- Jun 21, 2023
-
-
Mark de Wever authored
All supported compilers support the -std=year except Clang < 17 which doesn't support C++23. This fallback is marked for removal once we no longer support Clang 16. Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D152106
-
Mark de Wever authored
The mask used to check whether a code unit is a valid continuation was incorrect and accepts non-continuation code points. This fixes the issue. Reviewed By: ldionne, tahonermann, #libc Differential Revision: https://reviews.llvm.org/D149672
-
Mark de Wever authored
This version allowed testing the std module in C++26 mode. Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D153227
-
Anna Thomas authored
{mini|maxi}mum intrinsics are different from {min|max}num intrinsics in the propagation of NaN and signed zero. Also, the minnum/maxnum intrinsics require the presence of nsz flags to be valid reductions in vectorizer. In this regard, we introduce a new recurrence kind and also add support for identifying reduction patterns using these intrinsics. The reduction intrinsics and lowering was introduced here: 26bfbec5. There are tests added which show how this interacts across chains of min/max patterns. Differential Revision: https://reviews.llvm.org/D151482 -
Nico Weber authored
2700da5f added lld/test/Unit/lit.site.cfg.py.in in a state that half-supports relocatable lld lit tests. Make them fully relocatable. See description of fb80b6b2 for background. Differential Revision: https://reviews.llvm.org/D152885
-
Matt Arsenault authored
The optimizing, non-broken features have all been moved to AMDGPUAttributor. The only remaining piece of functionality was the broken propagation of the wavesize features. This was fundamentally broken and a hack for device library linking. It doesn't matter when the device libraries are correctly linked and internalized. In case of linked-as-normal-bitcode (as comgr still does), we're reliant on the global subtarget anyway. If we can get away without forcing target-cpu, we should just as well be able to get away without propagating target-features.
-
Amy Kwan authored
This patch adds support for the TLS local-exec access model on AIX to allow for the ability to generate the 32-bit (specifically, non-optimized) code sequence. This work is a follow up of D149722. The particular sequence that is generated for this sequence is as follows: ``` .tc var[TC],var[TL]@le. // variable offset, with the le relocation specifier bla .__get_tpointer() // get the thread pointer, modifies r3 lwz reg1, var[TC](2) // load the variable offset add reg2, r3, reg1 // add the variable offset to the retrieved thread pointer ``` Differential Revision: https://reviews.llvm.org/D152669
-
Craig Topper authored
Unless I'm missing something we need to update the whole vector not just where OddMask is true. Reviewed By: luke Differential Revision: https://reviews.llvm.org/D153087
-
Simi Pallipurath authored
This reverts commit 8cf89568.
-
Matt Arsenault authored
This was looking for full copies produced by SplitKit, but SplitKit introduces copy bundles if not all lanes are live. The scan for uses needs to look at bundles, not individual instructions. This is a prerequisite to avoiding some redundant spills due to subregisters which will help avoid an allocation failure in a future patch.
-
Matt Arsenault authored
-
Matt Devereau authored
This patch renames load/store enums to be equivalent to the contiguous addressing modes with _STRIDED and _STRIDED_IMM suffixed.
-
- Jun 20, 2023
-
-
Nikita Popov authored
-
Nikita Popov authored
-
Nikita Popov authored
-
Owen Pan authored
Also, reformat all clang-format related files. Differential Revision: https://reviews.llvm.org/D153208
-
Louis Dionne authored
Otherwise, Clang complains about format_ctx being unused in the tests when exceptions are disabled in Freestanding mode. I don't know why it doesn't complain not in freestanding mode. Differential Revision: https://reviews.llvm.org/D153301
-
Nikita Popov authored
-
Nikita Popov authored
-
Mark de Wever authored
Changes to preconditions have no impact on the library. Implements - LWG3927 Unclear preconditions for operator[] for sequence containers Reviewed By: #libc, philnik Differential Revision: https://reviews.llvm.org/D153286
-
Mark de Wever authored
Note libc++ implemented this in its initial version. Implements - LWG3935 template<class X> constexpr complex& operator=(const complex<X>&) has no specification Reviewed By: #libc, philnik, ldionne Differential Revision: https://reviews.llvm.org/D153287
-
Nikita Popov authored
-
Simon Pilgrim authored
Wrap the getSExtOrTrunc/getZExtOrTrunc calls behind an IsSigned argument.
-
Nikita Popov authored
Commit the bitcode file instead of generating it dynamically, as this will no longer be possible in the future.
-
Louis Dionne authored
This fixes rdar://110330781, which asked for the feature-test macro for std::pmr to take into account the deployment target. It doesn't fix https://llvm.org/PR62212, though, because the availability markup itself must be disabled until some Clang bugs have been fixed. This is pretty vexing, however at least everything should work once those Clang bugs have been fixed. In the meantime, this patch at least adds the required markup (as disabled) and ensures that the feature-test macro for std::pmr is aware of the deployment target requirement. Differential Revision: https://reviews.llvm.org/D135813
-
Nikita Popov authored
-
Nikita Popov authored
These tests were testing both typed an opaque pointers. Only keep opaque pointers tests.
-
Nikita Popov authored
Commit the typed pointer bitcode file instead of producing it, as this will not be possible in the future anymore.
-
Simon Pilgrim authored
This is a generic DAG combine version of D151055 which recognizes when a signed ABDS can be safely replaced with a unsigned ABDU instruction if it is legal. Alive2: https://alive2.llvm.org/ce/z/pb5BjG Differential Revision: https://reviews.llvm.org/D153328
-
David Green authored
-
Jingu Kang authored
Add tablegen pattern for uaddlv(uaddlp(x)) ==> uaddlv(x). Differential Revision: https://reviews.llvm.org/D153323
-
Stephen Thomas authored
Differential Revision: https://reviews.llvm.org/D153322
-
Weining Lu authored
This patch optimizes code generation by leveraging the zeroing behavior of the `maskeqz`/`masknez` instructions. ``` int sel(int a, int b) { return (a < b) ? a : 0; } ``` ``` slt $a1,$a0,$a1 masknez $a2,$r0,$a1 maskeqz $a0,$a0,$a1 or $a0,$a0,$a2 ``` => ``` slt $a1,$a0,$a1 maskeqz $a0,$a0,$a1 ``` Reviewed By: SixWeining Differential Revision: https://reviews.llvm.org/D153193 -
Pravin Jagtap authored
Verifying dominator tree is expensive using intra-pass asserts. Asserts added during D147408 are increasing the build time of libc significantly. This change does the verification after the atomic optimizer pass and should fix the regression reported in D153232. Reviewed By: arsenm, #amdgpu Differential Revision: https://reviews.llvm.org/D153261
-
Weining Lu authored
-
Tue Ly authored
Re-organize special cases and add a special case when `|x| < 2^-5`. Reviewed By: michaelrj Differential Revision: https://reviews.llvm.org/D153134
-
Tue Ly authored
Re-order exceptional branches and slightly adjust the evaluation. Depends on https://reviews.llvm.org/D153026 . Reviewed By: michaelrj Differential Revision: https://reviews.llvm.org/D153062
-
Tue Ly authored
Re-order exceptional branches and slightly adjust the evaluation. Performance tested with the CORE-MATH project on AMD EPYC 7B12 (clocks/op) Reciprocal throughputs: ``` --- BEFORE --- $ CORE_MATH_PERF_MODE=rdtsc ./perf.sh tanhf [####################] 100 % (with -mavx2 -mfma) Ntrial = 20 ; Min = 7.794 + 0.102 clc/call; Median-Min = 0.066 clc/call; Max = 8.267 clc/call; [####################] 100 %. (with -msse4.2) Ntrial = 20 ; Min = 10.783 + 0.172 clc/call; Median-Min = 0.144 clc/call; Max = 11.446 clc/call; [####################] 100 %. (SSE2) Ntrial = 20 ; Min = 18.926 + 0.381 clc/call; Median-Min = 0.342 clc/call; Max = 19.623 clc/call; --- AFTER --- $ CORE_MATH_PERF_MODE=rdtsc ./perf.sh tanhf [####################] 100 % (with -mavx2 -mfma) Ntrial = 20 ; Min = 6.598 + 0.085 clc/call; Median-Min = 0.052 clc/call; Max = 6.868 clc/call; [####################] 100 % (with -msse4.2) Ntrial = 20 ; Min = 9.245 + 0.304 clc/call; Median-Min = 0.248 clc/call; Max = 10.675 clc/call; [####################] 100 %. (SSE2) Ntrial = 20 ; Min = 11.724 + 0.440 clc/call; Median-Min = 0.444 clc/call; Max = 12.262 clc/call; ``` Latency: ``` --- BEFORE --- $ PERF_ARGS="--latency" CORE_MATH_PERF_MODE=rdtsc ./perf.sh tanhf [####################] 100 % (with -mavx2 -mfma) Ntrial = 20 ; Min = 38.821 + 0.157 clc/call; Median-Min = 0.122 clc/call; Max = 39.539 clc/call; [####################] 100 %. (with -msse4.2) Ntrial = 20 ; Min = 44.767 + 0.766 clc/call; Median-Min = 0.681 clc/call; Max = 45.951 clc/call; [####################] 100 %. (SSE2) Ntrial = 20 ; Min = 55.055 + 1.512 clc/call; Median-Min = 1.571 clc/call; Max = 57.039 clc/call; --- AFTER --- $ PERF_ARGS="--latency" CORE_MATH_PERF_MODE=rdtsc ./perf.sh tanhf [####################] 100 % (with -mavx2 -mfma) Ntrial = 20 ; Min = 36.147 + 0.194 clc/call; Median-Min = 0.181 clc/call; Max = 36.536 clc/call; [####################] 100 % (with -msse4.2) Ntrial = 20 ; Min = 40.904 + 0.728 clc/call; Median-Min = 0.557 clc/call; Max = 42.231 clc/call; [####################] 100 %. (SSE2) Ntrial = 20 ; Min = 55.776 + 0.557 clc/call; Median-Min = 0.542 clc/call; Max = 56.551 clc/call; ``` Reviewed By: michaelrj Differential Revision: https://reviews.llvm.org/D153026
-
Weining Lu authored
[LoongArch] Add missing chains and remove unnecessary `SDNPSideEffect` property for some intrinsic nodes
-
David Green authored
Similar to the existing f32 pattern, this adds a tablegen pattern for the fp16 fcvtn2.
-