- Jun 21, 2022
-
-
Simon Pilgrim authored
If the LHS op has a single use then using the more general AND op is likely to allow commutation, load folding, generic folds etc.
-
David Green authored
D127680 added some unnecessary funnel shift costs for AArch64 to "match the legacy behaviour". The default costs are closer to the correct values and line up with the scalar/neon costs better. Remove the lines again to clean up the code, they can be added back at a later date with better values if needed.
-
Nikita Popov authored
Tests were updated with this script: https://gist.github.com/nikic/98357b71fd67756b0f064c9517b62a34 However, in this case a lot of fixup was required, due to many minor, but ultimately immaterial differences in results. In particular, the GEP representation changes slightly in many cases, either because we now use an i8 GEP, or because we now leave a GEP alone, using it's original index types and (lack of) inbounds. basictest-opaque-ptrs.ll has been dropped, because it was an opaque pointers duplicate of basictest.ll.
-
Simon Pilgrim authored
This requires us to override the isTargetCanonicalConstantNode callback introduced in D128144, so we can recognise the various cases where a VBROADCAST_LOAD constant is being reused at different vector widths to prevent infinite loops.
-
Martin Storsjö authored
This reverts commit 9ffeaaa0. This fixes debugging large executables with lldb and gdb. When StringTableBuilder is used, the string offsets for any string can point anywhere in the string table - while previously, all strings were inserted in order (without deduplication and tail merging). For symbols, there's no complications in encoding the string offset; the offset is encoded as a raw 32 bit binary number in half of the symbol name field. For sections, the string table offset is written as "/<decimaloffset>", but if the decimal offset would be larger than 7 digits, it's instead written as "//<base64offset>". Tools that operate on object files can handle the base64 offset format, but apparently neither lldb nor gdb expect that syntax when locating the debug information section. Prior to the reverted commit, all long section names were located at the start of the string table, so their offset never exceeded the range for the decimal syntax. Just reverting this change for now, as the actual benefit from it was fairly modest. Longer term, lld could write all long section names unoptimized at the start of the string table, followed by all the strings for symbol names, with deduplication and tail merging. And lldb and gdb could be fixed to handle sections with the base64 offset syntax. This fixes https://github.com/mstorsjo/llvm-mingw/issues/289.
-
Nikita Popov authored
-
Shraiysh Vaishay authored
[mlir][OpenMP][NFC] Parameter refers to single args and hence changing description for taskgroup allocate clause.
-
Florian Hahn authored
-
Balazs Benics authored
There was a copy-paste mistake at the embedded link: `clang-tidy/checks/cppcoreguidelines-virtual-class-destructor` -> `clang-tidy/checks/cppcoreguidelines/virtual-class-destructor` Sphinx error: /home/zbebnal/git/llvm-project/clang-tools-extra/docs/ReleaseNotes.rst:168:unknown document: clang-tidy/checks/cppcoreguidelines-virtual-class-destructor Build bot: https://lab.llvm.org/buildbot#builders/115/builds/29805 Differential Revision: https://reviews.llvm.org/D126891
-
Balazs Benics authored
The `cppcoreguidelines-virtual-class-destructor` supposed to enforce http://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines#c35-a-base-class-destructor-should-be-either-public-and-virtual-or-protected-and-non-virtual Quote: > A **base** class destructor should be either public and virtual, or > protected and non-virtual [emphasis mine] However, this check still rules the following case: class MostDerived final : public Base { public: MostDerived() = default; ~MostDerived() = default; void func() final; }; Even though `MostDerived` class is marked `final`, thus it should not be considered as a **base** class. Consequently, the rule is satisfied, yet the check still flags this code. In this patch, I'm proposing to ignore `final` classes since they cannot be //base// classes. Reviewed By: whisperity Differential Revision: https://reviews.llvm.org/D126891
-
David Green authored
As a followup to D128144, this adds extract(DUP(C)) as a canonical constant to prevent it being transformed back into a BUILD_VECTOR, leading to an infinite loop.
-
Carl Ritson authored
If shader only has depth exports use MRTZ otherwise use MRT0. Differential Revision: https://reviews.llvm.org/D128185
-
Nicolas Vasilache authored
Differential Revision: https://reviews.llvm.org/D128247
-
Matthias Springer authored
Differential Revision: https://reviews.llvm.org/D128160
-
Nicolas Vasilache authored
This revision adds the necessary plumbing for canonicalizing scf::ForeachThread with the `AffineOpSCFCanonicalizationPattern`. In the process the `loopMatcher` helper is updated to take OpFoldResult instead of just values. This allows composing various scenarios without the need for an artificial builder. Differential Revision: https://reviews.llvm.org/D128244
-
Balázs Kéri authored
This updates StdLibraryFunctionsChecker to set the state of 'errno' by using the new errno_modeling functionality. The errno value is set in the PostCall callback. Setting it in call::Eval did not work for some reason and then every function should be EvalCallAsPure which may be bad to do. Now the errno value and state is not allowed to be checked in any PostCall checker callback because it is unspecified if the errno was set already or will be set later by this checker. Reviewed By: martong, steakhal Differential Revision: https://reviews.llvm.org/D125400
-
Kazu Hirata authored
-
Nikolas Klauser authored
Reviewed By: ldionne, Mordante, #libc Spies: jwakely, libcxx-commits Differential Revision: https://reviews.llvm.org/D127387
-
Kazu Hirata authored
-
Markus Lavin authored
Do not include debug instructions when comparing block sizes with thresholds. Differential Revision: https://reviews.llvm.org/D127208
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Chen Zheng authored
-
Argyrios Kyrtzidis authored
-
Shraiysh Vaishay authored
This patch adds omp.taskgroup operation according to OpenMP 5.0 2.17.6. Also added tests for the same. Reviewed By: kiranchandramohan, peixin Differential Revision: https://reviews.llvm.org/D127250
-
Shraiysh Vaishay authored
This patch fixes the unintentional data race in firstprivate implementation. There is a Read-Write race when one thread tries to copy the value inside the omp.parallel region while other thread modifies it from inside the region (using pointers or some other form of indirect access). For detailed discussion please refer to [[ https://discourse.llvm.org/t/issues-with-the-current-implementation-of-privatization-in-openmp-with-fortran/62335 | discourse ]]. Reviewed By: kiranchandramohan, peixin, NimishMishra Differential Revision: https://reviews.llvm.org/D125689
-
Argyrios Kyrtzidis authored
To accomodate macOS universal configuration include the assembly files and `blake3_neon.c` without a CMake check but instead guard their source with architecture "#ifdef" checks. Differential Revision: https://reviews.llvm.org/D128132
-
Craig Topper authored
Use it along with RISCVISD::HI and ADD_LO to avoid emitting MachineSDNodes during lowering.
-
Craig Topper authored
The failure that caused the previous revert has been fixed by https://reviews.llvm.org/D126048 Original commit message: RVV makes heavy use of subregisters due to LMUL>1 and segment load/store tuples. Enabling subregister liveness tracking improves the quality of the register allocation. I've added a command line that can be used to turn it off if it causes compile time or functional issues. I used the command line to keep the old behavior for one interesting test case that was testing register allocation. Reviewed By: kito-cheng Differential Revision: https://reviews.llvm.org/D128016
-
Serguei Katkov authored
There is no instruction to fold NZCV, so, just do not do it. Without the fix the added test case crashes with an assert "Mismatched register size in non subreg COPY" Reviewed By: danilaml Subscribers: llvm-commits Differential Revision: https://reviews.llvm.org/D127294
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
LLVM GN Syncbot authored
-
Chen Zheng authored
This patch implements a new way to generate the CTR loops. Now the intrinsics inserted in hardware loop pass will be mapped to pseudo instructions and these pseudo instructions will be expanded to CTR loop or normal compare+branch loop in this post ISEL pass. Reviewed By: lkail Differential Revision: https://reviews.llvm.org/D122125
-
Craig Topper authored
Use it in place of VSELECT_VL+VRGATHER*_VL. This simplifies the isel patterns. Overall, I think trying to match select+op to create masked instructions in isel doesn't scale. We either need to do it in DAG combine, pre-isel peepole, or post-isel peephole. I don't yet know which is the right answer, but for this case it seemed best to be able to request the masked form directly from lowering. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D128023
-
chenglin.bi authored
When already have (op N0, N2), reassociate (op (op N0, N1), N2) to (op (op N0, N2), N1) to reuse the exist (op N0, N2) Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D122539
-
Luo, Yuanke authored
For below case, virtual register is defined twice in the self loop. We don't need to spill %0 after the third instruction `%0 = def (tied %0)`, because it is defined in the second instruction `%0 = def`. 1 bb.1 2 %0 = def 3 %0 = def (tied %0) 4 ... 5 jmp bb.1 Reviewed By: MatzeB Differential Revision: https://reviews.llvm.org/D125079
-
Mogball authored
Depends on D127373 Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D127375
-
Phoebe Wang authored
This fixes issue #56103. Reviewed By: mingmingl Differential Revision: https://reviews.llvm.org/D128122
-