- Jun 14, 2023
-
-
Krzysztof Parzyszek authored
-
Krzysztof Parzyszek authored
-
Francesco Petrogalli authored
As reported in https://github.com/llvm/llvm-project/issues/63225, we need to make sure we can use the `&operator<<` on instances of the `ResourceSegments` class for builds that set `NDEBUG`. Reviewed By: sylvestre.ledru Differential Revision: https://reviews.llvm.org/D152817
-
Amaury Séchet authored
-
Nikita Popov authored
Fix the infinite loop reported on https://reviews.llvm.org/D151807#4420467. collectShuffleElements() will widen vectors and replace extracts via replaceExtractElements(), to allow the next call of collectShuffleElements() to fold. However, it's possible for another fold to run first, and break the expected sequence again. To ensure this does not happen, directly rerun the collectShuffleElements() fold if we have adjusted extracts.
-
Kristina Bessonova authored
RFC https://discourse.llvm.org/t/rfc-dwarfdebug-fix-and-improve-handling-imported-entities-types-and-static-local-in-subprogram-and-lexical-block-scopes/68544 !Note! Extracted from the following patch for review purpose only, should be squashed with the next patch (D144004) before committing. Currently the back-end emits imported entities in `DwarfDebug::beginModule()`. However in case an imported declaration is a function, it must point to an abstract subprogram if it exists (see PR51501). But in `DwarfDebug::beginModule()` the DWARF generator doesn't have information to identify if an abstract subprogram needs to be created. Only by entering `DwarfDebug::endModule()` all subprograms are processed, so it's clear which subprogram DIE should be referred to. Hence, the patch moves the emission there. The patch is need to fix PR51501, but it only does the preliminary work. Since it changes the order of debug entities ...
-
Takuya Shimizu authored
BEFORE this patch, unused const-qualified variable templates such as `template <typename T> const double var_t = 0;` were diagnosed as `unused variable 'var_t'` This patch fixes this message to `unused variable template 'var_t'` Differential Revision: https://reviews.llvm.org/D152796
-
Krishna Narayanan authored
The tautological comparison warning was not properly looking through parenthesized expressions, which is now fixed. Fixes https://github.com/llvm/llvm-project/issues/42992 Differential Revision: https://reviews.llvm.org/D149000
-
Jay Foad authored
D135579 added support for fixedStack, but did not cope with the output of -simplify-mir which does not include the fixedStack section by default. Differential Revision: https://reviews.llvm.org/D152896
-
Simon Pilgrim authored
-
Matt Arsenault authored
For cases where we cannot insert an addrspacecast, we can still expand like a memcpy if we know the address spaces cannot alias. Normally non-aliasing memmoves are optimized to memcpy, but we cannot rely on that for lowering. If a target has aliasing address spaces that cannot be casted between, we still have to give up lowering this.
-
Simon Pilgrim authored
-
Simon Pilgrim authored
[X86] X86FixupVectorConstantsPass - attempt to replace full width integer vector constant loads with broadcasts on AVX2+ targets (REAPPLIED) lowerBuildVectorAsBroadcast will not broadcast splat constants in all cases, resulting in a lot of situations where a full width vector load that has failed to fold but is loading splat constant values could use a broadcast load instruction just as cheaply, and save constant pool space. This is an updated commit of ab4b9248 after being reverted at 78de45fd
-
Nemanja Ivanovic authored
The commmit added clang-tidy checks without adding the required library to the link step. Caused failures with shared library builds.
-
Ivan Kosarev authored
Removes the need to add and remove them manually depending on whether they are used in cvt*() functions. Also removes the compiler warnings about unused handlers when it happens to be the case. Part of <https://github.com/llvm/llvm-project/issues/62629>. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D151688
-
Jay Foad authored
-
Mikael Holmen authored
We don't want the existence of debug instructions affect codegen so we now ignore debug instructions and other "isAssumeLikeIntrinsics in the "extend schedule region" search loop in BoUpSLP::BlockScheduling::extendSchedulingRegion. Differential Revision: https://reviews.llvm.org/D152441
-
Ivan Kosarev authored
Reviewed By: Joe_Nash Differential Revision: https://reviews.llvm.org/D152715
-
Simon Pilgrim authored
-
Guillaume Chatelet authored
Failing buildbot https://lab.llvm.org/buildbot/#/builders/73/builds/49707 This reverts commit 9a7b4c93.
-
Guillaume Chatelet authored
This patch mimics the behavior of Google Test and allow users to log custom messages after all flavors of ASSERT_ / EXPECT_. Reviewed By: sivachandra, lntue Differential Revision: https://reviews.llvm.org/D152630
-
Simon Pilgrim authored
It looks like we were trying to account for SLM costs, which are actually handled separately Fixes #62969
-
Simon Pilgrim authored
Addresses part of Issue #62969 - if the upper 32-bits of the vXi64 elements are known to be zero, then a multiply simplifies to a single (fast) PMULUDQ instruction We still have the problem that minRequiredElementSize can't determine that the upper bits are zero for the test case from Issue #62969 - I'll take a look at that next.
-
Théo Degioanni authored
This revision introduces support for memset intrinsics in SROA and mem2reg for the LLVM dialect. This is achieved for SROA by breaking memsets of aggregates into multiple memsets of scalars, and for mem2reg by promoting memsets of single integer slots into the value the memset operation would yield. The SROA logic supports breaking memsets of static size operating at the start of a memory slot. The intended most common case is for memsets covering the entirety of a struct, most often as a way to initialize it to 0. The mem2reg logic supports dynamic values and static sizes as input to promotable memsets. This is achieved by lowering memsets into `ceil(log_2(n))` LeftShift operations, `ceil(log_2(n))` Or operations and up to one ZExt operation (for n the byte width of the integer), computing in registers the integer value the memset would create. Only byte-aligned integers are supported, more types could easily be added afterwards. Reviewed By: gysit Dif...
-
Cullen Rhodes authored
Apologies I shouldn't have comitted this, need to wait until the planned MLIR ODM: https://discourse.llvm.org/t/rfc-creating-a-armsme-dialect/67208/76 This reverts commit a48fe898.
-
Cullen Rhodes authored
This patch adds support for lowering a `vector.transfer_write` of zeroes and type `vector<[16x16]xi8>` to the SME `zero {za}` instruction [1], which zeroes the entire accumulator. This contributes to supporting a path from `linalg.fill` to SME. [1] https://developer.arm.com/documentation/ddi0602/2022-06/SME-Instructions/ZERO--Zero-a-list-of-64-bit-element-ZA-tiles- Reviewed By: awarzynski, dcaballe Differential Revision: https://reviews.llvm.org/D152508 -
Guillaume Chatelet authored
Most of the time `memmove` is called on buffers that are disjoint, in that case we can use `memcpy` which is faster. The additional test is branchless on x86, aarch64 and RISCV with the zbb extension (bitmanip). On x86 this patch adds a latency of 2 to 3 cycles. Before ``` -------------------------------------------------------------------------------- Benchmark Time CPU Iterations UserCounters... -------------------------------------------------------------------------------- BM_Memmove/0/0_median 5.00 ns 5.00 ns 10 bytes_per_cycle=1.25477/s bytes_per_second=2.62933G/s items_per_second=199.87M/s __llvm_libc::memmove,memmove Google A BM_Memmove/1/0_median 6.21 ns 6.21 ns 10 bytes_per_cycle=3.22173/s bytes_per_second=6.75106G/s items_per_second=160.955M/s __llvm_libc::memmove,memmove Google B BM_Memmove/2/0_median 8.09 ns 8.09 ns 10 bytes_per_cycle=5.31462/s bytes_per_second=11.1366G/s items_per_second=123.603M/s __llvm_libc::memmove,memmove Google D BM_Memmove/3/0_median 5.95 ns 5.95 ns 10 bytes_per_cycle=2.71865/s bytes_per_second=5.69687G/s items_per_second=167.967M/s __llvm_libc::memmove,memmove Google L BM_Memmove/4/0_median 5.63 ns 5.63 ns 10 bytes_per_cycle=2.28294/s bytes_per_second=4.78383G/s items_per_second=177.615M/s __llvm_libc::memmove,memmove Google M BM_Memmove/5/0_median 5.68 ns 5.68 ns 10 bytes_per_cycle=2.16798/s bytes_per_second=4.54295G/s items_per_second=176.015M/s __llvm_libc::memmove,memmove Google Q BM_Memmove/6/0_median 7.46 ns 7.46 ns 10 bytes_per_cycle=3.97619/s bytes_per_second=8.332G/s items_per_second=134.044M/s __llvm_libc::memmove,memmove Google S BM_Memmove/7/0_median 5.40 ns 5.40 ns 10 bytes_per_cycle=1.79695/s bytes_per_second=3.76546G/s items_per_second=185.211M/s __llvm_libc::memmove,memmove Google U BM_Memmove/8/0_median 5.62 ns 5.62 ns 10 bytes_per_cycle=3.18747/s bytes_per_second=6.67927G/s items_per_second=177.983M/s __llvm_libc::memmove,memmove Google W BM_Memmove/9/0_median 101 ns 101 ns 10 bytes_per_cycle=9.77359/s bytes_per_second=20.4803G/s items_per_second=9.9333M/s __llvm_libc::memmove,uniform 384 to 4096 ``` After ``` BM_Memmove/0/0_median 3.57 ns 3.57 ns 10 bytes_per_cycle=1.71375/s bytes_per_second=3.59112G/s items_per_second=280.411M/s __llvm_libc::memmove,memmove Google A BM_Memmove/1/0_median 4.52 ns 4.52 ns 10 bytes_per_cycle=4.47557/s bytes_per_second=9.37843G/s items_per_second=221.427M/s __llvm_libc::memmove,memmove Google B BM_Memmove/2/0_median 5.70 ns 5.70 ns 10 bytes_per_cycle=7.37396/s bytes_per_second=15.4519G/s items_per_second=175.399M/s __llvm_libc::memmove,memmove Google D BM_Memmove/3/0_median 4.47 ns 4.47 ns 10 bytes_per_cycle=3.4148/s bytes_per_second=7.15563G/s items_per_second=223.743M/s __llvm_libc::memmove,memmove Google L BM_Memmove/4/0_median 4.53 ns 4.53 ns 10 bytes_per_cycle=2.86071/s bytes_per_second=5.99454G/s items_per_second=220.69M/s __llvm_libc::memmove,memmove Google M BM_Memmove/5/0_median 4.19 ns 4.19 ns 10 bytes_per_cycle=2.5484/s bytes_per_second=5.3401G/s items_per_second=238.924M/s __llvm_libc::memmove,memmove Google Q BM_Memmove/6/0_median 5.02 ns 5.02 ns 10 bytes_per_cycle=5.94164/s bytes_per_second=12.4505G/s items_per_second=199.14M/s __llvm_libc::memmove,memmove Google S BM_Memmove/7/0_median 4.03 ns 4.03 ns 10 bytes_per_cycle=2.47028/s bytes_per_second=5.17641G/s items_per_second=247.906M/s __llvm_libc::memmove,memmove Google U BM_Memmove/8/0_median 4.70 ns 4.70 ns 10 bytes_per_cycle=3.84975/s bytes_per_second=8.06706G/s items_per_second=212.72M/s __llvm_libc::memmove,memmove Google W BM_Memmove/9/0_median 90.7 ns 90.7 ns 10 bytes_per_cycle=10.8681/s bytes_per_second=22.7739G/s items_per_second=11.02M/s __llvm_libc::memmove,uniform 384 to 4096 ``` Reviewed By: courbet Differential Revision: https://reviews.llvm.org/D152811
-
Vitaly Buka authored
-
Carl Ritson authored
-
Nikita Popov authored
Just naming changes.
-
Nikita Popov authored
These two helpers also decrement the use count of the replaced operand, so give them the same treatment as eraseInstruction().
-
Vitaly Buka authored
-
Nikita Popov authored
This no longer creates a bitcast, just changes the element type of the ConstantAddress.
-
Joshua Cao authored
If a select's condition is a AND/OR, we can unswitch invariant operands. This patch uses existing logic from unswitching AND/OR's for branch conditions. This patch fixes the Cost computation for unswitching selects to have the cost of the entire loop, since unswitching selects do not remove branches. This is required for this patch because otherwise, there are cases where unswitching selects of AND/OR is beating out unswitching of branches. This patch also prevents unswitching of logical AND/OR selects. This should instead be done by unswitching of AND/OR branch conditions. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D151677
-
Joshua Cao authored
-
Matthias Springer authored
The destination operand does not bufferize to a memory read if it is completely overwritten. Differential Revision: https://reviews.llvm.org/D152823
-
Michael Platings authored
-
Nikita Popov authored
Many folds in InstCombine are limited to one-use instructions. For that reason, if the use-count of an instruction drops to one, it makes sense to revisit that one user. This is one of the most common reasons why InstCombine fails to finish in a single iteration. Doing this revisit actually slightly improves compile-time, because we save an extra InstCombine iteration in enough cases to make a visible difference. This is conceptually NFC, but not NFC in practice, because differences in worklist order can result in slightly different folding behavior. The regressed tests in or-shifted-masks.ll now require a sequence of instcombine,early-cse,instcombine to fold fully. D152876 would make these fold in a single instcombine run again. Differential Revision: https://reviews.llvm.org/D151807
-
eopXD authored
This is the 11th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152078. This patch also fixes the suffix for non-overloaded variants for vset on tuple types. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152079
-
eopXD authored
This is the 10th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152077. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152078
-