- Jun 14, 2023
-
-
Simon Pilgrim authored
It looks like we were trying to account for SLM costs, which are actually handled separately Fixes #62969
-
Simon Pilgrim authored
Addresses part of Issue #62969 - if the upper 32-bits of the vXi64 elements are known to be zero, then a multiply simplifies to a single (fast) PMULUDQ instruction We still have the problem that minRequiredElementSize can't determine that the upper bits are zero for the test case from Issue #62969 - I'll take a look at that next.
-
Théo Degioanni authored
This revision introduces support for memset intrinsics in SROA and mem2reg for the LLVM dialect. This is achieved for SROA by breaking memsets of aggregates into multiple memsets of scalars, and for mem2reg by promoting memsets of single integer slots into the value the memset operation would yield. The SROA logic supports breaking memsets of static size operating at the start of a memory slot. The intended most common case is for memsets covering the entirety of a struct, most often as a way to initialize it to 0. The mem2reg logic supports dynamic values and static sizes as input to promotable memsets. This is achieved by lowering memsets into `ceil(log_2(n))` LeftShift operations, `ceil(log_2(n))` Or operations and up to one ZExt operation (for n the byte width of the integer), computing in registers the integer value the memset would create. Only byte-aligned integers are supported, more types could easily be added afterwards. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D152367
-
Cullen Rhodes authored
Apologies I shouldn't have comitted this, need to wait until the planned MLIR ODM: https://discourse.llvm.org/t/rfc-creating-a-armsme-dialect/67208/76 This reverts commit a48fe898.
-
Cullen Rhodes authored
This patch adds support for lowering a `vector.transfer_write` of zeroes and type `vector<[16x16]xi8>` to the SME `zero {za}` instruction [1], which zeroes the entire accumulator. This contributes to supporting a path from `linalg.fill` to SME. [1] https://developer.arm.com/documentation/ddi0602/2022-06/SME-Instructions/ZERO--Zero-a-list-of-64-bit-element-ZA-tiles- Reviewed By: awarzynski, dcaballe Differential Revision: https://reviews.llvm.org/D152508 -
Guillaume Chatelet authored
Most of the time `memmove` is called on buffers that are disjoint, in that case we can use `memcpy` which is faster. The additional test is branchless on x86, aarch64 and RISCV with the zbb extension (bitmanip). On x86 this patch adds a latency of 2 to 3 cycles. Before ``` -------------------------------------------------------------------------------- Benchmark Time CPU Iterations UserCounters... -------------------------------------------------------------------------------- BM_Memmove/0/0_median 5.00 ns 5.00 ns 10 bytes_per_cycle=1.25477/s bytes_per_second=2.62933G/s items_per_second=199.87M/s __llvm_libc::memmove,memmove Google A BM_Memmove/1/0_median 6.21 ns 6.21 ns 10 bytes_per_cycle=3.22173/s bytes_per_second=6.75106G/s items_per_second=160.955M/s __llvm_libc::memmove,memmove Google B BM_Memmove/2/0_median 8.09 ns 8.09 ns 10 bytes_per_cycle=5.31462/s bytes_per_second=11.1366G/s items_per_second=123.603M/s __llvm_libc::memmove,memmove Google D BM_Memmove/3/0_median 5.95 ns 5.95 ns 10 bytes_per_cycle=2.71865/s bytes_per_second=5.69687G/s items_per_second=167.967M/s __llvm_libc::memmove,memmove Google L BM_Memmove/4/0_median 5.63 ns 5.63 ns 10 bytes_per_cycle=2.28294/s bytes_per_second=4.78383G/s items_per_second=177.615M/s __llvm_libc::memmove,memmove Google M BM_Memmove/5/0_median 5.68 ns 5.68 ns 10 bytes_per_cycle=2.16798/s bytes_per_second=4.54295G/s items_per_second=176.015M/s __llvm_libc::memmove,memmove Google Q BM_Memmove/6/0_median 7.46 ns 7.46 ns 10 bytes_per_cycle=3.97619/s bytes_per_second=8.332G/s items_per_second=134.044M/s __llvm_libc::memmove,memmove Google S BM_Memmove/7/0_median 5.40 ns 5.40 ns 10 bytes_per_cycle=1.79695/s bytes_per_second=3.76546G/s items_per_second=185.211M/s __llvm_libc::memmove,memmove Google U BM_Memmove/8/0_median 5.62 ns 5.62 ns 10 bytes_per_cycle=3.18747/s bytes_per_second=6.67927G/s items_per_second=177.983M/s __llvm_libc::memmove,memmove Google W BM_Memmove/9/0_median 101 ns 101 ns 10 bytes_per_cycle=9.77359/s bytes_per_second=20.4803G/s items_per_second=9.9333M/s __llvm_libc::memmove,uniform 384 to 4096 ``` After ``` BM_Memmove/0/0_median 3.57 ns 3.57 ns 10 bytes_per_cycle=1.71375/s bytes_per_second=3.59112G/s items_per_second=280.411M/s __llvm_libc::memmove,memmove Google A BM_Memmove/1/0_median 4.52 ns 4.52 ns 10 bytes_per_cycle=4.47557/s bytes_per_second=9.37843G/s items_per_second=221.427M/s __llvm_libc::memmove,memmove Google B BM_Memmove/2/0_median 5.70 ns 5.70 ns 10 bytes_per_cycle=7.37396/s bytes_per_second=15.4519G/s items_per_second=175.399M/s __llvm_libc::memmove,memmove Google D BM_Memmove/3/0_median 4.47 ns 4.47 ns 10 bytes_per_cycle=3.4148/s bytes_per_second=7.15563G/s items_per_second=223.743M/s __llvm_libc::memmove,memmove Google L BM_Memmove/4/0_median 4.53 ns 4.53 ns 10 bytes_per_cycle=2.86071/s bytes_per_second=5.99454G/s items_per_second=220.69M/s __llvm_libc::memmove,memmove Google M BM_Memmove/5/0_median 4.19 ns 4.19 ns 10 bytes_per_cycle=2.5484/s bytes_per_second=5.3401G/s items_per_second=238.924M/s __llvm_libc::memmove,memmove Google Q BM_Memmove/6/0_median 5.02 ns 5.02 ns 10 bytes_per_cycle=5.94164/s bytes_per_second=12.4505G/s items_per_second=199.14M/s __llvm_libc::memmove,memmove Google S BM_Memmove/7/0_median 4.03 ns 4.03 ns 10 bytes_per_cycle=2.47028/s bytes_per_second=5.17641G/s items_per_second=247.906M/s __llvm_libc::memmove,memmove Google U BM_Memmove/8/0_median 4.70 ns 4.70 ns 10 bytes_per_cycle=3.84975/s bytes_per_second=8.06706G/s items_per_second=212.72M/s __llvm_libc::memmove,memmove Google W BM_Memmove/9/0_median 90.7 ns 90.7 ns 10 bytes_per_cycle=10.8681/s bytes_per_second=22.7739G/s items_per_second=11.02M/s __llvm_libc::memmove,uniform 384 to 4096 ``` Reviewed By: courbet Differential Revision: https://reviews.llvm.org/D152811
-
Vitaly Buka authored
-
Carl Ritson authored
-
Nikita Popov authored
Just naming changes.
-
Nikita Popov authored
These two helpers also decrement the use count of the replaced operand, so give them the same treatment as eraseInstruction().
-
Vitaly Buka authored
-
Nikita Popov authored
This no longer creates a bitcast, just changes the element type of the ConstantAddress.
-
Joshua Cao authored
If a select's condition is a AND/OR, we can unswitch invariant operands. This patch uses existing logic from unswitching AND/OR's for branch conditions. This patch fixes the Cost computation for unswitching selects to have the cost of the entire loop, since unswitching selects do not remove branches. This is required for this patch because otherwise, there are cases where unswitching selects of AND/OR is beating out unswitching of branches. This patch also prevents unswitching of logical AND/OR selects. This should instead be done by unswitching of AND/OR branch conditions. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D151677
-
Joshua Cao authored
-
Matthias Springer authored
The destination operand does not bufferize to a memory read if it is completely overwritten. Differential Revision: https://reviews.llvm.org/D152823
-
Michael Platings authored
-
Nikita Popov authored
Many folds in InstCombine are limited to one-use instructions. For that reason, if the use-count of an instruction drops to one, it makes sense to revisit that one user. This is one of the most common reasons why InstCombine fails to finish in a single iteration. Doing this revisit actually slightly improves compile-time, because we save an extra InstCombine iteration in enough cases to make a visible difference. This is conceptually NFC, but not NFC in practice, because differences in worklist order can result in slightly different folding behavior. The regressed tests in or-shifted-masks.ll now require a sequence of instcombine,early-cse,instcombine to fold fully. D152876 would make these fold in a single instcombine run again. Differential Revision: https://reviews.llvm.org/D151807
-
eopXD authored
This is the 11th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152078. This patch also fixes the suffix for non-overloaded variants for vset on tuple types. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152079
-
eopXD authored
This is the 10th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152077. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152078
-
eopXD authored
This is the 9th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152076. This patch expands all variants of indexed strided segment store. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vsuxseg2ei32, vsoxseg2ei32 is added for now. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152077 -
eopXD authored
This is the 8th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152075. This patch expands all variants of indexed strided segment load, including the policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vluxseg2ei32, vloxseg2ei32 is added for now. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152076 -
eopXD authored
This is the 7th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152074. This patch expands all variants for strided segment store. The store intrinsics does not have any policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vssseg2e32 is added for now. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152075 -
eopXD authored
This is the 6th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152073. This patch expands all variants of strided segment load, including the policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vlsseg2e32 is added. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152074 -
Chuanqi Xu authored
The new added test has problems on windows since the patch is about ABI and MSVC ABI is not covered. Skip the test on windows to make the CI green.
-
eopXD authored
This is the 5th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152072. This patch expands all variants of unit stride fault-first segment load, including the policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vlseg2e32ff is added. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152073 -
Zi Xuan Wu (Zeson) authored
Complete fp16 support by ensuring that load extension / truncate store operations are properly expanded.
-
Piyou Chen authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152506
-
Matthias Springer authored
`RewriterBase::Listener::notifyOperationReplaced` notifies observers that an op is about to be replaced with a range of values. This notification is not very useful for ops without results, because it does not specify the replacement op (and it cannot be deduced from the replacement values). It provides no additional information over the `notifyOperationRemoved` notification. This revision adds an additional notification when a rewriter replaces an op with another op. By default, this notification triggers the original "op replaced with values" notification, so there is no functional change for existing code. This new API is useful for the transform dialect, which needs to track op replacements. (Updated in a subsequent revision.) Also includes minor documentation improvements. Differential Revision: https://reviews.llvm.org/D152814
-
eopXD authored
This is the 4th patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152071. This patch expands all variants for unit stride segment store. The store intrinsics does not have any policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. For the same reason mentioned in [3/11], only full test case for vsseg2e32 is added. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152072 -
eopXD authored
This is the 3rd patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152070. This patch expands all variants of unit stride segment load, including the policy variants. This patch also fixes the trailing suffix in the intrinsics' function name that representing the return type, adding `x{NF}`. Currently the tuple type co-exists with the non-tuple type intrinsics. Since the co-existance is temporary, this patch only adds test cases of all variants for vlseg2e32 to show the capability done. Test cases of other data type and NF will be added in the patch-set when the replacement happens. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152071 -
eopXD authored
This is the 2nd patch of the patch-set. For the cover letter, please checkout D152069. Depends on D152069. This patch also removes redundant checks related to tuples and dedicate the check to happen in `RVVType::verifyType`. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D152070
-
Valentin Clement authored
Lower gang dim from the parse tree to the new MLIR representation. Depends on D151972 Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D151973
-
River Riddle authored
We're getting asserts for duplicate section registration during linking which stems back to these sections. From previous discussions, it seems like these are metadata sections that can be dropped. See the discussion in D116474 and https://bugs.llvm.org/show_bug.cgi?id=45111. Differential Revision: https://reviews.llvm.org/D152574
-
Michael Platings authored
Reviewed By: peter.smith, MaskRay Differential Revision: https://reviews.llvm.org/D143587
-
Michael Platings authored
This enables layering baremetal multilibs on top of each other. For example a multilib containing only a no-exceptions libc++ could be layered on top of a multilib containing C libs. This avoids the need to duplicate the C library for every libc++ variant. Differential Revision: https://reviews.llvm.org/D143075
-
Michael Platings authored
This will enable layering multilibs on top of each other. For example a multilib containing only a no-exceptions libc++ could be layered on top of a multilib containing C libs. This avoids the need to duplicate the C library for every libc++ variant. This change doesn't expose the functionality externally, it only opens the functionality up to be potentially used by ToolChain classes. Differential Revision: https://reviews.llvm.org/D143059
-
Michael Platings authored
The default location for multilib.yaml is lib/clang-runtimes, without any target-specific suffix. This will allow multilibs for different architectures to share a common include directory. To avoid breaking the arm-execute-only.c CHECK-NO-EXECUTE-ONLY-ASM test, add a ForMultilib argument to getARMTargetFeatures. Since the presence of multilib.yaml can change the exact location of a library, relax the baremetal.cpp test. Differential Revision: https://reviews.llvm.org/D142986
-
Michael Platings authored
This option causes the flags used for selecting multilibs to be printed. This is an experimental feature that is documented in detail in D143587. Differential Revision: https://reviews.llvm.org/D142933
-
Michael Platings authored
The format includes a ClangMinimumVersion entry to avoid a potential source of subtle errors if an older version of Clang were to be used with a multilib.yaml that requires a newer Clang to work correctly. This feature is comparable to CMake's cmake_minimum_required. Reviewed By: peter.smith Differential Revision: https://reviews.llvm.org/D142932
-
Kirill Stoimenov authored
Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D152763
-