- Jun 15, 2022
-
-
Martin Boehme authored
-
Ilya Biryukov authored
When compiled with `-D_LIBCPP_ENABLE_CXX20_REMOVED_ALLOCATOR_MEMBERS` uses of `allocator<void>::pointer` resulted in compiler errors after D104323. If we instantiate the primary template, `allocator<void>::reference` produces an error 'cannot form references to void'. To workaround this, allow to bring back the `allocator<void>` specialization by defining the new `_LIBCPP_ENABLE_CXX20_REMOVED_ALLOCATOR_VOID_SPECIALIZATION` macro. To make sure the code that uses `allocator<void>` and the removed members does not break, both `_LIBCPP_ENABLE_CXX20_REMOVED_ALLOCATOR_MEMBERS` and `_LIBCPP_ENABLE_CXX20_REMOVED_ALLOCATOR_MEMBERS` have to be defined. Reviewed By: ldionne, #libc, philnik Differential Revision: https://reviews.llvm.org/D126210
-
Martin Boehme authored
-
David Sherwood authored
We can remove the MatrixZADRegisterTable table of tile registers and just calculate the register index directly. Differential Revision: https://reviews.llvm.org/D127757
-
Kadir Cetinkaya authored
This has been tested on a large set of c++ developers for a long while, without any crashes or complaints. Differential Revision: https://reviews.llvm.org/D127833
-
owenca authored
-
Benjamin Kramer authored
Just short-circuit when a change was made, the erased value is invalid after that. Found by asan. This pass looks like it could use rewrite patterns instead which don't have this issue, but let's fix the asan build first.
-
Kito Cheng authored
RISC-V expand register tuple spilling into series of register spilling after register allocation phase by the pseudo instruction expansion, however part of register tuple might be still undefined during spilling, machine verifier will complain the spill instruction is using an undefined physical register. Optimal solution should be doing liveness analysis and do not emit spill and reload for those undefined parts, but accurate liveness info at that point is not so easy to get. So the suboptimal solution is still spill and reload those undefined parts, but adding implicit-use of super register to spill function, then machine verifier will only report report using undefined physical register if the when whole super register is undefined, and this behavior are also documented in MachineVerifier::checkLiveness[1]. Example for demo what happend: ``` v10m2 = xxx # v12m2 not define yet PseudoVSPILL2_M2 v10m2_v12m2 ... ``` After expansion: ``` v10m2 = xxx # v12m2 not define yet # Expand PseudoVSPILL2_M2 v10m2_v12m2 to 2 vs2r VS2R_V v10m2 VS2R_V v12m2 # Use undef reg! ``` What this patch did: ``` v10m2 = xxx # v12m2 not define yet # Expand PseudoVSPILL2_M2 v10m2_v12m2 to 2 vs2r VS2R_V v10m2 implicit v10m2_v12m2 # Use undef reg (v12m2), but v10m2_v12m2 ins't totally undef, so # that's OK. VS2R_V v12m2 implicit v10m2_v12m2 ``` [1] https://github.com/llvm-mirror/llvm/blob/master/lib/CodeGen/MachineVerifier.cpp#L2016-L2019 Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D127642
-
Matthias Springer authored
If `create-deallocs=0`, mark all bufferization.alloc_tensor ops as escaping. (Unless they already have an `escape` attribute.) In the absence of analysis information, check SSA use-def chains to see if the value may be yielded. Differential Revision: https://reviews.llvm.org/D127302
-
Siva Chandra Reddy authored
-
Matthias Springer authored
`AnalysisState` now has default implementations of all virtual functions. Differential Revision: https://reviews.llvm.org/D127301
-
Heejin Ahn authored
Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D127810
-
Matthias Springer authored
Bufferization of the func dialect must go through `OneShotModuleBufferize`. With this change, the analysis interface methods of the BufferizableOpInterface of func dialect ops can be used together with the normal `OneShotBufferize`. (In the absence of analysis information, they will return conservative results.) Differential Revision: https://reviews.llvm.org/D127299
-
Peixin-Qiao authored
As OpenMP 5.0, for firstprivate, lastprivate, copyin, and copyprivate clauses, if the list item is a polymorphic variable with the allocatable attribute, the behavior is unspecified. Reviewed By: kiranchandramohan Differential Revision: https://reviews.llvm.org/D127601
-
Matthias Springer authored
This flag was introduced for a use case in IREE, but it is no longer needed. Differential Revision: https://reviews.llvm.org/D126965
-
Martin Boehme authored
This is an analog to the `annotate` attribute but for types. The intent is to allow adding arbitrary annotations to types for use in static analysis tools. For details, see this RFC: https://discourse.llvm.org/t/rfc-new-attribute-annotate-type-iteration-2/61378 Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D111548
-
Peixin-Qiao authored
This constraint is used in OMP2012 benchmark, and other compilers do not enforce it. Change it into one warning. This addresses the issue https://github.com/llvm/llvm-project/issues/56003. Reviewed By: klausler, kiranchandramohan Differential Revision: https://reviews.llvm.org/D127740
-
Nikita Popov authored
The same condition already exists inside optimizeMemCmpConstantSize().
-
Austin Kerbow authored
Some buildbots (lto, windows) were failing due to some function reference variables being improperly initialized.
-
Siva Chandra Reddy authored
-
Petr Hosek authored
Rather than invoking the linker directly, let the compiler driver handle it. This ensures that we use the correct linker in the case of cross-compiling. Differential Revision: https://reviews.llvm.org/D127828
-
Matthias Springer authored
scf::ForOp and scf::WhileOp must insert buffer copies not only for out-of-place bufferizations, but also to enforce additional invariants wrt. to buffer aliasing behavior. This is currently happening in the respective `bufferize` methods. With this change, the tensor copy insertion pass will also enforce these invariants by inserting copies. The `bufferize` methods can then be simplified and made independent of the `AnalysisState` data structure in a subsequent change. Differential Revision: https://reviews.llvm.org/D126822
-
owenca authored
-
chenglin.bi authored
#53877
-
Siva Chandra Reddy authored
Before this change, they were unconditionally added, irrespective of the availability of the architecture specific pieces.
-
Kadir Cetinkaya authored
Differential Revision: https://reviews.llvm.org/D127749
-
Yeting Kuo authored
VSETVLIInfos right after VLEFF/VLSEGFF are currently unknown since they modify VL. Unknown VSETVLIInfos make next vector operations needed to be inserted VSET(I)VLI. Actually the next vector operation of VLEFF/VLSEGFF may not need to be inserted VSET(I)VLI if it uses same VTYPE and the resulted vl of VLEFF/VLSEGFF. Take the below C code as an example, vint8m4_t vec_src1 = vle8ff_v_i8m4(str1, &new_vl, vl); vbool2_t mask1 = vmseq_vx_i8m4_b2(vec_src1, 0, new_vl); vsetvli insertion adds a redundant vsetvli for that, Assembly result: vsetvli a2,a2,e8,m4,ta,mu vle8ff.v v28,(a0) csrr a3,vl ; redundant vsetvli zero,a3,e8,m4,ta,mu ; redundant vmseq.vi v25,v28,0 After D126794, VLEFF/VLSEGFF has a define having value of VL. The patch consider there is a ghost vsetvli right after VLEFF/VLSEGFF. The ghost VSET(I)LIs use the vl output of the VLEFF/VLSEGFF as its AVL and same VTYPE of the VLEFF/VLSEGFF. The ghost vsetvli must be redundant, and we could use it to get the VSETVLIInfo right after VLEFF/VLSEGFF. Reviewed By: reames Differential Revision: https://reviews.llvm.org/D127576
-
Ping Deng authored
Reviewed By: craig.topper, spatel Differential Revision: https://reviews.llvm.org/D127474
-
Siva Chandra Reddy authored
Futexes are 32 bits in size on all platforms, including 64-bit systems.
-
owenca authored
Turn off RemoveBracesLLVM while analyzing InsertBraces and vice versa to avoid potential interference of each other and better the performance. Differential Revision: https://reviews.llvm.org/D127685
-
LLVM GN Syncbot authored
-
Austin Kerbow authored
The sched_barrier builtin allow the scheduler's behavior to be shaped by users when very specific codegen is needed in order to create highly optimized code. This patch adds more granular control over the types of instructions that are allowed to be reordered with respect to one or multiple sched_barriers. A mask is used to specify groups of instructions that should be allowed to be scheduled around a sched_barrier. The details about this mask may be used can be found in llvm/include/llvm/IR/IntrinsicsAMDGPU.td. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D127123
-
Austin Kerbow authored
Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D127124
-
Fangrui Song authored
switchSection should be used instead.
-
Peter S. Housel authored
This change adds test cases targeting the AArch64 Linux platform to the ORC runtime integration test suite. Reviewed By: lhames, sunho Differential Revision: https://reviews.llvm.org/D127720
-
Ping Deng authored
precommit tests for D127474 Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D127475
-
Joe Loser authored
Several span constructors use `enable_if` which is verbose. Replace these with concepts or requires expressions.
-
Venkata Ramanaiah Nalamothu authored
For the 'thread until' command, the selected thread ID, to perform the operation on, could be of the current thread or the specified thread. Reviewed By: jingham Differential Revision: https://reviews.llvm.org/D48865
-
Lei Zhang authored
Per GLSL Pow extended instruction spec: "Result is undefined if x < 0. Result is undefined if x = 0 and y <= 0." So we need to handle negative `x` values specifically. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D127816
-
wangpc authored
Since almost all pseudos have the same form of BaseInstr, we can just set it as default value to reduce some lines. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D127632
-