- Jul 12, 2023
-
-
Philip Reames authored
This change continues with the line of work discussed in https://discourse.llvm.org/t/riscv-transition-in-vector-pseudo-structure-policy-variants/71295. This change handles most of the binary pseudos. I excluded pseudos which _TIED variants, and those that produce mask results. Both a bit different in functionality, and deserve their own change and review. As with previous changes in the series, we replace the existing TA and TU forms with a single unified pseudo with a passthru (which may be implicit_def) and a policy operand. As before, we see codegen changes (some improvements and some regressions) due to scheduling differences caused by the extra implicit_def instructions. Differential Revision: https://reviews.llvm.org/D154245
-
David Green authored
See D153632 and D154063
-
Craig Topper authored
The register being replaced might have a more restrictive register class due to requirements of the using instruction. We should constrain the register class to preserve any restrictions. This was found in our downstream on a custom instruction. I don't have a test case for upstream currently. Differential Revision: https://reviews.llvm.org/D154920
-
Aleksandr Popov authored
This reverts commit 4c6f95be and relands e16c5c09 https://reviews.llvm.org/D154069
-
Alex Zinenko authored
This is the counterpart to the forward dense dataflow analysis and integrates into the dataflow framework. The implementation follows the structure of existing dataflow analyses. Reviewed By: Mogball, phisiart Differential Revision: https://reviews.llvm.org/D154713
-
Slava Zakharin authored
When an initializer value is missing for an allocatable component in a structure constructor, the RHS is NULL() expression. We should just skip this part of the initializer, since the component must become unallocated (as it is from the initialization). Runtime detected rank mismatch when we tried to pass NULL() box RHS for assigning it to the unallocated component of rank 1, 2, etc. Reviewed By: tblah Differential Revision: https://reviews.llvm.org/D154906
-
Slava Zakharin authored
In the context of elemental operation a dynamically optional intrinsic argument must be lowered such that the elemental designator is generated under isPresent check. Reviewed By: tblah Differential Revision: https://reviews.llvm.org/D154897
-
Shoaib Meenai authored
We need to explicitly mark DWARFUnitInfo as non-copyable since MSVC's STL has a `noexcept(false)` move constructor for `unordered_map`; see the added comment for more details. An alternative might be using SmallVector instead of std::vector, since that never tries to copy elements [1]. That would result in a bunch of API changes though, so I figured a smaller targeted fix was better. [1] https://llvm.org/docs/ProgrammersManual.html#llvm-adt-smallvector-h Reviewed By: ayermolo, maksfb Differential Revision: https://reviews.llvm.org/D154924
-
Nick Desaulniers authored
There is no need to print the entire function after a transform via LLVM_DEBUG statements. These can be emulated via: $ llc -print-after=consthoist -filter-print-funcs=<function name> Otherwise, this makes the output of $ llc -debug-only=consthoist too verbose. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D154904
-
Wael Yehia authored
Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D154239
-
Nick Desaulniers authored
A follow up to commit 6bad76c7 ("[Demangle] fix windows tests") based on @thakis' report. Fixes: #63740 Reviewed By: thakis Differential Revision: https://reviews.llvm.org/D154875
-
Bryan Chan authored
-
Simon Pilgrim authored
Building on the support for wider input vector types from D154592, try to more aggressively widen inputs instead of scalarizing them.
-
- Jul 11, 2023
-
-
Valentin Clement authored
Add support for `ieor` reduction operator in OpenACC lowering. Depends on D154887 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D154888
-
Valentin Clement authored
Add support for `ior` reduction operator in OpenACC lowering. Depends on D154886 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D154887
-
Joseph Huber authored
The 'RPCHandleTy' was intended to capture the intention that a specific device owns its slot in the RPC server. However, this required creating a temporary store to hold these pointers. This was causing really weird spurious failure due to undefined behaviour in the order of library teardown. For example, the x64 plugin would be torn down, set this to some invalid memory, and then the CUDA plugin would crash. Rather than spend the time to fully diagnose this problem I found it pertinent to simply remove the failure mode. This patch removes this indirection so now the usage of the RPC server must always be done with the intended device. This just requires some extra handling for the AMDGPU indirection where we need to store a reference to the device. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D154971
-
Zarko Todorovski authored
On PowerPC, the vec_ct* builtin function take the form of eg. d=vec_cts(a,b) LLVM (llc) will crash when a user specifies a number out of the allowed range (0-31) for b.This patch truncates b so that we avoid the backend crash in some cases. Further documentation for the builtins can be found here: https://www.ibm.com/docs/en/xl-c-and-cpp-linux/16.1.0?topic=functions-vec-ctf https://www.ibm.com/docs/en/xl-c-and-cpp-linux/16.1.0?topic=functions-vec-cts Reviewed By: nemanjai, #powerpc Differential Revision: https://reviews.llvm.org/D106409
-
Tuan Chuong Goh authored
Differential Revision: https://reviews.llvm.org/D154835
-
Valentin Clement authored
Add support for `iand` reduction operator in OpenACC lowering. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D154886
-
Richard Smith authored
We were accidentally profiling the fabricated second argument (`0`), resulting in overloaded dependent `a++` and non-overloaded dependent `a++` having different hashes.
-
Fangrui Song authored
Port D69671 (llvm-readobj) to llvm-objdump. Add a class llvm::objdump::Dumper and move some free functions into Dumper so that they can call reportUniqueWarning. Warnings seems preferable in these cases as the issue is localized and we can continue dumping other information. Differential Revision: https://reviews.llvm.org/D154754
-
Viktoriia Bakalova authored
Differential Revision: https://reviews.llvm.org/D154962
-
Fangrui Song authored
-
Guray Ozen authored
`mbarrier` is a barrier created in shared memory that supports different flavors of synchronizing threads other than `__syncthreads`, for more information see below. https://docs.nvidia.com/cuda/parallel-thread-execution/#parallel-synchronization-and-communication-instructions-mbarrier This work adds initial Ops wrt `mbarrier` to nvgpu dialect. First, it introduces to two types: `mbarrier.barrier` that is barrier object in shared memory `mbarrier.barrier.token` that is token It introduces following Ops: `mbarrier.create` creates `mbarrier.barrier` `mbarrier.init` initializes `mbarrier.barrier` `mbarrier.arrive` performs arrive-on `mbarrier.barrier` returns `mbarrier.barrier.token` `mbarrier.arrive.nocomplete` performs arrive-on (non-blocking) `mbarrier.barrier` returns `mbarrier.barrier.token` `mbarrier.test_wait` waits on `mbarrier.barrier` and `mbarrier.barrier.token` Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D154090
-
Aliia Khasanova authored
Differential Revision: https://reviews.llvm.org/D154976
-
Petr Hosek authored
This reverts commit dae9d1b5 since it caused https://github.com/llvm/llvm-project/issues/63799.
-
Luke Lau authored
It no longer defaults to false as of 63336795 Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D154973
-
Juan Manuel MARTINEZ CAAMAÑO authored
This reverts commit 125b9074.
-
Juan Manuel MARTINEZ CAAMAÑO authored
Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D154970
-
Matthias Springer authored
Add a new option that allows users to specify a memcpy op: "memref.tensor_store", "memref.copy" or "linalg.copy". Differential Revision: https://reviews.llvm.org/D154968
-
Matthias Springer authored
This unit attribute indicates to the bufferization that the resulting buffer will not be written to by another op. Differential Revision: https://reviews.llvm.org/D154967
-
Serge Pavlov authored
Builtin floating-point number classification functions: - __builtin_isnan, - __builtin_isinf, - __builtin_finite, and - __builtin_isnormal now are implemented using `llvm.is_fpclass`. This change makes the target callback `TargetCodeGenInfo::testFPKind` unneeded. It is preserved in this change and should be removed later. Differential Revision: https://reviews.llvm.org/D112932 -
Matthias Springer authored
Return all ops that were generated as part of the bufferization, so that users do not have to match them in the enclosing op. Differential Revision: https://reviews.llvm.org/D154966
-
David Mo authored
For inline WebAssembly, passing a numeric operand to global.get is unsupported. This causes encodeInstruction to reach an llvm_unreachable call, leading to undefined behaviors. This patch fixes the issue for this invalid instruction encoding, making it report an error by adding an MCContext field in class WebAssemblyMCCodeEmitter. Reviewed By: sbc100, bryanpkc Differential Revision: https://reviews.llvm.org/D154734
-
Phoebe Wang authored
The combination was designed to combine a negative imaginary value rather then a full negative complex value. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D154213
-
Matthias Springer authored
This transform op can be used to select all payload ops with a given name from a handle. Differential Revision: https://reviews.llvm.org/D154956
-
gilsaia authored
Added a series of optimizations to the Intersect function of PresburgerRelation, referring to the ISL implementation. Tested it on a simple Benchmark implemented by myself to see that it can speed up the Intersect operation The Benchmark can be found here:https://github.com/gilsaia/llvm-project-test-fpl/blob/develop_benchmark/mlir/benchmark/presburger/Benchmark.cpp The overall results for Intersect are as follows {F28191553} The results for each case are as follows {F28191556} Reviewed By: Groverkss Differential Revision: https://reviews.llvm.org/D154771
-
NAKAMURA Takumi authored
-
Juan Manuel MARTINEZ CAAMAÑO authored
Moving out some changes not related to the bugfix in https://reviews.llvm.org/D154946 Reviewed By: JonChesterfield, arsenm Differential Revision: https://reviews.llvm.org/D154959
-