- Sep 13, 2023
-
-
Martin Erhart authored
Revert "[mlir][bufferization] Update linalg integration tests to lower ops created by bufferization-to-memref pass" This reverts commit 6f35401f. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit 1bebb60a. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit 29d86175. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit 89117f18. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit 1356e853. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit f0c46639. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Martin Erhart authored
This reverts commit cb5fe6ce. This caused problems in downstream projects. We are reverting to give them more time for integration.
-
Mikhail R. Gadelha authored
This patch changes the size of time_t to be an int64_t. This still follows the POSIX standard which only requires time_t to be an integer. Making time_t a 64-bit integer also fixes two cases in 32 bits platforms that use SYS_clock_nanosleep_time64 and SYS_clock_gettime64, as the name of these calls implies, they require a 64-bit time_t. For instance, in rv32, the 32-bit version of these syscalls is not available. We also follow glibc here, where time_t is still a 32-bit integer in arm32. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D159125
-
Joseph Huber authored
Summary: Currently, there is an assertion that prevents us from emitting an AMDGPU global with a non-target specific address space (i.e. numerical attribute). I'm unsure what the original intentions of this assertion were, but we should be able to use OpenCL address spaces when compiling directly to AMDGPU from C++. This is permitted on NVPTX so I'm unsure what this assertion is guarding. The patch simply removes the assertion and adds a test to ensure that these emit the expected address spaces. Fixes https://github.com/llvm/llvm-project/issues/65069
-
Guray Ozen authored
clang was used for local testing. The PR changes it to `mlir-cpu-runner`
-
Joseph Huber authored
Summary: We use the `llvm.amgcn.abi.version` varaible to control code generation. This is emitted in every module now to indicate what should be used when compiling. Previously, the logic caused us to emit an external reference to this variable when creating the code for the `none` type. This would then cause us not to emit the actual definition. This patch refines the logic to create the external reference, and then update it if it is found unset by the time we emit the global. I had to remove the reference to `GetOrCreateLLVmGlobal` because it did not accept the proper address space.
-
Guray Ozen authored
An integration test for the 128b Swizzling TMA. TMA with 128B Swizzle loads data as follows (each numbered cell is 16 bytes). The program tests this pattern for `128x64xf16` type. ``` |-------------------------------| | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | | 1 | 0 | 3 | 2 | 5 | 4 | 7 | 6 | | 2 | 3 | 0 | 1 | 6 | 7 | 4 | 5 | | 3 | 2 | 1 | 0 | 7 | 6 | 5 | 4 | | 4 | 5 | 6 | 7 | 0 | 1 | 2 | 3 | | 5 | 4 | 7 | 6 | 1 | 0 | 3 | 2 | | 6 | 7 | 4 | 5 | 2 | 3 | 0 | 1 | |-------------------------------| | ... pattern repeats ... | |-------------------------------| ```
-
martinboehme authored
Otherwise, the test doesn't actually do anything.
-
Guray Ozen authored
The register number of predicate is calculated incorrectly. This PR fixes that.
-
Timm Bäder authored
-
Luke Drummond authored
We were using a mix of unsigned and signed ints in the various PTX asm printers. All calls from tablgen use a non-negative immediate, so either will work, but when doing arithmetic on the return value from `getNumOperands`, or calling `getOperand`, it makes sense to keep everything unsigned.
-
Paul T Robinson authored
Adds a bunch of stuff overlooked in the original setup. Steps back a little on the llvm/lib/CodeGen/AsmPrinter paths.
-
Michael Maitland authored
VPIntrinsics with VP_PROPERTY_BINARYOP property should have the ability to be queried with with VPBinOpIntrinsic::isVPBinOp, similiar to how intrinsics with the VP_PROPERTY_REDUCTION property can be queried with VPReductionIntrinsic::isVPReduction. This will be used in #65706. In that PR the usage of this class is tested.
-
Luke Lau authored
This adds a helper method to get the ID of the functionally equivalent intrinsic, similar to the existing getFunctionalOpcodeForVP and getConstrainedIntrinsicIDForVP methods. Not sure if it's notable or not, but I can't find any existing uses of VP_PROPERTY_FUNCTIONAL_INTRINSIC? It could potentially be used in #65706 to scalarize VP intrinsics.
-
Timm Bäder authored
-
Simon Pilgrim authored
Fixes #66194
-
David Truby authored
This adds a new pass to add an Any comdat to each linkonce and linkonce_odr function in the LLVM dialect. These comdats are necessary on Windows to allow the default system linker to link binaries containing these functions.
-
Matthias Springer authored
This was an oversight in 0ac21e65.
-
Simon Pilgrim authored
Followup to D59363 which failed to handle the icmp(X,undef) -> isTrueWhenEqual case - similar to llvm::ConstantFoldCompareInstruction As discussed on the review, this is affecting some previously reduced test cases, but will also prevent reductions from relying on this inconsistent behaviour in the future. Reapplied after reversion at e1e3c75c with a tweak to the pseudo-probe-peep.ll test Differential Revision: https://reviews.llvm.org/D158068
-
Benjamin Kramer authored
This reverts commit c6a33ff4. Makes clang segfault. // clang t.cc class a; class c { public: [[clang::annotate("")]] c(const c *) {} }; class d { d(const c *, a *, a *); c e; }; d::d(const c *f, a *, a *) : e(f) {}
-
Simon Pilgrim authored
[AMDGPU] Remove constexpr from getNumUserSGPRForField/getMaxNumPreloadedSGPRs to appease older gcc builds Older versions of gcc wouldn't accept the constexpr getNumUserSGPRForField (introduced in D159439 / 343be513) as it couldn't treat the llvm_unreachable call as constexpr
-
Mikhail Goncharov authored
ninja takes every target as a separate argument
-
Ben Shi authored
The newly added tests are all about scalable vector types.
-
Simon Pilgrim authored
Revert rG6c56cf71 "[DAG] FoldSetCC - add missing icmp(X,undef) -> isTrueWhenEqual case" Need to address a missed test change
-
Matthias Springer authored
This commit generalizes the special tensor.extract_slice/tensor.insert_slice bufferization rules to tensor subset ops. Ops that insert a tensor into a tensor at a specified subset (e.g., tensor.insert_slice, tensor.scatter) can implement the `SubsetInsertionOpInterface`. Apart from adding a new op interface (extending the API), this change is NFC. The only ops that currently implement the new interface are tensor.insert_slice and tensor.parallel_insert_slice, and those ops were are supported by One-Shot Bufferize.
-
Sander de Smalen authored
-
Simon Pilgrim authored
Followup to D59363 which failed to handle the icmp(X,undef) -> isTrueWhenEqual case - similar to llvm::ConstantFoldCompareInstruction As discussed on the review, this is affecting some previously reduced test cases, but will also prevent reductions from relying on this inconsistent behaviour in the future. Differential Revision: https://reviews.llvm.org/D158068
-
David Spickett authored
So folks at least know what we are differing from.
-
Martin Erhart authored
-
Martin Erhart authored
Since buffer deallocation requires a few passes to be run in a somewhat fixed sequence, it makes sense to have a pipeline for convenience (and to reduce the number of transform ops to represent default deallocation). Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D159432
-
Martin Erhart authored
The scf.forall.in_parallel terminator operation has a nested graph region with the NoTerminator trait. Such regions are not supported by the default implementations. Therefore, this commit adds a specialized implementation for this operation which only covers the case where the nested region is empty. This is because after bufferization, ops like tensor.parallel_insert_slice were already converted to memref operations residing int the scf.forall only and the nested region of scf.forall.in_parallel ends up empty. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158979
-
Martin Erhart authored
Add a method to the BufferDeallocationOpInterface that allows operations to implement the interface and provide custom logic to compute the ownership indicators of values it defines. As a demonstrating example, this new method is implemented by the `arith.select` operation. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158828
-
Martin Erhart authored
This new interface allows operations to implement custom handling of ownership values and insertion of dealloc operations which is useful when an op cannot implement the interfaces supported by default by the buffer deallocation pass (e.g., because they are not exactly compatible or because there are some additional semantics to it that would render the default implementations in buffer deallocation invalid, or because no interfaces exist for this kind of behavior and it's not worth introducing one plus a default implementation in buffer deallocation). Additionally, it can also be used to provide more efficient handling for a specific op than the interface based default implementations can. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158756
-
Martin Erhart authored
Add a new Buffer Deallocation pass replacing the old one with the goal of inserting fewer clone operations and supporting additional use-cases. Please refer to the Buffer Deallocation section in the updated Bufferization.md file for more information on how this new pass works. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158421
-
Martin Erhart authored
[mlir][bufferization] Update linalg integration tests to lower ops created by bufferization-to-memref pass This commit prepares the linalg integration tests to be run with the new BufferDeallocation pass which requires the bufferization-to-memref pass to be run afterwards. The bufferization-to-memref pass may create ops of the SCF, Func, Arith, and MemRef dialects. Currently, not all integration tests execute all the conversion passes necessary to lower these dialects to LLVM. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D156663
-