- Sep 13, 2023
-
-
Timm Bäder authored
-
Luke Drummond authored
We were using a mix of unsigned and signed ints in the various PTX asm printers. All calls from tablgen use a non-negative immediate, so either will work, but when doing arithmetic on the return value from `getNumOperands`, or calling `getOperand`, it makes sense to keep everything unsigned.
-
Paul T Robinson authored
Adds a bunch of stuff overlooked in the original setup. Steps back a little on the llvm/lib/CodeGen/AsmPrinter paths.
-
Michael Maitland authored
VPIntrinsics with VP_PROPERTY_BINARYOP property should have the ability to be queried with with VPBinOpIntrinsic::isVPBinOp, similiar to how intrinsics with the VP_PROPERTY_REDUCTION property can be queried with VPReductionIntrinsic::isVPReduction. This will be used in #65706. In that PR the usage of this class is tested.
-
Luke Lau authored
This adds a helper method to get the ID of the functionally equivalent intrinsic, similar to the existing getFunctionalOpcodeForVP and getConstrainedIntrinsicIDForVP methods. Not sure if it's notable or not, but I can't find any existing uses of VP_PROPERTY_FUNCTIONAL_INTRINSIC? It could potentially be used in #65706 to scalarize VP intrinsics.
-
Timm Bäder authored
-
Simon Pilgrim authored
Fixes #66194
-
David Truby authored
This adds a new pass to add an Any comdat to each linkonce and linkonce_odr function in the LLVM dialect. These comdats are necessary on Windows to allow the default system linker to link binaries containing these functions.
-
Matthias Springer authored
This was an oversight in 0ac21e65.
-
Simon Pilgrim authored
Followup to D59363 which failed to handle the icmp(X,undef) -> isTrueWhenEqual case - similar to llvm::ConstantFoldCompareInstruction As discussed on the review, this is affecting some previously reduced test cases, but will also prevent reductions from relying on this inconsistent behaviour in the future. Reapplied after reversion at e1e3c75c with a tweak to the pseudo-probe-peep.ll test Differential Revision: https://reviews.llvm.org/D158068
-
Benjamin Kramer authored
This reverts commit c6a33ff4. Makes clang segfault. // clang t.cc class a; class c { public: [[clang::annotate("")]] c(const c *) {} }; class d { d(const c *, a *, a *); c e; }; d::d(const c *f, a *, a *) : e(f) {}
-
Simon Pilgrim authored
[AMDGPU] Remove constexpr from getNumUserSGPRForField/getMaxNumPreloadedSGPRs to appease older gcc builds Older versions of gcc wouldn't accept the constexpr getNumUserSGPRForField (introduced in D159439 / 343be513) as it couldn't treat the llvm_unreachable call as constexpr
-
Mikhail Goncharov authored
ninja takes every target as a separate argument
-
Ben Shi authored
The newly added tests are all about scalable vector types.
-
Simon Pilgrim authored
Revert rG6c56cf71 "[DAG] FoldSetCC - add missing icmp(X,undef) -> isTrueWhenEqual case" Need to address a missed test change
-
Matthias Springer authored
This commit generalizes the special tensor.extract_slice/tensor.insert_slice bufferization rules to tensor subset ops. Ops that insert a tensor into a tensor at a specified subset (e.g., tensor.insert_slice, tensor.scatter) can implement the `SubsetInsertionOpInterface`. Apart from adding a new op interface (extending the API), this change is NFC. The only ops that currently implement the new interface are tensor.insert_slice and tensor.parallel_insert_slice, and those ops were are supported by One-Shot Bufferize.
-
Sander de Smalen authored
-
Simon Pilgrim authored
Followup to D59363 which failed to handle the icmp(X,undef) -> isTrueWhenEqual case - similar to llvm::ConstantFoldCompareInstruction As discussed on the review, this is affecting some previously reduced test cases, but will also prevent reductions from relying on this inconsistent behaviour in the future. Differential Revision: https://reviews.llvm.org/D158068
-
David Spickett authored
So folks at least know what we are differing from.
-
Martin Erhart authored
-
Martin Erhart authored
Since buffer deallocation requires a few passes to be run in a somewhat fixed sequence, it makes sense to have a pipeline for convenience (and to reduce the number of transform ops to represent default deallocation). Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D159432
-
Martin Erhart authored
The scf.forall.in_parallel terminator operation has a nested graph region with the NoTerminator trait. Such regions are not supported by the default implementations. Therefore, this commit adds a specialized implementation for this operation which only covers the case where the nested region is empty. This is because after bufferization, ops like tensor.parallel_insert_slice were already converted to memref operations residing int the scf.forall only and the nested region of scf.forall.in_parallel ends up empty. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158979
-
Martin Erhart authored
Add a method to the BufferDeallocationOpInterface that allows operations to implement the interface and provide custom logic to compute the ownership indicators of values it defines. As a demonstrating example, this new method is implemented by the `arith.select` operation. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158828
-
Martin Erhart authored
This new interface allows operations to implement custom handling of ownership values and insertion of dealloc operations which is useful when an op cannot implement the interfaces supported by default by the buffer deallocation pass (e.g., because they are not exactly compatible or because there are some additional semantics to it that would render the default implementations in buffer deallocation invalid, or because no interfaces exist for this kind of behavior and it's not worth introducing one plus a default implementation in buffer deallocation). Additionally, it can also be used to provide more efficient handling for a specific op than the interface based default implementations can. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158756
-
Martin Erhart authored
Add a new Buffer Deallocation pass replacing the old one with the goal of inserting fewer clone operations and supporting additional use-cases. Please refer to the Buffer Deallocation section in the updated Bufferization.md file for more information on how this new pass works. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D158421
-
Martin Erhart authored
[mlir][bufferization] Update linalg integration tests to lower ops created by bufferization-to-memref pass This commit prepares the linalg integration tests to be run with the new BufferDeallocation pass which requires the bufferization-to-memref pass to be run afterwards. The bufferization-to-memref pass may create ops of the SCF, Func, Arith, and MemRef dialects. Currently, not all integration tests execute all the conversion passes necessary to lower these dialects to LLVM. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D156663
-
Martin Erhart authored
[mlir][bufferization] Remove allow-return-allocs and create-deallocs pass options, remove bufferization.escape attribute This is the first commit in a series with the goal to rework the BufferDeallocation pass. Currently, this pass heavily relies on copies to perform correct deallocations, which leads to very slow code and potentially high memory usage. Additionally, there are unsupported cases such as returning memrefs which this series of commits aims to add support for as well. This first commit removes the deallocation capabilities of one-shot-bufferization.One-shot-bufferization should never deallocate any memrefs as this should be entirely handled by the buffer-deallocation pass going forward. This means the allow-return-allocs pass option will default to true now, create-deallocs defaults to false and they, as well as the escape attribute indicating whether a memref escapes the current region, will be removed. The documentation should w.r.t. these pass option changes should also be updated in this commit. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D156662
-
Zhangyin authored
Reviewed By: #libc, philnik Differential Revision: https://reviews.llvm.org/D159509
-
David Spickett authored
Previously we would check all built-ins first for suggestions, then check built-ins and aliases. This meant that if you had an alias brkpt -> breakpoint, "br" would complete to "breakpoint". Instead of giving you the choice of "brkpt" or "breakpoint".
-
David Spickett authored
``` $ ./bin/clang --target=arm-linux-gnueabihf --print-supported-extensions <...> All available -march extensions for ARM crc crypto sha2 aes dotprod <...> ``` This follows the format set by RISC-V and AArch64. As for AArch64, ARM doesn't have versioned extensions like RISC-V does. So there is only 1 column, which contains the name. Any extension without a "feature" is hidden as these cannot be used with -march. -
David Spickett authored
SME reuses SVE's register state but adds new modes to it. Therefore we can't check all those in the same test as the existing SVE checks. SME's ZA, SVG and SVCR register checks will be added to this test in later patches. Prior to this we didn't have any testing of writing streaming mode SVE registers from lldb, only writing SVE registers in normal (non-streaming) SVE mode. Reviewed By: omjavaid Differential Revision: https://reviews.llvm.org/D157846
-
Timm Baeder authored
I don't now squat about Objective C, but the extra variable (and cast) seemed unnecessary and this is a good opportunity to re-format that ugly parameter list.
-
Sergey Kachkov authored
findDominatingValue has a search limit, and when it is reached, optimization is not applied. This patch fixes the issue that this limit also takes into account debug intrinsics, so the result of optimization can depend from the presence of debug info.
-
Qiu Chaofan authored
Reviewed By: shchenz Differential Revision: https://reviews.llvm.org/D158704
-
Matt Arsenault authored
Try to avoid expensive checks failures from reporting no changes when some dead instructions were introduced.
-
Konstantin Varlamov authored
Differential Revision: https://reviews.llvm.org/D159065
-
Matt Arsenault authored
This reverts commit d9333e36.
-
Matt Arsenault authored
Ported from old amdgcn intrinsic which will soon be deleted. https://reviews.llvm.org/D149587
-
Kai Luo authored
In Orc runtime, we use `dlopen(nullptr, ...)` to open current executable and use `dlsym` to find addresses of symbols, this requires `-rdynamic` flag. As `llvm/CMakeLists.txt` suggests ``` # Make sure we don't get -rdynamic in every binary. For those that need it, # use export_executable_symbols(target). ``` This patch exports symbols in `ClangReplInterpreterExceptionTests`. This also fixes `ClangReplInterpreterExceptionTests` is skipped on ppc64 when jitlink is used. Reviewed By: v.g.vassilev Differential Revision: https://reviews.llvm.org/D159167
-