- Apr 24, 2024
-
-
jeanPerier authored
The pass assumed that all fun.func symbol usages could be safely replaced by undef, that is not true after #87796 that added a back link from internal procedure back to the parent procedure. This caused the internal procedures to be erased and then processed (segfault). Also set visibility of such internal procedures so that MLIR do not remove them before the target function is generated for the target region.
-
Mike Rice authored
Add an initializer for StreamSym, which is a pointer. The pointers in this class are set in the Init function, but all should be initialized in the constructor to avoid confusion and static verifier hits.
-
Luke Lau authored
Seems to cause an address sanitizer failure on one of the buildbots related to live intervals.
-
Benjamin Maxwell authored
These tests basically were integration tests as unit tests, checking too many passes at once to be useful, and brittle to any changes. This patch moves (non-duplicated) tests to `vector-to-arm-sme.mlir` (which only tests `-convert-vector-to-arm-sme`). The lowering after that e.g. `ArmSME -> SCF` and `ArmSME -> LLVM` already have their own set of tests covering these cases.
-
Alexander M authored
ae389b24 change doesn't cover "_d" suffix for Debug build on Windows. Fixed #87381.
-
Antonio Frighetto authored
A logic issue arose when inlining via `CloneAndPruneFunctionInto`, which, besides cloning, performs instruction simplification as well. By the time a new cloned instruction is being simplified, phi-nodes are not remapped yet as the whole CFG needs to be processed first. As `VMap` state at this stage is incomplete, `threadCmpOverPHI` and variants could lead to unsound optimizations. This issue has been addressed by performing basic constant folding while cloning, and postponing instruction simplification once phi-nodes are revisited. Fixes: https://github.com/llvm/llvm-project/issues/87534.
-
Antonio Frighetto authored
-
Philip Reames authored
We can expand these as the three instruction sequence: (sub (shl X, C1), (shXadd X, x)).
-
Tom Stellard authored
This target will be used to generate the release binary package for uploading to GitHub.
-
Brandon Wu authored
Since the requirement is EEW=32, it's impossible that EGW=128 needs LMUL=8.
-
pvanhout authored
-
yingopq authored
About unsigned max/min, ANDi is available for all ISA revisions in extend before slt insn. So that we can reduce one instruction.
-
Matt Arsenault authored
-
Mircea Trofin authored
-
Emma Pilkington authored
On gfx11 shaders run with PRIV=1, which causes `s_trap 2` to be treated as a nop, which means it isn't a correct lowering for the trap intrinsic. As a workaround, this commit instead lowers the trap intrinsic to instructions that simulate the behavior of s_trap 2. Fixes: SWDEV-438421
-
smanna12 authored
In the lambda function within clang::Sema::InstantiateFunctionDefinition, the return value of a function that may return null is now checked before dereferencing to avoid potential null pointer dereference issues which can lead to crashes or undefined behavior in the program.
-
David Sherwood authored
When SVE is available we can lower calls to get.active.lane.mask using the SVE whilelo instruction, however in practice since vXi1 types are not legal for NEON we often end up expanding the predicate into a vector of integers, e.g. v4i1 -> v4i32. This usually happens when we have to keep the predicate live out of the block, for example when the predicate is the incoming value to a PHI node in a tail-folded vector loop. Currently in such cases the intrinsic call has a cost of 1, which is far too low when considering the extra instructions required to expand the predicate. This patch fixes that by basing the cost on the number of lane moves required for expansion. This is required for a follow-on patch that adds the cost of the intrinsic call to the vectorisation cost model, so that we can teach the vectoriser to make better choices.
-
jeanPerier authored
Semantics usually fold SHAPE into an array constructor, but sometimes it cannot (like when the source is a function result that cannot be duplicated in expression analysis). Add lowering handling for shape.
-
Sergio Afonso authored
This patch updates lowering from PFT to MLIR of workshare loops to follow the loop wrapper approach. Unit tests impacted by this change are also updated. As the last patch of the stack, this should compile and pass unit tests.
-
Sergio Afonso authored
This patch introduces minimal changes to the MLIR to LLVM IR translation of `omp.wsloop` to support the loop wrapper approach. There is `omp.loop_nest` related translation code that should be extracted and shared among all loop operations (e.g. `omp.simd`). This would possibly also help in the addition of support for compound constructs later on. This first approach is only intended to keep things running after the transition to loop wrappers and not to add support for other use cases enabled by that transition. This PR on its own will not pass premerge tests. All patches in the stack are needed before it can be compiled and passes tests.
-
Sergio Afonso authored
This patch makes changes to the `scf.parallel` to `omp.parallel` + `omp.wsloop` lowering pass in order to introduce a nested `omp.loop_nest` as well, and to follow the new loop wrapper role for `omp.wsloop`. This PR on its own will not pass premerge tests. All patches in the stack are needed before it can be compiled and passes tests.
-
Sergio Afonso authored
This patch updates verifiers for `omp.ordered`, `omp.ordered.region`, `omp.cancel` and `omp.cancellation_point`, which check for a parent `omp.wsloop`. After transitioning to a loop wrapper-based approach, the expected direct parent will become `omp.loop_nest` instead, so verifiers need to take this into account. This PR on its own will not pass premerge tests. All patches in the stack are needed before it can be compiled and passes tests.
-
Hans Wennborg authored
to reflect that there are three variants.
-
Xu Zhang authored
Fixes #82659 There are some functions, such as `findRegisterDefOperandIdx` and `findRegisterDefOperand`, that have too many default parameters. As a result, we have encountered some issues due to the lack of TRI parameters, as shown in issue #82411. Following @RKSimon 's suggestion, this patch refactors 9 functions, including `{reads, kills, defines, modifies}Register`, `registerDefIsDead`, and `findRegister{UseOperandIdx, UseOperand, DefOperandIdx, DefOperand}`, adjusting the order of the TRI parameter and making it required. In addition, all the places that call these functions have also been updated correctly to ensure no additional impact. After this, the caller of these functions should explicitly know whether to pass the `TargetRegisterInfo` or just a `nullptr`. -
Sergio Afonso authored
This patch updates the definition of `omp.wsloop` to enforce the restrictions of a loop wrapper operation. Related tests are updated but this PR on its own will not pass premerge tests. All patches in the stack are needed before it can be compiled and passes tests.
-
Dinar Temirbulatov authored
[Clang][AArch64] Extend diagnostics when warning non/streaming about vector size difference (#88380) Add separate messages about passing arguments or returning parameters with scalable types. --------- Co-authored-by:Sander de Smalen <sander.desmalen@arm.com>
-
Fabio D'Urso authored
This can occur if the virtual address space is (almost) entirely mapped or heavily fragmented.
-
Krzysztof Parzyszek authored
This function will break up a construct into constituent leaf and composite constructs, e.g. if OMPD_c_d_e and OMPD_d_e are composite constructs, then OMPD_a_b_c_d_e will be broken up into the list {OMPD_a, OMPD_b, OMPD_c_d_e}. -
Nico Weber authored
Reverts d3f6c2c5, since ARMTargetDefEmitter.cpp has to be in llvm-min-tblgen too.
-
Hans Wennborg authored
-
Daniel Grumberg authored
This changes the handling of anonymous TagDecls to the following rules: - If the TagDecl is embedded in the declaration for some VarDecl (this is the only possibility for RecordDecls), then pretend the child decls belong to the VarDecl - If it's an EnumDecl proceed as we did previously, i.e., embed it in the enclosing DeclContext. Additionally this fixes a few issues with declaration fragments not consistently including "{ ... }" for anonymous TagDecls. To make testing these additions easier this patch fixes some text declaration fragments merging issues and updates tests accordingly. rdar://121436298 -
Nico Weber authored
-
Maksim Levental authored
Add bindings for LLVM pointer type.
-
Christian Ulmann authored
This commit enhances the LLVM dialect's Mem2Reg interfaces to support partial stores to memory slots. To achieve this support, the `getStored` interface method has to be extended with a parameter of the reaching definition, which is now necessary to produce the resulting value after this store.
-
Simon Pilgrim authored
Simplify callers which don't have their own DemandedElts mask. Noticed while reviewing #88801
-
Joseph Huber authored
-
Joseph Huber authored
Summary: The AMDGPU toolchain simply took the short name to get the link job instead of using the common utilities that respect options like `-fuse-ld`. Any linker that isn't `ld.lld` will fail, however we should be able to override it.
-
Allen authored
On AArch64, rdvl can accept a nagative value, while cntd/cntw/cnth can't. As we do support VScale with a negative multiply value, so we did not limit the negative value and instead took the hit of having the extra patterns according PR88108. Also add NoUseScalarIncVL to avoid affecting patterns works for -mattr=+use-scalar-inc-vl Fix https://github.com/llvm/llvm-project/issues/84620 -
Dmitry Chernenkov authored
-
Florian Hahn authored
Precommit tests for https://github.com/llvm/llvm-project/pull/83860.
-