- May 28, 2024
-
-
Artem Kroviakov authored
Building on top of [#88204](https://github.com/llvm/llvm-project/pull/88204), this PR adds support for converting `vector.insert` into an equivalent `vector.shuffle` operation that operates on linearized (1-D) vectors.
-
Kelvin Li authored
-
Matt Arsenault authored
-
josel-amd authored
-
David Green authored
Other than some additional checks needed for compare predicates and selects with scalar condition operands, these are relatively simple additions to what already exists.
-
Shengchen Kan authored
The generated table will be used in #93508
-
Stefan Gränitz authored
-
Eymen Ünay authored
In ARM mode, the Program Counter (PC) points to the current instruction's address + 8 instead of + 4. An offset is added to RuntimeDyldChecker to use `next_pc` expression in JITLink tests with both Thumb and Arm.
-
Tom Eccles authored
The pass constructor can be generated automatically. This pass is module-level and then runs on all of the relevant HLFIR operations inside of the module, no matter what top level operation they are inside of.
-
Adrian Kuegel authored
It removed the dependency from the wrong target. Also, we need to remove the header include to be able to remove the dependency from VectorToSPIRV.
-
Abid Qadeer authored
The fortran arrays use 'dataLocation', 'rank', 'allocated' and 'associated' fields of the DICompositeType. These were not available in 'DICompositeTypeAttr'. This PR adds the missing fields. --------- Co-authored-by:Tobias Gysi <tobias.gysi@nextsilicon.com>
-
Louis Dionne authored
This is a first step towards splitting up the <__config> header. The <__config> header is large and rather disorganized at this point, leading to confusion and subtle mistakes. For example, we never noticed that the string layout used on arm64 was only enabled for the Clang compiler, as the setting being in the compiler == clang block was probably never intentional. The danger of splitting up the <__config> header is to implicitly use undefined macros that should have been defined prior to their usage, however this can be remediated with -Wundef and we've started moving towards -Wundef enforceable macros.
-
Lukacma authored
Reverts llvm/llvm-project#88251
-
Vlad Serebrennikov authored
A follow up for #93318. Discussion happened at https://github.com/llvm/llvm-project/pull/93318#discussion_r1616281934
-
Ralender authored
-
Lukacma authored
According to the specification in https://github.com/ARM-software/acle/pull/309 this adds the intrinsics ``` svbfloat16x2_t svclamp[_single_bf16_x2](svbfloat16x2_t zd, svbfloat16_t zn, svbfloat16_t zm) __arm_streaming; svbfloat16x4_t svclamp[_single_bf16_x4](svbfloat16x4_t zd, svbfloat16_t zn, svbfloat16_t zm) __arm_streaming; ``` These are available only if __ARM_FEATURE_SME_B16B16 is enabled.
-
Kunwar Grover authored
Reverts llvm/llvm-project#93488 Buildbot failure: https://lab.llvm.org/buildbot/#/builders/220/builds/39911
-
Timm Bäder authored
-
Kunwar Grover authored
These passes have been depreciated for a long time and replaced by one-shot bufferization. These passes are also unsafe because they do not check for read-after-write conflicts.
-
David Green authored
This just adds splat constants, which can be treated like any other splat which hopefully makes them very simple. It does not try to handle more complex constant vectors yet, just the more common splats.
-
Simon Pilgrim authored
[X86] isHorizontalBinOp - always create HADD/SUB if it will be merged with another existing HADD/SUB Fixes some more cases from #34072 where undemanded vector elements prevent HADD/SUB being matched on slow targets
-
Stefan Gränitz authored
Until now the IncrExecutor was created lazily on the first execution request. In order to process the PTUs that come from initialization, we have to do it upfront implicitly.
-
Nikita Popov authored
gep inbounds of undef can only be folded to poison if we know that the offset is non-zero. I don't think precise handling here is important, so just drop the inbounds special case. This matches what InstSimplify does.
-
Nikita Popov authored
If the offset is zero, then returning poison here is not correct.
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
David Spickett authored
DumpValueObjectOptions can only be created and modified from C++. This means it's currently only testable from Python by calling some command that happens to use one, and even so, you can't pick which options get chosen. So we have decent coverage for the major options that way, but I want to add more niche options that will be harder to test from Python (register field options). So this change adds some "unit tests", though it's stretching the definition to the point it's more "test written in C++". So we can test future options in isolation. Since I want to add options specific to enums, that's all it covers. There is a test class that sets up the type system so it will be easy to test other types in future (e.g. structs, which register fields also use).
-
JP Lehr authored
This broke several buildbots: https://lab.llvm.org/buildbot/#/builders/193 https://lab.llvm.org/buildbot/#/builders/259 https://lab.llvm.org/staging/#/builders/185 https://lab.llvm.org/staging/#/builders/140 This reverts commit 435ea21c.
-
Cullen Rhodes authored
-force-streaming-compatible-sve was renamed in #92774 but this test was missed, no longer required so removing.
-
David Green authored
VPTState was holding static state, and acting as both the info in a VPTBlock and the overall state of all the blocks in the loop. This has been split up into a class (VPTBlock) to hold the instructions of one block, and VPTState that holds the overall state. The PredicatedInsts is also made into a map<MachineInstr *, SetVector<MachineInstr *>>, as the double-storing of MI inside a unique pointer is unneeded.
-
donald chen authored
This patch add more precise memory effect to linalg op. Including the following points: 1. Remove the read side effects for operands that are not used. 2. Set the effect for all side effects to "full".
-
Chuanqi Xu authored
-
Pierre van Houtryve authored
This is just something I noticed while going over this pass logic one more time and didn't cause issues (yet). If we find an indirect call, we stop looking assuming we added all functions to the list, but if not all functions in the module were indirectly callable, some may still be missing. Just to be safe, keep looking until we did everything we could to find dependencies, so we don't accidentally miss one.
-
Pierre van Houtryve authored
When I rewrote this, I made a mistake in the control flow. I thought we could just stop promoting if an alloca is too big to vectorize, but we can't. Other allocas in the list may be promotable and fit within the budget. Fixes SWDEV-455343
-
Shengchen Kan authored
1. Merge apx/ccmp-flags-copy-lowering.mir into apx/flags-copy-lowering.mir 2. Update check lines for flags-copy-lowering.mir by script This is for the coming NF (no flags update) support in flag copy lowering.
-
Shengchen Kan authored
-
Yingwei Zheng authored
This patch extends the transform `(icmp pred iM (shl iM %v, N), C) -> (icmp pred i(M-N) (trunc %v iM to i(M-N)), (trunc (C>>N))` to handle icmps with the flipped strictness of predicate. See the following case: ``` icmp ult i64 (shl X, 32), 8589934593 -> icmp ule i64 (shl X, 32), 8589934592 -> icmp ule i32 (trunc X, i32), 2 -> icmp ult i32 (trunc X, i32), 3 ``` Fixes the regression introduced by https://github.com/llvm/llvm-project/pull/86111#issuecomment-2098203152. Alive2 proofs: https://alive2.llvm.org/ce/z/-sp5n3 `nuw` cannot be propagated as we always use `ashr` here. I don't see the value of fixing this (see the test `test_icmp_shl_nuw`).
-
Ricky Zhou authored
xray instruments tail call function exits by inserting a nop sled before the tail call. When tracing is enabled, the nop sled is replaced with a call to `__xray_FunctionTailExit()`. This currently does not work for conditional tail calls, as the instrumentation assumes that the tail call will be unconditional. This causes two issues: - `__xray_FunctionTailExit()` is inappropately called even when the tail call is not taken. - `__xray_FunctionTailExit()`'s prologue/epilogue adjusts the stack pointer with add/sub instructions. This clobbers condition flags, which can flip the condition used for the tail call, leading to incorrect program behavior. Fix this by rewriting conditional calls when lowering patchable tail calls. With this change, a conditional patchable tail call like: ``` je target ``` Will be lowered to: ``` jne .fallthrough .p2align 1, .. .Lxray_sled_N: SLED_CODE jmp target .fallthrough: ```
-
Ricky Zhou authored
Calls to @llvm.xray.{custom,typed}event are lowered to an nop sled that can be patched with a call to an instrumentation function at runtime. Prior to this change, x86 codegen did not guarantee that these patched calls run with the appropriate stack alignment (particularly in leaf functions, where normal stack alignment assumptions may not hold). This leads to crashes on x86, as the custom event hook can end up using SSE instructions that assume a 16-byte aligned stack. Fix this by wrapping custom event hooks in CALLSEQ_START/CALLSEQ_END as done for regular function calls. One downside of this approach is that on functions whose stacks aren't already aligned, we may end up running stack alignment fixup instructions even when instrumentation is disabled. An alternative could be to make the custom event assembly trampolines stack-alignment-agnostic. This was the case in the past, but https://github.com/llvm/llvm-project/commit/b46c89892fe25bec197fd30f09b3a312da126422 removed this due to complexity in maintaining CFI directives for these stack adjustments. Since we are already willing to pay the call argument setup cost for custom hooks when instrumentation is disabled, I am hoping that an extra push/pop in this hopefully uncommon unaligned stack case is tolerable. -
Fangrui Song authored
-