- Sep 19, 2022
-
-
Nicolas Vasilache authored
Given an opOperand uniquely determined by the operation `%op` and the operand number `num`, the `transform.get_producer_of_operand %op[num]` returns the handle to the unique operation that produced the SSA value used as opOperand. The transform fails if the operand is a block argument. Differential Revision: https://reviews.llvm.org/D134171
-
Nicolas Vasilache authored
[mlir][Linalg] NFC - Cleanup internal transform APIs and produce better messages on failure to apply.
-
Max Kazantsev authored
This is required because if there is a pure loop-invariant instruction, Loop Rotation may decide to not clone it and just hoist it instead. If SCEV has previously cached that it was loop-variant (not being smart enough to prove invariance), we may end up with inconsistent cache state (which may later trigger false-negative assertion failures checking that something was invariant). This is a conservative fix that unconditionally drops the dispositions. We could only drop it if the hoisting has actually happened, but it should take some time understanding whether it's safe with all other things this function does. Differential Revision: https://reviews.llvm.org/D134167 Reviewed By: fhahn
-
Nuno Lopes authored
Alive2 doesn't support verification of optimizations that use inter-procedural analyses. Right now, clang uses GlobalsAA by default and there's no way to disable it. This leads to Alive2 producing false positives. The added flag allows us to skip global analyses altogether. Differential Revision: https://reviews.llvm.org/D134139
-
LLVM GN Syncbot authored
-
Simon Pilgrim authored
Add tests for the core static shuffle pattern match helpers
-
Lorenzo Chelini authored
The batch-reduce GEMM kernel essentially multiplies a sequence of input tensor blocks (which form a batch) and the partial multiplication results are reduced into a single output tensor block. See: https://ieeexplore.ieee.org/document/9139809 for more details. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D134163
-
Max Kazantsev authored
This bug was found by recent improvement in SCEV verifier. The code in LoopFuse directly reassigns blocks to be a part of a different loop, which should automatically invalidate all related cached loop dispositions. Differential Revision: https://reviews.llvm.org/D134173 Reviewed By: nikic
-
Max Kazantsev authored
It seems that it is sometimes broken. Initial motivation for this was investigation of https://github.com/llvm/llvm-project/issues/56260, but it also seems that we have found an unrelated bug in LoopFusion that leaves broken caches. Differential Revision: https://reviews.llvm.org/D134158 Reviewed By: nikic
-
David Green authored
Previously only using the UnsafeFPMath option, this now looks for the fast moth flags on the instructions, using the same flag flags as other backends.
-
Lorenzo Chelini authored
This reverts commit f381768a.
-
lorenzo chelini authored
The batch-reduce GEMM kernel essentially multiplies a sequence of input tensor blocks (which form a batch) and the partial multiplication results are reduced into a single output tensor block. See: https://ieeexplore.ieee.org/document/9139809 for more details. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D134163
-
Simon Pilgrim authored
-
Simon Pilgrim authored
DOS can't handle -passes='default<O3>' correctly
-
Nuno Lopes authored
As discussed in https://reviews.llvm.org/D133967
-
Valentin Clement authored
-
Max Kazantsev authored
Let's be honest about it, we don't drop loop dispositions for particular loops. Remove the parameter that misleadingly makes it apparent that we do.
-
Fangrui Song authored
-
Zi Xuan Wu (Zeson) authored
CodeGenSchedModels::hasReadOfWrite tries to predicate whether the WriteDef is contained in the list of ValidWrites of someone ProcReadAdvance, so that WriteID of WriteDef can be compressed and reusable. It tries to iterate all ProcReadAdvance entry, but not all ProcReadAdvance defs also inherit from SchedRead. Some ProcReadAdvances are defined by ReadAdvance.So it's not complete to enumerate all ProcReadAdvances if just iterate all SchedReads. Differential Revision: https://reviews.llvm.org/D132205
-
LiaoChunyu authored
shrinkdemandedconstant does some optimizations, but is not very friendly to riscv, targetShrinkDemandedConstant to limit the damage. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D134155
-
Kazu Hirata authored
These files don't seem to use StringSwitch.
-
Mingming Liu authored
- There is an outer while-loop and an inner for-loop in the test case. Inner-loop has `llvm.loop.unroll.enable` metadata that is not preserved. This happens around [1], when the loop metadata of outer loop overrides the inner loop metadata directly, without looking at whether inner-loop itself has loop metadata. [1] https://github.com/llvm/llvm-project/blob/ab755e65629ea098cb6faa77b13ac087849ffc67/llvm/lib/Transforms/Utils/Local.cpp#L1146 Differential Revision: https://reviews.llvm.org/D134014
-
Christian Sigg authored
-
Kazu Hirata authored
These files don't seem to use StringSwitch.
-
Kazu Hirata authored
This patch deprecates llvm::empty as I've migrated all known uses of llvm::empty(x) to x.empty(). Differential Revision: https://reviews.llvm.org/D134141
-
bixia1 authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D134062
-
Weining Lu authored
Reuse most of RISCV's implementation with several exceptions: 1. Assign signext/zeroext attribute to args passed in stack. On RISCV, integer scalars passed in registers have signext/zeroext when promoted, but are anyext if passed on the stack. This is defined in early RISCV ABI specification. But after this change [1], integers should also be signext/zeroext if passed on the stack. So I think RISCV's ABI lowering should be updated [2]. While in LoongArch ABI spec, we can see that integer scalars narrower than GRLEN bits are zero/sign-extended no matter passed in registers or on the stack. 2. Zero-width bit fields are ignored. This matches GCC's behavior but it hasn't been documented in ABI sepc. See https://gcc.gnu.org/r12-8294. 3. `char` is signed by default. There is another difference worth mentioning is that `char` is signed by default on LoongArch while it is unsigned on RISCV. This patch also adds `_BitInt` type support to LoongArch and handle it in LoongArchABIInfo::classifyArgumentType. [1] https://github.com/riscv-non-isa/riscv-elf-psabi-doc/commit/cec39a064ee0e5b0129973fffab7e3ad1710498f [2] https://github.com/llvm/llvm-project/issues/57261 Differential Revision: https://reviews.llvm.org/D132285
-
Chuanqi Xu authored
module unit directly Previously we lack a test which ensures that the module unit will generate initializer if it is compiled directly (instead of from a pcm file). Now we add the test back.
-
jacquesguan authored
This patch adds constant folder for ErfOp by using erf/erff of libm. Reviewed By: ftynse, Mogball Differential Revision: https://reviews.llvm.org/D134017
-
Kazu Hirata authored
-
wanglian authored
Reviewed By: tra Differential Revision: https://reviews.llvm.org/D134007
-
Chuanqi Xu authored
Previsouly the module-initializer*.cpp lives in the CodeGen dir instead of CodeGenCXX dir, which is not consistency with other tests since modules are features for C++.
-
LiaoChunyu authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D134154
-
Kazu Hirata authored
-
Emilia Dreamer authored
There already exists logic to disallow requires *expressions* to be treated as function declarations, but this expands it to include requires *clauses*, when they happen to also be parenthesized. Previously, in the following case: ``` template <typename T> requires(Foo<T>) T foo(); ``` The line with the requires clause was actually being considered as the line with the function declaration due to the parentheses, and the *real* function declaration on the next line became a trailing annotation (Together with https://reviews.llvm.org/D134049) Fixes https://github.com/llvm/llvm-project/issues/56213 Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D134052
-
Emilia Dreamer authored
In the following construction: `template <typename T> requires Foo<T> || Bar<T> auto func() -> int;` The `->` of the trailing return type was actually considered as an operator as part of the binary operation in the requires clause, with the precedence level of `PrecedenceArrowAndPeriod`, leading to fake parens being inserted in strange locations, that would never be closed. Fixes one part of https://github.com/llvm/llvm-project/issues/56213 (the rest will probably be in a separate patch) Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D134049
-
Lang Hames authored
Serialized calls to void-wrapper-functions should have zero bytes of argument data, but accessing ArgData[0] may (and will, in the case of SmallVector) fail if the argument data buffer is empty. This commit fixes the issue by adding a check for empty argument buffers.
-
Kazu Hirata authored
The LEA optimization pass visits each basic block of a given machine function. In each basic block, for each pair of LEAs that differ only in their displacement fields, we replace all uses of the second LEA with the first LEA while adjusting the displacement. Now, without this patch, after all the replacements are made, the following assert triggers: assert(MRI->use_empty(LastVReg) && "The LEA's def register must have no uses"); The replacement loop uses: for (MachineOperand &MO : llvm::make_early_inc_range(MRI->use_operands(LastVReg))) { which is equivalent to: for (auto UI = MRI->use_begin(LastVReg), UE = MRI->use_end(); UI != UE;) { MachineOperand &MO = *UI++; // <-- Look! That is, immediately after the post increment, make_early_inc_range already has the iterator for the next iteration in its mind. The problem is that in one iteration of the loop, we could replace two uses in a debug instruction like: DBG_VALUE_LIST !"r", !DIExpression(DW_OP_LLVM_arg, 0), %0:gr64, %0:gr64, ... So, the iterator for the next iteration becomes invalid. We end up traversing a garbage use list from that point on. In turn, we don't get to visit remaining uses. The patch fixes the problem by switching to a "draining" while loop: while (!MRI->use_empty(LastVReg)) { MachineOperand &MO = *MRI->use_begin(LastVReg); MachineInstr &MI = *MO.getParent(); The credit goes to Simon Pilgrim for reducing the test case. Fixes https://github.com/llvm/llvm-project/issues/57673 Differential Revision: https://reviews.llvm.org/D133631 -
Kazu Hirata authored
-
Yaxun (Sam) Liu authored
SimplifyCFG folds bool foo() { if (cond1) return false; if (cond2) return false; return true; } as bool foo() { if (cond1 | cond2) return false return true; } 'cond2' is called 'bonus insts' in branch folding since they introduce overhead since the original CFG could do early exit but the folded CFG always executes them. SimplifyCFG calculates the costs of 'bonus insts' of a folding a BB into its predecessor BB which shares the destination. If it is below bonus-inst-threshold, SimplifyCFG will fold that BB into its predecessor and cond2 will always be executed. When SimplifyCFG calculates the cost of 'bonus insts', it only consider 'bonus' insts in the current BB to be considered for folding. This causes issue for unrolled loops which share destinations, e.g. bool foo(int *a) { for (int i = 0; i < 32; i++) if (a[i] > 0) return false; return true; } After unrolling, it becomes bool foo(int *a) { if(a[0]>0) return false if(a[1]>0) return false; //... if(a[31]>0) return false; return true; } SimplifyCFG will merge each BB with its predecessor BB, and ends up with 32 'bonus insts' which are always executed, which is much slower than the original CFG. The root cause is that SimplifyCFG does not consider the accumulated cost of 'bonus insts' which are folded from different BB's. This patch fixes that by introducing a ValueMap to track costs of 'bonus insts' coming from different BB's into the same BB, and cuts off if the accumulated cost exceeds a threshold. Reviewed by: Artem Belevich, Florian Hahn, Nikita Popov, Matt Arsenault Differential Revision: https://reviews.llvm.org/D132408
-