- Sep 19, 2022
-
-
Lorenzo Chelini authored
This reverts commit f381768a.
-
lorenzo chelini authored
The batch-reduce GEMM kernel essentially multiplies a sequence of input tensor blocks (which form a batch) and the partial multiplication results are reduced into a single output tensor block. See: https://ieeexplore.ieee.org/document/9139809 for more details. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D134163
-
Simon Pilgrim authored
-
Simon Pilgrim authored
DOS can't handle -passes='default<O3>' correctly
-
Nuno Lopes authored
As discussed in https://reviews.llvm.org/D133967
-
Valentin Clement authored
-
Max Kazantsev authored
Let's be honest about it, we don't drop loop dispositions for particular loops. Remove the parameter that misleadingly makes it apparent that we do.
-
Fangrui Song authored
-
Zi Xuan Wu (Zeson) authored
CodeGenSchedModels::hasReadOfWrite tries to predicate whether the WriteDef is contained in the list of ValidWrites of someone ProcReadAdvance, so that WriteID of WriteDef can be compressed and reusable. It tries to iterate all ProcReadAdvance entry, but not all ProcReadAdvance defs also inherit from SchedRead. Some ProcReadAdvances are defined by ReadAdvance.So it's not complete to enumerate all ProcReadAdvances if just iterate all SchedReads. Differential Revision: https://reviews.llvm.org/D132205
-
LiaoChunyu authored
shrinkdemandedconstant does some optimizations, but is not very friendly to riscv, targetShrinkDemandedConstant to limit the damage. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D134155
-
Kazu Hirata authored
These files don't seem to use StringSwitch.
-
Mingming Liu authored
- There is an outer while-loop and an inner for-loop in the test case. Inner-loop has `llvm.loop.unroll.enable` metadata that is not preserved. This happens around [1], when the loop metadata of outer loop overrides the inner loop metadata directly, without looking at whether inner-loop itself has loop metadata. [1] https://github.com/llvm/llvm-project/blob/ab755e65629ea098cb6faa77b13ac087849ffc67/llvm/lib/Transforms/Utils/Local.cpp#L1146 Differential Revision: https://reviews.llvm.org/D134014
-
Christian Sigg authored
-
Kazu Hirata authored
These files don't seem to use StringSwitch.
-
Kazu Hirata authored
This patch deprecates llvm::empty as I've migrated all known uses of llvm::empty(x) to x.empty(). Differential Revision: https://reviews.llvm.org/D134141
-
bixia1 authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D134062
-
Weining Lu authored
Reuse most of RISCV's implementation with several exceptions: 1. Assign signext/zeroext attribute to args passed in stack. On RISCV, integer scalars passed in registers have signext/zeroext when promoted, but are anyext if passed on the stack. This is defined in early RISCV ABI specification. But after this change [1], integers should also be signext/zeroext if passed on the stack. So I think RISCV's ABI lowering should be updated [2]. While in LoongArch ABI spec, we can see that integer scalars narrower than GRLEN bits are zero/sign-extended no matter passed in registers or on the stack. 2. Zero-width bit fields are ignored. This matches GCC's behavior but it hasn't been documented in ABI sepc. See https://gcc.gnu.org/r12-8294. 3. `char` is signed by default. There is another difference worth mentioning is that `char` is signed by default on LoongArch while it is unsigned on RISCV. This patch also adds `_BitInt` type support to LoongArch and handle it in LoongArchABIInfo::classifyArgumentType. [1] https://github.com/riscv-non-isa/riscv-elf-psabi-doc/commit/cec39a064ee0e5b0129973fffab7e3ad1710498f [2] https://github.com/llvm/llvm-project/issues/57261 Differential Revision: https://reviews.llvm.org/D132285
-
Chuanqi Xu authored
module unit directly Previously we lack a test which ensures that the module unit will generate initializer if it is compiled directly (instead of from a pcm file). Now we add the test back.
-
jacquesguan authored
This patch adds constant folder for ErfOp by using erf/erff of libm. Reviewed By: ftynse, Mogball Differential Revision: https://reviews.llvm.org/D134017
-
Kazu Hirata authored
-
wanglian authored
Reviewed By: tra Differential Revision: https://reviews.llvm.org/D134007
-
Chuanqi Xu authored
Previsouly the module-initializer*.cpp lives in the CodeGen dir instead of CodeGenCXX dir, which is not consistency with other tests since modules are features for C++.
-
LiaoChunyu authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D134154
-
Kazu Hirata authored
-
Emilia Dreamer authored
There already exists logic to disallow requires *expressions* to be treated as function declarations, but this expands it to include requires *clauses*, when they happen to also be parenthesized. Previously, in the following case: ``` template <typename T> requires(Foo<T>) T foo(); ``` The line with the requires clause was actually being considered as the line with the function declaration due to the parentheses, and the *real* function declaration on the next line became a trailing annotation (Together with https://reviews.llvm.org/D134049) Fixes https://github.com/llvm/llvm-project/issues/56213 Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D134052
-
Emilia Dreamer authored
In the following construction: `template <typename T> requires Foo<T> || Bar<T> auto func() -> int;` The `->` of the trailing return type was actually considered as an operator as part of the binary operation in the requires clause, with the precedence level of `PrecedenceArrowAndPeriod`, leading to fake parens being inserted in strange locations, that would never be closed. Fixes one part of https://github.com/llvm/llvm-project/issues/56213 (the rest will probably be in a separate patch) Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D134049
-
Lang Hames authored
Serialized calls to void-wrapper-functions should have zero bytes of argument data, but accessing ArgData[0] may (and will, in the case of SmallVector) fail if the argument data buffer is empty. This commit fixes the issue by adding a check for empty argument buffers.
-
Kazu Hirata authored
The LEA optimization pass visits each basic block of a given machine function. In each basic block, for each pair of LEAs that differ only in their displacement fields, we replace all uses of the second LEA with the first LEA while adjusting the displacement. Now, without this patch, after all the replacements are made, the following assert triggers: assert(MRI->use_empty(LastVReg) && "The LEA's def register must have no uses"); The replacement loop uses: for (MachineOperand &MO : llvm::make_early_inc_range(MRI->use_operands(LastVReg))) { which is equivalent to: for (auto UI = MRI->use_begin(LastVReg), UE = MRI->use_end(); UI != UE;) { MachineOperand &MO = *UI++; // <-- Look! That is, immediately after the post increment, make_early_inc_range already has the iterator for the next iteration in its mind. The problem is that in one iteration of the loop, we could replace two uses in a debug instruction like: DBG_VALUE_LIST !"r", !DIExpression(DW_OP_LLVM_arg, 0), %0:gr64, %0:gr64, ... So, the iterator for the next iteration becomes invalid. We end up traversing a garbage use list from that point on. In turn, we don't get to visit remaining uses. The patch fixes the problem by switching to a "draining" while loop: while (!MRI->use_empty(LastVReg)) { MachineOperand &MO = *MRI->use_begin(LastVReg); MachineInstr &MI = *MO.getParent(); The credit goes to Simon Pilgrim for reducing the test case. Fixes https://github.com/llvm/llvm-project/issues/57673 Differential Revision: https://reviews.llvm.org/D133631 -
Kazu Hirata authored
-
Yaxun (Sam) Liu authored
SimplifyCFG folds bool foo() { if (cond1) return false; if (cond2) return false; return true; } as bool foo() { if (cond1 | cond2) return false return true; } 'cond2' is called 'bonus insts' in branch folding since they introduce overhead since the original CFG could do early exit but the folded CFG always executes them. SimplifyCFG calculates the costs of 'bonus insts' of a folding a BB into its predecessor BB which shares the destination. If it is below bonus-inst-threshold, SimplifyCFG will fold that BB into its predecessor and cond2 will always be executed. When SimplifyCFG calculates the cost of 'bonus insts', it only consider 'bonus' insts in the current BB to be considered for folding. This causes issue for unrolled loops which share destinations, e.g. bool foo(int *a) { for (int i = 0; i < 32; i++) if (a[i] > 0) return false; return true; } After unrolling, it becomes bool foo(int *a) { if(a[0]>0) return false if(a[1]>0) return false; //... if(a[31]>0) return false; return true; } SimplifyCFG will merge each BB with its predecessor BB, and ends up with 32 'bonus insts' which are always executed, which is much slower than the original CFG. The root cause is that SimplifyCFG does not consider the accumulated cost of 'bonus insts' which are folded from different BB's. This patch fixes that by introducing a ValueMap to track costs of 'bonus insts' coming from different BB's into the same BB, and cuts off if the accumulated cost exceeds a threshold. Reviewed by: Artem Belevich, Florian Hahn, Nikita Popov, Matt Arsenault Differential Revision: https://reviews.llvm.org/D132408 -
Carl Ritson authored
Special registers, e.g. MODE, do not have register classes so will cause null pointer exception if passed to isSGPRReg. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D134025
-
Kazu Hirata authored
This patch, a follow-up for 588628de, removes duplicate types like T and PointerT in favor of reference and pointer, respectively.
-
Xiang Li authored
Add vector version of abs as ``` __attribute__((clang_builtin_alias(__builtin_elementwise_abs))) int2 abs (int2 ); __attribute__((clang_builtin_alias(__builtin_elementwise_abs))) int3 abs (int3 ); ``` To make this work. Allowed custom type checking builtins to be recelareable. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D133737
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
This patch moves getInlineCostWrapper to an anonymous namespace. While I am at it, I'm moving the function closer to the beginning of the file so that I can use it elsewhere in the file without a forward declaration.
-
Shafik Yaghmour authored
[Clang] Fix compat diagnostic to detect a nontype template parameter has a placeholder type using getContainedAutoType() Based on the changes introduced by 15361a21 it looks like C++17 compatibility diagnostic should have been checking getContainedAutoType(). This fixes: https://github.com/llvm/llvm-project/issues/57369 https://github.com/llvm/llvm-project/issues/57643 https://github.com/llvm/llvm-project/issues/57793 Differential Revision: https://reviews.llvm.org/D132990
-
Sanjay Patel authored
This was originally part of D133788. There are no visible regressions. All of the diffs show a large unsigned constant becoming a small negative constant. This should be better for analysis (and slightly less compile-time) and codegen.
-
Kazu Hirata authored
I'm planning to deprecate and eventually remove llvm::empty. Note that no use of llvm::empty requires the ability of llvm::empty to determine the emptiness from begin/end only.
-
owenca authored
Differential Revision: https://reviews.llvm.org/D134103
-