- Feb 03, 2024
-
-
Simon Pilgrim authored
-
Simon Pilgrim authored
Shift down the value so the active bits are at the lsb
-
Krystian Stasiowski authored
Consider the following: ``` namespace N0 { namespace N1 { template<typename T> int x1 = 0; } using namespace N1; } template<> int N0::x1<int>; ``` According to [dcl.meaning.general] p3.3: > - If the _declarator_ declares an explicit instantiation or a partial or explicit specialization, the _declarator_ does not bind a name. If it declares a class member, the terminal name of the _declarator-id_ is not looked up; otherwise, **only those lookup results that are nominable in `S` are considered when identifying any function template specialization being declared**. In particular, the requirement for lookup results to be nominal in the lookup context of the terminal name of the _declarator-id_ only applies to function template specializations -- not variable template specializations. We currently reject the above declaration, but we do (correctly) accept it if the using-directive is replaced with a `using` declaration naming `N0::N1::x1`. This patch makes it so the above specialization is (correctly) accepted. -
Nathan Gauër authored
This new analysis returns a hierarchical view of the convergence regions in the given function. This will allow our passes to query which basic block belongs to which convergence region, and structurize the code in consequence. Definition ---------- A convergence region is a CFG with: - a single entry node. - one or multiple exit nodes (different from LLVM's regions). - one back-edge - zero or more subregions. Excluding sub-regions nodes, the nodes of a region can only reference a single convergence token. A subregion uses a different convergence token. Algorithm --------- This algorithm assumes all loops are in the Simplify form. Create an initial convergence region for the whole function. - the convergence token is the function entry token. - the entry is the function entrypoint. - Exits are all the basic blocks terminating with a return instruction. Take the function CFG, and process it in DAG order (ignoring back-edges). If a basic block is a loop header: - Create a new region. - The parent region is the parent's loop region if any, otherwise, the top level region. - The region blocks are all the blocks belonging to this loop. - For each loop exit: - visit the rest of the CFG in DAG order (ignore back-edges). - if the region's convergence token is found, add all the blocks dominated by the exit from which the token is reachable to the region. - continue the algorithm with the loop headers successors.
-
Manish Kausik H authored
Prior to this patch, SelectionDAG generated aligned move onto stacks for AVX registers when the function was marked as a no-realign-stack function. This lead to misalignment between the stack and the instruction generated. This patch fixes the issue. Fixes #77730
-
Philip Reames authored
The primary motivation of this patch is to add testing infrastructure atop the recently landed 8ad14b6d, so that we can separate the costing aspects of strided memory operations from the SLP implementation details. I want to be clear that I am *not* proposing that we use the vp.strided.* forms as our canonical IR representation. I'm merely using them as a testing vehicle to exercise the costing machinery. The canonical IR form remains a masked.gather or masked.scatter. I do want to explore adding a non-vp strided load/store intrinsic, but that's a separate line of work. There is one costing change included in this. As I wrote my test, I discovered that the default implementation was scalarized (if invoked via generic routines such as getInstructionCost), and when adding the call into the strided specific costing discovered that we hadn't modeled the fallback to scalarization properly in the initial patch. After fixing that, there is a minor difference in scalarization cost reported for the unaligned case but I believe that to be uninteresting. For the record, I did confirm that vp.strided.store is lowered to a strided store on RISCV. :)
-
Rushi Bhamani authored
-
Philip Reames authored
Follow up to https://github.com/llvm/llvm-project/pull/74747 This change extends the previously added fixed expansion threshold by scaling down the cost allowed for an expansion for a loop with either a small known trip count or a profile which indicates the trip count is likely small. The goal here is to improve code generation for a loop nest where the outer loop has a high trip count, and the inner loop runs only a handful of iterations. --------- Co-authored-by:
Nikita Popov <github@npopov.com>
-
Guillaume Chatelet authored
-
LLVM GN Syncbot authored
-
Nikolas Klauser authored
This patch introduces a new trait to represent whether a type is trivially relocatable, and uses that trait to optimize the growth of a std::vector of trivially relocatable objects. ``` -------------------------------------------------- Benchmark old new -------------------------------------------------- bm_grow<int> 1354 ns 1301 ns bm_grow<std::string> 5584 ns 3370 ns bm_grow<std::unique_ptr<int>> 3506 ns 1994 ns bm_grow<std::deque<int>> 27114 ns 27209 ns ``` This also changes to order of moving and destroying the objects when growing the vector. This should not affect our conformance.
-
Natalie Chouinard authored
-
- Feb 02, 2024
-
-
Nikita Popov authored
To allow reusing it in IndVars.
-
Yaxun (Sam) Liu authored
AMDGPU does not support unaligned atomics, therefore make the warning an error. This patch is transferred from https://reviews.llvm.org/D99201
-
Simon Pilgrim authored
-
alexfh authored
Implements the fix proposed by Evgeny Eltsin on https://github.com/llvm/llvm-project/pull/66514#issuecomment-1924039038. No test case provided, since the bug is extremely sensitive to the preprocessor state (headers, macros, including the ones defined on command line), and it turned out to be non-trivial to create an isolated test.
-
Nikita Popov authored
-
Simon Pilgrim authored
Extend #79989 slightly to use KnownBits on the CTPOP input - this should make it easier to add additional cases identified in #79823
-
lifengxiang1025 authored
Fix one corner case when `CallStackTrie` has a single chain to leaf with multi alloc type. This will cause stackIds in function summary is empty.
-
Harald van Dijk authored
Currently, half operations can be promoted in one of two ways. * If softPromoteHalfType() returns false, fp16 values are passed around in fp32 registers, and whole chains of fp16 operations are promoted to fp32 in one go. * If softPromoteHalfType() returns true, fp16 values are passed around in i16 registers, and individual fp16 operations are promoted to fp32 and the result truncated to fp16 right away. The softPromoteHalfType behavior is necessary for correctness, but changing this for an existing target breaks the ABI. Therefore, this commit adds a third option: * If softPromoteHalfType() returns true and useFPRegsForHalfType() returns true as well, fp16 values are passed around in fp32 registers, but individual fp16 operations are promoted to fp32 and the result truncated to fp16 right away. This change does not yet update any target to make use of it.
-
Timm Bäder authored
Instead of asserting that it's wrong, assert the correct value. See the discussion in https://github.com/llvm/llvm-project/commit/a8b5994b337cf1d461202a65204a4ee6c5eae341
-
Konstantin Zhuravlyov authored
-
Simon Pilgrim authored
[X86] FP<->INT helpers - share the same SDLoc argument instead of recreating it over and over again.
-
Sergio Afonso authored
This patch removes the omp.target module attribute, since the information it held on the target CPU and features is available through the fir.target_cpu and fir.target_features module attributes. Target outlining during the MLIR to LLVM IR translation stage is updated, so that these attributes, at that point available as llvm.func attributes, are passed along to the newly created function.
-
jeanPerier authored
In fortran, it is possible to give a negative "i" in "character(i)" in which case the standard says the length is zero. So the length must be sanitized as max(0, user_input) in lowering. This is already done when lowering specification parts, but was not done when "character(i)" appears in array constructors. Sanitize the length when lowering SetLength in lowering. Fixes https://github.com/llvm/llvm-project/issues/80270
-
Florian Hahn authored
This patch adds a set of tests taken from/llvm/test/Transforms/IndVarSimplify/iv-poison.ll with multiple congruent IVs but different set of flags on the increments. Extra tests for https://github.com/llvm/llvm-project/pull/80430.
-
Valery Pykhtin authored
Reapply #71556 with added lit test constraint: `REQUIRES: amdgpu-registered-target`. This reverts commit 9791e541.
-
Sander de Smalen authored
When a function F has ZA and ZT0 state, calls another function G that only shares ZT0 state with its caller, F will have to save ZA before the call to G, and restore it afterwards (rather than setting up a lazy-sve). This is not yet implemented in LLVM and does not result in a compile-time error either. So instead of silently generating incorrect code, it's better to emit an error saying this is not yet implemented.
-
Jie Fu authored
llvm-project/llvm/lib/Target/X86/X86MCInstLower.cpp:1588:48: error: comparison of integers of different signs: 'unsigned int' and 'int' [-Werror,-Wsign-compare] if (C && C->getType()->getScalarSizeInBits() == SrcEltBits) { ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ^ ~~~~~~~~~~ 1 error generated. -
NAKAMURA Takumi authored
The current implementation (D138849) assumes `Branch`(es) would follow after the corresponding `Decision`. It is not true if `Branch`(es) are forwarded to expanded file ID. As a result, consecutive `Decision`(s) would be confused with insufficient number of `Branch`(es). `Expansion` will point `Branch`(es) in other file IDs if `Expansion` is included in the range of `Decision`. Fixes #77871 --------- Co-authored-by:Alan Phipps <a-phipps@ti.com>
-
Jie Fu authored
llvm-project/mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp:2763:23: error: comparison of integers of different signs: 'int64_t' (aka 'long') and 'size_type' (aka 'unsigned long') [-Werror,-Wsign-compare] assert(t1.getRank() == droppedDims.size() && "incorrect number of bits"); ~~~~~~~~~~~~ ^ ~~~~~~~~~~~~~~~~~~ llvm-project/mlir/lib/Dialect/MemRef/IR/MemRefOps.cpp:2764:38: error: comparison of integers of different signs: 'int64_t' (aka 'long') and 'size_type' (aka 'unsigned long') [-Werror,-Wsign-compare] assert(t1.getRank() - t2.getRank() == droppedDims.count() && ~~~~~~~~~~~~~~~~~~~~~~~~~~~ ^ ~~~~~~~~~~~~~~~~~~~ -
Simon Pilgrim authored
[X86] X86FixupVectorConstants - load+sign-extend vector constants that can be stored in a truncated form (#79815) Reduce the size of the vector constant by storing it in the constant pool in a truncated form, and sign-extend it as part of the load. I've extended the existing FixupConstant functionality to support these sext constant rebuilds - we still select the smallest stored constant entry and prefer vzload/broadcast/vextload for same bitwidth to avoid domain flips. I intend to add the matching load+zero-extend handling in a future PR, but that requires some alterations to the existing MC shuffle comments handling first.
-
Simon Pilgrim authored
[X86] LowerBuildVector* - share the same SDLoc argument instead of recreating it over and over again.
-
Nikolas Klauser authored
-
Jay Foad authored
Use !tolower instead of repeating the name when defining a renamed Real instruction.
-
Benjamin Kramer authored
-
chuongg3 authored
Add support of i16 vector operation for BSWAP and change to TableGen to select instructions Handle vector types that are smaller/larger than legal for BSWAP
-
Simon Pilgrim authored
This is the first basic proposal in #79823 - we can investigate improving support for other widths if we can find further use cases.
-
Nikolas Klauser authored
This reduces the time to include `<vector>` from 468ms to 367ms.
-
Peter Smith authored
-