- May 17, 2022
-
-
Utkarsh Saxena authored
Reduces time spent in findRef by 66%. Differential Revision: https://reviews.llvm.org/D125675
-
Robert Suderman authored
We were custom counting per bit for the clz instruction. Math dialect now has an intrinsic to do this in one instruction. Migrated to this instruction and fixed a minor bug math-to-llvm for the intrinsic. Reviewed By: mravishankar Differential Revision: https://reviews.llvm.org/D125592
-
zhijian authored
Summary: llvm-ar can not read empty big archive correctly. it output error as error: unable to load 'empty.a': truncated or malformed archive (characters in size field in archive member header are not all decimal numbers: '<bigaf>' Reviewers: James Henderson Differential Revision: https://reviews.llvm.org/D124017
-
Stanislav Mekhanoshin authored
This reverts ffbee7ac, see also bug 37653 which it was fixing. The bug claims this is an undocumented feature which actually works. In the reality it is documented as not working for a good reason. It likely does something, but it is useless anyway. These instructions write into the LDS. The LDS address is: M0 + inst_offset + (TIDinWave * 4). For a store wider than a DWORD neighboring lanes will overwrite each other. Differential Revision: https://reviews.llvm.org/D125409
-
Utkarsh Saxena authored
Differential Revision: https://reviews.llvm.org/D125682
-
Alex Bradbury authored
The logic around IsCanonical previously used getAsString and compared to "1". Just using getValueAsBit is simpler.
-
Jakub Kuderski authored
In this instance, the trailing return type does not improve readability as it repeats what is returned in the same line. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125697
-
Fangrui Song authored
The former is a bit misleading and a user may try installing zlib which will not help.
-
Fangrui Song authored
-
Rahman Lavaee authored
[llvm-objdump] Let --symbolize-operands symbolize basic block addresses based on the SHT_LLVM_BB_ADDR_MAP section. `--symbolize-operands` already symbolizes branch targets based on the disassembly. When the object file is created with `-fbasic-block-sections=labels` (ELF-only) it will include a SHT_LLVM_BB_ADDR_MAP section which maps basic blocks to their addresses. In such case `llvm-objdump` can annotate the disassembly based on labels inferred on this section. In contrast to the current labels, SHT_LLVM_BB_ADDR_MAP-based labels are created for every machine basic block including empty blocks and those which are not branched into (fallthrough blocks). The old logic is still executed even when the SHT_LLVM_BB_ADDR_MAP section is present to handle functions which have not been received an entry in this section. Reviewed By: jhenderson, MaskRay Differential Revision: https://reviews.llvm.org/D124560
-
Adrian Prantl authored
-
David Green authored
There have been some patterns in the AArch64 backend to optimize code of the form: ldrsh w8, [x0] scvtf s0, w8 to: ldr h0, [x0] sshll v0.4s, v0.4h, #0 scvtf s0, s0 The idea is to remove the GRP->FPR move, but in reality is making code larger and slower (or the same) on all the cpus I tried. This patch adds the UseAlternateSExtLoadCVTF32 predicate similar to nearby related pattern. Differential Revision: https://reviews.llvm.org/D125470
-
Sanjay Patel authored
The existing transform was wrong in 3 ways: 1. It created an extra instruction when the source and dest types don't match. 2. It did not account for an extra use of the icmp, so could create 2 extra insts. 3. It favored bit hacks over icmp (icmp generally has better analysis). This fixes #54692 (modeled by the PhaseOrdering tests). This is a minimal step to fix the bug, but we should likely invert the sibling transform for the "is negative" pattern too. The backend should be able to invert this back to a shift if that leads to better codegen.
-
Sanjay Patel authored
-
Craig Topper authored
This bug is in generic DAG combine and easily reproducible on many targets. Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D125640
-
Craig Topper authored
If we use multiply it would be with 0x0101 which is 1 more than a power of 2. On some targets we would expand this to shl+add. By avoiding the multiply earlier, we can generate better code. Note, PowerPC doesn't do the shl+add expansion of multiply so one of the tests increased in instruction count. Limiting to scalars because it almost always increased the number of instructions in vector tests. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D125638
-
Craig Topper authored
-
Philip Reames authored
-
Hongtao Yu authored
Current profile generation caculcates callsite body samples and call target samples separately. The former is done based on LBR range samples while the latter is done based on branch samples. Note that there's a subtle difference. LBR ranges is formed from two consecutive branch samples. Therefore the last entry in a LBR record will not be counted towards body samples while there's still a chance for it to be counted towards call targets if it is a function call. I'm making sense of the call body samples by updating it to the aggregation of call targets. Reviewed By: wenlei Differential Revision: https://reviews.llvm.org/D122609
-
Matthias Springer authored
This changes replaces the `fully-dynamic-layout-maps` options (which was badly named) with two new options: * `unknown-type-conversion` controls the layout maps on buffer types for which no layout map can be inferred. * `function-boundary-type-conversion` controls the layout maps on buffer types inside of function signatures. Differential Revision: https://reviews.llvm.org/D125615
-
- May 16, 2022
-
-
Ellis Hoag authored
Add a map from functions to load instructions that compute the profile bias. Previously we assumed that if the first instruction in the function was a load instruction, then it must be computing the bias. This was likely to work out because functions usually start with the `llvm.instrprof.increment` instruction, but optimizations could change this. For example, inlining into a non-profiled function. Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D114319
-
Sanjay Patel authored
-
Philip Reames authored
-
Simon Pilgrim authored
As mentioned on D123678 this appears to be causing namespace resolution issues on some versions of gcc.
-
David Green authored
-
Sanjay Patel authored
The tests (see C++ source in #54692) have multiple potential optimizations/canonicalizations, but we should be consistent since they are logically identical.
-
Sanjay Patel authored
The bug number was typo'd when it was added for D86243.
-
Alexey Bataev authored
The root of the buildvector can have only one use, otherwise it can be treated only as a final element of the previous buildvector sequence.
-
Nikita Popov authored
isImpliedCondition() currently handles and/or on the LHS, but not on the RHS, resulting in asymmetric behavior. This patch adds two new implication rules: * LHS ==> (RHS1 || RHS2) if LHS ==> RHS1 or LHS ==> RHS2 * LHS ==> !(RHS1 && RHS2) if LHS ==> !RHS1 or LHS ==> !RHS2 Differential Revision: https://reviews.llvm.org/D125551
-
Florian Hahn authored
This patch adds initial support for a pointer diff based runtime check scheme for vectorization. This scheme requires fewer computations and checks than the existing full overlap checking, if it is applicable. The main idea is to only check if source and sink of a dependency are far enough apart so the accesses won't overlap in the vector loop. To do so, it is sufficient to compute the difference and compare it to the `VF * UF * AccessSize`. It is sufficient to check `(Sink - Src) <u VF * UF * AccessSize` to rule out a backwards dependence in the vector loop with the given VF and UF. If Src >=u Sink, there is not dependence preventing vectorization, hence the overflow should not matter and using the ULT should be sufficient. Note that the initial version is restricted in multiple ways: 1. Pointers must only either be read or written, by a single instruction (this allows re-constructing source/sink for dependences with the available information) 2. Source and sink pointers must be add-recs, with matching steps 3. The step must be a constant. 3. abs(step) == AccessSize. Most of those restrictions can be relaxed in the future. See https://github.com/llvm/llvm-project/issues/53590. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D119078
-
Joe Nash authored
Address the FIXME by marking the sendmsg instructions with hasSideEffects. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D125569
-
Jay Foad authored
On GFX10 VOP3 instructions can have a literal operand, so the conversion from VOP3 MAD/FMA to VOP2 MADAK/MADMK/FMAAK/FMAMK will not happen in SIFoldOperands. The only benefit of the VOP2 form is code size, so do it in SIShrinkInstructions instead. Differential Revision: https://reviews.llvm.org/D125567
-
Nikita Popov authored
Add toKnownBits() method to mirror fromKnownBits(). We know the top bits that are constant between min and max. The return value for an empty range is chosen to be conservative.
-
Joe Nash authored
Includes MachineCode layer support and tests, and MIR tests not requiring CodeGen pass changes. Includes a small change in SMInstructions.td to correct encoded bits. Contributors: Petar Avramovic <Petar.Avramovic@amd.com> Dmitry Preobrazhensky <dmitry.preobrazhensky@amd.com> Depends on D125316 Patch 6/N for upstreaming of AMDGPU gfx11 architecture. Reviewed By: dp, Petar.Avramovic Differential Revision: https://reviews.llvm.org/D125319
-
Stephen Long authored
`#pragma alloc_text` is a MSVC pragma that names the code section where functions should be placed. It only applies to functions with C linkage. https://docs.microsoft.com/en-us/cpp/preprocessor/alloc-text?view=msvc-170 Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D125011
-
Mehdi Amini authored
-
Mehdi Amini authored
-
Nathan James authored
Reimplement the matching logic using Visitors instead of matchers. Benchmarks from running the check over SemaCodeComplete.cpp Before 0.20s, After 0.04s Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D125026
-
Liqin.Weng authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D123656
-
Louis Dionne authored
This way, we could use it for LLVM_ENABLE_PROJECTS too if desired. Differential Revision: https://reviews.llvm.org/D125121
-