- Apr 11, 2023
-
-
Diana Picus authored
The GFX11 NGG Streamout Instructions perform atomic operations on dedicated registers. At the moment, they lack machine memory operands, which causes the si-memory-legalizer pass to treat them conservatively and introduce several unnecessary waits and cache invalidations. This patch introduces a new address space to represent these special registers and teaches instruction selection to add memory operands with this new address space to DS_ADD/SUB_GS_REG_RTN. Since this address space is meant to be compiler-internal, we move it up a bit from the other address spaces and give it the number 128. According to the LLVM Language Reference, address space numbers can go all the way up to 2^24, but I'm not sure how well this is supported in practice [1], so using a smaller number seems safer. [1] https://github.com/llvm/llvm-project/blob/0107513fe79da7670e37c29c0862794a2213a89c/llvm/utils/TableGen/IntrinsicEmitter.cpp#L401 Differential Revision: https://reviews.llvm.org/D146031
-
Diana Picus authored
-
Heejin Ahn authored
When we encounter an `else`, `catch`, or `catch_all`, we currently just push the structure `NestingType` and don't preserve the original `if` and `try`'s signature. So after we pass `else`/`catch`/`catch_all`, we can't check if the values on stack have the correct types when we encounter `end_if` or `end_try`. This CL fixes the issue, and modifies the existing test to be correct (some of them had `try` without `catch`). Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D147881
-
Heejin Ahn authored
We disable type check in unreachable code, but when the unreachable code is enclosed within a block-like structure, the block as a whole has a valid type and we should continue type checking after the block. But it looks we currently only do that for blocks and not other block-like structures (`loop`s, `try`s, and `if`s). Also unreachable code within `if`'s true body shouldn't disable type checking in `else` body, and that in `try` body shouldn't disable type checking in `catch/catch_all` body. This also causes the values/types on the stack to be correctly checked when encounterint `catch`, `catch_all`, and `delegate`. Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D147852
-
Heejin Ahn authored
The current code is ``` ExpectBlockType = false; TC.setLastSig(*Signature.get()); if (ExpectBlockType) NestingStack.back().Sig = *Signature.get(); ``` Because of the first line, the third line's `if (ExpectBlockType)` is always false and we don't get to update `NestingStack.back().Sig`. This results in not correctly erroring out when the types of remaining values on the stack do not match the block type if the block type is written in the form of a function type. We should set `ExpectBlockType` to false after the `if`. Reviewed By: sbc100 Differential Revision: https://reviews.llvm.org/D147837 -
Sjoerd Meijer authored
This lowers the cost for FADD, FSUB, and FNEG. The motivation is to avoid over-eager SLP vectorisation, that makes it look like SLP vectorisation is profitable but results in significant slow downs. Lowering the cost for scalar FADD/FSUB costs helps the profitability decision to favour the scalar version where vectorisation isn't beneficial. Lowering the cost for these floating point operations makes sense because a lot of other instructions including many shuffles have only a cost of 1; these FADD/FSUB/FNEG instructions should not be twice the cost. Performance results show a 7% improvement for Imagick from SPEC FP 2017, a small improvement in Blender, and unchanged results for the other apps in SPEC. RAJAPerf is neutral and mostly shows no changes. Differential Revision: https://reviews.llvm.org/D146033
-
wanglei authored
-
Nikita Popov authored
Reassociate gep (gep ptr, idx1), idx2 to gep (gep ptr, idx2), idx1 if this would make the inner GEP loop invariant and thus hoistable. This is intended to replace an InstCombine fold that does this (in https://github.com/llvm/llvm-project/blob/04f61fb73dc6a994ab267d431f2fdaedc67430ff/llvm/lib/Transforms/InstCombine/InstructionCombining.cpp#L2006). The problem with the InstCombine fold is that LoopInfo is an optional dependency, so it is not performed reliably. Differential Revision: https://reviews.llvm.org/D146813
-
Nikita Popov authored
When converting this test to opaque pointers, we get a register move between the call and the inline asm. However, the test comment specifically says that there should be nothing between them. As far as I can tell, this is fine, both in that the inline asm doesn't use the relevant registers, but also more generally because the inline asm doesn't declare any clobbers, so really LLVM can do whatever, side effects or not. The test was added by 618ce3e8 with only a reference to Apple's internal issue tracker. Differential Revision: https://reviews.llvm.org/D147512
-
Vlad Serebrennikov authored
[[https://wg21.link/p1787 | P1787]]: CWG1822 is resolved by specifying that the body of a lambda remains in the surrounding (function parameter) scope. Wording: A parameter-declaration-clause P introduces a function parameter scope that includes P. <...> If P is associated with a lambda-declarator, its scope extends to the end of the compound-statement in the lambda-expression. ([basic.scope.param]) Reviewed By: #clang-language-wg, shafik Differential Revision: https://reviews.llvm.org/D147836
-
John McIver authored
This change is made to enable conversion of a masked icmp splat vector containing poison/undef to an equality expression. llvm::decomposeBitTestICmp Alive2 correctness examples using splat/masking vectors: SLT < https://alive2.llvm.org/ce/z/pPTTHh SLE <= https://alive2.llvm.org/ce/z/qQhAmU SGT > https://alive2.llvm.org/ce/z/koFHzF SGE >= https://alive2.llvm.org/ce/z/3SNz2S ULT <u https://alive2.llvm.org/ce/z/W8ktzQ ULE <=u https://alive2.llvm.org/ce/z/G5SdUY UGT >u https://alive2.llvm.org/ce/z/WFwYxq UGE >=u https://alive2.llvm.org/ce/z/DzJszP Tests have been verified using Alive2: icmp-logical.ll: @nomask_splat_and_B_allones https://alive2.llvm.org/ce/z/zmJwQU icmp-logical.ll: @nomask_splat_and_B_mixed https://alive2.llvm.org/ce/z/ktzgzd signed-truncation-check.ll: @positive_vec_undef0 https://alive2.llvm.org/ce/z/-sTRLD Differential Revision: https:/... -
John McIver authored
Add tests to verifying future support for splat vectors containing poison/undef in llvm::decomposeBitTestICmp. Differential Revision: https://reviews.llvm.org/D143031
-
Max Kazantsev authored
-
Shivam Gupta authored
This reverts commit e08af170.
-
Job Noorman authored
As discussed in 1d1b3c49, instruction flags set in the *.td files are under-approximations. For C_JR, isBranch and isConditionalBranch are set even though it is used for for returns which are not considered branches. This patch proposes to remove those flags from C_JR. More detailed analysis can be implemented in RISCVMCInstrAnalysis. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D147784
-
pvanhout authored
Use an `OptimizationRemark` for them even though it's not really an optimization. It just integrates better with the other diagnostics (enabling is easy with `-pass-remark`). Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D147703
-
Max Kazantsev authored
-
Max Kazantsev authored
We already hoist min/max functions and want to do more of this kind. Some refactoring to make growth points for it.
-
Max Kazantsev authored
-
Craig Topper authored
-
Douglas Chen authored
[clang-tidy] Fix hungarian notation failed to indicate the number of asterisks in check-clang-extra-clang-tidy-checkers-readability Fix hungarian notation failed to indicate the number of asterisks for the pointers of multiple word types. - WRONG: `unsigned char* value` : `value` --> `ucValue` - RIGHT: `unsigned cahr* value` : `value` --> `pucValue` - RIGHT: `unsigned char** value` : `value` --> `ppucValue` Reviewed By: amurzeau, PiotrZSL Differential Revision: https://reviews.llvm.org/D147779
-
Bing1 Yu authored
Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D147771
-
Tobias Gysi authored
The revision ensures the newly introduced argument and result handlers cannot be used for type conversion. Instead use the existing materializeCallConversion hook to perform type conversions. Reviewed By: Dinistro Differential Revision: https://reviews.llvm.org/D147605
-
Kazu Hirata authored
-
sgokhale authored
This patch splits a restore point to allow it to only post-dominate blocks reachable by use or def of CSRs(Callee Saved Registers)/FI(Frame Index). Benchmarking this on SPEC2017, this gives around 4% improvement on povray and no significant change for others. Co-authored-by: junbuml Differential Revision: https://reviews.llvm.org/D42600
-
Mehdi Amini authored
Caught by ASAN on a bot: https://lab.llvm.org/buildbot/#/builders/168/builds/12872/steps/14/logs/stdio
-
Shengchen Kan authored
-
Craig Topper authored
Instead of copying a vector and clearing the original, we can swap with an empty vector.
-
Joshua Cao authored
* combine zext and sext into the one switch case * combine vscale and udiv into one switch case * renames according to LLVM style
-
Max Kazantsev authored
-
Bing1 Yu authored
Reviewed By: xiangzhangllvm Differential Revision: https://reviews.llvm.org/D147921
-
Shengchen Kan authored
Avoid the LIT fail when only arm/aarch64 is built.
-
Alex Brachet authored
-
Alex Brachet authored
Differential Revision: https://reviews.llvm.org/D147464
-
Owen Shepherd authored
Fixes error about missing reference to __atomic_fetch_add_8 when linking LLVMOrcJit on ARMv6. Reviewed By: lhames Differential Revision: https://reviews.llvm.org/D147937
-
Shengchen Kan authored
Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D147835
-
Joshua Cao authored
SCEV determines that loops with trip count >=2^32 have a trip multiple of 1 to guard against huge multiples. This patch stregthens this to instead find the greatest power of 2 divisor that is less than the threshold. Differential Revision: https://reviews.llvm.org/D147868
-
Joshua Cao authored
-
Joshua Cao authored
-
Joshua Cao authored
This patch improves on https://reviews.llvm.org/D110587. To summarize the patch, given backedge-taken count BC, trip count TC is `BC + 1`. However, we don't know if BC we might overflow. So the patch modifies TC computation to `1 + zext(BC)`. This patch only adds the zext if necessary by looking at the constant range. If we can determine that BC cannot be the max value for its bitwidth, then we know adding 1 will not overflow, and the zext is not needed. We apply loop guards before computing TC to get more data. The primary motivation is to support my work on more precise trip multiples in https://reviews.llvm.org/D141823. For example: ``` void test(unsigned n) __builtin_assume(n % 6 == 0); for (unsigned i = 0; i < n; ++i) foo(); ``` Prior to this patch, we had `TC = 1 + zext(-1 + 6 * ((6 umax %n) /u 6))<nuw>`. SCEV range computation is able to determine that the BC cannot be the max value, so the zext is not needed. The result is `TC -> (6 * ((6 umax %n) /u 6))<nuw>`. From here, we would be able to determine that %n is a multiple of 6. There was one change in LoopCacheAnalysis/LoopInterchange required. Before this patch, if a loop has BC = false, it would compute `TC -> 1 + zext(false) -> 1`, which was fine. After this patch, it computes `TC -> 1 + false = true`. CacheAnalysis would then sign extend the `true`, which was not the intended the behavior. I modified CacheAnalysis such that it would only zero extend trip counts. This patch is not NFC, but also does not change any SCEV outputs. I would like to get this patch out first to make work with trip multiples easier. Differential Revision: https://reviews.llvm.org/D147117
-