- Dec 25, 2021
-
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Fangrui Song authored
This decreases the 0.2% time (no debug info) to nearly no.
-
Fangrui Song authored
This avoid repeated load of the unique_ptr in hot paths.
-
Fangrui Song authored
This mainly avoid `relsOrRelas` cost in `InputSectionBase::relocate`. `llvm::object::ELFFile::sections()` has redundant and expensive checks.
-
William S. Moses authored
LLVM Dialect in MLIR doesn't have a memmove op. This adds one. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D116274
-
Fangrui Song authored
This reverts 8cb7876c and follow-ups. GNU ld/gold/ld.lld -O has nothing to do with any code related linker optimizations. It has very small benefit (save 144Ki (.hash, .gnu_hash) with GNU ld, save 0.7% .debug_str with gold/ld.lld) while it makes gold/ld.lld significantly slower when linking RelWithDebInfo clang (gold: 16.437 vs 19.488; ld.lld: 1.882 vs 4.881).
-
Markus Böck authored
-
Fangrui Song authored
-
Fangrui Song authored
-
Fangrui Song authored
It is fairly easy to forget SectionBase::repl after ICF. Let ICF rewrite a Defined symbol's `section` field to avoid references to SectionBase::repl in subsequent passes. This slightly improves the --icf=none performance due to less indirection (maybe for --icf={safe,all} as well if most symbols are Defined). With this change, there is only one reference to `repl` (--gdb-index D89751). We can undo f4fb5fd7 (`Move Repl to SectionBase.`) but move `repl` to `InputSection` instead. Reviewed By: ikudrin Differential Revision: https://reviews.llvm.org/D116093 -
Gabriel Smith authored
Single-variant enums were still getting placed on a single line even when AllowShortEnumsOnASingleLine was false. This fixes that by checking that setting when looking to merge lines. Differential Revision: https://reviews.llvm.org/D116188
-
Groverkss authored
This patch moves some static functions from AffineStructures.cpp to Presburger/Utils.cpp and some to be private members of FlatAffineConstraints (which will later be moved to IntegerPolyhedron) to allow for a smoother transition for moving FlatAffineConstraints math functionality to Presburger/IntegerPolyhedron. This patch is part of a series of patches for moving math functionality to Presburger directory. Reviewed By: arjunp, bondhugula Differential Revision: https://reviews.llvm.org/D115869
-
Anastasia Stulova authored
-
- Dec 24, 2021
-
-
Alexey Zhikhartsev authored
Otherwise, it is possible that the state defined in the determinator block defines the state for the next iteration of the loop, rather than for the current one. Fixes llvm-test-suite's SingleSource/Regression/C/gcc-c-torture/execute/pr80421.c Differential Revision: https://reviews.llvm.org/D115832
-
alex-t authored
In 'trunc' i16/32/64 to i1 pattern the 'and $src, 1' node supply operand to 'setcc'. The latter is selected to S_CMP_EQ/V_CMP_EQ dependent on the divergence. In case the 'and' is scalar and 'setcc' is divergent, we need VGPR to SGPR copy to adjust input operand for V_CMP_EQ. This patch changes the S_AND_B32 to V_AND_B32_e64 in the 'trunc to i1' divergent patterns. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D116241
-
Matt Arsenault authored
Officially this is currently required to always use the datalayout's alloca address space. This may change in the future, and it's cleaner to propagate the existing alloca's addrspace anyway. This is a triple fix. Initially the change in simplifyAllocaArraySize would drop the address space, but produce output. Fixing this hit an assertion in the cast combine. This patch also makes the changes to handle this situation from a33e1280 dead, so eliminate it. InstCombine should not take it upon itself to introduce addrspacecasts, and preserve the original address space instead.
-
Simon Pilgrim authored
-
Shilei Tian authored
This patch adds the support for `atomic compare` in parser. The support in Sema and CodeGen will come soon. For now, it simply eimits an error when it is encountered. Reviewed By: ABataev Differential Revision: https://reviews.llvm.org/D115561
-
Alexandros Lamprineas authored
Converts RSHRN/RSHRN2 to RADDHN/RADDHN2 when the shift amount is half the width of the vector element. The latter has twice the throughput and half the latency on Arm out-of-order cores. Setting up the zero register adds no latency. Differential Revision: https://reviews.llvm.org/D116166
-
LLVM GN Syncbot authored
-
Krasimir Georgiev authored
We need some internal updates for this, shared directly with the author. This reverts commit 71b3bfde.
-
Nikita Popov authored
This fixes a typo in 81d69e1b. Of course we should only skip the particular store if it isn't removable, not bail out of the whole loop. Add a test to cover this case.
-
Nikita Popov authored
MemoryLocation::getForDest() checks this itself, call it directly.
-
Nikita Popov authored
We used to have both getLocForWrite() and getLocForWriteEx(). Now that we only have a single method, the "ex" suffix no longer makes sense.
-
Phoebe Wang authored
This reverts commit a954558e. Thanks Yuanfang's help. I think I found the root cause of the buildbot fail. The failed test has both Memory and Immediate X86Operand. All data of different operand kinds share the same memory space by a union definition. So it has chance we get the wrong result if we don't check the operand kind. It's probably it happen to be the correct value in my local environment so that I can't reproduce the fail. Differential Revision: https://reviews.llvm.org/D116090
-
Nikita Popov authored
As requested on D116210. The function is not necessarily well-defined without this precondition.
-
Nikita Popov authored
At this point the instruction may either have an analyzable write or be a terminator. For terminators, isRemovable() is not necessarily well-defined. Move the check until after we have ensured that it is not a terminator.
-
Nikita Popov authored
The only non-trivial change here is that the isReadClobber() check for redundant stores is now on the DefLoc, not the UpperLoc. This is semantically the right location to use, though in practice it makes no difference (the locations are either the same, or the def inst does not read).
-
Fangrui Song authored
-
Nikita Popov authored
We have Value->stripInBoundsConstantOffsets() which does what we want here, but the inbounds requirement isn't actually necessary. We should probably add Value->stripConstantOffsets() as well.
-
Nikita Popov authored
Remove the special casing for intrinsics in MemoryLocation::getForDest() and handle them through the general attribute based code. On the DSE side, this means that isRemovable() now needs to handle more than a hardcoded list of intrinsics. We consider everything apart from volatile memory intrinsics and lifetime markers to be removable. This allows us to perform DSE on intrinsics that DSE has not been specially taught about, using a matrix store as an example here. There is an interesting test change for invariant.start, but I believe that optimization is correct. It only looks a bit odd because the code is immediate UB anyway. Differential Revision: https://reviews.llvm.org/D116210
-
Nikita Popov authored
78d15a11 has been reverted, but the test not deleted, so it is failing now.
-
Nikita Popov authored
Instead of using the ArgumentPromotion implementation, we now walk call sites using checkForAllCallSites() and directly call areTypesABICompatible() using the replacement types. I believe that resolves the TODO in the code. Differential Revision: https://reviews.llvm.org/D116033
-
Nikita Popov authored
The reduction initialization code creates a "naturally aligned null pointer to void lvalue", which I found somewhat odd, even though it works out in the end because it is not actually used. It doesn't look like this code actually needs an LValue for anything though, and we can use an invalid Address to represent this case instead. Differential Revision: https://reviews.llvm.org/D116214
-
Fangrui Song authored
Avoid repeated load of global pointer (symtab) / members (sections.size(), firstGlobal) in the hot paths. And remove some unneeded this->
-
Chuanqi Xu authored
According to [dcl.fct.def.coroutine]p6, the promise_type is allowed to not define return_void nor return_value: > If searches for the names return_void and return_value in the scope > of the promise type each find any declarations, the program is > ill-formed. > [Note 1: If return_void is found, flowing off the end of a coroutine is > equivalent to a co_return with no operand. Otherwise, flowing off the > end of a coroutine results in > undefined behavior ([stmt.return.coroutine]). — end note] So the program isn't ill-formed if the promise_type doesn't define return_void nor return_value. It is just a potential UB. So the program should be allowed to compile. Reviewed By: urnathan Differential Revision: https://reviews.llvm.org/D116204
-