- Jan 25, 2023
-
-
Benjamin Kramer authored
-
Ilya Tokar authored
AVX/AVX512 instructions may cause frequency drop on e.g. Skylake. The magnitude of frequency/performance drop depends on instruction (multiplication vs load/store) and vector width. Currently users, that want to avoid this drop can specify -mprefer-vector-width=128. However this also prevents generations of 256-bit wide instructions, that have no associated frequency drop (mainly load/stores). Add a tuning flag that allows generations of 256-bit AVX load/stores, even when -mprefer-vector-width=128 is set, to speed-up memcpy&co. Verified that running memcpy loop on all cores has no frequency impact and zero CORE_POWER:LVL[12]_TURBO_LICENSE perf counters. Makes coping memory faster e.g.: BM_memcpy_aligned/256 80.7GB/s ± 3% 96.3GB/s ± 9% +19.33% (p=0.000 n=9+9) Differential Revision: https://reviews.llvm.org/D134982
-
Shilei Tian authored
GCC doesn't support `-fopenmp-version`, causing test failure if the compiler used for testing is GCC. GCC's OpenMP 5.2 support is very limited yet. Disable those tests requiring 5.2 feature for GCC as well. We might want to take a look at all `libomp` tests and mark those tests that don't support GCC yet. Reviewed By: ABataev Differential Revision: https://reviews.llvm.org/D142173
-
Nick Desaulniers authored
As pointed out by @arsenm in https://reviews.llvm.org/D141451#4045099, we don't handle ConstantExpressions for dontcall-{warn|error} IR Fn Attrs. Use CallBase::getCalledOperand() and Value::stripPointerCasts() should the call to CallBase::getCalledFunction return nullptr. I don't know how to express the IR test case in C, otherwise I'd add a clang test, too. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D142058
-
Matt Arsenault authored
These are essentially add/sub 1 with a clamping value. AMDGPU has instructions for these. CUDA/HIP expose these as atomicInc/atomicDec. Currently we use target intrinsics for these, but those do no carry the ordering and syncscope. Add these to atomicrmw so we can carry these and benefit from the regular legalization processes.
-
Sanjay Patel authored
https://alive2.llvm.org/ce/z/2iC4oB This is similar to changes made for zext + lshr: 21d3871b 6c39a3aa The existing fold did not account for extra uses, so we see some instruction count reductions in the test diffs. This is intended to improve analysis (icmp likely has more transforms than any other opcode), make other transforms more symmetric with zext/lshr, and it can be inverted in codegen if profitable. As with the earlier changes, there is potential to uncover infinite combine loops, but I have not found any yet.
-
Slava Zakharin authored
It looks like a flaky issue that sometimes breaks the buildbot: https://lab.llvm.org/buildbot/#/builders/181/builds/13475 Reviewed By: clementval Differential Revision: https://reviews.llvm.org/D142081
-
Jay Foad authored
This will allow an entry in the table to access data that is stored immediately after the end of the table, by adding its opcode value to its address. Differential Revision: https://reviews.llvm.org/D142217
-
Philipp Tomsich authored
The Ampere1A core improves on the Ampere1 with key differences being: * memory tagging is supported * SM3/SM4 are supported * adds a new fusion pair for (A+B+1 and A-B-1) (added in a later commit) Depends on D142395 Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D142396
-
Jay Foad authored
Combine the implicit uses and defs lists into a single list of uses followed by defs. Instead of 0-terminating the list, store the number of uses and defs. This avoids having to scan the whole list to find the length and removes one pointer from MCInstrDesc (although it does not get any smaller due to alignment issues). Remove the old accessor methods getImplicitUses, getNumImplicitUses, getImplicitDefs and getNumImplicitDefs as all clients are using the new implicit_uses and implicit_defs. Differential Revision: https://reviews.llvm.org/D142216
-
Johannes Doerfert authored
-
Philipp Tomsich authored
The original enablement for the Ampere1 core inadvertently had omitted that FEAT_RAND is support and errorously claimed that FEAT_MTE was available. Adjust the definition of Ampere1 to match reality. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D142395
-
Kevin Sala authored
-
Philipp Tomsich authored
This change fixes the following warnings: llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:727:55: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 727 | TYPED_TEST_SUITE(RangeSetCastToNoopTest, NoopCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:728:65: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 728 | TYPED_TEST_SUITE(RangeSetCastToPromotionTest, PromotionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:729:67: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 729 | TYPED_TEST_SUITE(RangeSetCastToTruncationTest, TruncationCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:730:67: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 730 | TYPED_TEST_SUITE(RangeSetCastToConversionTest, ConversionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:732:46: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 732 | PromotionConversionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:734:47: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 734 | TruncationConversionCastTypes); | ^ Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D142439 -
Joseph Huber authored
Currently, we embed device code into the host to perform multi-architecture linking and handling of device code. If the user specified `-S -emit-llvm` then the embedded output will be textual LLVM-IR. This is a problem because it can't be used by the LTO backend and it makes reading the file confusing. This patch changes the behaviour to only emit textual device IR if we are in device only mode, that is, if the device code is presented directly to the user instead of being embedded. Otherwise we should always embed device bitcode instead. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D141717
-
Manas authored
A couple of packages were out-dated while building satest docker image. This patch updates those. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D142454
-
Manas authored
This patch fixes certain cases where solver was not able to infer disequality due to overlapping of values in rangeset. This case was casting from lower signed type to bigger unsigned type. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D140086
-
Douglas Yung authored
This reverts commit 2b807336. This change is causing Windows builds to hang and out of memory errors with clang-15: - https://lab.llvm.org/buildbot/#/builders/17/builds/33129 - https://lab.llvm.org/buildbot/#/builders/174/builds/17069 - https://lab.llvm.org/buildbot/#/builders/83/builds/28484 - https://lab.llvm.org/buildbot/#/builders/172/builds/22803 - https://lab.llvm.org/buildbot/#/builders/216/builds/16210
-
Florian Hahn authored
This patch updates SCCP to use the value ranges of AddInst operands to try to prove the AddInst does not overflow in the unsigned sense and adds the NUW flag. The reasoning is done with makeGuaranteedNoWrapRegion (thanks @nikic for point it out!). Follow-ups will include adding NSW and extension to more OverflowingBinaryOperators. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D142387
-
Sanjay Patel authored
not (bitcast (sext i1 X)) --> bitcast (sext (not i1 X)) https://alive2.llvm.org/ce/z/-6Ygkd This shows up as a potential regression if we change canonicalization of ashr+not to icmp+sext.
-
Sanjay Patel authored
-
Aaron Ballman authored
The test currently expects to be run in a directory named 'clang' but that's not valid for our release tarballs. We don't actually care what base directory the test is run from, so this removes the path component entirely.
-
Aart Bik authored
Note that I did not track why this started failing exactly, which is why I CC Matthias on this fix. But at least we run asan clean again for the whole suite after this change. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D142496
-
Paul Robinson authored
Found by the Rotten Green Tests project.
-
Philip Reames authored
This matches the behavior from a number of other targets, including e.g. X86. This does have the effect of increasing register pressure slightly, but we have a relative abundance of registers in the ISA compared to other targets which use the same heuristic. The motivation here is that our current cost heuristic treats number of registers as the dominant cost. As a result, an extra use outside of a loop can radically change the LSR result. As an example consider test4 from the recently added test/Transforms/LoopStrengthReduce/RISCV/lsr-cost-compare.ll. Without a use outside the loop (see test3), we convert the IV into a pointer increment. With one, we leave the gep in place. The pointer increment version both decreases number of instructions in some loops, and creates parallel chains of computation (i.e. decreases critical path depth). Both are generally profitable. Arguably, we should really be using a more sophisticated model here - such as e.g. using profile information or explicitly modeling parallelism gains. However, as a practical matter starting with the same mild hack that other targets have used seems reasonable. Differential Revision: https://reviews.llvm.org/D142227
-
Siva Chandra Reddy authored
This is the first of patches doing similar cleanup. A section in the code style doc has been added explaining where and how LIBC_INLINE is to be used. Reviewed By: jeffbailey, lntue Differential Revision: https://reviews.llvm.org/D142434
-
Guilherme Valarini authored
-
Sanjay Patel authored
Value name propagation improved.
-
Sanjay Patel authored
-
Sanjay Patel authored
There's no reason to use "CI" (cast instruction) when we know that the value is a more specific (exact) type of instruction (although we might want to common-ize some of this code to eliminate duplication or logic diffs). It's also visually difficult to distinguish between "CI", "ICI", and "IC" acronyms (and those could change meaning depending on context). This was partially changed in earlier commits, so this makes this pair of functions consistent.
-
Stanislav Mekhanoshin authored
Differential Revision: https://reviews.llvm.org/D142407
-
Shilei Tian authored
-
Stanislav Mekhanoshin authored
These are unsupported. Differential Revision: https://reviews.llvm.org/D142493
-
Han Zhu authored
-
Valentin Clement authored
When referencing a single component from a polymorphic array in an expression, the rebox operation should output a boxed array of that component type and not a polymorphic boxed array as it was done. Reviewed By: PeteSteinfeld Differential Revision: https://reviews.llvm.org/D142462
-
Alexandros Lamprineas authored
Found here https://github.com/llvm/llvm-project/issues/60191 The compiler would crash when specializing a function based on a function pointer whose call sites may expect less parameters than those of the function we are replacing the pointer with. Differential Revision: https://reviews.llvm.org/D142444
-
Joseph Huber authored
-
Vassil Vassilev authored
Patch by Lang Hames and Sunho Kim! Differential revision: https://reviews.llvm.org/D138264
-
Pratik Sharma authored
There were some dead links in Suppressing Undesired Diagnostics which I replaced with the working links. Fixes #60023 Differential Revision: https://reviews.llvm.org/D142377
-
Slava Zakharin authored
OpenMP buildbots are failing: https://lab.llvm.org/buildbot/#/builders/193/builds/25434 https://lab.llvm.org/buildbot/#/builders/193/builds/25420 This reverts commit 7fbf1221.
-