- Jan 25, 2023
-
-
Tom Stellard authored
The stage1 and stage2 builds aren't packaged, so we only need to build enough of the toolchain to build the next phase. Reviewed By: thieta, amyk Differential Revision: https://reviews.llvm.org/D141552
-
Muhammad Omair Javaid authored
This patch remove xfail decorator from builtins/Unit/trampoline_setup_test.c as it is passing on Windows/AArch64 nowz. It is being skipped in code with __clang__ not defined. https://lab.llvm.org/buildbot/#/builders/120/builds/3873
-
Craig Topper authored
If we're extracting an element and inserting into a undef vector with the same number of elements, we can use the original vector. This pattern occurs around reductions that have been cascaded together. This can be generalized to wider/narrow vectors by using insert_subvector/extract_subvector, but we don't have lit tests for that case currently. We can also support non-undef before by using a slide or vmv.v.v Reviewed By: reames Differential Revision: https://reviews.llvm.org/D142264
-
yronglin authored
Remove __unexpected namespace. Reviewed By: philnik, #libc, ldionne Differential Revision: https://reviews.llvm.org/D141947
-
Jez Ng authored
This is what ld64 does, and also what we already do for most of the other load commands. I'm not aware of a good way to test this, but I don't think it really matters. Differential Revision: https://reviews.llvm.org/D141462
-
Benjamin Kramer authored
-
Kirill Stoimenov authored
Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D142504
-
usama hameed authored
-
Younan Zhang authored
2db6b34e introduces circular dependency on llvm::ArrayRef. By inspecting commit history, it appears that we have some issue using deduction guide on std::array. Why don't we try std::array with explicit template arguments? Differential revision: https://reviews.llvm.org/D141352
-
Ben Langmuir authored
The check for "no SOURCE_DATE_EPOCH" wasn't especially interesting, and I am not aware of a _portable_ way to unset and environment variable in a lit test. So remove it since it can fail if the build environment has SOURCE_DATE_EPOCH set globally. Differential Revision: https://reviews.llvm.org/D142511
-
Alexander Yermolovich authored
In some binaries produced with ThinLTO there are CUs that share entry in .debug_addr. Before we would generate a new entry for each. Which lead to binary size increase. This changes the behavior so that we re-use entries in .debug_addr. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D142425
-
Luke Hutton authored
Adds the RFFT2d TOSA operation and supporting shape inference function. Signed-off-by:
Luke Hutton <luke.hutton@arm.com> Change-Id: I7e49c47cdd846cdc1b187545ef76d5cda2d5d9ad Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D142336
-
usama hameed authored
[ASan] Introduce a flag -asan-constructor-kind to control the generation of the Asan module constructor. By default, ASan generates an asan.module_ctor function that initializes asan and registers the globals in the module. This function is added to the @llvm.global_ctors array. Previously, there was no way to control the generation of this function. This patch adds a way to control the generation of this function. The flag -asan-constructor-kind has two options: global: This is the default option and the default behavior of ASan. It generates an asan.module_ctor function. none: This skips the generation of the asan.module_ctor function. rdar://104448572 Differential revision: https://reviews.llvm.org/D142505
-
usama hameed authored
rdar://103570533 Differential Revision: https://reviews.llvm.org/D142243
-
Kevin Sala authored
This patch implements the memory lock/unlock API, introduced in patch https://reviews.llvm.org/D139208, in the NextGen plugins. Locked buffers feature reference counting and we allow certain overlapping. Given an already locked buffer A, other buffers that are fully contained inside A can be locked again, even if they are smaller than A. In this case, the reference count of locked buffer A will be incremented. However, extending an existing locked buffer is not allowed. The original buffer is actually unlocked once all its users have released the locked buffer and sub-buffers (i.e., the reference counter becomes zero). Differential Revision: https://reviews.llvm.org/D141227
-
Nick Desaulniers authored
Very similar to https://reviews.llvm.org/D111272. We very often can evaluate calls to llvm.objectsize.* regardless of inlining. Don't count calls to llvm.objectsize.* against the InlineCost when we can evaluate the call to a constant. Link: https://github.com/ClangBuiltLinux/linux/issues/1302 Reviewed By: manojgupta Differential Revision: https://reviews.llvm.org/D111456
-
Joseph Huber authored
Summary: Forgot to add this.
-
Joseph Huber authored
The AMDGPU target can only emit LLVM-IR, so we can always rely on LTO to link the static version of the runtime optimally. Using the static library only has a few advantages. Namely, it avoids several known bugs and allows us to optimize out more functions. This is legal since the changes in D142486 and D142484 Depends on D142486 D142484 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142491
-
Joseph Huber authored
Currently we have two versions of the static library. One is built as individual bitcode files and linked via `-mlink-builtin-bitcode`. The other is built as a single static archive `omptarget.devicertl.a` and is linked via `-lomptarget.devicertl` and handled by the linker wrapper during LTO. We use the former in the case that we are not performing LTO, because linking the library late wouldn't allow us to optimize the runtime library effectively. The support in D142484 allows us to unconditionally link this library, so it will only be pulled in if needed. That is, if we linked already via `-mlink-builtin-bitcode` then we will not pull in the static library even if it's linked on the command line. Depends on D142484 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142486
-
Joseph Huber authored
Currently, we pull in every single static archive member as long as we have an offloading architecture that requires it. This goes against the standard sematnics of static libraries that only pull in symbols that define currently undefined symbols. In order to support this we roll some custom symbol resolution logic to check if a static library is needed. Because of offloading semantics, this requires an extra check for externally visibile symbols. E.g. if a static member defines a kernel we should import it. The main benefit to this is that we can now link against the `libomptarget.devicertl.a` library unconditionally. This removes the requirement for users to specify LTO on the link command. This will also allow us to stop using the `amdgcn` bitcode versions of the libraries. ``` clang foo.c -fopenmp --offload-arch=gfx1030 -foffload-lto -c clang foo.o -fopenmp --offload-arch=gfx1030 -foffload-lto ``` Reviewed By: tra Differential Revis...
-
Giorgis Georgakoudis authored
Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142492
-
Benjamin Kramer authored
-
Ilya Tokar authored
AVX/AVX512 instructions may cause frequency drop on e.g. Skylake. The magnitude of frequency/performance drop depends on instruction (multiplication vs load/store) and vector width. Currently users, that want to avoid this drop can specify -mprefer-vector-width=128. However this also prevents generations of 256-bit wide instructions, that have no associated frequency drop (mainly load/stores). Add a tuning flag that allows generations of 256-bit AVX load/stores, even when -mprefer-vector-width=128 is set, to speed-up memcpy&co. Verified that running memcpy loop on all cores has no frequency impact and zero CORE_POWER:LVL[12]_TURBO_LICENSE perf counters. Makes coping memory faster e.g.: BM_memcpy_aligned/256 80.7GB/s ± 3% 96.3GB/s ± 9% +19.33% (p=0.000 n=9+9) Differential Revision: https://reviews.llvm.org/D134982
-
Shilei Tian authored
GCC doesn't support `-fopenmp-version`, causing test failure if the compiler used for testing is GCC. GCC's OpenMP 5.2 support is very limited yet. Disable those tests requiring 5.2 feature for GCC as well. We might want to take a look at all `libomp` tests and mark those tests that don't support GCC yet. Reviewed By: ABataev Differential Revision: https://reviews.llvm.org/D142173
-
Nick Desaulniers authored
As pointed out by @arsenm in https://reviews.llvm.org/D141451#4045099, we don't handle ConstantExpressions for dontcall-{warn|error} IR Fn Attrs. Use CallBase::getCalledOperand() and Value::stripPointerCasts() should the call to CallBase::getCalledFunction return nullptr. I don't know how to express the IR test case in C, otherwise I'd add a clang test, too. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D142058
-
Matt Arsenault authored
These are essentially add/sub 1 with a clamping value. AMDGPU has instructions for these. CUDA/HIP expose these as atomicInc/atomicDec. Currently we use target intrinsics for these, but those do no carry the ordering and syncscope. Add these to atomicrmw so we can carry these and benefit from the regular legalization processes.
-
Sanjay Patel authored
https://alive2.llvm.org/ce/z/2iC4oB This is similar to changes made for zext + lshr: 21d3871b 6c39a3aa The existing fold did not account for extra uses, so we see some instruction count reductions in the test diffs. This is intended to improve analysis (icmp likely has more transforms than any other opcode), make other transforms more symmetric with zext/lshr, and it can be inverted in codegen if profitable. As with the earlier changes, there is potential to uncover infinite combine loops, but I have not found any yet.
-
Slava Zakharin authored
It looks like a flaky issue that sometimes breaks the buildbot: https://lab.llvm.org/buildbot/#/builders/181/builds/13475 Reviewed By: clementval Differential Revision: https://reviews.llvm.org/D142081
-
Jay Foad authored
This will allow an entry in the table to access data that is stored immediately after the end of the table, by adding its opcode value to its address. Differential Revision: https://reviews.llvm.org/D142217
-
Philipp Tomsich authored
The Ampere1A core improves on the Ampere1 with key differences being: * memory tagging is supported * SM3/SM4 are supported * adds a new fusion pair for (A+B+1 and A-B-1) (added in a later commit) Depends on D142395 Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D142396
-
Jay Foad authored
Combine the implicit uses and defs lists into a single list of uses followed by defs. Instead of 0-terminating the list, store the number of uses and defs. This avoids having to scan the whole list to find the length and removes one pointer from MCInstrDesc (although it does not get any smaller due to alignment issues). Remove the old accessor methods getImplicitUses, getNumImplicitUses, getImplicitDefs and getNumImplicitDefs as all clients are using the new implicit_uses and implicit_defs. Differential Revision: https://reviews.llvm.org/D142216
-
Johannes Doerfert authored
-
Philipp Tomsich authored
The original enablement for the Ampere1 core inadvertently had omitted that FEAT_RAND is support and errorously claimed that FEAT_MTE was available. Adjust the definition of Ampere1 to match reality. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D142395
-
Kevin Sala authored
-
Philipp Tomsich authored
This change fixes the following warnings: llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:727:55: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 727 | TYPED_TEST_SUITE(RangeSetCastToNoopTest, NoopCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:728:65: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 728 | TYPED_TEST_SUITE(RangeSetCastToPromotionTest, PromotionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:729:67: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 729 | TYPED_TEST_SUITE(RangeSetCastToTruncationTest, TruncationCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:730:67: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 730 | TYPED_TEST_SUITE(RangeSetCastToConversionTest, ConversionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:732:46: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 732 | PromotionConversionCastTypes); | ^ llvm/clang/unittests/StaticAnalyzer/RangeSetTest.cpp:734:47: warning: ISO C++11 requires at least one argument for the "..." in a variadic macro 734 | TruncationConversionCastTypes); | ^ Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D142439 -
Joseph Huber authored
Currently, we embed device code into the host to perform multi-architecture linking and handling of device code. If the user specified `-S -emit-llvm` then the embedded output will be textual LLVM-IR. This is a problem because it can't be used by the LTO backend and it makes reading the file confusing. This patch changes the behaviour to only emit textual device IR if we are in device only mode, that is, if the device code is presented directly to the user instead of being embedded. Otherwise we should always embed device bitcode instead. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D141717
-
Manas authored
A couple of packages were out-dated while building satest docker image. This patch updates those. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D142454
-
Manas authored
This patch fixes certain cases where solver was not able to infer disequality due to overlapping of values in rangeset. This case was casting from lower signed type to bigger unsigned type. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D140086
-
Douglas Yung authored
This reverts commit 2b807336. This change is causing Windows builds to hang and out of memory errors with clang-15: - https://lab.llvm.org/buildbot/#/builders/17/builds/33129 - https://lab.llvm.org/buildbot/#/builders/174/builds/17069 - https://lab.llvm.org/buildbot/#/builders/83/builds/28484 - https://lab.llvm.org/buildbot/#/builders/172/builds/22803 - https://lab.llvm.org/buildbot/#/builders/216/builds/16210
-
Florian Hahn authored
This patch updates SCCP to use the value ranges of AddInst operands to try to prove the AddInst does not overflow in the unsigned sense and adds the NUW flag. The reasoning is done with makeGuaranteedNoWrapRegion (thanks @nikic for point it out!). Follow-ups will include adding NSW and extension to more OverflowingBinaryOperators. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D142387
-