- Jul 06, 2023
-
-
Nikita Popov authored
Currently, our GEP specification has a special case that makes gep inbounds (null, 0) legal. This patch proposes to expand this special case to all gep inbounds (ptr, 0), where ptr is no longer required to point to an allocated object. This was previously discussed in some detail at https://discourse.llvm.org/t/question-about-getelementptr-inbounds-with-offset-0/62533. The motivation for this change is twofold: * Rust relies on getelementptr inbounds with zero offset to be legal for arbitrary pointers to support zero-sized types. The current rules are unclear on whether this is legal or not (saying that there is a zero-size "allocated object" at every address may be consistent with our current rules, but more clarity is desired here). * The current semantics require us to drop the inbounds flag when materializing zero-index GEPs, which is done by some InstCombine transforms. Preserving the inbounds flag can substantially improve optimization quality in some cases, as illustrated in D154055. As far as I know, the only analysis/transforms affected by this semantics change are: * A special-case for comparisons with null in CaptureTracking, which is fixed by D154054. As far as I can tell, that special case is not particularly valuable and should be recovered by other transforms. * Folding gep inbounds undef, idx to poison. We now need to fold to undef instead (D154215). Differential Revision: https://reviews.llvm.org/D154051
-
Louis Dionne authored
Based on the comment in https://reviews.llvm.org/D54290#4418958, these attributes need to be on the top-level functions in order to work properly. Also, add tests. Fixes http://llvm.org/PR57035. Differential Revision: https://reviews.llvm.org/D154354
-
Alex Bradbury authored
Thanks to D154555, these intrinsics no longer crash when used with a soft float ABI.
-
Florian Hahn authored
When a scalar epilogue is required, at least one iteration of the scalar loop has to execute. Adjust ConstTripCount accordingly to avoid picking a max VF that results in a dead vector loop. Reviewed By: Ayal Differential Revision: https://reviews.llvm.org/D154261
-
Felipe de Azevedo Piovezan authored
Two identical loops were iterating over different ranges, leading to code duplication. We replace this by a loop over the concatenation of the ranges. We also use early returns to avoid deeply nested code and explicitly check for a condition mentioned in comments. Differential Revision: https://reviews.llvm.org/D154505
-
John Brawn authored
This test is failing to compile when LLVM_ENABLE_MODULES=ON due to NamedDecl being multiply defined. Fix this by avoiding declaring our own NamedDecl in the test and instead cast a struct of appropriate size and alignment to NamedDecl.
-
Eddie Phillips authored
Matches behaviour in update_llc_test_checks.py etc. Fixes #63112 Differential Revision: https://reviews.llvm.org/D152333
-
Ivan Kosarev authored
Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D154527
-
Marco Elver authored
Current FreeBSD has increased size of cpuset. Match it to not break the build on newer FreeBSD. Patch by John F. Carr Fixes: https://github.com/llvm/llvm-project/issues/63485
-
Amilendra Kodithuwakku authored
This commit provides linker support for Cortex-M Security Extensions (CMSE). The specification for this feature can be found in ARM v8-M Security Extensions: Requirements on Development Tools. The linker synthesizes a security gateway veneer in a special section; `.gnu.sgstubs`, when it finds non-local symbols `__acle_se_<entry>` and `<entry>`, defined relative to the same text section and having the same address. The address of `<entry>` is retargeted to the starting address of the linker-synthesized security gateway veneer in section `.gnu.sgstubs`. In summary, the linker translates input: ``` .text entry: __acle_se_entry: [entry_code] ``` into: ``` .section .gnu.sgstubs entry: SG B.W __acle_se_entry .text __acle_se_entry: [entry_code] ``` If addresses of `__acle_se_<entry>` and `<entry>` are not equal, the linker considers that `<entry>` already defines a secure gateway veneer so does not synthesize one. If `--out-implib=<out.lib>` is specified, the linker writes the list of secure gateway veneers into a CMSE import library `<out.lib>`. The CMSE import library will have 3 sections: `.symtab`, `.strtab`, `.shstrtab`. For every secure gateway veneer <entry> at address `<addr>`, `.symtab` contains a `SHN_ABS` symbol `<entry>` with value `<addr>`. If `--in-implib=<in.lib>` is specified, the linker reads the existing CMSE import library `<in.lib>` and preserves the entry function addresses in the resulting executable and new import library. Reviewed By: MaskRay, peter.smith Differential Revision: https://reviews.llvm.org/D139092 -
Haojian Wu authored
We're in favor of the llvm::writeToOutput API, and all writeFileAtomically usages have been migrated to writeToOutput. Differential Revision: https://reviews.llvm.org/D153740
-
Matthias Springer authored
This should have been part of D154585.
-
Lorenzo Chelini authored
Make the transformation accessible to other drivers (i.e., passes).
-
Simon Pilgrim authored
Fold allsignbits pack patterns to make better use of cheap (and commutable) logic ops Reapplied after a32d14fd / 156913cb with bitcast fix
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
Matthias Springer authored
The TileOp builders did not set `scalable_sizes`, which produces invalid ops. `scalable_sizes` must contain as any booleans as there are sizes. Differential Revision: https://reviews.llvm.org/D154585
-
Haojian Wu authored
Reviewed By: kadircet Differential Revision: https://reviews.llvm.org/D153340
-
Gedare Bloom authored
Fixes a bug that prevents alignment from proceeding through a function pointer in a list of declarations. Fixes #63451. Differential Revision: https://reviews.llvm.org/D153585
-
Dmitri Gribenko authored
-
Gedare Bloom authored
Fixes a bug with the handling of right aligned references with left/middle alignment pointers. Fixes #63452. Differential Revision: https://reviews.llvm.org/D153579
-
David Spickett authored
Previously the following would crash: (lldb) run Process 2594053 launched: '/tmp/test.o' (aarch64) Process 2594053 exited with status = 0 (0x00000000) (lldb) register read <tab> As the completer assumed that the execution context would always have a register context. After a program has finished, it does not. Split out the generic parts of the test from the x86 specific tests, and added "register info" to both. Reviewed By: JDevlieghere Differential Revision: https://reviews.llvm.org/D154413
-
Craig Topper authored
This matches the data type of the intrinsics. This case be seen from the removal of sext and trunc instructions from the IR. Reviewed By: kito-cheng Differential Revision: https://reviews.llvm.org/D154572
-
Craig Topper authored
These builtins were recently changed to return 'int' like the similar __builtin_clz/__builtin_ctz builtins, but the IR generation was not updated to use a truncate.
-
Job Noorman authored
`JITLinkContext` is notified (using `notifyResolved`) of the final symbol addresses after allocating memory and running the post-allocation passes. However, linker relaxation, which can cause symbol addresses to change, was run during the pre-fixup passes. This causes users of JITLink (e.g., ORC) to pick-up wrong symbol addresses when linker relaxation was enabled. This patch fixes this by running relaxation during the post-allocation passes. Fixes #63671 Reviewed By: lhames Differential Revision: https://reviews.llvm.org/D154501
-
Andrzej Warzynski authored
This patch adds the missing logic to vectorise `tensor.extract` for 0-d tensors. Fixes #63688 Differential Revision: https://reviews.llvm.org/D154518
-
Craig Topper authored
This matches the definition for the underlying builtins and what is done in the Zbb test.
-
esmeyi authored
Summary: Clang uses LLVM's integrated assembler by default on most targets, however non-integrated-as mode is default on AIX. Currently integrated-as mode on AIX has passed tests of LLVM test-suite, bootstrap and Spec2017, therefore this patch sets integrated-as as the default assembler mode on AIX. Reviewed By: DiggerLin Differential Revision: https://reviews.llvm.org/D150758
-
Nikita Popov authored
As reported at https://reviews.llvm.org/D153305#4475840.
-
Serge Pavlov authored
The issue: https://github.com/llvm/llvm-project/issues/63704
-
XChy authored
[InstCombine] Transform (A > 0) | (A < 0) -> zext (A != 0) fold This extends **foldCastedBitwiseLogic** to handle the similar cases. Actually, for `(A > B) | (A < B)`, when B != 0, it can be optimized to `zext( A != B )` by **foldAndOrOfICmpsUsingRanges**. However, when B = 0, **transformZExtICmp** will transform `zext(A < 0) to i32` into `A << 31`, which cannot be optimized by **foldAndOrOfICmpsUsingRanges**. Because I'm new to LLVM and has no concise knowledge about how LLVM decides the order of optimization, I choose to extend **foldCastedBitwiseLogic** to fold `( A << (X - 1) ) | ((A > 0) zext to iX) -> (A != 0) zext to iX`. And the equivalent fold follows: ``` A << (X - 1) ) | ((A > 0) zext to iX -> A < 0 | A > 0 -> (A != 0) zext to iX ``` It's proved by [[https://alive2.llvm.org/ce/z/33HzjE|alive-tv]] Related issue: [[https://github.com/llvm/llvm-project/issues/62586 | (a > b) | (a < b) is not simplified only for the case b=0 ]] Reviewed By: goldstein.w.n Differential Revision: https://reviews.llvm.org/D154126
-
XChy authored
Tests for an upcoming (A > 0) | (A < 0) -> zext (A != 0) fold. Related issue: [[ https://github.com/llvm/llvm-project/issues/62586 | (a > b) | (a < b) is not simplified only for the case b=0 ]] Differential Revision: https://reviews.llvm.org/D154089
-
Martin Braenne authored
See comments in the code for details. Reviewed By: xazax.hun Differential Revision: https://reviews.llvm.org/D154479
-
Valery Pykhtin authored
1. Improved code that deduces register class from instruction definitions. Previously if some instruction didn't contain a reg class for an operand it was considered as no information on register class even if other instructions specified the class. 2. Added check on required size of resulting register because in some cases classes with smaller registers had been selected (for example VReg_1). Reviewed By: arsenm, #amdgpu Differential Revision: https://reviews.llvm.org/D152832
-
Jim Lin authored
[LibCallsShrinkWrap] Set IsFPConstrained is true for creating quiet floating comparision if function has strictfp attribute Create a quiet floating-point comparision if function has strictfp attribute. Avoid unexpected FP exception raised during libcall domain error checking. It raises an FP exception only in case where an input is a signaling NaN. Reviewed By: efriedma Differential Revision: https://reviews.llvm.org/D152776
-
Fangrui Song authored
fdr-thread-order.cpp can be very slow when the thread contention is large. Enable it for AArch64 and x86-64 for now. fdr-mode.cpp fails on a ppc64le machine. Unsupport it on ppc64le for now. The remaining modified tests pass on AArch64, ppc64le, and x86-64.
-
Fangrui Song authored
-
Fangrui Song authored
-
Jianjian GUAN authored
Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D154487
-
Tom Stellard authored
A project that bundles the llvm source code may have their own PACKAGE_VERSION variable, so only use this to compute the CLANG_RESOURCE_DIR if CLANG_VERSION_MAJOR is undefined. Reviewed By: sebastian-ne Differential Revision: https://reviews.llvm.org/D152608
-