- Jun 01, 2023
-
-
Haojian Wu authored
-
Nikita Popov authored
-
Igor Kirillov authored
This patch allows us to gain all the benefits provided by LoopLoadElimination pass to descending loops. Differential Revision: https://reviews.llvm.org/D151448
-
Joseph Huber authored
The linker wrapper performs its own very basic symbol resolution for the purpose of supporting standard static library semantics. We do this here because the Nvidia `nvlink` wrapper does not support static linking and we have some offloading specific extensions. Currently, we always place symbols in the "table" even if they aren't extracted. This caused the logic to fail when many files were used that referenced the same undefined variable. This patch changes the pass to only add the symbols to the global "table" if the file is actually extracted. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D151839
-
Simon Pilgrim authored
Only uses port2+3 for agen, and was missing port4 for the actual store Noticed while investigating the skylake vs icelake diffs for Issue #62602
-
Nimish Mishra authored
Verification of support for lowering private/firstprivate clauses on unstructured sections. Differential Revision: https://reviews.llvm.org/D145352 Reviewed By: TIFitis
-
Ritanya B Bharadwaj authored
Initial support for OpenMP 5.0 declare target "as if" behavior for "initializer expressions". OpenMP 5.0, 2.12.7 declare target. Reviewed By: Alexey Differential Revision: https://reviews.llvm.org/D146418
-
David Green authored
i1 inserts will need an extra cset, and i1 extracts need a cmp (or tst) in order to be used. This increase the cost of them a little to account for those extra instructions. https://godbolt.org/z/3c5z4G7Mh Differential Revision: https://reviews.llvm.org/D151189
-
Nikita Popov authored
We need to add the replaced instruction itself to the worklist as well. We want to remove the old instructions, but can't easily do so directly, as the icmp is also one of the users and we need to retain it until the fold has finished.
-
Antonio Abbatangelo authored
Adds a dynamic stack alignment to functions under the interrupt call convention on x86-32. This fixes the issue where the stack can be misaligned on entry, since x86-32 makes no guarantees about the stack pointer position when the interrupt service routine is called. The alignment is done by overriding X86RegisterInfo::shouldRealignStack, and by setting the correct alignment in X86FrameLowering::calculateMaxStackAlign. This forces the interrupt handler to be dynamically aligned, generating the appropriate `and` instruction in the prologue and `lea` in the epilogue. The `no-realign-stack` attribute can be used as an opt-out. Fixes #26851 Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D151400
-
Timm Bäder authored
Our comparison opcodes always produce a Boolean value and push it on the stack. However, the result of such a comparison in C is int, so the later code expects an integer value on the stack. Work around this problem by casting the boolean value to int in those cases. This is not ideal for C however. The comparison is usually wrapped in a IntegerToBool cast anyway. Differential Revision: https://reviews.llvm.org/D149645
-
Nikita Popov authored
Use replaceInstUsesWith() rather than plain RAUW to make sure the old instructions are added back to the worklist for DCE.
-
zhuna authored
Now, if the offset overflow happens, we just silently ignore it. We will generate a bad dwp file, which will crash the gdb or make it undefined behavior, and hard to address the root cause. So, we need to produce some messages if overflow happens. Reviewed By: ayermolo, dblaikie, steven.zhang Differential Revision: https://reviews.llvm.org/D144565
-
David Green authored
This expands the reduction cost of i1 and/or/xor, so that larger type sizes get handled by the existing code. For i1 reductions - and will use maxv, or will use minv and xor will use addv, plus the cost of legalizing the type for larger vectors using and/or/xor. The i1 vectors will be legalized to higher width integers (say v16i8), which this overrides the cost of. As with all i1 vectors there is a chance that the types the i1 vector is created with and how it is used will not match, introducing extra extends that are not necessarily costmodelled. https://godbolt.org/z/6Gc9K6b7T Differential Revision: https://reviews.llvm.org/D151184
-
Andrzej Warzynski authored
This patch enables specifying scalable tile sizes when using the Transform dialect to drive tiling, e.g.: ``` %1, %loop = transform.structured.tile %0 [[4]] ``` This is implemented by extending the TileOp with a dedicated attribute for "scalability" and by updating various parsing hooks. At the moment, only the trailing tile size can be scalable. The following is not yet supported: ``` %1, %loop = transform.structured.tile %0 [[4], [4]] ``` This change is a part of larger effort to enable scalable vectorisation in Linalg. See this RFC for more context: * https://discourse.llvm.org/t/rfc-scalable-vectorisation-in-linalg/ Differential Revision: https://reviews.llvm.org/D150944
-
Nikita Popov authored
Make sure the old operand is added back to the worklist for DCE.
-
Petr Hosek authored
This reverts commit f99a7d3e since it broke the bolt-aarch64-ubuntu-clang-shared bot.
-
Balázs Kéri authored
[clang][analyzer] Merge apiModeling.StdCLibraryFunctions and StdCLibraryFunctionArgs checkers into one. Main reason for this change is that these checkers were implemented in the same class but had different dependency ordering. (NonNullParamChecker should run before StdCLibraryFunctionArgs to get more special warning about null arguments, but the apiModeling.StdCLibraryFunctions was a modeling checker that should run before other non-modeling checkers. The modeling checker changes state in a way that makes it impossible to detect a null argument by NonNullParamChecker.) To make it more simple, the modeling part is removed as separate checker and can be only used if checker StdCLibraryFunctions is turned on, that produces the warnings too. Modeling the functions without bug detection (for invalid argument) is not possible. The modeling of standard functions does not happen by default from this change on. Reviewed By: Szelethus Differential Revision: https://reviews.llvm.org/D151225
-
Nikita Popov authored
Make ValueTracking directly call the KnownBits shift helpers, which provides more precise results. Unfortunately, ValueTracking has a special case where sometimes we determine non-zero shift amounts using isKnownNonZero(). I have my doubts about the usefulness of that special-case (it is only tested in a single unit test), but I've reproduced the special-case via an extra parameter to the KnownBits methods. Differential Revision: https://reviews.llvm.org/D151816
-
Matthias Springer authored
Certain ExtractSliceOps, that do extract all elements from the destination, are treated like casts when looking for replacement ops. Such ExtractSliceOps are typically rank expansions. Differential Revision: https://reviews.llvm.org/D151804
-
Matthias Springer authored
Drop insert_slice rank expansions if they are directly followed by an inverse rank reduction. Differential Revision: https://reviews.llvm.org/D151800
-
Manas authored
Zero ranked tensor (say tensor<i1>) when used for arith.select's condition, crashes optimizer during bufferization. This patch puts a constraint on condition to be either scalar or of matching shape as to its result. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D151270
-
Petr Hosek authored
This reverts commit 80614e16.
-
Petr Hosek authored
The existing BOLT install targets are broken on Windows becase they don't properly handle the output extension. We cannot use the existing LLVM macros since those make assumptions that don't hold for BOLT. This change instead implements custom macros following the approach used by Clang and LLD. Differential Revision: https://reviews.llvm.org/D151595
-
Phoebe Wang authored
Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D151808
-
Guillaume Chatelet authored
Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D151798
-
wangpc authored
So that we can remove `SchedSEWSetF` and simplify some code. Reviewed By: michaelmaitland Differential Revision: https://reviews.llvm.org/D151790
-
Piyou Chen authored
Encountered ASAN crash and found it dereference without check pointer. Reviewed By: kito-cheng, eklepilkina Differential Revision: https://reviews.llvm.org/D151716
-
Joshua Cao authored
Before this patch, we can only use the MaxBECount for an AddRec's range computation if the MaxBECount has <= bit width of the AddRec. This patch reasons that if a MaxBECount has > bit width, and is <= the max value of AddRec's bit width, we can still use the MaxBECount. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D151698
-
Joshua Cao authored
-
Joshua Cao authored
-
Craig Topper authored
Support for this was removed previously. Change them to "supervisor" since they were testing generic "interrupt" things.
-
zhanglimin authored
In LoongArch ABI spec, we can see that in the LP64D ABI, unsigned 32-bit types, such as unsigned int, are stored in general-purpose registers as proper sign extensions of their 32-bit values. Reference: https://loongson.github.io/LoongArch-Documentation/LoongArch-ELF-ABI-EN.html#_abi_lp64d Reviewed By: SixWeining, xen0n Differential Revision: https://reviews.llvm.org/D151794
-
Kevin Gleason authored
Currently desired bytecode version is clamped to the maximum. This allows requesting bytecode versions that do not exist. We have added callsite validation for this in StableHLO to ensure we don't pass an invalid version number, probably better if this is managed upstream. If a user wants to use the current version, then omitting `setDesiredBytecodeVersion` is the best way to do that (as opposed to providing a large number). Adding this check will also properly error on older version numbers as we increment the minimum supported version. Silently claming on minimum version would likely lead to unintentional forward incompatibilities. Separately, due to bytecode version being `int64_t` and using methods to read/write uints, we can generate payloads with invalid version numbers: ``` mlir-opt file.mlir --emit-bytecode --emit-bytecode-version=-1 | mlir-opt <stdin>:0:0: error: bytecode version 18446744073709551615 is newer than the current version 5 ``` This is fixed with version bounds checking as well. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D151838
-
Jason Molenda authored
On AArch64, it is possible to have a program that accesses both low (0x000...) and high (0xfff...) memory, and with pointer authentication, you can have different numbers of bits used for pointer authentication depending on whether the address is in high or low memory. This adds a new target.process.highmem-virtual-addressable-bits setting which the AArch64 Mac ABI plugin will use, when set, to always set those unaddressable high bits for high memory addresses, and will use the existing target.process.virtual-addressable-bits setting for low memory addresses. This patch does not change the existing behavior when only target.process.virtual-addressable-bits is set. In that case, the value will apply to all addresses. Not yet done is recognizing metadata in a live process connection (gdb-remote qHostInfo) or a Mach-O corefile LC_NOTE to set the correct number of addressing bits for both memory ranges. That will be a future change. Differential Revisio...
-
Ellis Hoag authored
The tests introduced by https://reviews.llvm.org/D151589 were failing because I guess some test platforms don't have `lld`. Similar tests add `-B%S/Inputs/lld` to the clang commands so lets try this here to fix the tests. ``` clang: error: invalid linker name in argument '-fuse-ld=lld' ```
-
LLVM GN Syncbot authored
-
Nikolas Klauser authored
``` --------------------------------------------------- Benchmark old new --------------------------------------------------- bm_for_each/1 3.00 ns 2.98 ns bm_for_each/2 4.53 ns 4.57 ns bm_for_each/3 5.82 ns 5.82 ns bm_for_each/4 6.94 ns 6.91 ns bm_for_each/5 7.55 ns 7.75 ns bm_for_each/6 7.06 ns 7.45 ns bm_for_each/7 6.69 ns 7.14 ns bm_for_each/8 6.86 ns 4.06 ns bm_for_each/16 11.5 ns 5.73 ns bm_for_each/64 43.7 ns 4.06 ns bm_for_each/512 356 ns 7.98 ns bm_for_each/4096 2787 ns 53.6 ns bm_for_each/32768 20836 ns 438 ns bm_for_each/262144 195362 ns 4945 ns bm_for_each/1048576 685482 ns 19822 ns ``` Reviewed By: ldionne, Mordante, #libc Spies: arichardson, libcxx-commits Differential Revision: https://reviews.llvm.org/D151274
-
Nikolas Klauser authored
This simplifies the code inside copy/move and makes it easier to apply the optimization to other algorithms. Reviewed By: ldionne, Mordante, #libc Spies: arichardson, libcxx-commits Differential Revision: https://reviews.llvm.org/D151265
-
Ellis Hoag authored
Enable support for CSPGO for lld MachO targets. Since lld MachO does not support `-plugin-opt=`, we need to create the `--cs-profile-generate` and `--cs-profile-path=` options and propagate them in `Darwin.cpp`. These flags are not supported by ld64. Also outline code into `getLastCSProfileGenerateArg()` to share between `CommonArgs.cpp` and `Darwin.cpp`. CSPGO is already implemented for ELF (https://reviews.llvm.org/D56675) and COFF (https://reviews.llvm.org/D98763). Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D151589
-