- Jan 20, 2023
-
-
Frederik Gossen authored
This reverts commit 399b8ee7.
-
Paul Kirth authored
We were iterating over a SmallPtrSet when outputting slot variables. This is still correct but made the test fail under reverse iteration. This patch replaces the SmallPtrSet with a SmallVector. Also remove the "Stack Frame Layout" lines from arm64-opt-remarks-lazy-bfi test, since those also break under reverse iteration. Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D142127
-
Jim Ingham authored
This is processed by hand in CommandObjectMultiword, and is undiscoverable, it doesn't work in all cases. Because it is a bare word, it can't really be extended w/o introducing the possibility of collisions as well. If we did want to do something like this we should add a --help flag to CommandObject. That way the feature would be consistent and documented. Differential Revision: https://reviews.llvm.org/D142067
-
Mark de Wever authored
Implements the range-default-formatter specialization range_format::map. Implements parts of - P2286R8 Formatting Ranges - P2585R0 Improving default container formatting Depends on D140653 Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D140801
-
Mehdi Amini authored
-
Arvind Sudarsanam authored
This is an issue reported inside the NewPMDriver module. Static analyzer reported that Null pointer 'P' may be dereferenced at line 371 and two more sites. Proposed change guards this use. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D142047
-
Noah Goldstein authored
According to https://uops.info/ ICL and newer have fast 3-term LEA. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D141974
-
Noah Goldstein authored
Definitionally a non-zero power of 2 will only have 1 bit set so this is a freebee. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D141990
-
Noah Goldstein authored
Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D141989
-
Frederik Gossen authored
-
LLVM GN Syncbot authored
-
Nikolas Klauser authored
This has multiple benefits: - The optimizations are also performed for the `ranges::` versions of the algorithms - Code duplication is reduced - it is simpler to add this optimization for other segmented iterators, like `ranges::join_view::iterator` - Algorithm code is removed from `<deque>` Reviewed By: ldionne, huixie90, #libc Spies: mstorsjo, sstefan1, EricWF, libcxx-commits, mgorny Differential Revision: https://reviews.llvm.org/D132505
-
Stanislav Mekhanoshin authored
MFMA and WMMA essentially the same thing, but apear on different ASICs. Differential Revision: https://reviews.llvm.org/D142062
-
Stanislav Mekhanoshin authored
Current implementation abuses ErrorMargin to apply an additional bias to VGPR and SGPR limits under a high register pressure. The ErrorMargin exists to account for inaccuracies of the RP tracker and not to tackle an excess pressure. Introduce separate bias for this purpose and also make it different for SGPRs and VGPRs as we may want to use different values in the future. This is supposed to be NFC, however there is a subtle difference when subtracting a margin overflows the limit. Doing two subtractions makes it less probable, although manifests only in mir tests with an artificially small register budget. Differential Revision: https://reviews.llvm.org/D142051
-
Joseph Huber authored
Summary: There shouldn't be an extra newline in these messages.
-
Joseph Huber authored
Right now in the linker wrapper we manually invoke a lot of the toolchain programs. This reproduces a lot of logic that is already handled in clang. Since D140158 we can now target all supported toolchains directly via cross-compilation. This patch changes the linker wrapper to consolidate all the alternate linking and assembler steps into a generic call to `clang` and let clang handle the argument handling. This heavily simplifies the interface. Reviewed By: tra, JonChesterfield Differential Revision: https://reviews.llvm.org/D142133
-
Krzysztof Drewniak authored
This reverts commit 45530562. Linker error, unbreak build while I work out how to fix it. Differential Revision: https://reviews.llvm.org/D142142
-
Gulfem Savrun Yeniceri authored
This patch replaces CallInstr with CallBase to cover InvokeInstr besides CallInstr while removing nocallback attribute on a call site. It also extends drop-attribute.ll test to include a case for an invoke instruction. Differential Revision: https://reviews.llvm.org/D141740
-
Arthur Eubanks authored
This reverts commit da5a8d14. Causes more duplicate symbol errors, see https://bugs.chromium.org/p/chromium/issues/detail?id=1408161.
-
Frederik Gossen authored
Differential Revision: https://reviews.llvm.org/D142049
-
Erich Keane authored
As reported in https://github.com/llvm/llvm-project/issues/54524, and later in https://github.com/llvm/llvm-project/issues/60038, we were not properly implmenting temp.constr.atomic P3. This patch stops implicitly converting constraints to bool, and ensures the Rvalue conversion takes place as needed. Differential Revision: https://reviews.llvm.org/D141954
-
Florian Hahn authored
The scope of DT updates are very limited when unrolling loops: the DT should only need updating for * new blocks added * exiting blocks we simplified branches This can be done manually without too much extra work. MergeBlockIntoPredecessor also needs to be updated to support direct DT updates. This fixes excessive time spent in DTU for same cases. In an internal example, time spent in LoopUnroll with this patch goes from ~200s to 2s. It also is slightly positive for CTMark: * NewPM-O3: -0.13% * NewPM-ReleaseThinLTO: -0.11% * NewPM-ReleaseLTO-g: -0.13% Notable improvements are mafft (~ -0.50%) and lencod (~ -0.30%), with no workload regressed. https://llvm-compile-time-tracker.com/compare.php?from=78a9ee7834331fb4360457cc565fa36f5452f7e0&to=687e08d011b0dc6d3edd223612761e44225c7537&stat=instructions:u Reviewed By: kuhar Differential Revision: https://reviews.llvm.org/D141487
-
Matthias Springer authored
Upper bound and step size should be symbols instead of dims. Differential Revision: https://reviews.llvm.org/D142136
-
David Carlier authored
Reviewers: dvyukov Reviewed-By: dvyukov Differental Revision: https://reviews.llvm.org/D140688
-
Krzysztof Drewniak authored
Implement InferIntRangeInterface for all operations in the Index dialect. The inference implementation, unlike the one for Arith, accounts for the fact that Index can be either 64 or 32 bits long by evaluating both cases. Bounds are stored as if index were i64, but when inferring new bounds, we compute both f(...) and f(trunc(...)). We then compare trunc(f(...)) to f(trunc(...)). If they are equal in the relevant range components, we use the 64-bit range computation, otherwise we give the range ext(f(trunc(...))) union f(...). Note that this can cause surprising behavior as seen in the tests, where, for example, the order of min and max operations impacts the behavior of the inference. The inference could perhaps be made more precise in the future (ex. by tracking 32 and 64-bit results separately and having them influence each other somehow) butt, since my project targets an index=i32 platform and doesn't see index-valued values > uint32_max, I'm not too concerned about it. Depends on https://reviews.llvm.org/D141299 Depends on https://reviews.llvm.org/D141296 Reviewed By: Mogball Differential Revision: https://reviews.llvm.org/D140899
-
Xing Xue authored
Summary: This patch adds OpenMP runtime to the linker command line if -fopenmp is specifed for AIX. Reviewed by: daltenty Differential Revision: https://reviews.llvm.org/D141862
-
v1nh1shungry authored
Current version there is a fix-it for template <class> constexpr int x = 0; template <> constexpr int x<int>; // fix-it here but it will cause template <> constexpr int x = 0<int>; Differential Revision: https://reviews.llvm.org/D139705
-
Paul Robinson authored
This reverts commit a0f8bdbb. Several bots are failing in shtest-format.py, likely because of this.
-
Slava Zakharin authored
Reviewed By: jeanPerier, PeteSteinfeld Differential Revision: https://reviews.llvm.org/D142070
-
Aaron Ballman authored
The std::optional implementation in MSVC causes this code to produce a sign comparison warning. This ensures the types are the same sign.
-
Michael Jones authored
This patch adds the %f/F/e/E/g/G/a/A conversions for scanf, as well as accompanying tests. This implementation matches the definition set forth in the standard, which may conflict with some other implementations. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D141091
-
Michael Jones authored
The scanf implementation needs a dynamically resizing string class. This patch adds a minimal version of that class along with tests to check the current functionality. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D141162
-
Kiran Chandramohan authored
-> Use file pathname from the Flang frontend. It is the frontend that is in-charge of finding the files and is hence the canonical source for paths. -> Convert pathname to absolute pathname while creating the moduleOp. Co-authored-by:
Peter Klausler <pklausler@nvidia.com> Reviewed By: PeteSteinfeld, vzakhari, jeanPerier, awarzynski Differential Revision: https://reviews.llvm.org/D141674
-
Mark de Wever authored
Implements parts of - P2286R8 Formatting Ranges Depends on D140653 Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D141761
-
LLVM GN Syncbot authored
-
Zino Benaissa authored
shuffle_vector instructions are serialized targeting SVE fixed vectors, see https://reviews.llvm.org/D139111. This patch disables optimizeExtendOrTruncateConversion peepholes that generates shuffle_vector. Differential Revision: https://reviews.llvm.org/D141439
-
Sjoerd Meijer authored
-
Mark de Wever authored
Implements parts of - P2286R8 Formatting Ranges Depends on D140653 Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D141290
-
Jonas Paulsson authored
Only allow replacements of nodes that have a single user. This is better as simple instructions (e.g. XGRK) are one cycle faster, and it helps in cases where both inputs share a common node. Review: Ulrich Weigand
-
Jordan Rupprecht authored
-