- Jun 23, 2023
-
-
Ties Stuij authored
[ARM] generate armv6m eXecute Only (XO) code for immediates, globals Previously eXecute Only (XO) support was implemented for targets that support MOVW/MOVT (~armv7+). See: https://reviews.llvm.org/D27449 XO prevents the compiler from generating data accesses to code sections. This patch implements XO codegen for armv6-M, which does not support MOVW/MOVT, and must resort to the following general pattern to avoid loads: movs r3, :upper8_15:foo lsls r3, #8 adds r3, :upper0_7:foo lsls r3, #8 adds r3, :lower8_15:foo lsls r3, #8 adds r3, :lower0_7:foo ldr r3, [r3] This is equivalent to the code pattern generated by GCC. The above relocations are new to LLVM and have been implemented in a parent patch: https://reviews.llvm.org/D149443. This patch limits itself to implementing codegen for this pattern and enabling XO for armv6-M in the backend. Separate patches will follow for: - switch tables - replacing specific loads from constant islands which are spread out over the ARM backend codebase. Amongst others: FastISel, call lowering, stack frames. Reviewed By: john.brawn Differential Revision: https://reviews.llvm.org/D152795
-
pvanhout authored
It was overdue for a clang-format run, and it avoids unrelated formatting changes sneaking into diffs.
-
pvanhout authored
Accidentally copy-pasted them into the .cpp while refactoring the file in D151432 Those functions are currently only used in the .cpp so it didn't cause an issue, but it causes an undefined reference if another file attempts to use them.
-
Jolanta Jensen authored
This patch removes DAG combines that are no longer relevant because equivalent IR combines have been added. Differential Revision: https://reviews.llvm.org/D153445
-
Jeremy Furtek authored
This diff adds APFloat support for a semantic that matches the TF32 data type used by some accelerators (most notably GPUs from both NVIDIA and AMD). For more information on the TF32 data type, see https://blogs.nvidia.com/blog/2020/05/14/tensorfloat-32-precision-format/. Some intrinsics that support the TF32 data type were added in https://reviews.llvm.org/D122044. For some discussion on supporting common semantics in `APFloat`, see similar efforts for 8-bit formats at https://reviews.llvm.org/D146441, as well as https://discourse.llvm.org/t/rfc-adding-the-amd-graphcore-maybe-others-float8-formats-to-apfloat/67969. A subsequent diff will extend MLIR to use this data type. (Those changes are not part of this diff to simplify the review process.) Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D151923
-
Kazu Hirata authored
Differential Revision: https://reviews.llvm.org/D153615
-
Kazu Hirata authored
The corresponding function definition was removed by: commit 773d663e Author: Arthur Eubanks <aeubanks@google.com> Date: Mon Feb 27 19:00:37 2023 -0800
-
Tamás Danyluk authored
If a value is already the last element of the worklist, then I think that we don't have to add it again, it is not needed to process it repeatedly. For some long Triton-generated LLVM IR, this can cause a ~100x speedup. Differential Revision: https://reviews.llvm.org/D153561
-
Alex Zinenko authored
Wrapping a warning into a silenceable failure will result in the warning being interpreted as an error, which it is not. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D153546
-
Alex Zinenko authored
When exiting the scope of a region attached to a transform op, clean up the handle invalidation checks assocaited with handles defined in this region. Otherwise, these checks may trigger on the next entry to the region while there is no incorrect usage. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D153545
-
Balázs Kéri authored
Fix for issue #62770. Reviewed By: donat.nagy Differential Revision: https://reviews.llvm.org/D153424
-
Dhruv Chawla authored
When the sign of either of the operands is known, it is possible to determine what the saturating value will be without having to compute it using the sign bits. Differential Revision: https://reviews.llvm.org/D153575
-
Dhruv Chawla authored
-
Kazu Hirata authored
Differential Revision: https://reviews.llvm.org/D153610
-
Ulrich Weigand authored
This makes the bytecode reader/writer work on big-endian platforms. The only problem was related to encoding of multi-byte integers, where both reader and writer code make implicit assumptions about endianness of the host platform. This fixes the current test failures on s390x, and in addition allows to remove the UNSUPPORTED markers from all other bytecode-related test cases - they now also all pass on s390x. Also adding a GFAIL_SKIP to the MultiModuleWithResource unit test, as this still fails due to an unrelated endian bug regarding decoding of external resources. Differential Revision: https://reviews.llvm.org/D153567 Reviewed By: mehdi_amini, jpienaar, rriddle
-
Haojian Wu authored
Remove the existing `Rng` field. From the review comment: https://reviews.llvm.org/D147034 Reviewed By: kadircet Differential Revision: https://reviews.llvm.org/D153259
-
Jean Perier authored
Use hlfir::loadTrivialScalars to dereference pointer, allocatables, and load numerical and logical scalars. This has a small fallout on tests: - load is done on the HLFIR entity (#0 of hlfir.declare) and not the FIR one (#1). This makes no difference at the FIR level (#1 and #0 only differs to account for assumed and explicit shape lower bounds). - loadTrivialScalars get rids of allocatable fir.box for monomoprhic scalars (it is not needed). This exposed a bug in lowering of MERGE with a polymorphic and a monomorphic argument: when the monomorphic is not a fir.box, the polymorphic fir.class should not be reboxed but its address should be read. Reviewed By: tblah Differential Revision: https://reviews.llvm.org/D153252
-
Martin Braenne authored
- The AST of the function we're currently analyzing - The CFG - The CFG element we're currently processing Reviewed By: ymandel Differential Revision: https://reviews.llvm.org/D153549
-
Kazu Hirata authored
The declaration was added without a corresponding class definition by: commit 13bb8f49 Author: Stella Laurenzo <laurenzo@google.com> Date: Wed Apr 3 11:16:32 2019 -0700
-
Kazu Hirata authored
The declaration was added without a corresponding class definition by: commit a84064bc Author: Florian Hahn <flo@fhahn.com> Date: Wed Dec 21 22:02:31 2022 +0000 It is most likely a misspelling of PredicatedScalarEvolution.
-
Kazu Hirata authored
-
Jacques Pienaar authored
Allows for single op nested regions.
-
Kazu Hirata authored
The declaration was added without a corresponding function by: commit cc3bb855 Author: James Nagurne <j-nagurne@ti.com> Date: Fri Oct 22 17:08:16 2021 -0500
-
Jon Chesterfield authored
NFC. Simplifies process slightly, gives more options for testing it. Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D153604
-
Jon Chesterfield authored
Private member variable minimises scope of access to Process Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D153603
-
Oleksii Lozovskyi authored
Instead of dumping all sources into RTXray object library with a weird special case for x86, handle multiarch builds better. Build a separate object library for each arch with its arch-specific sources, then link in all those libraries. This fixes the build on platforms that produce fat binaries, such as new macOS which expects both x86_64 and aarch64 objects in the same library since Apple Silicon is a thing. This only enables building XRay support for Apple Silicon. It does not actually work yet on macOS, neither on Intel nor on Apple Silicon CPUs. Thus the tests are still disabled. Reviewed By: MaskRay, phosek Differential Revision: https://reviews.llvm.org/D153221
-
Jon Chesterfield authored
CMake plumbing cargo culted from other tests. Minor changes to Process to allow statically allocating a buffer. Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D153594
-
Amaury Séchet authored
This fixes a regression in D127115 Depends on D127115 Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D151916
-
Sheng authored
`TargetGlobalTLSAddress` is not considered and handled correctly when matching addressing mode, which leads to an incorrect result of instruction selection. fixes #63162. Reviewed By: myhsu Differential Revision: https://reviews.llvm.org/D153103
-
Vitaly Buka authored
-
Vitaly Buka authored
-
Vitaly Buka authored
-
Vitaly Buka authored
-
Alex Langford authored
The Objective-C runtime and the shared cache has changed slightly. Given a class_ro_t, the baseMethods ivar is now a pointer union and may either be a method_list_t pointer or a pointer to a relative list of lists. The entries of this relative list of lists are indexes that refer to a specific image in the shared cache in addition to a pointer offset to find the accompanying method_list_t. We have to go over each of these entries, parse it, and then if the relevant image is loaded in the process, we add those methods to the relevant clang Decl. In order to determine if an image is loaded, the Objective-C runtime exposes a symbol that lets us determine if a particular image is loaded. We maintain a data structure SharedCacheImageHeaders to keep track of that information. There is a known issue where if an image is loaded after we create a Decl for a class, the Decl will not have the relevant methods from that image (i.e. for Categories). rdar://107957209 Differential Revision: https://reviews.llvm.org/D153597
-
Matt Arsenault authored
This reverts commit aa7e09eb.
-
Matt Arsenault authored
-
Daniel Hoekwater authored
On AArch64, object files may be greater than 2^32 bytes. If an offset is greater than the max value of a 32-bit unsigned integer, LLVM silently truncates the offset. Instead, make it return an error. Differential Revision: https://reviews.llvm.org/D153494
-
Fangrui Song authored
-
Philip Reames authored
I tried to give a rough overview of our current pseudo structure. I'm mostly focused on the policy handling bits - since that's what I'm in the process of changing - but touched on the other dimensions in the process of framing it. Differential Revision: https://reviews.llvm.org/D152937
-
Shatian Wang authored
Order code sections with names in the form of ".text.cold.i" based on the value of i [Context] SplitFunctions.cpp implements splitting strategies that can potentially split each function into maximum N>2 fragments. When such N-way splitting happens, new code sections with names ".text.cold.1", ..., ".text.cold.i", ... "text.cold.N-2" will be created A section with name ".text.cold.i" contains the the (i+2)th fragment of each function. As an example, if each function is splitted into N=3 fragments: hot, warm, cold, then code sections will now include - a section with name ".text" containing hot fragments - a section with name ".text.cold" containing warm fragments - a section with name ".text.cold.1" containing cold fragments The order of these new sections in the output binary currently depends on the order in which they are encountered by the emitter. For example, under N=3-way splitting, if the first function is 2-way ...
-