- Jun 29, 2023
-
-
Paul Kirth authored
Fat LTO objects contain both LTO compatible IR, as well as generated object code. This allows users to defer the choice of whether to use LTO or not to link-time. This is a feature available in GCC for some time, and makes the existing -ffat-lto-objects flag functional in the same way as GCC's. Within LLVM, we add a new EmbedBitcodePass that serializes the module to the object file, and expose a new pass pipeline for compiling fat objects. The new pipeline initially clones the module and runs the selected (Thin)LTOPrelink pipeline, after which it will serialize the module into a `.llvm.lto` section of an ELF file. When compiling for (Thin)LTO, this normally the point at which the compiler would emit a object file containing the bitcode and metadata. After that point we compile the original module using the PerModuleDefaultPipeline used for non-LTO compilation. We generate standard object files at the end of this pipeline, which contain machine code and the new `.llvm.lto` section containing bitcode. Since the two pipelines operate on different copies of the module, we can be sure that the bitcode in the `.llvm.lto` section and object code in `.text` are congruent with the existing output produced by the default and LTO pipelines. Original RFC: https://discourse.llvm.org/t/rfc-ffat-lto-objects-support/63977 Earlier versions of this patch were missing REQUIRES lines for llc related tests in Transforms/EmbedBitcode. Those tests are now under CodeGen/X86, which should avoid running the check on unsupported platforms. The EmbedbBitcodePass also returned PreservedAnalyses::all when adding a metadata section, which failed expensive checks, since it modified the module. This is now corrected. Reviewed By: tejohnson, MaskRay, nikic Differential Revision: https://reviews.llvm.org/D146776
-
Peiming Liu authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D153998
-
Fangrui Song authored
clang -ffat-lto-objects can use this new ELF section type for the .llvm.lto section for fat LTO support (D146776). Original RFC: https://discourse.llvm.org/t/rfc-ffat-lto-objects-support/63977 Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D153215
-
Matt Arsenault authored
-
Alexey Bataev authored
If the buildvector node is a full match of another node, need to correctly build the mask for the original vector value and build common mask for the emitted node.
-
Wenlei He authored
Exposing a non-const accessor for clearing CallsiteSamples during flattening is a big of an overkill. Replace the non-const accessor with removeAllCallsiteSamples. Differential Revision: https://reviews.llvm.org/D153995
-
Nikolas Klauser authored
Fixes #63192 Reviewed By: cor3ntin Spies: cfe-commits Differential Revision: https://reviews.llvm.org/D153890
-
Ethan Luis McDonough authored
Flang currently supports offloading for AMD GPUs. This patch establishes a test structure for Fortran offloading tests in libomptarget. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D148778
-
Florian Hahn authored
Extra tests for D152730 with different GEP step sizes and the end pointer being an argument.
-
David Green authored
Similar to the other code that costs main/alt instructions, the cmp should be using the VecTy for the costs, not the ScalarTy. One of the tests look like it gets worse just because it is not simplified to 0. Differential Revision: https://reviews.llvm.org/D153507
-
Serge Pavlov authored
This reverts commit 98390ccb. It caused issue #63542.
-
Matt Arsenault authored
-
Matt Arsenault authored
-
root authored
Currently, bf16 has been scatteredly added to the PTX codegen. This patch aims to complete the set of instructions and code path required to support bf16 data type. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D144911 Co-authored-by:
Artem Belevich <tra@google.com>
-
Matt Arsenault authored
-
Matt Arsenault authored
Add an intrinsic which returns the two pieces as multiple return values. Alternatively could introduce a pair of intrinsics to separately return the fractional and exponent parts. AMDGPU has native instructions to return the two halves, but could use some generic legalization and optimization handling. For example, we should be able to handle legalization of f16 on older targets, and for bf16. Additionally antique targets need a hardware workaround which would be better handled in the backend rather than in library code where it is now.
-
Caroline Tice authored
In two calls to ReadMemory in DWARFExpression.cpp, the buffer size passed to ReadMemory is not actually the size of the buffer (I suspect a copy/paste error where the variable name was not properly updated). This caused a buffer overflow bug, which we found throuth Address Sanitizer. This patch fixes the problem by passing the correct buffer size to the calls to ReadMemory (and to the DataExtractor). Differential Revision: https://reviews.llvm.org/D153840
-
Nikolas Klauser authored
-
Tue Ly authored
Implement correctly rounded `erff` functions. For `x >= 4`, `erff(x) = 1` for `FE_TONEAREST` or `FE_UPWARD`, `0x1.ffffep-1` for `FE_DOWNWARD` or `FE_TOWARDZERO`. For `0 <= x < 4`, we divide into 32 sub-intervals of length `1/8`, and use a degree-15 odd polynomial to approximate `erff(x)` in each sub-interval: ``` erff(x) ~ x * (c0 + c1 * x^2 + c2 * x^4 + ... + c7 * x^14). ``` For `x < 0`, we can use the same formula as above, since the odd part is factored out. Performance tested with `perf.sh` tool from the CORE-MATH project on AMD Ryzen 9 5900X: Reciprocal throughput (clock cycles / op) ``` $ ./perf.sh erff --path2 GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH reciprocal throughput -- with -march=native (with FMA instructions) [####################] 100 % Ntrial = 20 ; Min = 11.790 + 0.182 clc/call; Median-Min = 0.154 clc/call; Max = 12.255 clc/call; -- CORE-MATH reciprocal throughput -- with -marc...
-
Daniel Thornburgh authored
The symbolizer markup syntax is structured such that fields require only previous fields for their interpretation; this was originally intended to make adding new fields a natural extension mechanism for existing elements. This codifies this into the spec and makes the behavior of the llvm-symbolizer match. Extra fields are now warned about, but ignored, rather than ignoring the whole element. Reviewed By: mcgrathr Differential Revision: https://reviews.llvm.org/D153821
-
Jon Roelofs authored
In https://reviews.llvm.org/D149445, it was lowered from 32 to 16bits, which broke an internal project of ours. The relevant code being compiled is a fairly large nested switch that results in a PHI node with 65k+ operands, which can't easily be turned into a table for perf reasons. This change unifies `NumOperands`, `Flags`, and `AsmPrinterFlags` into a packed 7-byte struct, which `CapOperands` can follow as the 8th byte, rounding it up to a nice alignment before the `Info` field. rdar://111217742&109362033 Differential revision: https://reviews.llvm.org/D153791
-
Matt Arsenault authored
Move it up with other module passes. It's a higher level optimization that should probably be done before hacking up the IR for codegen. It should really be done earlier than this. We could possibly move this with other IPO passes, but we'd have to stop inferring the lack of lds.kernel.id calls and have the LDS module pass mark functions which don't need the ID. The one test change is because that pass is relying on the backend run of SROA (which we ideally wouldn't have).
-
Philip Reames authored
-
Snehasish Kumar authored
Add an overload for InstrProfWriter::write so that users can emit the buffer to a string. Also use this new overload for existing unit test usecases. Reviewed By: tejohnson Differential Revision: https://reviews.llvm.org/D153904
-
David Green authored
See D153507. The existing test is over-simplified, as written it should have been simpified prior to SLP vectorization. I have left it as-is to ensure the crash it was protecting against doesn't arise again. A new test with valid inputs is also added to show the incorrect costs of alt cmp vectorization.
-
Fraser Cormack authored
D143505 fixed/simplified folding of operations with SNaN operands. In doing so it introduced a crash when handling scalable vector types, wherein the scalable-vector ConstantVector was cast to a ConstantFP. Since we know by that point in the code that if we've found a NaN, we're dealing with a scalable-vector splat (as there are no other kinds of scalable-vector constant for which that holds), we can grab the splatted value and re-use the existing code, which will automatically splat the new NaN back to a scalable vector for us. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153566
-
Valentin Clement authored
Some symbols were not resolved in the device, host and self clause resulting in an `Internal: no symbol found` error. This patch adds symbol resolution for these clauses. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D153919
-
Valentin Clement authored
Some compiler treat `acc routine` without a parallelism clause as if seq is present. Relax the parser rule to allow acc routine without clause. The default clause will be handled in lowering. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D153896
-
- Jun 28, 2023
-
-
Paul Robinson authored
-
LLVM GN Syncbot authored
-
Yusra Syeda authored
Revert "[SystemZ][z/OS] This patch adds support for the ADA (associated data area), doing the following:" This reverts commit 9df0f66a.
-
Jeffrey Byrnes authored
[AMDGPU] NFC: Add schedule-relaxed-occupancy to relax occupancy targets for wave-limited/membound kernels Default scheduling behavior for these types of kernels is to chase high occupancy goals with scheduling heuristics, but allow occupancy drops if we are unable to reach the target. This (experimental, off-by-default) feature relaxes occupancy target from the beginning, which enables scheduler to produce better ILP schedules. Differential Revision: https://reviews.llvm.org/D153925 Change-Id: I112833214e2db869704591f4df3c4574d0fcbb1b
-
Shilei Tian authored
the color of finished task
-
Craig Topper authored
The functions are identical except for the opcode of the node. We can have a single function and use N->getOpcode(). Reviewed By: luke, paulwalker-arm Differential Revision: https://reviews.llvm.org/D153929
-
Paul Robinson authored
Differential Revision: https://reviews.llvm.org/D153884
-
LLVM GN Syncbot authored
-
Yusra Syeda authored
- Creates the ADA table to handle displacements - Emits the ADA section in the SystemZAsmPrinter - Lowers the ADA_ENTRY node into the appropriate load instruction Differential Revision: https://reviews.llvm.org/D153788
-
LLVM GN Syncbot authored
-
Felipe de Azevedo Piovezan authored
This concludes the migration of accelerator tables from LLDB code to LLVM code. Differential Revision: https://reviews.llvm.org/D153868
-
David Green authored
-