- Jun 24, 2023
-
-
Matt Arsenault authored
-
Matt Arsenault authored
These are basically the same thing and only differ for strictfp, so add both for future proofing. Note all the elementwise functions are currently broken for strictfp, and use non-constrained ops. Add a test that demonstrates this, but doesn't attempt to fix it.
-
Hongtao Yu authored
Compiler-generated static symbols, such as the global initializers, can shared the same name and can coexist in the binary. As a result, their pseudo probes are all kept in the binary too. This could cause multiple call probes decoded against one callsite, as probes are decoded against there owning functions by name. I'm temporarily disabling an assert to keep the debug build green until we have a better fix. Reviewed By: wenlei Differential Revision: https://reviews.llvm.org/D153588
-
Matt Arsenault authored
-
Johannes Doerfert authored
If the user wants to avoid running additional passes, they can now initialize the AnalysisGetter accordingly.
-
Johannes Doerfert authored
It was never really useful to track #iterations, though it helped during the initial development. What we should track, in a follow up, are potentially #updates. That is also what we should restrict instead of the #iterations.
-
Matt Arsenault authored
Pass all arguments so now assumes work.
-
Matt Arsenault authored
Introduce a full featured wrapper around computeKnownFPClass to start replacing the uses with.
-
Paul Kirth authored
Fat LTO objects contain both LTO compatible IR, as well as generated object code. This allows users to defer the choice of whether to use LTO or not to link-time. This is a feature available in GCC for some time, and makes the existing -ffat-lto-objects flag functional in the same way as GCC's. Within LLVM, we add a new EmbedBitcodePass that serializes the module to the object file, and expose a new pass pipeline for compiling fat objects. The new pipeline initially clones the module and runs the selected (Thin)LTOPrelink pipeline, after which it will serialize the module into a `.llvm.lto` section of an ELF file. When compiling for (Thin)LTO, this normally the point at which the compiler would emit a object file containing the bitcode and metadata. After that point we compile the original module using the PerModuleDefaultPipeline used for non-LTO compilation. We generate standard object files at the end of this pipeline, which contain machine code and the new `.llvm.lto` section containing bitcode. Since the two pipelines operate on different copies of the module, we can be sure that the bitcode in the `.llvm.lto` section and object code in `.text` are congruent with the existing output produced by the default and LTO pipelines. Original RFC: https://discourse.llvm.org/t/rfc-ffat-lto-objects-support/63977 Earlier versions of this patch were missing REQUIRES lines for llc related tests in Transforms/EmbedBitcode. Those tests are now under CodeGen/X86, which should avoid running the check on unsupported platforms. Reviewed By: tejohnson, MaskRay, nikic Differential Revision: https://reviews.llvm.org/D146776
-
Sami Tolvanen authored
With `-fsanitize=kcfi` (Kernel Control-Flow Integrity), Clang emits "kcfi" operand bundles to indirect call instructions. Similarly to the target-specific lowering added in D119296, implement KCFI operand bundle lowering for RISC-V. This patch disables the generic KCFI pass for RISC-V in Clang, and adds the KCFI machine function pass in `RISCVPassConfig::addPreSched` to emit target-specific `KCFI_CHECK` pseudo instructions before calls that have KCFI operand bundles. The machine function pass also bundles the instructions to ensure we emit the checks immediately before the calls, which is not possible with the generic pass. `KCFI_CHECK` instructions are lowered in `RISCVAsmPrinter` to a contiguous code sequence that traps if the expected hash in the operand bundle doesn't match the hash before the target function address. This patch emits an `ebreak` instruction for error handling to match the Linux kernel's `BUG()` implementation. Just like for X86, we also emit trap locations to a `.kcfi_traps` section to support error handling, as we cannot embed additional information to the trap instruction itself. Relands commit 62fa708c with fixed tests. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D148385
-
Mehdi Amini authored
This is a rarely used API: `matchAndRewrite()` is the usual way of using patterns.
-
Matt Arsenault authored
Move cannotBeOrderedLessThanZero logic into computeKnownFPClass.
-
Kazu Hirata authored
This patch fixes -Wsign-compare warnings from: llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:117:3: note: in instantiation of function template specialization 'testing::internal::EqHelper::Compare<unsigned long, int, nullptr>' requested here llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:134:3: note: in instantiation of function template specialization 'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>' requested here llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:144:3: note: in instantiation of function template specialization 'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>' requested here llvm/unittests/tools/llvm-profdata/MD5CollisionTest.cpp:160:3: note: in instantiation of function template specialization 'testing::internal::EqHelper::Compare<unsigned int, int, nullptr>' requested here
-
Vitaly Buka authored
-
Vitaly Buka authored
Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D153599
-
LLVM GN Syncbot authored
-
William Huang authored
[llvm-profdata] Refactoring Sample Profile Reader to increase FDO build speed using MD5 as key to Sample Profile map This is phase 1 of multiple planned improvements on the sample profile loader. The major change is to use MD5 hash code ((instead of the function itself) as the key to look up the function offset table and the profiles, which significantly reduce the time it takes to construct the map. The optimization is based on the fact that many practical sample profiles are using MD5 values for function names to reduce profile size, so we shouldn't need to convert the MD5 to a string and then to a SampleContext and use it as the map's key, because it's extremely slow. Several changes to note: (1) For non-CS SampleContext, if it is already MD5 string, the hash value will be its integral value, instead of hashing the MD5 again. In phase 2 this is going to be optimized further using a union to represent MD5 function (without converting it to string) and regular function names. (2) The SampleProfileMap is a wrapper to *map<uint64_t, FunctionSamples>, while providing interface allowing using SampleContext as key, so that existing code still work. It will check for MD5 collision (unlikely but not too unlikely, since we only takes the lower 64 bits) and handle it to at least guarantee compilation correctness (conflicting old profile is dropped, instead of returning an old profile with inconsistent context). Other code should not try to use MD5 as key to access the map directly, because it will not be able to handle MD5 collision at all. (see exception at (5) ) (3) Any SampleProfileMap::emplace() followed by SampleContext assignment if newly inserted, should be replaced with SampleProfileMap::Create(), which does the same thing. (4) Previously we ensure an invariant that in SampleProfileMap, the key is equal to the Context of the value, for profile map that is eventually being used for output (as in llvm-profdata/llvm-profgen). Since the key became MD5 hash, only the value keeps the context now, in several places where an intermediate SampleProfileMap is created, each new FunctionSample's context is set immediately after insertion, which is necessary to "remember" the context otherwise irretrievable. (5) When reading a profile, we cache the MD5 values of all functions, because they are used at least twice (one to index into FuncOffsetTable, the other into SampleProfileMap, more if there are additional sections), in this case the SampleProfileMap is directly accessed with MD5 value so that we don't recalculate it each time (expensive) Performance impact: When reading a ~1GB extbinary profile (fixed length MD5, not compressed) with 10 million function names and 2.5 million top level functions (non CS functions, each function has varying nesting level from 0 to 20), this patch improves the function offset table loading time by 20%, and improves full profile read by 5%. Reviewed By: davidxl, snehasish Differential Revision: https://reviews.llvm.org/D147740
-
Sami Tolvanen authored
This reverts commit 62fa708c. Reverting to investigate -verify-machineinstrs errors in MIR tests.
-
xortoast authored
WebAssembly doesn't expose inexact exceptions, so frint can be mapped to fnearbyint. Likewise, WebAssembly always rounds ties-to-even, so froundeven can be mapped to fnearbyint. Differential Revision: https://reviews.llvm.org/D153451
-
Amara Emerson authored
e018cbf7 changed the default behaviour for Darwin, and this breaks some existing software. rdar://110350601
-
Joseph Huber authored
This patch adds the utilities for the clocks on the GPU. This is done prior to exporting it via some other interface and is mainly just done so they are availible if we wish to do internal testing. Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D153388
-
Joseph Huber authored
Currently we simply overwrite the output file if we get muliple matches in the fatbinary. This patch introduces the `--archive` option which allows us to combine all of the files into a static archive instead. This is usefuly for creating a device specific static archive library from a fatbinary. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D153568
-
Florian Hahn authored
Additional test coverage for D151799.
-
Matt Arsenault authored
This reverts commit e429fdd0.
-
Joseph Huber authored
Summary: We include things off of `libc/include` so we need to use the current binary dir when setting up the directory.
-
Robert Suderman authored
The existing lowering has lower precision for certain use cases, e.g. tanh. Improved version should demonstrate an overall higher level of precision. Reviewed By: cota, jpienaar Differential Revision: https://reviews.llvm.org/D153592
-
Matt Arsenault authored
IMPLICIT_DEPENDS doesn't actually work with ninja and this does.
-
Matt Arsenault authored
-
Zahira Ammarguellat authored
a '#pragma clang fp eval_method, it can lead to ABI breakage. See https://godbolt.org/z/56zG4Wo91 This patch prevents this. Differential Revision: https://reviews.llvm.org/D153590
-
LLVM GN Syncbot authored
-
Jonas Devlieghere authored
Addresses Jason's post-commit feedback in D153644.
-
Chia-hung Duan authored
Ensure the thread that refills freelist will get the Batch without contending the lock in SizeClassAllocator64. Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D152419
-
Chia-hung Duan authored
Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D152420
-
Joseph Huber authored
The patch in D152592 changed the logic for this. We could never check if we were on the GPU as this was before the variable was defined so I moved it later. Secondly, we cannot use the `LLVM_BINARY_DIR` here, and I do not know if that works in general. The problem is that it will isntall the headers under a normal path outside of the `LLVM_ENABLE_RUNTIMES` build. I don't know if that's correct for the other targets, but for the GPU I need to set it back to the CMAKE_BINARY_DIR so it works. Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D153637
-
Paul Kirth authored
There seems to be a problem on arm buildbots. Reverting until I can investigate. https://lab.llvm.org/buildbot#builders/245/builds/10184 This reverts commit a67208e1 and dependent commit e54a3112.
-
Sami Tolvanen authored
With `-fsanitize=kcfi` (Kernel Control-Flow Integrity), Clang emits "kcfi" operand bundles to indirect call instructions. Similarly to the target-specific lowering added in D119296, implement KCFI operand bundle lowering for RISC-V. This patch disables the generic KCFI pass for RISC-V in Clang, and adds the KCFI machine function pass in `RISCVPassConfig::addPreSched` to emit target-specific `KCFI_CHECK` pseudo instructions before calls that have KCFI operand bundles. The machine function pass also bundles the instructions to ensure we emit the checks immediately before the calls, which is not possible with the generic pass. `KCFI_CHECK` instructions are lowered in `RISCVAsmPrinter` to a contiguous code sequence that traps if the expected hash in the operand bundle doesn't match the hash before the target function address. This patch emits an `ebreak` instruction for error handling to match the Linux kernel's `BUG()` implementation. Just like for X86, we also emit trap locations to a `.kcfi_traps` section to support error handling, as we cannot embed additional information to the trap instruction itself. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D148385
-
LLVM GN Syncbot authored
-
Benjamin Kramer authored
-
Jonas Devlieghere authored
When specifying the C-string format for dumping memory, we treat unprintable characters as signed. Whether a character is signed or not is implementation defined, but all printable characters are signed. Therefore it's fair to assume that unprintable characters are unsigned. Before this patch, "\xcf\xfa\xed\xfe\f" would be printed as "\xffffffcf\xfffffffa\xffffffed\xfffffffe\f". Now we correctly print the original string. rdar://111126134 Differential revision: https://reviews.llvm.org/D153644
-
Artem Belevich authored
This avoids unnecessary vector splitting that was needed for vectorized store instruction. Differential Revision: https://reviews.llvm.org/D152593
-