- Aug 02, 2022
-
-
jacquesguan authored
This patch adds constant folder for TanOp which only supports single and double precision floating-point. Differential Revision: https://reviews.llvm.org/D130873
-
Chuanqi Xu authored
This reverts commit db6152ad. This commit fails in ppc64. Since we want to backport it to 15.x. So revert it now to keep the patch complete.
-
Fangrui Song authored
-
Fangrui Song authored
Also update -ftime-trace='s help to fix a recommonmark error.
-
Chuanqi Xu authored
Previously when we add module initializer, we forget to handle header units. This results that we couldn't compile a Hello World Example with Header Units. This patch tries to fix this. Reviewed By: iains Differential Revision: https://reviews.llvm.org/D130871
-
Jeff Niu authored
Previously, DenseArrayAttr used VectorType for its shaped type. VectorType is problematic for arrays because it doesn't support zero dimensions, meaning that an empty array would have `vector<i32>` as its type. ElementsAttr would think that an empty dense array is size 1, not 0. This patch switches over to TensorType, which does support zero dimensions. Fixes #56860 Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D130921
-
Fangrui Song authored
-
Siva Chandra Reddy authored
Reviewed By: michaelrj, lntue Differential Revision: https://reviews.llvm.org/D130872
-
Fangrui Song authored
-
Alex Brachet authored
-
Sotiris Apostolakis authored
Reviewed By: davidxl Differential Revision: https://reviews.llvm.org/D129817
-
Craig Topper authored
-
Sunho Kim authored
Fixes compilation error. Differential Revision: https://reviews.llvm.org/D130898
-
Manish Gupta authored
Adds optional attribute to support tensor cores on F32 datatype by lowering to `mma.sync` with TF32 operands. Since, TF32 is not a native datatype in LLVM we are adding `tf32Enabled` as an attribute to allow the IR to be aware of `MmaSyncOp` datatype. Additionally, this patch adds placeholders for nvgpu-to-nvgpu transformation targeting higher precision tf32x3. For mma.sync on f32 input using tensor cores there are two possibilites: (a) tf32 (1 `mma.sync` per warp-level matrix-multiply-accumulate) (b) tf32x3 (3 `mma.sync` per warp-level matrix-multiply-accumulate) Typically, tf32 tensor core acceleration comes at a cost of accuracy from missing precision bits. While f32 has 23 precision bits, tf32 has only 10 precision bits. tf32x3 aims to recover the precision bits by splitting each operand into two tf32 values and issue three `mma.sync` tensor core operations. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D130294
-
Martin Sebor authored
Reflect in the pointer's offset the length of the leading part of the consumed string preceding the first converted digit. Reviewed By: efriedma Differential Revision: https://reviews.llvm.org/D130912
-
Ben Langmuir authored
As progress towards having FileEntryRef contain the requested name of the file, this commit narrows the "remap" hack to only apply to paths that were remapped to an external contents path by a VFS. That was always the original intent of this code, and the fact it was making relative paths absolute was an unintended side effect. Differential Revision: https://reviews.llvm.org/D130935
-
Craig Topper authored
It's possible we have: lui a0, %hi(sym) addi a0, %lo(sym) addi a0, <offset1> lw a0, <offset2>(a0) We want to arrive at lui a0, %hi(sym+offset1+offset2) lw a0, %lo(sym+offset1+offset2) We currently fail to do this because we only consider loads/stores if we didn't find any arithmetic. This patch splits arithmetic folding and load/store folding into two separate phases. The load/store folding can no longer assume the offset in hi/lo is 0 so we must combine the offsets. I've applied the same simm32 limit that we applied in the arithmetic folding. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D130931
-
Ilia Diachkov authored
The patch replaces SPIRVBaseInfo.* previously created using macros by the tablegen approach. There are many small changes in other files due to differences in namespaces. Also, functions in SPIRVUtils are moved to the llvm namespace. Differential Revision: https://reviews.llvm.org/D130518 Co-authored-by:
Aleksandr Bezzubikov <zuban32s@gmail.com> Co-authored-by:
Michal Paszkowski <michal.paszkowski@outlook.com> Co-authored-by:
Andrey Tretyakov <andrey1.tretyakov@intel.com> Co-authored-by:
Konrad Trifunovic <konrad.trifunovic@intel.com>
-
Alex Brachet authored
-
Alex Brachet authored
This test had to be disabled because ps4 targets don't support -fuse-ld. Preferably, this should just be unsupported for ps4 targets. However no such lit feature exists so I have just gone ahead and set the target explicitly. Moreover, this needs to create a terminal link step, either an executable or shared object to get the link error. With the change to the explicit target I've had to also add -nostartfiles -nostdlib so that clang doesn't pull crt files into the link which may not be present. Again, this would likely be solved if this test was unsupported for the one platform that disables -fuse-ld
-
River Riddle authored
This attribute is technical debt from the early stages of MLIR, before ElementsAttr was an interface and when it was more difficult for dialects to define their own types of attributes. At present it isn't used at all in tree (aside from being convenient for eliding other ElementsAttr), and has had little to no evolution in the past three years. Differential Revision: https://reviews.llvm.org/D129917
-
David Blaikie authored
Fixes #56724
-
Ben Langmuir authored
It's an accident that we started return asbolute paths from FileEntry::getName for all relative paths. Prepare for getName to get (closer to) return the requested path. Note: conceptually it might make sense for the dependency scanner to allow relative paths and have the DependencyConsumer decide if it wants to make them absolute, but we currently document that it's absolute and I didn't want to change behaviour here. Differential Revision: https://reviews.llvm.org/D130934
-
Tue Ly authored
-
Slava Gurevich authored
Fix incorrect null-check logic, likely cause by copy-paste Differential Revision: https://reviews.llvm.org/D130937
-
Alexander Yermolovich authored
We were not handling correclty multiple DW_OP_addrx in the location expression. This was exposed by clang-15 build in release mode with debug information. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D130812
-
Mehdi Amini authored
We noticed this failing depending on the platform, checking the last few digit isn't necessary for this test anyway.
-
Alex Brachet authored
-
Alex Brachet authored
-
Kirill Okhotnikov authored
-
Jakob Johnson authored
The use of `std::unique_ptr` with `TraceCursor` adds unnecessary complexity to adding `SBTraceCursor` bindings Specifically, since `TraceCursor` is an abstract class there's no clean way to provide "deep clone" semantics for `TraceCursorUP` short of creating a pure virtual `clone()` method (afaict). After discussing with @wallace, we decided there is no strong reason to favor wrapping `TraceCursor` with `std::unique_ptr` over `std::shared_ptr`, thus this diff replaces all usages of `std::unique_ptr<TraceCursor>` with `std::shared_ptr<TraceCursor>`. This sets the stage for future diffs to introduce `SBTraceCursor` bindings in a more clean fashion. Test Plan: Differential Revision: https://reviews.llvm.org/D130925
-
Craig Topper authored
addMachineSSAOptimization is skipped for -O0, but this pass is required for -O0.
-
Kirill Okhotnikov authored
Correct rounding function. Performance ~2x faster than glibc analog. Performance (llvm 12 intel): ``` CORE_MATH_PERF_MODE=rdtsc PERF_ARGS='' ./perf.sh tanhf GNU libc version: 2.31 GNU libc release: stable 13.279 37.492 18.145 CORE_MATH_PERF_MODE=rdtsc PERF_ARGS='--latency' ./perf.sh tanhf GNU libc version: 2.31 GNU libc release: stable 40.658 109.582 66.568 ``` Differential Revision: https://reviews.llvm.org/D130780
-
Alex Brachet authored
-fuse-ld is not available for ps4 targets
-
Joseph Huber authored
Summary: This file is no longer used, get rid of it.
-
Markus Böck authored
GCC and that specific build bot issued warnings turned errors, due to a narrowing conversion from `unsigned` to `int32_t`. Silence these via a static_cast.
-
Vasileios Porpodas authored
2xi64 is the legalized type for wide reductions (like 16xi64) and setting the cost to 2 makes `load-reduce` and `load-zext-reduce` patterns profitable. The few performance measurments that I did on an aarch64 machine confirm that these patterns are actually faster when vectorized. Differential Revision: https://reviews.llvm.org/D130740
-
Alex Brachet authored
This was discussed on https://discourse.llvm.org/t/rfc-generating-lld-reproducers-on-crashes/58071/12 When lld crashes, or errors when -gen-reproducer=error and -fcrash-diagnostics=all clang will re-run lld with --reproduce=$temp_file for easily reproducing the crash/error. Differential Revision: https://reviews.llvm.org/D120175
-
Joseph Huber authored
The runtime makes some use of `std::vector` data structures. We should be able to replace these trivially with `llvm::SmallVector` instead. This should allow us to avoid heap allocations in the majority of cases now. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D130927
-
Craig Topper authored
-