- Jul 14, 2023
-
-
Sean Fertile authored
Followup to D101178 - peephole optimization that converts a load address instruction and a consuming load/store into just the load/store when its safe to do so. eg: converts the 2 instruction code sequence la 4, i[TD](2) stw 3, 0(4) to stw 3, i[TD](2) Differential Revision: https://reviews.llvm.org/D101470
-
Noah Goldstein authored
We can do this if `Y*C` doesn't overflow. This is trivial if `C` is 0/1. Otherwise we actually generate a `mul` instruction iff the `div` has one use. Alive2 Links: udiv: https://alive2.llvm.org/ce/z/GWPW67 sdiv: https://alive2.llvm.org/ce/z/bUoX9h Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D150091 -
Noah Goldstein authored
Differential Revision: https://reviews.llvm.org/D150090
-
Jon Roelofs authored
-
boxu.zhang authored
I'm using clang to compile CUDA code. And just found that clang doesn't support the per-thread stream option for NV CUDA. I don't know if there is another solution. Reviewed By: tra Differential Revision: https://reviews.llvm.org/D154822
-
Pavel Iliin authored
The patch fixes second argument of Function Multi Versioning resolvers, it is pointer to an extendible struct containing hwcap and hwcap2 not a unsigned long hwcap2. Also fixes FMV features caching in resolver. Differential Revision: https://reviews.llvm.org/D155026
-
Jon Chesterfield authored
Requires D155190 Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D155238
-
Sergei Barannikov authored
ParseStatus is slightly more convenient to use due to implicit conversion from bool, which allows to do something like: ``` return Error(L, "msg"); ``` when with MatchOperandResultTy it had to be: ``` Error(L, "msg"); return MatchOperand_ParseFail; ``` It also has more appropriate name since parse* methods are not only for parsing operands. Reviewed By: olista01 Differential Revision: https://reviews.llvm.org/D154304
-
Sergei Barannikov authored
ParseStatus is slightly more convenient to use due to implicit conversion from bool, which allows to do something like: ``` return Error(L, "msg"); ``` when with MatchOperandResultTy it had to be: ``` Error(L, "msg"); return MatchOperand_ParseFail; ``` It also has more appropriate name since parse* methods are not only for parsing operands. Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D154292
-
Aart Bik authored
Reviewed By: K-Wu Differential Revision: https://reviews.llvm.org/D155244
-
Jon Chesterfield authored
Do the LDS frame calculation once, in the IR pass, instead of repeating the work in the backend. Prior to this patch: The IR lowering pass sets up a per-kernel LDS frame and annotates the variables with absolute_symbol metadata so that the assembler can build lookup tables out of it. There is a fragile association between kernel functions and named structs which is used to recompute the frame layout in the backend, with fatal_errors catching inconsistencies in the second calculation. After this patch: The IR lowering pass additionally sets a frame size attribute on kernels. The backend uses the same absolute_symbol metadata that the assembler uses to place objects within that frame size. Deleted the now dead allocation code from the backend. Left for a later cleanup: - enabling lowering for anonymous functions - removing the elide-module-lds attribute (test churn, it's not used by llc any more) - adjusting the dynamic alignment check to not use symbol names Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D155190
-
Po-yao Chang authored
Remove spaces between operator"" and identifier to suppress -Wdeprecated-literal-operator, and between operator and "" like how they are written in [string.view.literals] and [basic.string.literals]. Differential Revision: https://reviews.llvm.org/D155200
-
Dave Lee authored
These synthetic providers use expression evaluation and fail in some cases. Examples: ``` llvm::PointerIntPair<llvm::PointerUnion<const Type *, const ExtQuals *>, Qualifiers::FastWidth> Value; ``` and ``` typedef llvm::PointerUnion<const ValueDecl *, const Expr *, TypeInfoLValue, DynamicAllocLValue> PtrTy; ``` Original contribution: D117779 rdar://110791233 rdar://112195543 Differential Revision: https://reviews.llvm.org/D155219 -
Jonas Devlieghere authored
Make sure TestCTF only run on Darwin when ctfconvert and llvm-objdump are available.
-
Alexey Bataev authored
-
Jonas Devlieghere authored
Add support for compressed CTF data. The flags in the header can indicate whether the CTF body is compressed with zlib deflate. This patch supports inflating the data before parsing. Differential revision: https://reviews.llvm.org/D155221
-
Jan Svoboda authored
In D114095, `HeaderFileInfo::NumIncludes` was moved into `Preprocessor`. This still makes sense, because we want to track this on the granularity of submodules (D112915, D114173), but the way this information is serialized is not ideal. In `ASTWriter`, the set of included files gets deserialized eagerly, issuing lots of calls to `FileManager::getFile()` for input files the PCM consumer might not be interested in. This patch makes the information part of the header file info table, taking advantage of its lazy deserialization which typically happens when a file is about to be included. Reviewed By: benlangmuir Differential Revision: https://reviews.llvm.org/D155131
-
Fangrui Song authored
To fix undefined errors like to_float. This tool is often not built as LLVM_ENABLE_HTTPLIB defaults to off (and the external dependency cpp-httplib is difficult to set up due to a dependency on brotli) and LLVM_TOOL_LLVM_DEBUGINFOD_BUILD disabling logic in D147185.
-
Hanhan Wang authored
It also unifies the computation of StridedLayoutAttr. If the stride is static known value, we can just use it. Differential Revision: https://reviews.llvm.org/D155017
-
Fangrui Song authored
-
Krzysztof Drewniak authored
During a conversion to MLIR_ENABLE_EXECUTION_ENGINE from checking for the native target, the ROCm conversion passes (--serialize-to-hsaco) were mistakenly flagged for being disabled if the execution ending is not being built. These passes use LLVM to build binaries for AMD GPUs, and so require that backend to be enabled. However, they do not produce native code, nor do they interact with the JIT or any of the execution engine support libraries. When building MLIR into a compiler library that's intended to produce GPU binaries, we want to build only the AMDGPU backend and have the binary serialization passes available. This change makes that possible. It looks like the CUDA path might currently require a native target, it's hard to tell, so this commit leaves that if statement untouched. Reviewed By: fmorac Differential Revision: https://reviews.llvm.org/D155227
-
Yaxun (Sam) Liu authored
When -mcpu=native is specified, try detecting GPU on the system by using amdgpu-arch tool. If it fails to detect GPU, emit an error about GPU not detected. If multiple GPUs are detected, use the first GPU and emit a warning. Reviewed by: Matt Arsenault, Fangrui Song Differential Revision: https://reviews.llvm.org/D154531
-
Elliot Goodrich authored
Move the implementation of the `toString` function from `llvm/Support/Error.h` to the source file, which allows us to move `#include "llvm/ADT/StringExtras.h"` to the source file as well. As `Error.h` is present in a large number of translation units this means we are unnecessarily bringing in the contents of `StringExtras.h` - itself a large file with lots of includes - and slowing down compilation. Also move the `#include "llvm/ADT/SmallVector.h"` directive to the source file as it's no longer needed, but this does not give as much of a benefit. This reduces the total number of preprocessing tokens across the LLVM source files in lib from (roughly) 1,920,413,050 to 1,903,629,230 - a reduction of ~0.87%. This should result in a small improvement in compilation time. Differential Revision: https://reviews.llvm.org/D155178
-
Elliot Goodrich authored
In preparation for removing the #include "llvm/ADT/StringExtras.h" from the header to source file of llvm/Support/Error.h, first add in all the missing includes that were previously included transitively through this header. This is fixing all files missed in b0abd489, 39d8e6e2, a11efd49, 5551657b, and 90bfe2df. Differential Revision: https://reviews.llvm.org/D155178
-
Stanislav Mekhanoshin authored
I need this for future patch in the MC, while TII is not available in the llvm-mc. Besides this is not a first time I want it there. Differential Revision: https://reviews.llvm.org/D155228
-
Craig Topper authored
According to the spec, Zce is an alias for Zca, Zcb, Zcmp, and Zcmt. If F is enabled on RV32 it also includes Zcf. This patch adds the Zce and the implication rule which unfortunately requires custom handling for adding Zcf. I've also made all the Zc* extensions imply Zca. I've also added an error for Zcf without RV32. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D153742
-
Nemanja Ivanovic authored
Improve codegen for vectors modulo additions. Reviewed By: nemanjai Differential Revision: https://reviews.llvm.org/D154447
-
Guray Ozen authored
When targeting NVIDIA GPUs, seeing the generated PTX is important. Currently, we don't have simple way to do it. This work adds dump-ptx to gpu-to-cubin pass. One can use it like `gpu-to-cubin{chip=sm_90 features=+ptx80 dump-ptx}`. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D155166 -
Jeffrey Byrnes authored
This adds the IGLP strategy for single-wave gemms. The SchedGroup pipeline is laid out in multiple phases, with each phase corresponding to a distinct pattern present in gemm kernels. The resilience of the optimization is dependent upon IR (as seen by pre-RA scheduling) continuing to have these patterns (as defined by instruction class and dependencies) in their current relative ordering. The kernels of interest have these specific phases: NT: 1, 2a, 2c NN: 1, 2a, 2b TT: 1, 2b, 2c TN: 1, 2b The general approach taken was to have a long SchedGroup pipeline. In this way the scheduler will have less capability of doing the wrong thing. In order to resolve the challenge of correctly fitting these long pipelines, we leverage the rules infrastructure to help the solver. Differential Revision: https://reviews.llvm.org/D149773 Change-Id: I1a35962a95b4bdf740602b8f110d3297c6fb9d96
-
Slava Zakharin authored
I changed the set of files that are built for experimental CUDA/OMP builds, i.e. the files with enabled device support are built as such and the rest of the files are built just for the host target. With this change we can build Flang runtime library that is fully functional on the host target, so in-tree targets like check-flang become operational. Reviewed By: klausler, PeteSteinfeld Differential Revision: https://reviews.llvm.org/D155029
-
Ivan Kosarev authored
Part of <https://github.com/llvm/llvm-project/issues/62629>. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D155061
-
Ivan Kosarev authored
The patch adds the support for 'noa16' operands in non-A16 variants of the instructions, fixes validation of A16 operands and eliminates the custom conversion to MCInst. Part of <https://github.com/llvm/llvm-project/issues/62629>. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D155057
-
Luke Lau authored
This patch adds pseudos and SDNode patterns for vbrev.v, vrev8.v, vclz.v, vctz.v and vcpop.v. I've only added them for integer element types so far since we're lacking tests for floats. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D155216
-
Craig Topper authored
The instructions produce DLEN bits per cycle. The vsetvli LMUL for these instructions is the output EMUL. The input EMUL is scaled down by the vector factor suffix on the instruction name. So for LMUL=1 there are 2*DLEN bits of result produced over 2 cycles. This makes SiFive7GetCyclesDefault the correct resource cycles. Reviewed By: monkchiang Differential Revision: https://reviews.llvm.org/D155010
-
Ivan Kosarev authored
The added instructions are incorrectly encoded as a16 ones despite the 'noa16' modifiers. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D155059
-
Alex Langford authored
In an attempt to make it easier to catch errors when parsing the debug_abbrev section, we should force users to call `parse` before calling `begin`. In a follow-up change, I will change the return type of `parse` from `void` to `Error`. I also explored using the fallible_iterator pattern instead of forcing users to parse everything up front. I think it would be a useful and interesting pattern to implement, but it would require more extensive changes to both DWARFDebugAbbrev and its users. Because my top priority is improving the safety around parsing debug_abbrev, I'm opting to preserve existing behavior until I or somebody else has time to refactor to be able to implement a fallible_iterator. Differential Revision: https://reviews.llvm.org/D154655
-
Jonas Devlieghere authored
Add support for the Compact C Type Format (CTF) in LLDB. The format describes the layout and sizes of C types. It is most commonly consumed by dtrace. We generate CTF for the XNU kernel and want to be able to use this in LLDB to debug kernels for which we don't have dSYMs (anymore). CTF is a much more limited debug format than DWARF which allows is to be an order of magnitude smaller: a 1GB dSYM can be converted to a handful of megabytes of CTF. For XNU, the goal is not to replace DWARF, but rather to have CTF serve as a "better than nothing" debug info format when DWARF is not available. It's worth noting that the LLVM toolchain does not support emitting CTF. XNU uses ctfconvert to generate CTF from DWARF which is used for testing. Differential revision: https://reviews.llvm.org/D154862
-
Jonas Devlieghere authored
Teach LLDB about the ctf (Compact C Type Format) section. Differential revision: https://reviews.llvm.org/D154668
-
Philip Reames authored
We can share the code for both the unmasked and masked cases, and add a missing consistency assert in the process. This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct. This is the last major piece.
-
Alex Langford authored
This has an `lldb_private` type in its parameter, it should be in `lldb-private-types.h` Differential Revision: https://reviews.llvm.org/D155129
-