- Jul 14, 2023
-
-
Craig Topper authored
According to the spec, Zce is an alias for Zca, Zcb, Zcmp, and Zcmt. If F is enabled on RV32 it also includes Zcf. This patch adds the Zce and the implication rule which unfortunately requires custom handling for adding Zcf. I've also made all the Zc* extensions imply Zca. I've also added an error for Zcf without RV32. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D153742
-
Nemanja Ivanovic authored
Improve codegen for vectors modulo additions. Reviewed By: nemanjai Differential Revision: https://reviews.llvm.org/D154447
-
Guray Ozen authored
When targeting NVIDIA GPUs, seeing the generated PTX is important. Currently, we don't have simple way to do it. This work adds dump-ptx to gpu-to-cubin pass. One can use it like `gpu-to-cubin{chip=sm_90 features=+ptx80 dump-ptx}`. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D155166 -
Jeffrey Byrnes authored
This adds the IGLP strategy for single-wave gemms. The SchedGroup pipeline is laid out in multiple phases, with each phase corresponding to a distinct pattern present in gemm kernels. The resilience of the optimization is dependent upon IR (as seen by pre-RA scheduling) continuing to have these patterns (as defined by instruction class and dependencies) in their current relative ordering. The kernels of interest have these specific phases: NT: 1, 2a, 2c NN: 1, 2a, 2b TT: 1, 2b, 2c TN: 1, 2b The general approach taken was to have a long SchedGroup pipeline. In this way the scheduler will have less capability of doing the wrong thing. In order to resolve the challenge of correctly fitting these long pipelines, we leverage the rules infrastructure to help the solver. Differential Revision: https://reviews.llvm.org/D149773 Change-Id: I1a35962a95b4bdf740602b8f110d3297c6fb9d96
-
Slava Zakharin authored
I changed the set of files that are built for experimental CUDA/OMP builds, i.e. the files with enabled device support are built as such and the rest of the files are built just for the host target. With this change we can build Flang runtime library that is fully functional on the host target, so in-tree targets like check-flang become operational. Reviewed By: klausler, PeteSteinfeld Differential Revision: https://reviews.llvm.org/D155029
-
Ivan Kosarev authored
Part of <https://github.com/llvm/llvm-project/issues/62629>. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D155061
-
Ivan Kosarev authored
The patch adds the support for 'noa16' operands in non-A16 variants of the instructions, fixes validation of A16 operands and eliminates the custom conversion to MCInst. Part of <https://github.com/llvm/llvm-project/issues/62629>. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D155057
-
Luke Lau authored
This patch adds pseudos and SDNode patterns for vbrev.v, vrev8.v, vclz.v, vctz.v and vcpop.v. I've only added them for integer element types so far since we're lacking tests for floats. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D155216
-
Craig Topper authored
The instructions produce DLEN bits per cycle. The vsetvli LMUL for these instructions is the output EMUL. The input EMUL is scaled down by the vector factor suffix on the instruction name. So for LMUL=1 there are 2*DLEN bits of result produced over 2 cycles. This makes SiFive7GetCyclesDefault the correct resource cycles. Reviewed By: monkchiang Differential Revision: https://reviews.llvm.org/D155010
-
Ivan Kosarev authored
The added instructions are incorrectly encoded as a16 ones despite the 'noa16' modifiers. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D155059
-
Alex Langford authored
In an attempt to make it easier to catch errors when parsing the debug_abbrev section, we should force users to call `parse` before calling `begin`. In a follow-up change, I will change the return type of `parse` from `void` to `Error`. I also explored using the fallible_iterator pattern instead of forcing users to parse everything up front. I think it would be a useful and interesting pattern to implement, but it would require more extensive changes to both DWARFDebugAbbrev and its users. Because my top priority is improving the safety around parsing debug_abbrev, I'm opting to preserve existing behavior until I or somebody else has time to refactor to be able to implement a fallible_iterator. Differential Revision: https://reviews.llvm.org/D154655
-
Jonas Devlieghere authored
Add support for the Compact C Type Format (CTF) in LLDB. The format describes the layout and sizes of C types. It is most commonly consumed by dtrace. We generate CTF for the XNU kernel and want to be able to use this in LLDB to debug kernels for which we don't have dSYMs (anymore). CTF is a much more limited debug format than DWARF which allows is to be an order of magnitude smaller: a 1GB dSYM can be converted to a handful of megabytes of CTF. For XNU, the goal is not to replace DWARF, but rather to have CTF serve as a "better than nothing" debug info format when DWARF is not available. It's worth noting that the LLVM toolchain does not support emitting CTF. XNU uses ctfconvert to generate CTF from DWARF which is used for testing. Differential revision: https://reviews.llvm.org/D154862
-
Jonas Devlieghere authored
Teach LLDB about the ctf (Compact C Type Format) section. Differential revision: https://reviews.llvm.org/D154668
-
Philip Reames authored
We can share the code for both the unmasked and masked cases, and add a missing consistency assert in the process. This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct. This is the last major piece.
-
Alex Langford authored
This has an `lldb_private` type in its parameter, it should be in `lldb-private-types.h` Differential Revision: https://reviews.llvm.org/D155129
-
Valentin Clement authored
Add support to lower reduction with the multiply operator and complex type. Depends on D155007 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D155014
-
Lang Hames authored
Thanks to Simon Pilgrim for letting me know about these in https://reviews.llvm.org/rG9d701c8a8d65.
-
Maksim Panchenko authored
Propagate Linux Kernel ORC information read from the file to the whole function CFG once the graph has been built. We have a choice to either attach ORC state annotation to every instruction, or to the first instruction in the basic block to conserve processing memory. I chose to attach to every instruction under --print-orc option which is currently on by default. Depends on D155153, D154815 Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D155156
-
Maksim Panchenko authored
Add MetadataRewriter::postCFGInitializer(). Reviewed By: jobnoorman Differential Revision: https://reviews.llvm.org/D155153
-
Philip Reames authored
This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct. The code structure here is overly verbose. I'm landing this staging change with the code structure exactly matching the non-masked case to make the following cleanup that commons this all obviously correct.
-
Maksim Panchenko authored
Read ORC (oops rewind capability) info used for unwinding the stack by Linux Kernel. The info is stored in .orc_unwind and .orc_unwind_ip sections. There is also a related .orc_lookup section that is being populated by the kernel during runtime. Contents of the sections are sorted for quicker lookup by a post-link objtool. Unless we modify stack access instructions, we don't have to change ORC info attributed to instructions in the binary. However, we need to update instruction addresses and sort both sections based on the new layout. For pretty printing, we add "--print-orc" option that prints ORC info next to instructions in code dumps. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D154815
-
Matthew Voss authored
This test has been failing on sanitizer-x86_64-linux-bootstrap-asan since it was commited. Removing this test while I work on reproducing this. Example: https://lab.llvm.org/buildbot/#/builders/168/builds/14579
-
Louis Dionne authored
This patch moves a few tests that were still using std::fprintf to using TEST_REQUIRE instead, which provides a single point to tweak for platforms that don't implement fprintf. As a fly-by fix, it also avoids including `time_utils.h` in filesystem_clock.cpp when it is not required, since that header makes some pretty large assumptions about the platform it is on. Differential Revision: https://reviews.llvm.org/D155019
-
Alexander Yermolovich authored
There are cases in DWARF4 when Skeleton CU has ranges, but dwo CU doesn't. Bug was introduced in new DWARFRewriter where for DWARF4 it would fall through to DWARF5 case. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D155033
-
Alex Langford authored
Differential Revision: https://reviews.llvm.org/D155137
-
Alexander Yermolovich authored
The DWO Unit DIE, doesn't have low_pc/high_pc, so we were printing this error for valid cases. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D155032
-
Amara Emerson authored
GCC and existing codebases allow the use of integral values to be used with this constraint. A recent change D133914 in this area started causing asserts. Removing the assert is enough as the rest of the code works fine. rdar://109675485 Differential Revision: https://reviews.llvm.org/D155023
-
Craig Topper authored
The greediness of the operand matching regular expressions made the test pass even though an operand is missing.
-
Aart Bik authored
Also makes some minor consistency edits in the cuSparseLt wrapper lib. Reviewed By: Peiming, K-Wu Differential Revision: https://reviews.llvm.org/D155139
-
Alexander Yermolovich authored
Setting initial offset of DIE to input DIE. This is to make "printf" debugging easier. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D155031
-
Fangrui Song authored
ENABLE_X86_RELAX_RELOCATIONS has defaulted to on (c41a18cf) for nearly 3 years. As a clean-up, remove overrides from some early adopters. Change OHOS to use true as agreed by the patch author D145227.
-
Philip Reames authored
This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct. This particular change involves a slightly ugly bit of code to match the glue to the mask. I'm staging it this way as I ran into a bit of weirdness when commoning mask operands, and wanted to isolate the complexity.
-
Nick Desaulniers authored
To fix expensive check builds that were failing when using MSVC's std::string_view::iterator::operator*, I added a few expressions like &*std::string_view::begin. @nico pointed out that this is literally the same thing and more clearly expressed as std::string_view::data. Link: https://github.com/llvm/llvm-project/issues/63740 Reviewed By: #libc_abi, ldionne, philnik, MaskRay Differential Revision: https://reviews.llvm.org/D154876
-
Andrew Gozillon authored
This is an attempt at mimicing the method in which threadprivate handles the following type of variables: program main integer :: i !$omp declare target to(i) end Which essentially generates a GlobalOp for the variable (which would normally only be an alloca) when it's instantiated. The main difference is there is no operation generated within the function, instead the declare target attribute is appended later within handleDeclareTarget. Reviewers: kiranchandramohan Differential Revision: https://reviews.llvm.org/D152037
-
Philip Reames authored
Very minor change, just making sure each step is obvious and easy to follow. This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct.
-
Jan Sjodin authored
The early outlining pass was erasing target functions that need to be kept. It should only erase functions that contain target ops.
-
Philip Reames authored
We have the SEW operand access repeating in all paths, common it up to make the code easier to read. This is a subset of Luke's D155063. I'm splitting pieces and landing them in the process of convincing myself all the individual transforms are in fact correct.
-
Slava Zakharin authored
The problem appeared as a segfault for case like this: ``` type t character(11), allocatable :: c end type character(12), alloctable :: x type(t) y y = t(x) ``` The frontend representes `y = t(x)` as `y=t(c=%SET_LENGTH(x,11_8))`. When 'x' is unallocated the hlfir.set_length lowering results in segfault. It could probably be handled in hlfir.set_length lowering by using NULL base for the hlfir.declare depending on the allocation status of 'x', but I am not sure if !hlfir.expr, in general, is supposed to represent an expression created from unallocated allocatable. I believe in Fortran that would mean referencing an unallocated allocatable, which is not allowed. I decided to special case `SET_LENGTH` in structure constructor, so that we use its 'x' operand as the RHS for the assign operation implying the isAllocatable check for cases when 'x' is allocatable. This requires setting keep_lhs_length_if_realloc flag for the assign operation. Note that when the component being intialized has deferred length the frontend does not produce `SET_LENGTH`. Differential Revision: https://reviews.llvm.org/D155151
-
Christian Walther authored
This seems to match https://gcc.gnu.org/install/specific.html#powerpc-x-eabi It seems that anything with OS `none` (although that doesn’t seem to be distinguished from `unknown`) or with environment `eabi` should be treated as bare-metal. Since this seems to have been handled on a case-by-case basis in the past ([arm](https://reviews.llvm.org/D33259), [riscv](https://reviews.llvm.org/D91442), [aarch64](https://reviews.llvm.org/D111134)), what I am proposing here is to add another case to the list to also handle `powerpc[64][le]-unknown-unknown-eabi` using the `BareMetal` toolchain, following the example of the existing cases. (We don’t care about powerpc64 and powerpc[64]le, but it seemed appropriate to lump them in.) At Indel, we have been building bare-metal embedded applications that run on custom PowerPC and ARM systems with Clang and LLD for a couple of years now, using target triples `powerpc-indel-eabi`, `powerpc-indel-eabi750`, `arm-indel-eabi`, `aarch64-indel-eabi` (which I just learned from D153430 is wrong and should be `aarch64-indel-elf` instead, but that’s a different matter). This has worked fine for ARM, but for PowerPC we have been unable to call the linker (LLD) through the Clang driver, because it would insist on calling GCC as the linker, even when told `-fuse-ld=lld`. That does not work for us, there is no GCC around. Instead we had to call `ld.lld` directly, introducing some special cases in our build system to translate between linker-via-driver and linker-called-directly command line arguments. I have now dug into why that is, and found that the difference between ARM and PowerPC is that `arm-indel-eabi` hits a special case that causes the Clang driver to instantiate a `BareMetal` toolchain that is able to call LLD and works the way we need, whereas `powerpc-indel-eabi` lands in the default case of a `Generic_ELF` (subclass of `Generic_GCC`) toolchain which expects GCC. Reviewed By: MaskRay, michaelplatings, #powerpc, nemanjai Differential Revision: https://reviews.llvm.org/D154357
-
- Jul 13, 2023
-
-
Simon Pilgrim authored
Removing the x86-specific node helps further folding and improves commutativity
-