- Aug 24, 2022
-
-
Michal Terepeta authored
Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D115742
-
Joseph Huber authored
This patch replaces uses of `dlopen` and `dlsym` with LLVM's support with `loadPermanentLibrary` and `getSymbolAddress`. This allows us to remove the explicit dependency on the `dl` libraries in the CMake. This removes another explicit dependency and solves an issue encountered while building on Windows platforms. The one downside to this is that the LLVM library does not currently support `dlclose` functionality, but this could be added in the future. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D131507
-
Joseph Huber authored
We use the offloading entires array to determine the relative names and addressed of device-side kernel functions. The x86_64 plugin previously derived the device-side entry table by first identifying the `omp_offloading_entries` section offset in the loaded elf. Then we would use the base offset of the loaded dyanmic library to identify the entries array within the loaded image. This relied on some more unconventional methods which prevented us from using the LLVM dynamic library loader for this plugin. This patch simplifies this by instead copying the host-side entry and replacing its address with the device-side address looked up through `dlsym`. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D131516
-
Kito Cheng authored
The only use of TM is checking result of TargetMachine::getFunctionSections, check that directly instead of introdce a local variable.
-
Sanjay Patel authored
Added/removed braces, reduced indents, and renamed a variable.
-
Sanjay Patel authored
-
Mircea Trofin authored
New logic works for both `tensorflow` and `tf-nightly`.
-
Simon Pilgrim authored
Revert rGc360955c "[InstCombine] Canonicalize ((X & -X) - 1) --> (~X & (X - 1)) (PR51784)" The test changes are failing on some buildbots (but not others.....).
-
Phoebe Wang authored
Fixes #57340 Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D132563
-
Louis Dionne authored
This will make the following patches to migrate projects off of the LLVM_ENABLE_PROJECTS build onto the LLVM_ENABLE_RUNTIMES build much easier to comprehend. This patch should be a NFC since it keeps the same set of runtimes being built by default. Differential Revision: https://reviews.llvm.org/D132478
-
Muhammad Omair Javaid authored
This test fails on buildbot while passes on standalone builds. I am marking it as skipped until actual problem is found and resolved.
-
Valentin Clement authored
This patch creates a temporary of the appropriate length while lowering SetLength. The corresponding character can be truncated or padded if necessary. This fix issue with array constructor in argument and also with statement function. ``` character(7) :: str = "1234567" call s(str(1:1)) contains subroutine s(a) character(*) :: a call s2([Character(3)::a]) end subroutine subroutine s2(c) character(3) :: c(1) print "(4a)", c(1), "end" end subroutine end ``` The example prior the patch prints `123end` instead of `1. end` Reviewed By: PeteSteinfeld, jeanPerier Differential Revision: https://reviews.llvm.org/D132464
-
Zain Jaffal authored
Following the work on `D131672` we do the same optimisations for integer products. We add tests to check if a loop gets removed if we repeatdly multiply an array elements with an accumulator initalised to zero Reviewed By: fhahn Differential Revision: https://reviews.llvm.org/D132553
-
Simon Pilgrim authored
As confirmed on D132520 - this should always return true
-
Simon Tatham authored
This test contained some data tables that llvm-objdump was disassembling as code, so the test was recovering the 32-bit values in the table from the instruction encoding column of the disassembly. D131589 changed how llvm-objdump decides what to disassemble as code or as data. As a result, these data tables are now being disassembled as data, which I think is actually more sensible -- but the test wasn't expecting it, and got confused.
-
Pierre van Houtryve authored
Allows things like `(G_PTR_ADD (G_PTR_ADD a, b), c)` to be simplified into a single ADD3 instruction instead of two adds. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D131254
-
Jakub Kuderski authored
The intention is to have this op lowered to `llvm.intr.uadd.with.overflow` or `spv.IAddCarry`. LLVM has a second intrinsic for signed add-with-overflow, `llvm.intr.sadd.with.overflow`, with different semantics. Therefore we should have 2 ops with `arith`, and be explicit about signed/unsigned semantics. Rename `arith.addi_carry` to `arith.addui_carry` before we introduce a signed version of this op: `arith.addsi_carry`. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D132491
-
Michele Scuttari authored
-
Louis Dionne authored
Otherwise, we would end up passing `-lNOTFOUND` to the compiler, which caused various compiler checks to fail and ended up breaking the build in the most obscure ways. For example, checks for -faligned-allocation would fail because the compiler would complain about an unknown library called NOTFOUND, and we would end up not passing -faligned-allocation anywhere in our build. This is madness. An even better alternative would be to simply FATAL_ERROR if we don't find the builtins library. However, it seems like our build has been working fine without finding it for a while, so instead of making a bunch of builds fail, we can figure out why linking against compiler-rt doesn't actually seem to be required in a follow-up, and perhaps relax that.
-
Simon Pilgrim authored
Enables the ctpop((x & -x ) - 1) -> cttz(x, false) fold Alive2: https://alive2.llvm.org/ce/z/EDk4h7 (((X & -X) - 1) --> (~X & (X - 1)) ) Alive2: https://alive2.llvm.org/ce/z/8Yr3XG (CTPOP -> CTTZ) Fixes #51126 Differential Revision: https://reviews.llvm.org/D110488
-
Stephen Tozer authored
Reverting due to reported errors when running Linux kernel builds with KMSAN -gdwarf-4. This reverts commit 2cb9e1ac.
-
Alex Richardson authored
While working on https://reviews.llvm.org/D131429, I got a test diff in one of the VE tests and running update_llc_test_checks.py deleted all the code for that function. This updates the regex to handle this new output. Reviewed By: kaz7 Differential Revision: https://reviews.llvm.org/D131431
-
Alex Richardson authored
While working on https://reviews.llvm.org/D131429, I got a test diff in one of the VE tests and running update_llc_test_checks.py deleted all the code for that function. This is a baseline test for this bug (incorrect regex for VE when .Lfoo$local symbols are used). Reviewed By: kaz7 Differential Revision: https://reviews.llvm.org/D131434
-
Alex Richardson authored
This commit moves the information on whether a register is constant into the Tablegen files to allow generating the implementaiton of isConstantPhysReg(). I've marked isConstantPhysReg() as final in this generated file to ensure that changes are made to tablegen instead of overriding this function, but if that turns out to be too restrictive, we can remove the qualifier. This should be pretty much NFC, but I did notice that e.g. the AMDGPU generated file also includes the LO16/HI16 registers now. The new isConstant flag will also be used by D131958 to ensure that constant registers are marked as call-preserved. Differential Revision: https://reviews.llvm.org/D131962
-
John Ericson authored
A simple sed doing these substitutions: - `${LLVM_BINARY_DIR}/(\$\{CMAKE_CFG_INTDIR}/)?lib(${LLVM_LIBDIR_SUFFIX})?\>` -> `${LLVM_LIBRARY_DIR}` - `${LLVM_BINARY_DIR}/(\$\{CMAKE_CFG_INTDIR}/)?bin\>` -> `${LLVM_TOOLS_BINARY_DIR}` where `\>` means "word boundary". The only manual modifications were reverting changes in - `compiler-rt/cmake/Modules/CompilerRTUtils.cmake - `runtimes/CMakeLists.txt` because these were "entry points" where we wanted to tread carefully not not introduce a "loop" which would end with an undefined variable being expanded to nothing. This hopefully increases readability overall, and also decreases the usages of `LLVM_LIBDIR_SUFFIX`, preparing us for D130586. Reviewed By: sebastian-ne Differential Revision: https://reviews.llvm.org/D132316 -
Simon Tatham authored
The main disassembly loop in llvm-objdump works by iterating through the symbols in a code section, and for each one, dumping the range of the section from that symbol to the next. If there's another symbol defined at the same location, then that range will have length 0, and llvm-objdump will skip over the symbol entirely. As a result, llvm-objdump will only show the last of the symbols defined at that address. Not only that, but the other symbols won't even be checked against the `--disassemble-symbol` list. So if you have two symbols `foo` and `bar` defined in the same place, then one of `--disassemble-symbol=foo` and `--disassemble-symbol=bar` will generate an error message and no disassembly. I think a better approach in that situation is to prioritise display of the symbol the user actually asked for. Also, if the user specifically asks for disassembly of //both// of two symbols defined at the same address, the best response I can think of is to disassemble the code once, preceded by both symbol names. This involves teaching llvm-objdump to be able to display more than one symbol name at the head of a disassembled section, which also makes it possible to implement a `--show-all-symbols` option to display //every// symbol defined in the code, not just the most preferred one at each address. This change also turns out to fix a bug in which `--disassemble-all` on a mixed Arm/Thumb ELF file would fail to switch disassembly states between Arm and Thumb functions, because the mapping symbols were accidentally ignored. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D131589
-
Tarun Prabhu authored
Lower F08 parity intrinsic. This largely follows the implementation of the ANY and ALL intrinsics which are related. Differential Revision: https://reviews.llvm.org/D129788
-
Joseph Huber authored
The new driver supports device-only compilation for the offloading device. The way this is handlded is a little different from the old offloading driver. The old driver would put all the outputs in the final action list akin to a linker job. The new driver however generated these in the middle of the host's job so we instead put them all in a single offloading action. However, we only handled these kinds of offloading actions correctly when there was only a single input. When we had multiple inputs we would instead attempt to get the host job, which didn't exist, and crash. This patch simply adds some extra logic to generate the jobs for all dependencies if there is not host action. Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D132248
-
Kito Cheng authored
This issue is found by build llvm-testsuite with `-Oz`, linker will complain `dangerous relocation: %pcrel_lo missing matching %pcrel_hi` and that turn out cause by we outlined pcrel-lo, but leave pcrel-hi there, that's not problem in general, but the problem is they put into different section, they pcrel-hi and pcrel-lo pair (e.g. AUIPC+ADDI) *MUST* put be present in same section due to the implementation. Outlined function will put into .text name, but the source functions will put in .text.<function-name> if function-section is enabled or the function has `comdat` attribute. There are few solutions for this issue: 1. Always disallow instructions with pcrel-lo flags. 2. Only disallow instructions with pcrel-lo flags that when function-section is enabled or this function has `comdat` attribute. 3. Check the corresponding instruction with pcrel-high also included in the outlining candidate sequence or not, and allow that only when pcrel-high is included in the outlining candidate. First one is most conservative, that might lose some optimization opportunities, and second one could save those opportunities, and last one is hard to implement, and don't have any benefits since pcrel-high are using different label even accessing same symbol. Use custom section name might also cause this problem, but that already filtered by RISCVInstrInfo::isFunctionSafeToOutlineFrom. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D132528
-
Simon Pilgrim authored
Noticed by @spatel in D110488
-
Kito Cheng authored
Differential Revision: https://reviews.llvm.org/D132527
-
Louis Dionne authored
For the time being, we are still building libc++ and libc++abi's headers during stage 2 builds. Encode that in the cache file so that CI jobs don't have to manually specify LLVM_ENABLE_RUNTIMES when doing a stage 2 build.
-
Mats Petersson authored
Add simplifcation pass for MAXVAL intrinsic function This refactors some of the code to allow variation on the initialization value and operation performed within the loop, reusing the majority of code for both SUM and MAXVAL. Adding tests for the test-cases that produce different output than the SUM function. Reviewed By: vzakhari Differential Revision: https://reviews.llvm.org/D132234
-
Louis Dionne authored
-
Kiran Chandramohan authored
-
Simon Pilgrim authored
As originally suggested on D110488
-
David Green authored
The existing cost model for fixed-order recurrences models the phi as an extract shuffle of a v1 vector. The shuffle produced should be a splice, as they take two vectors inputs are extracting from a subset of the lanes. On certain architectures the existing cost model can drastically under-estimate the correct cost for the shuffle, so this changes it to a SK_Splice and passes a correct Mask through to the getShuffleCost call. I believe this might be the first use of a SK_Splice shuffle cost model outside of scalable vectors, and some targets may require additions to the cost-model to correctly account for them. In tree targets appear to all have been updated where needed. Differential Revision: https://reviews.llvm.org/D132308
-
Felipe de Azevedo Piovezan authored
The -fdebug-types-section flag is not supported on Apple platforms. Reviewed By: Michael137 Differential Revision: https://reviews.llvm.org/D132410
-
Keith Randall authored
Reviewed By: dvyukov Differential Revision: https://reviews.llvm.org/D131927
-
Markus Böck authored
A bitwise and with the bitwise negate of itself is always 0, regardless of the integer type. This patch adds detection of such a pattern in `arith.andi`s `fold` method. Differential Revision: https://reviews.llvm.org/D131860
-