- Mar 24, 2023
-
-
Nikita Popov authored
-
Dmitry Makogon authored
-
Dmitry Makogon authored
Currently it fails as LSR doesn't preserve LCSSA in some cases.
-
Nicolas Vasilache authored
This revision adds vector transform operations that allow us to better inspect the composition of various lowerings that were previously very opaque. This commit is NFC in that it does not change patterns beyond adding `rewriter.notifyFailure` messages and it does not change the tests beyond breaking them into pieces and using transforms instead of throwaway opaque test passes. Reviewed By: ftynse, springerm Co-authored-by:
Alex Zinenko <zinenko@google.com> Differential Revision: https://reviews.llvm.org/D146755
-
Nikita Popov authored
Don't merge invokes if this replaces constant operands with phis in a place where this is not legal. This also disallows converting operand bundles from constant to non-constant, in line with the restriction we use in other transforms. Fixes https://github.com/llvm/llvm-project/issues/61265. Differential Revision: https://reviews.llvm.org/D146723
-
Rainer Orth authored
The `Flang :: Driver/lto-flags.f90` test `FAIL`s on Solaris: /vol/llvm/src/llvm-project/dist/flang/test/Driver/lto-flags.f90:30:13: error: THIN-LTO: expected string not found in input ! THIN-LTO: "-plugin-opt=thinlto" ^ This is no wonder since the native Solaris `ld` doesn't support the linker plugin interface at all, so this patch marks the test as `UNSUPPORTED`. Tested on `amd64-pc-solaris2.11` and `x86_64-pc-linux-gnu`. Differential Revision: https://reviews.llvm.org/D146807 -
Nikita Popov authored
Rather than cleanup up dead constant expressions as we go along, do this once at the end. This aligns it with the CleanupConstantGlobalUsers() implementation and avoids any invalidation issues. Fixes https://github.com/llvm/llvm-project/issues/61674.
-
khei4 authored
Differential Revision: https://reviews.llvm.org/D146798 Reviewed By: nikic
-
Mariusz Sikora authored
Introducing Subtarget Features for instructions: - ds_pk_add_bf16 - ds_pk_add_f16 - ds_pk_add_rtn_bf16 - ds_pk_add_rtn_f16 - flat_atomic_pk_add_f16 - flat_atomic_pk_add_bf16 - global_atomic_pk_add_f16 - global_atomic_pk_add_bf16 - buffer_atomic_pk_add_f16 Differential Revision: https://reviews.llvm.org/D146701
-
Nico Weber authored
-
Nikita Popov authored
Peculiarly, the non-permissive variant handled this gracefully, but the permissive one did not.
-
Dmitry Makogon authored
-
Aaron Ballman authored
This addresses the issues found in: https://lab.llvm.org/buildbot/#/builders/92/builds/41772
-
Kiran Chandramohan authored
This is a temporary message until the feature is implemented and merged. Note: There is a proposed patch (https://reviews.llvm.org/D127215) Reviewed By: peixin Differential Revision: https://reviews.llvm.org/D146768
-
Stefan Gränitz authored
-
Sven van Haastregt authored
-
Kadir Cetinkaya authored
Fixes https://github.com/llvm/llvm-project/issues/61652 Differential Revision: https://reviews.llvm.org/D146732
-
Karl-Johan Karlsson authored
When compiling compiler-rt with -fsanitize=undefined and running testcases you end up with the following warning: UBSan:/repo/uabkaka/llvm-project/compiler-rt/lib/builtins/int_mulo_impl.inc:24:23: signed integer overflow: -1 * -2147483648 cannot be represented in type 'si_int' (aka 'long') This can be avoided by doing the multiplication in a matching unsigned variant of the type. This was found in an out of tree target. Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D146623
-
Ben Shi authored
The 'LPM' instruction has three forms: ------------------------ | form | feature | | ---------- | --------| | LPM | hasLPM | | LPM Rd, Z | hasLPMX | | LPM Rd, Z+ | hasLPMX | ------------------------ The second form is always selected in ISelDAGToDAG, even on devices without FeatureLPMX. This patch emits "LPM + MOV" on devices with only FeatureLPM. Reviewed By: jacquesguan Differential Revision: https://reviews.llvm.org/D141246
-
David Sherwood authored
Quite a few vectoriser tests were using a trip count of 1024, which meant: 1. For fixed-length VFs we would never actually tail-fold, e.g. see Transforms/LoopVectorize/RISCV/uniform-load-store.ll. This is because we can prove at compile-time there will never be a scalar tail. 2. As of D146199 the same optimisation mentioned above will also apply to scalable VFs too. I've changed all such trip counts to be 1025 instead. Differential Revision: https://reviews.llvm.org/D146219
-
Akshay Khadse authored
Fixes some usages of the "auto" keyword to avoid creation of copies. Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D146694
-
mydeveloperday authored
Regenerate the style documentation, requires some minor sphinx changes to avoid warnings Reviewed By: klimek Differential Revision: https://reviews.llvm.org/D146704
-
Stefan Gränitz authored
This first version lays the foundations for AArch32 support in JITLink. ELFLinkGraphBuilder_aarch32 processes REL-type relocations and populates LinkGraphs from ELF object files for both big- and little-endian systems. The ArmCfg member controls subarchitecture-specific details throughout the linking process (i.e. it's passed to ELFJITLinker_aarch32). Relocation types follow the ABI documentation's division into classes: Data (endian-sensitive), Arm (32-bit little-endian) and Thumb (2x 16-bit little-endian, "Thumb32" in the docs). The implementation of instruction encoding/decoding for relocation resolution is implemented symmetrically and is testable in isolation (see AArch32 category in JITLinkTests). Callable Thumb functions are marked with a ThumbSymbol target-flag and stored in the LinkGraph with their real addresses. The thumb-bit is added back in when the owning JITDylib requests the address for such a symbol. The StubsManager can generate (absolute) Thumb-state stubs for branch range extensions on v7+ targets. Proper GOT/PLT handling is not yet implemented. This patch is based on the backend implementation in ez-clang and has just enough functionality to model the infrastructure and link a Thumb function `main()` that calls `printf()` to dump "Hello Arm!" on Armv7a. It was tested on Raspberry Pi with 32-bit Raspbian OS. Reviewed By: lhames Differential Revision: https://reviews.llvm.org/D144083
-
Luke Lau authored
Reviewed By: reames Differential Revision: https://reviews.llvm.org/D146710
-
Andrzej Warzynski authored
This patch extends the `createConst` method so that it can generate constant vectors (it can already generate scalars). This change is required to be able to apply the converter for `arith.floordivsi` (i.e. `FloorDivSIOpConverter`) to vectors. While `arith.floordivsi` is my main motivation for this change, this patch should also allow other Arith ops to be converted in vector cases. In my example, the Linalg vectorizer updates `arith.floordivsi` to operate on vectors and hence the need for this change. Differential Revision: https://reviews.llvm.org/D146741
-
luxufan authored
-
Max Kazantsev authored
-
Martin Storsjö authored
When LLVM_NATIVE_TOOL_DIR was introduced in d3da9067 / D131052, it consisted of refactoring a couple cases of manual logic for tools in clang-tools-extra/clang-tidy, clang-tools-extra/pseudo/include and mlir/tools/mlir-linalg-ods-gen. The former two had the same consistent behaviour while the latter was slightly different, so the refactoring would end up slightly adjusting one or the other. The difference was that the clang-tools-extra tools respected the external variable for setting the tool name, regardless of the LLVM_USE_HOST_TOOLS variable, while mlir-linalg-ods-gen tool only checked its external variable if LLVM_USE_HOST_TOOLS was set. LLVM_USE_HOST_TOOLS is supposed to be enabled automatically whenever cross compiling, so this shouldn't have been an issue. In https://github.com/llvm/llvm-project/issues/60784, it seems like some users do cross compile LLVM, without CMake knowi...
-
Johannes de Fine Licht authored
This revealed a test case that wasn't hitting the intended branch because the inlinees had no function definition. Depends on D146628 Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D146633
-
Martin Storsjö authored
This test somewhat unconventionally assembles both aarch64 and x86 object files. This fixes test failures in build configurations with the aarch64 target enabled but x86 target disabled.
-
luxufan authored
-
Dmitry Chernenkov authored
This reverts commit 25557aa3.
-
Tobias Gysi authored
The revision switches the remaining LLVM dialect tests to use opaque pointers. Selected tests are copied to a postfixed test file for the time being. A number of tests disappear once we fully switch to opaque pointers. In particular, all tests that check verify a pointer element type matches another type as well as tests of recursive types. Part of https://discourse.llvm.org/t/rfc-switching-the-llvm-dialect-and-dialect-lowerings-to-opaque-pointers/68179 Reviewed By: Dinistro, zero9178 Differential Revision: https://reviews.llvm.org/D146726
-
Carlos Galvez authored
[clang-tidy][NFC] Improve naming convention in google-readability-avoid-underscore-in-googletest-name According to the Google docs, the convention is TEST(TestSuiteName, TestName). Apply that convention to the source code, test and documentation of the check. Differential Revision: https://reviews.llvm.org/D146713
-
Michael Platings authored
The new algorithm is: 1. Find all multilibs with flags that are a subset of the requested flags. 2. If more than one multilib matches, choose the last. In addition a new selection mechanism is permitted via an overload of MultilibSet::select() for which multiple multilibs are returned. This allows layering multilibs on top of each other. Since multilibs are now ordered within a list, they no longer need a Priority field. The new algorithm is different to the old algorithm, but in practise the old algorithm was always used in such a way that the effect is the same. The old algorithm was to find the set intersection of the requested flags (with the first character of each removed) with each multilib's flags (ditto), and for that intersection check whether the first character matched. However, ignoring the first characters, the requested flags were always a superset of all the multilibs flags. Therefore the new algorithm can be used as a drop-in replacement. The exception is Fuchsia, which needs adjusting slightly to set both fexceptions and fno-exceptions flags. Differential Revision: https://reviews.llvm.org/D142905
-
Kazu Hirata authored
This patch precommits a test for: https://github.com/llvm/llvm-project/issues/61365
-
Dave Lee authored
-
Xiang1 Zhang authored
Reviewed By: Luo yuanke Differential Revision: https://reviews.llvm.org/D146683
-
Kazu Hirata authored
This patch adds tests for umax(x, 1u). This patch fixes: https://github.com/llvm/llvm-project/issues/60233 It turns out that commit 86b4d864 on Feb 8, 2023 already performs the instcombine transformation proposed in the issue, so the issue requires no change on the codegen side.
-
Xiaodong Liu authored
Keep `EnableLoopDataPrefetch` option off for now because we need a few more TTIs and ISels. This patch is inspired by http://reviews.llvm.org/D17943. Reviewed By: SixWeining Differential Revision: https://reviews.llvm.org/D146600
-