- Oct 27, 2022
-
-
Kito Cheng authored
It splited into several zb* extensions, and `b` is dropped after 0.93, so it time to retired that as other non-ratified zb* extensions. Currntly clang can accept that with warning: $ clang -target riscv64-elf ~/hello.c -S -march=rv64gcb '+b' is not a recognized feature for this target (ignoring feature) '+b' is not a recognized feature for this target (ignoring feature) '+b' is not a recognized feature for this target (ignoring feature) Reviewed By: asb, luismarques Differential Revision: https://reviews.llvm.org/D136812
-
Dave Lee authored
-
Alexandros Lamprineas authored
When calculating the specialization bonus for a given function argument, we recursively traverse the chain of (certain) users, accumulating the instruction costs. Then we exponentially increase the bonus to account for loop nests. This is problematic for two reasons: (a) the users might not themselves be inside the loop nest, (b) if they are we are accounting for it multiple times. Instead we should be adjusting the bonus before traversing the user chain. This reduces the instruction count for CTMark (newPM-O3) when Function Specialization is enabled without actually reducing the amount of specializations performed (geomean: -0.001% non-LTO, -0.406% LTO). Differential Revision: https://reviews.llvm.org/D136692
-
Sanjay Patel authored
This is copying the code that was added for 'add' with D130075. (That patch removed a fallthrough in the cases, but we can probably still share at least some code again as a follow-up cleanup, but I didn't want to risk it here.) The reasoning is similar to the carry propagation for 'add': if we don't demand low bits of the subtraction and the subtrahend (aka RHS or operand 1) is known zero in those low bits, then there can't be any borrowing required from the higher bits of operand 0, so the low bits don't matter. Also, the no-wrap flags can be propagated (and I think that should be true for add too). Here's an attempt to prove that in Alive2: https://alive2.llvm.org/ce/z/xqh7Pa (can add nsw or nuw to src and tgt, and it should still pass) Differential Revision: https://reviews.llvm.org/D136788
-
Liming Liu authored
template-template parameters. Although it effects whether a template can be used as an argument for another template, the constraint seems not to be checked, nor other major implementations (GCC, MSVC, et al.) check it. Additionally, Part-A of the document seems to have been implemented. So mark P0857R0 as completed. Differential Revision: https://reviews.llvm.org/D134128
-
gonglingqin authored
Differential Revision: https://reviews.llvm.org/D135948
-
John Brawn authored
Currently MachineCSE forbids PRE when the instruction reads a physical register. Relax this so that it's allowed when the value being read is the same as what would be read in the place the instruction would be hoisted to. This is being done in preparation for adding FPCR handling to the AArch64 backend, in order to prevent it to from worsening the generated code, but for targets that already have a similar register it should improve things. This patch affects code generation in several tests. The new code looks better except for in Thumb2/LowOverheadLoops/memcall.ll where we perform PRE but the LowOverheadLoops transformation then undoes it. Also in AMDGPU/selectcc-opt.ll the CHECK makes things look worse, but actually the function as a whole is better (as a MOV is PRE'd). Differential Revision: https://reviews.llvm.org/D136675
-
LLVM GN Syncbot authored
-
LLVM GN Syncbot authored
-
Nico Weber authored
-
Michał Górny authored
This reverts commit 88d7508d. It's reported to break builds when symlinking other projects inside the `tools` directory.
-
Jakob Johnson authored
Update the Python tests (ie tests run via `lldb-dotest -p TestTrace`) to handle new error introduced in D136610. Test Plan: `lldb-dotest -p TestTrace` Differential Revision: https://reviews.llvm.org/D136801
-
Carlos Alberto Enciso authored
The fix for the unitest case introduced a dependency on the MC library causing a failure in: https://lab.llvm.org/buildbot/#/builders/121/builds/24567 clang-ppc64le-multistage/stage1 undefined reference to symbol 'llvm::TargetRegistry::lookupTarget' Added: - MC to the LLVM_LINK_COMPONENTS list. Reviewed By: jryans Differential Revision: https://reviews.llvm.org/D136837
-
Alexander Belyaev authored
Differential Revision: https://reviews.llvm.org/D136836
-
Momchil Velikov authored
[fixed test to work with reverse iteration] The `FunctionSpecialization` pass has support for specialising functions, which are called with literal arguments. This functionality is disabled by default and is enabled with the option `-function-specialization-for-literal-constant` . There are a few issues with the implementation, though: * even with the default, the pass will still specialise based on floating-point literals * even when it's enabled, the pass will specialise only for the `i1` type (or `i2` if all of the possible 4 values occur, or `i3` if all of the possible 8 values occur, etc) The reason for this is incorrect check of the lattice value of the function formal parameter. The lattice value is `overdefined` when the constant range of the possible arguments is the full set, and this is the reason for the specialisation to trigger. However, if the set of the possible arguments is not the full set, that must not prevent the specialisation. This patch changes the pass to NOT consider a formal parameter when specialising a function if the lattice value for that parameter is: * unknown or undef * a constant * a constant range with a single element on the basis that specialisation is pointless for those cases. Is also changes the criteria for picking up an actual argument to specialise if the argument is: * a LLVM IR constant * has `constant` lattice value has `constantrange` lattice value with a single element. Reviewed By: ChuanqiXu Differential Revision: https://reviews.llvm.org/D135893 Change-Id: Iea273423176082ec51339aa66a5fe9fea83557ee -
Michał Górny authored
Move `cmake_policy()` settings from `llvm/CMakeLists.txt` into a shared `cmake/modules/CMakePolicy.cmake`. Include it from all relevant projects that support standalone builds, in order to ensure that the policies are consistently set whether they are built in-tree or stand-alone. Differential Revision: https://reviews.llvm.org/D136572
-
Oleg Shyshkov authored
`AffineMap.dropResult` erases one result from the array and it changes indexing. Calling `dropResult` is a loop with increasing indexes does not produce a desired result. Differential Revision: https://reviews.llvm.org/D136833
-
Max Kazantsev authored
We should not delete block predecessors (via replacing successors of terminators) while iterating them, otherwise we may skip some of them. Instead, save predecessors to a separate vector and iterate over it.
-
Nikita Popov authored
-
Matthias Springer authored
This change adds memory space support to tensor.pad. (tensor.generate and tensor.from_elements do not support memory spaces yet.) The memory space is inferred from the buffer of the source tensor. Instead of lowering tensor.pad to tensor.generate + tensor.insert_slice, it is now lowered to bufferization.alloc_tensor (with the correct memory space) + linalg.map + tensor.insert_slice. Memory space support for the remaining two tensor ops is left for a later point, as this requires some more design discussions. Differential Revision: https://reviews.llvm.org/D136265
-
Phoebe Wang authored
-
Matthias Springer authored
This should have been part of D136767.
-
Matthias Springer authored
There is no memref equivalent of tensor.generate. The purpose of this change is to avoid creating scf.parallel loops during bufferization. Differential Revision: https://reviews.llvm.org/D136767
-
Nikita Popov authored
The writeonly attribute for memset_pattern16 (and other referenced libcalls) is being added by InferFunctionAttrs nowadays. No need to special-case it here.
-
Utkarsh Saxena authored
Fixes: https://github.com/llvm/llvm-project/issues/50886 **Adding requires clause to template head** or **constraining the template parameter type** is ineffective because, even though it creates a non-equivalent template head [temp.over.link#6](https://eel.is/c++draft/temp.over.link#6) and hence eligible for overload resolution, `Derived::foo` still [hides any previous using decl](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1283-L1301,). Clang diverges from gcc here and can be seen more clearly in this example: ``` struct base { template <int N, int M> int foo() { return 1; }; }; struct bar : public base { using base::foo; template <int N> int foo() { return 2; }; }; int main() { bar f; f.foo<10, 10>(); // clang previously errored while GCC does not. } ``` https://godbolt.org/z/v5hnh6czq. We see that `bar::foo` hides `base::foo` because it only differs in the head. Adding a trailing `requires` to the definition was a nice find. In this case, clang considers them [overloads](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1148-L1152) because of [mismatching requires clause.](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1390-L1405). So both of them make it to the overload resolution (where constrained Derived::foo is rejected then). --- In this patch, we do not ignore matching the template head (template parameters, type contraints and trailing requires) while considering whether the using decl of base member should be hidden. The return type of a templated function is still not considered as different return types would create ambiguous candidates. The changed tests looks reasonable and also matches GCC behaviour: https://godbolt.org/z/8KqPEThrY Note: We are now able to create an ambiguity in case where both base member and derived member specialisations satisfy the constraints (when the constraints are not same). Ideally using-decl should not create ambiguity. I plan to fix this later if it gathers more attention. Reviewed By: ilya-biryukov, #clang-language-wg Differential Revision: https://reviews.llvm.org/D136440
-
Matthias Springer authored
This change fixes a bug where a dialect is initialized multiple times. This triggers an assertion when the ops of the dialect are registered (`error: operation named ... is already registered`). This bug can be triggered as follows: 1. Dialect A depends on dialect B (as per ADialect.td). 2. Somewhere there is an extension of dialect B that depends on dialect A (e.g., it defines external models create ops from dialect A). E.g.: ``` registry.addExtension(+[](MLIRContext *ctx, BDialect *dialect) { BDialectOp::attachInterface ... ctx->loadDialect<ADialect>(); }); ``` 3. When dialect A is loaded, its `initialize` function is called twice: ``` ADialect::ADialect() | | | v | ADialect::initialize() v getOrLoadDialect<BDialect>() | v (load extension of BDialect) | v ctx->loadDialect<ADialect>() // user wrote this in the extension | v getOrLoadDialect<ADialect>() // the dialect is not "fully" loaded yet | v ADialect::ADialect() | v ADialect::initialize() ``` An example of a dialect extension that depends on other dialects is `Dialect/Tensor/Transforms/BufferizableOpInterfaceImpl.cpp`. That particular dialect extension does not trigger this bug. (It would trigger this bug if the SCF dialect would depend on the Tensor dialect.) This change introduces a new dialect state: dialects that are currently being loaded. Same as dialects that were already fully loaded (and initialized), dialects that are in the process of being loaded are not loaded a second time. Differential Revision: https://reviews.llvm.org/D136685 -
Phoebe Wang authored
For more details about these instructions, please refer to the latest ISE document: https://www.intel.com/content/www/us/en/develop/download/intel-architecture-instruction-set-extensions-programming-reference.html Initial authored by Liu Chen (@LiuChen3) Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D135951
-
Matthias Springer authored
This simplifies the BufferizableOpInterface implementation of vector.transfer_write. Differential Revision: https://reviews.llvm.org/D136348
-
Carlos Alberto Enciso authored
The unitest and test cases are platform dependent (x86_64) causing failures in: https://lab.llvm.org/buildbot/#/builders/245/builds/146 https://lab.llvm.org/buildbot/#/builders/188/builds/21397 No available targets are compatible with triple "x86_64-unknown-unknown". Added: - ';REQUIRES: x86-registered-target' to the LIT tests. - Code to check if the target 'Triple::x86_64' is supported to the unittest case.
-
eopXD authored
This is a pre-commit for the fix in D136784. Reviewed By: SjoerdMeijer Differential Revision: https://reviews.llvm.org/D136783
-
eopXD authored
Pre-commit test case for D126043 Reviewed By: Meinersbur, #loopoptwg Differential Revision: https://reviews.llvm.org/D134823
-
Matthias Springer authored
tensor.insert and tensor.insert_slice (as destination style ops) do no longer need to implement the entire BufferizableOpInterface. Differential Revision: https://reviews.llvm.org/D136347
-
Haojian Wu authored
-
David Spickett authored
GDB implemented data_bit_offset in https://sourceware.org/bugzilla/show_bug.cgi?id=12616 which has been present since GDB 8.0. GCC started using it at GCC 11. Reviewed By: dblaikie Differential Revision: https://reviews.llvm.org/D135583
-
wangpc authored
There are a lot of cases for pseudos of the same instruction, here we just use existed mapping table to map pseudos to real instructions to reduce cases. Reviewed By: kito-cheng Differential Revision: https://reviews.llvm.org/D128271
-
OCHyams authored
And update the unittest. Reviewed By: dblaikie Differential Revision: https://reviews.llvm.org/D136242
-
Chuanqi Xu authored
Rename module related things according to the consensus in https://discourse.llvm.org/t/rfc-unifying-the-terminology-about-modules-in-clang/66054/ to reduce further confusings. This only renames things I can make sure. It doesn't mean all the names in Preprocessor are correct now.
-
Guillaume Chatelet authored
This patch seems to introduce bugs on aarch64. Reverting while we investigate the root cause. This reverts commit 02841488.
-
Nikita Popov authored
Clang detects the GCC version from the libdir. However, modern GCC versions only include the major version in the libdir (something like lib/gcc/powerpc64le-linux-gnu/12/), not all version components. For this reason, even though the system has a supported libstdcxx, it will still fail the check against the 12.1.0 version requirement. Fix this by doing the same thing we do for patch versions: Assume that a missing minor version is larger than any specific version. To allow this to be tested, we need to fix two additional issues: First, the GCC toolchain directories used for testing need to contain a crtbegin.o file to be properly detected. The existing tests actually ended up using a 0.0.0 version, rather the intended one. Second, we also need to satisfy the glibc version check based on the dynamic linker. To do so, respect the --dyld-prefix argument and add the necessary file to the test toolchain directory. Differential Revision: https://reviews.llvm.org/D136258
-
Alexander Belyaev authored
This reverts commit c34de60e.
-