- Oct 28, 2022
-
-
Craig Topper authored
-
rkayaith authored
This adds a new function for creating pass managers that takes an argument for the anchor string. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D136404
-
Chris Bieneman authored
DXContainer files contain a part that has an MD5 of the generated shader. This adds support to the ObjectYAML tooling to expand the hash part data and hash iteself in preparation for adding hashing support to DirectX code generation. Reviewed By: python3kgae Differential Revision: https://reviews.llvm.org/D136632
-
Michael Jones authored
This adds the fgets function and its unit tests. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D136785
-
eopXD authored
The LSR may suggest less profitable transformation to the loop. This patch adds check to prevent LSR from generating worse code than what we already have. Since LSR affects nearly all targets, the patch is guarded by the option 'lsr-drop-solution' and default as disable for now. The next step should be extending an TTI interface to allow target(s) to enable this enhancememnt. Debug log is added to remind user of such choice to skip the LSR solution. Reviewed By: Meinersbur, #loopoptwg Differential Revision: https://reviews.llvm.org/D126043
-
rkayaith authored
Currently any errors during pipeline parsing are reported to stderr. This adds a new pipeline parsing function to the C api that reports errors through a callback, and updates the python bindings to use it. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D136402
-
Fangrui Song authored
readelf --section-details displays ch_type/ch_size/ch_addralign for a SHF_COMPRESSED section. Port the feature. There is a small difference that readelf doesn't display `[<corrupt>]` for an empty section while we do. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D136636
-
Alexander Belyaev authored
Differential Revision: https://reviews.llvm.org/D136852
-
-
Craig Topper authored
Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D136478
-
Dave Lee authored
-
- Oct 27, 2022
-
-
Kito Cheng authored
It splited into several zb* extensions, and `b` is dropped after 0.93, so it time to retired that as other non-ratified zb* extensions. Currntly clang can accept that with warning: $ clang -target riscv64-elf ~/hello.c -S -march=rv64gcb '+b' is not a recognized feature for this target (ignoring feature) '+b' is not a recognized feature for this target (ignoring feature) '+b' is not a recognized feature for this target (ignoring feature) Reviewed By: asb, luismarques Differential Revision: https://reviews.llvm.org/D136812
-
Dave Lee authored
-
Alexandros Lamprineas authored
When calculating the specialization bonus for a given function argument, we recursively traverse the chain of (certain) users, accumulating the instruction costs. Then we exponentially increase the bonus to account for loop nests. This is problematic for two reasons: (a) the users might not themselves be inside the loop nest, (b) if they are we are accounting for it multiple times. Instead we should be adjusting the bonus before traversing the user chain. This reduces the instruction count for CTMark (newPM-O3) when Function Specialization is enabled without actually reducing the amount of specializations performed (geomean: -0.001% non-LTO, -0.406% LTO). Differential Revision: https://reviews.llvm.org/D136692
-
Sanjay Patel authored
This is copying the code that was added for 'add' with D130075. (That patch removed a fallthrough in the cases, but we can probably still share at least some code again as a follow-up cleanup, but I didn't want to risk it here.) The reasoning is similar to the carry propagation for 'add': if we don't demand low bits of the subtraction and the subtrahend (aka RHS or operand 1) is known zero in those low bits, then there can't be any borrowing required from the higher bits of operand 0, so the low bits don't matter. Also, the no-wrap flags can be propagated (and I think that should be true for add too). Here's an attempt to prove that in Alive2: https://alive2.llvm.org/ce/z/xqh7Pa (can add nsw or nuw to src and tgt, and it should still pass) Differential Revision: https://reviews.llvm.org/D136788
-
Liming Liu authored
template-template parameters. Although it effects whether a template can be used as an argument for another template, the constraint seems not to be checked, nor other major implementations (GCC, MSVC, et al.) check it. Additionally, Part-A of the document seems to have been implemented. So mark P0857R0 as completed. Differential Revision: https://reviews.llvm.org/D134128
-
gonglingqin authored
Differential Revision: https://reviews.llvm.org/D135948
-
John Brawn authored
Currently MachineCSE forbids PRE when the instruction reads a physical register. Relax this so that it's allowed when the value being read is the same as what would be read in the place the instruction would be hoisted to. This is being done in preparation for adding FPCR handling to the AArch64 backend, in order to prevent it to from worsening the generated code, but for targets that already have a similar register it should improve things. This patch affects code generation in several tests. The new code looks better except for in Thumb2/LowOverheadLoops/memcall.ll where we perform PRE but the LowOverheadLoops transformation then undoes it. Also in AMDGPU/selectcc-opt.ll the CHECK makes things look worse, but actually the function as a whole is better (as a MOV is PRE'd). Differential Revision: https://reviews.llvm.org/D136675
-
LLVM GN Syncbot authored
-
LLVM GN Syncbot authored
-
Nico Weber authored
-
Michał Górny authored
This reverts commit 88d7508d. It's reported to break builds when symlinking other projects inside the `tools` directory.
-
Jakob Johnson authored
Update the Python tests (ie tests run via `lldb-dotest -p TestTrace`) to handle new error introduced in D136610. Test Plan: `lldb-dotest -p TestTrace` Differential Revision: https://reviews.llvm.org/D136801
-
Carlos Alberto Enciso authored
The fix for the unitest case introduced a dependency on the MC library causing a failure in: https://lab.llvm.org/buildbot/#/builders/121/builds/24567 clang-ppc64le-multistage/stage1 undefined reference to symbol 'llvm::TargetRegistry::lookupTarget' Added: - MC to the LLVM_LINK_COMPONENTS list. Reviewed By: jryans Differential Revision: https://reviews.llvm.org/D136837
-
Alexander Belyaev authored
Differential Revision: https://reviews.llvm.org/D136836
-
Momchil Velikov authored
[fixed test to work with reverse iteration] The `FunctionSpecialization` pass has support for specialising functions, which are called with literal arguments. This functionality is disabled by default and is enabled with the option `-function-specialization-for-literal-constant` . There are a few issues with the implementation, though: * even with the default, the pass will still specialise based on floating-point literals * even when it's enabled, the pass will specialise only for the `i1` type (or `i2` if all of the possible 4 values occur, or `i3` if all of the possible 8 values occur, etc) The reason for this is incorrect check of the lattice value of the function formal parameter. The lattice value is `overdefined` when the constant range of the possible arguments is the full set, and this is the reason for the specialisation to trigger. However, if the set of the possible arguments is not the full set, that must not prevent the specialisation. This patch changes the pass to NOT consider a formal parameter when specialising a function if the lattice value for that parameter is: * unknown or undef * a constant * a constant range with a single element on the basis that specialisation is pointless for those cases. Is also changes the criteria for picking up an actual argument to specialise if the argument is: * a LLVM IR constant * has `constant` lattice value has `constantrange` lattice value with a single element. Reviewed By: ChuanqiXu Differential Revision: https://reviews.llvm.org/D135893 Change-Id: Iea273423176082ec51339aa66a5fe9fea83557ee -
Michał Górny authored
Move `cmake_policy()` settings from `llvm/CMakeLists.txt` into a shared `cmake/modules/CMakePolicy.cmake`. Include it from all relevant projects that support standalone builds, in order to ensure that the policies are consistently set whether they are built in-tree or stand-alone. Differential Revision: https://reviews.llvm.org/D136572
-
Oleg Shyshkov authored
`AffineMap.dropResult` erases one result from the array and it changes indexing. Calling `dropResult` is a loop with increasing indexes does not produce a desired result. Differential Revision: https://reviews.llvm.org/D136833
-
Max Kazantsev authored
We should not delete block predecessors (via replacing successors of terminators) while iterating them, otherwise we may skip some of them. Instead, save predecessors to a separate vector and iterate over it.
-
Nikita Popov authored
-
Matthias Springer authored
This change adds memory space support to tensor.pad. (tensor.generate and tensor.from_elements do not support memory spaces yet.) The memory space is inferred from the buffer of the source tensor. Instead of lowering tensor.pad to tensor.generate + tensor.insert_slice, it is now lowered to bufferization.alloc_tensor (with the correct memory space) + linalg.map + tensor.insert_slice. Memory space support for the remaining two tensor ops is left for a later point, as this requires some more design discussions. Differential Revision: https://reviews.llvm.org/D136265
-
Phoebe Wang authored
-
Matthias Springer authored
This should have been part of D136767.
-
Matthias Springer authored
There is no memref equivalent of tensor.generate. The purpose of this change is to avoid creating scf.parallel loops during bufferization. Differential Revision: https://reviews.llvm.org/D136767
-
Nikita Popov authored
The writeonly attribute for memset_pattern16 (and other referenced libcalls) is being added by InferFunctionAttrs nowadays. No need to special-case it here.
-
Utkarsh Saxena authored
Fixes: https://github.com/llvm/llvm-project/issues/50886 **Adding requires clause to template head** or **constraining the template parameter type** is ineffective because, even though it creates a non-equivalent template head [temp.over.link#6](https://eel.is/c++draft/temp.over.link#6) and hence eligible for overload resolution, `Derived::foo` still [hides any previous using decl](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1283-L1301,). Clang diverges from gcc here and can be seen more clearly in this example: ``` struct base { template <int N, int M> int foo() { return 1; }; }; struct bar : public base { using base::foo; template <int N> int foo() { return 2; }; }; int main() { bar f; f.foo<10, 10>(); // clang previously errored while GCC does not. } ``` https://godbolt.org/z/v5hnh6czq. We see that `bar::foo` hides `base::foo` because it only differs in the head. Adding a trailing `requires` to the definition was a nice find. In this case, clang considers them [overloads](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1148-L1152) because of [mismatching requires clause.](https://github.com/llvm/llvm-project/blob/main/clang/lib/Sema/SemaOverload.cpp#L1390-L1405). So both of them make it to the overload resolution (where constrained Derived::foo is rejected then). --- In this patch, we do not ignore matching the template head (template parameters, type contraints and trailing requires) while considering whether the using decl of base member should be hidden. The return type of a templated function is still not considered as different return types would create ambiguous candidates. The changed tests looks reasonable and also matches GCC behaviour: https://godbolt.org/z/8KqPEThrY Note: We are now able to create an ambiguity in case where both base member and derived member specialisations satisfy the constraints (when the constraints are not same). Ideally using-decl should not create ambiguity. I plan to fix this later if it gathers more attention. Reviewed By: ilya-biryukov, #clang-language-wg Differential Revision: https://reviews.llvm.org/D136440
-
Matthias Springer authored
This change fixes a bug where a dialect is initialized multiple times. This triggers an assertion when the ops of the dialect are registered (`error: operation named ... is already registered`). This bug can be triggered as follows: 1. Dialect A depends on dialect B (as per ADialect.td). 2. Somewhere there is an extension of dialect B that depends on dialect A (e.g., it defines external models create ops from dialect A). E.g.: ``` registry.addExtension(+[](MLIRContext *ctx, BDialect *dialect) { BDialectOp::attachInterface ... ctx->loadDialect<ADialect>(); }); ``` 3. When dialect A is loaded, its `initialize` function is called twice: ``` ADialect::ADialect() | | | v | ADialect::initialize() v getOrLoadDialect<BDialect>() | v (load extension of BDialect) | v ctx->loadDialect<ADialect>() // user wrote this in the extension | v getOrLoadDialect<ADialect>() // the dialect is not "fully" loaded yet | v ADialect::ADialect() | v ADialect::initialize() ``` An example of a dialect extension that depends on other dialects is `Dialect/Tensor/Transforms/BufferizableOpInterfaceImpl.cpp`. That particular dialect extension does not trigger this bug. (It would trigger this bug if the SCF dialect would depend on the Tensor dialect.) This change introduces a new dialect state: dialects that are currently being loaded. Same as dialects that were already fully loaded (and initialized), dialects that are in the process of being loaded are not loaded a second time. Differential Revision: https://reviews.llvm.org/D136685 -
Phoebe Wang authored
For more details about these instructions, please refer to the latest ISE document: https://www.intel.com/content/www/us/en/develop/download/intel-architecture-instruction-set-extensions-programming-reference.html Initial authored by Liu Chen (@LiuChen3) Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D135951
-
Matthias Springer authored
This simplifies the BufferizableOpInterface implementation of vector.transfer_write. Differential Revision: https://reviews.llvm.org/D136348
-
Carlos Alberto Enciso authored
The unitest and test cases are platform dependent (x86_64) causing failures in: https://lab.llvm.org/buildbot/#/builders/245/builds/146 https://lab.llvm.org/buildbot/#/builders/188/builds/21397 No available targets are compatible with triple "x86_64-unknown-unknown". Added: - ';REQUIRES: x86-registered-target' to the LIT tests. - Code to check if the target 'Triple::x86_64' is supported to the unittest case.
-