- Sep 28, 2023
-
-
Nikita Popov authored
Use the constant folding API instead. In preparation for dropping zext constant expressions.
-
Krzysztof Drewniak authored
Define operations that wrap the gfx940's new operations for converting between f32 and registers containing packed sets of four 8-bit floats. Define rocdl operations for the intrinsics and an AMDGPU dialect wrapper around them (to account for the fact that MLIR distinguishes the two float formats at the type level but that the LLVM IR does not). Define an ArithToAMDGPU pass, meant to run before conversion to LLVM, that replaces relevant calls to arith.extf and arith.truncf with the packed operations in the AMDGPU dialect. Note that the conversion currently only handles scalars and vectors of rank <= 1, as we do not have a usecase for multi-dimensional vector support right now. Reviewed By: jsjodin Differential Revision: https://reviews.llvm.org/D152457
-
René Rebe authored
This addresses missing cmake files needed to build some sub-projects like libstdcxx. Co-authored-by:René Rebe <rene@exactcode.de>
-
Andrew Gozillon authored
Fix mistyped syntax in omptarget-region-parallel-llvm.mlir test added by b05d436e
-
Andrzej Warzynski authored
This patch updates `transform.loop.peel` so that this Op returns two rather than one handle: * one for the peeled loop, and * one for the remainder loop. Also, following this change this Op will fail if peeling fails. This is consistent with other similar Ops that also fail if no transformation takes place. Relands #67482 with an extra fix for transform_loop_ext.py
-
Fabio D'Urso authored
-
Nikita Popov authored
Instead work on APInt.
-
Nikita Popov authored
Work on APInt instead.
-
Nikita Popov authored
Use IRBuilder instead, which will either insert an instruction or constant fold.
-
Younan Zhang authored
From two aspects: - For function templates, emit additional template argument placeholders in the context where it can't be a call in order to specify an instantiation explicitly. - Consider expressions with base type specifier such as 'Derived().Base::foo^' a function call. Reviewed By: nridge Differential Revision: https://reviews.llvm.org/D156605
-
Nikita Popov authored
Let the IRBuilder constant fold instead.
-
Nikita Popov authored
We don't require a Constant here, so let IRBuilder fold this.
-
Sam McCall authored
(This fails if the input is not writable)
-
Jay Foad authored
This makes some tests robust against minor codegen differences that will be caused by PR #67038.
-
Goran Flegar authored
-
Nikita Popov authored
In preparation for removing these constant expressions.
-
Nikita Popov authored
Avoid an unnecessary use of ConstantExpr::getZExt() when APInt::zext() is sufficient.
-
Mel Chen authored
The vectorization of the FindLastIV reduction does not depend on the nocapture and readonly attributes.
-
Louis Dionne authored
We don't neeed to handle both spellings anymore since we don't support Clang 15 anymore.
-
Louis Dionne authored
The tests were a bit of a mess -- the testing coverage wasn't bad but it was extremely difficult to see what was being tested and where. I split up the tests to make them easier to audit for completeness and did such an audit, adding a few missing tests (e.g. the conditional noexcept-ness of std::cbegin and std::cend). I also audited the synopsis and adjusted it where it needed to be adjusted. This patch is in preparation of fixing #67471.
-
Nikita Popov authored
This fixes the bitcode upgrade failure reported in https://reviews.llvm.org/D155924#4616789. The expansion always happens in the entry block, so this may be inaccurate if there are trapping constant expressions.
-
Louis Dionne authored
Even though the underlying issue was fixed in #65177 when we made std::pointer_traits SFINAE-friendly, it is worth adding a regression test since it's so easy to do.
-
bipmis authored
-
Tuan Chuong Goh authored
Fix test since the review was created
-
martinboehme authored
Also, change the mouse cursor into a pointer instead of a text cursor. This makes it more discoverable that the element can be opened and closed.
-
Andrzej Warzyński authored
This conversion is identical to vector.broadcast when broadcasting a scalar.
-
Sam McCall authored
Currently it tries to call S.size() when preallocating the target string, which doesn't compile. vector<const char*> is used a bunch in e.g. clang driver.
-
Nikita Popov authored
We should check whether the element type is non-byte-sized, not the vector type. For types like <32 x i1> the whole type is byte-sized, but the individual elements (that we scalarize to) are not. Fixes https://github.com/llvm/llvm-project/issues/67060.
-
Nikita Popov authored
-
Jay Foad authored
-
Jay Foad authored
Previously this was relying on [[RESULT]] having been defined in an earlier function.
-
Cullen Rhodes authored
The following patterns - TransferReadToVectorLoadLowering - TransferWriteToVectorStoreLowering attempt to generate invalid vector.maskedload and vector.maskedstore ops for non rank-1 vector types. These ops operate on 1-D vectors. This patch adds a check to prevent this.
-
Muhammad Omair Javaid authored
Recently added TLS linux support fails on Arm/AArch64. I am skiping test for now and will investigate the issue later.
-
chuongg3 authored
G_VECREDUCE_ADD is now able to have v4i16 and v8i8 vector types as source registers
-
Timm Bäder authored
-
Timm Bäder authored
Pull the nested if statement into the outer one.
-
Cullen Rhodes authored
The vector.extract assembly format currently only contains the source type, for example: %1 = vector.extract %0[1] : vector<3x7x8xf32> it's not immediately obvious if this is the source or result type. This patch improves the assembly format to make this clearer, so the above becomes: %1 = vector.extract %0[1] : vector<7x8xf32> from vector<3x7x8xf32>
-
Nikita Popov authored
-
Mikael Holmen authored
Without the fix gcc warns about ../lib/Transforms/Vectorize/VPlanTransforms.cpp:968:42: warning: suggest parentheses around '&&' within '||' [-Wparentheses] 968 | UseActiveLaneMaskForControlFlow && | ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^~ 969 | "DataAndControlFlowWithoutRuntimeCheck implies " | ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 970 | "UseActiveLaneMaskForControlFlow"); | ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -
Cullen Rhodes authored
This patch adds support for lowering a vector.transfer_read with a transpose permutation map to a vertical tile load, for example: vector.transfer_read ... permutation_map: (d0, d1) -> (d1, d0) is converted to: arm_sme.tile_load ... <vertical> On SME the transpose can be done in-flight, rather than as a separate operation as in the TransferReadPermutationLowering, which would do the following: %0 = vector.transfer_read ... vector.transpose %0, [1, 0] ... The lowering doesn't support masking yet and the transfer_read must be in-bounds. It also intentionally doesn't handle simple loads as transfer_write currently does, as the generic TransferReadToVectorLoadLowering can lower these to simple vector.load ops, which can already be lowered to ArmSME. A subsequent patch will update the existing transfer_write lowering, this is a separate patch as there is currently no lowering for vector.transfer_read.
-