- Feb 11, 2021
-
-
Stephan Herhut authored
With the standard dialect being split up, the set of dialects that are used when converting to GPU is growing. This change modifies the SCFToGpu pass to allow all operations inside launch bodies. Differential Revision: https://reviews.llvm.org/D96480
-
Markus Lavin authored
Need to take endianness into account when doing vector to scalar casts such as %bc = bitcast <8 x i1> %v to i8 Companion commit for https://reviews.llvm.org/D94867 Upload in response to https://lists.llvm.org/pipermail/llvm-dev/2021-January/147862.html Attempting to document the actual memory layout rules for vectors in https://reviews.llvm.org/D94964 Differential Revision: https://reviews.llvm.org/D94765
-
David Green authored
We were storing predicate registers, such as a <8 x i1>, in the opposite order to how the rest of llvm expects. This actually turns out to be correct for the one place that usually uses it - the ScalarizeMaskedMemIntrin pass, but only because the pass was incorrect itself. This fixes the order so that bits are stored in the opposite order and bitcasts work as expected. This allows the Scalarization pass to be fixed, as in https://reviews.llvm.org/D94765. Differential Revision: https://reviews.llvm.org/D94867
-
Haojian Wu authored
The EndLoc of a type loc can be invalid for broken code. Also extend the existing test to support error code with `error-ok` annotation. Differential Revision: https://reviews.llvm.org/D96261
-
Haojian Wu authored
-
Sander de Smalen authored
This patch is NFC and changes occurrences of `unsigned Width` and `unsigned i` to work on type ElementCount instead. This patch is a preparatory patch with the ultimate goal of making `computeMaxVF()` return both a max fixed VF and a max scalable VF, so that `selectVectorizationFactor()` can pick the most cost-effective vectorization factor. Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D96019
-
Haojian Wu authored
It is useful for syntax-tree developement. Reviewed By: kbobyrev Differential Revision: https://reviews.llvm.org/D96017
-
Sander de Smalen authored
This patch fixes an issue in the implementation of DUP/CPY where certain immediates were not accepted. Immediates should be interpreted as a two's complement encoding of a value that fits the number of bits of the element type. mov z0.b, p0/z, #127 <=> mov z0.b, p0/z, #-129 <=> mov z0.b, p0/z, #0xffffffffffffff7f This behaviour is in line with the GNU assembler. Reviewed By: c-rhodes Differential Revision: https://reviews.llvm.org/D94776 -
Hanhan Wang authored
The dimension order of a filter in tensorflow is [filter_height, filter_width, in_channels, out_channels], which is different from current definition. The current definition follows TOSA spec. Add TF version conv ops to .tc, so we do not have to insert a transpose op around a conv op. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D96038
-
Arthur Eubanks authored
Some parameters were already part of the Config passed in.
-
Sanjoy Das authored
This should have gone in with a76761cf.
-
Sanjoy Das authored
- Remove leftover comment from de2568aa - Fix a typo in a comment
-
Max Kazantsev authored
Function `replaceMathCmpWithIntrinsic` artificially limits the scope of the optimization, setting a requirement of two instructions be in the same block, due to two reasons: - usage of DT for more general check is costly in terms of compile time; - risk of creating a new value that lives through multiple blocks. Because of this, two semantically equivalent tests may be or not be the subject of this opt depending on where the binary operation is located. See `test/CodeGen/X86/usub_inc_iv.ll` for motivation There is one important particular case where this limitation is too strict: it is when the binary operation is the increment of the induction variable. As result, the application of this opt becomes fragile and highly reliant on where other passes decide to place IV increment. In most cases, they place it in the end of the latch block, killing the opt opportunity (when in fact it does not matter where to insert the actual instruction). This patch handles this particular case separately. - The detector does not use dom tree and has constant cost; - The value of IV or IV.next lives through all loop in any case, so this should not create a new unexpected long-living value. As result, the transform becomes more robust. It also seems to lead to better code generation in some cases (see `test/CodeGen/X86/lsr-loop-exit-cond.ll`). Differential Revision: https://reviews.llvm.org/D96119 Reviewed By: spatel, reames
-
Max Kazantsev authored
-
Yang Fan authored
GCC warning: ``` /llvm-project/clang/lib/Frontend/TestModuleFileExtension.cpp:131:20: warning: ‘llvm::raw_ostream& clang::operator<<(llvm::raw_ostream&, const clang::TestModuleFileExtension&)’ has not been declared within ‘clang’ 131 | llvm::raw_ostream &clang::operator<<(llvm::raw_ostream &OS, | ^~~~~ In file included from /llvm-project/clang/lib/Frontend/TestModuleFileExtension.cpp:8: /llvm-project/clang/lib/Frontend/TestModuleFileExtension.h:75:3: note: only here as a ‘friend’ 75 | operator<<(llvm::raw_ostream &OS, const TestModuleFileExtension &Extension); | ^~~~~~~~ ``` -
Carl Ritson authored
Add mimgopc object to represent the opcode allowing different opcodes for different hardware variants. This enables image_atomic_fcmpswap, image_atomic_fmin, and image_atomic_fmax on GFX10 Reviewed By: foad, rampitec Differential Revision: https://reviews.llvm.org/D96309
-
Michael Kruse authored
Move SimplifiyVisitor from Simplify.h to Simplify.cpp. It is not relevant for applying the pass in either the NewPM or the legacyPM. Rename it to SimplifyImpl to account for that. This is possible due its state not being necessary to be preserved between runs and thefore SimplifyImpl not needed to be held in the pass object. Instead, SimplifyImpl is only instatiated for the current Scop. In the NewPM as a function-local variable, and in the legacy PM inside a llvm::Optional object because the state must be preserved between the printScop (invoked by opt -analyze) and the most recent runOnScop calls.
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
Identified with readability-const-return-type.
-
Craig Topper authored
This removes the commuted PatFrags that only existed to carry an SDNodeXForm in its OperandTransform field. We know all the places that need to use the commuted SDNodeXForm and there is one transform shared by signed and unsigned compares. So just hardcode the the SDNodeXForm where it is needed and use the non commuted PatFrag in the pattern. I think when I wrote this I thought the SDNodeXForm name had to match what is in the PatFrag that is being used. But that's not true. The OperandTransform is only used when the PatFrag is used in an instruction pattern and not a separate Pat pattern. All the commuted cases are Pat patterns.
-
Michael Kruse authored
"using namespace" pollutes the namespace of every file that includes such a header and universally considered a bad thing. Even the variant namespace polly { using namespace llvm; } (previously used by LoopGenerators.h) imports more symbols than the file is in control of. The header may include a fixed set of files from LLVM, but the header itself may by be included together with other headers from LLVM. For instance, LLVM's MemorySSA.h and Polly's ScopInfo.h both declare a class 'MemoryAccess' which may conflict. Instead of prefixing everything in Polly's header files, this patch adds 'using' statements to import only the symbols that are actually referenced in Polly. This approach is also used by MLIR to import commonly used symbols into the mlir namespace. This patch also puts the symbols declared in IslNodeBuilder.h into the Polly namespace to also be able to use the imported symbols. -
Valentin Clement authored
-
Daniel Hwang authored
Updates static analyzer to be able to generate both sarif and html output in a single run similar to plist-html. Differential Revision: https://reviews.llvm.org/D96389
-
Jessica Clarke authored
-
Jessica Clarke authored
-
Aart Bik authored
Rationale: BuiltinTypes.cpp observed overflow when computing size of tensor<100x200x300x400x500x600x700x800xf32>. Reviewed By: stellaraccident Differential Revision: https://reviews.llvm.org/D96475
-
Craig Topper authored
This binds the SDNodeXForm to the ImmLeaf so we only need to mention the ImmLeaf in both the input and output pattern.
-
Mehdi Amini authored
Differential Revision: https://reviews.llvm.org/D96474
-
Valentin Clement authored
This patch is a follow up of D96422 and move the ShapeShiftType to TableGen. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D96442
-
peter klausler authored
When accessing a specific procedure of a USE-associated generic interface, we need to allow for the case in which that specific procedure has the same name as the generic when testing for its availability in the current scope. Differential Revision: https://reviews.llvm.org/D96467
-
Adrian Prantl authored
-
Jianzhou Zhao authored
https://reviews.llvm.org/D95835 implements origin tracking for DFSan. It reuses the chained origin depot of MSan. This change moves the utility to sanitizer_common to share between MSan and DFSan. Reviewed-by: eugenis, morehouse Differential Revision: https://reviews.llvm.org/D96319
-
peter klausler authored
Some state in name resolution is stored in the DeclarationVisitor instance and processed at the end of the specification part. This state needs to accommodate nested specification parts, namely the ones that can be nested in a subroutine or function interface body. Differential Revision: https://reviews.llvm.org/D96466
-
xgupta authored
An error has occurred when I build libunwind with -DLLVM_BUILD_DOCS=ON. Reviewed By: #libunwind, compnerd Differential Revision: https://reviews.llvm.org/D96107
-
xgupta authored
-
Mehdi Amini authored
The CMake changes in 2aa1af9b to make it possible to build MLIR as a standalone project unfortunately disabled all unit-tests from the regular in-tree build.
-
Duncan P. N. Exon Smith authored
Rename the `RF_MoveDistinctMDs` flag passed into `MapValue` and `MapMetadata` to `RF_ReuseAndMutateDistinctMDs` in order to more precisely describe its effect and clarify the header documentation. Found this while helping to investigate PR48841, which pointed out an unsound use of the flag in `CloneModule()`. For now I've just added a FIXME there, but I'm hopeful that the new (more precise) name will prevent other similar errors.
-
Valentin Clement authored
This is the first patch of a serie to move FIR types to TableGen format as suggested in D96172. This patch is setting up the files for FIR types and move the ShapeType to TableGen. As discussed with @schweitz, I'm taking over this task to help the FIR upstreaming effort. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D96422
-
Vedant Kumar authored
This test started failing after https://reviews.llvm.org/D95849 defaulted --allow-unused-prefixes to false. Taking a look at the test, I didn't see an obvious need to add OS-specific check lines for each supported value of %os. rdar://74207657
-