- Jan 19, 2023
-
-
David Green authored
As Armv9-a implies SVE2 it implies SVE (added in D141411) and so it should also imply FP16, which this patch adds. This helps get the target features correct when using `target("arch=armv9-a")` attributes. There is also an adjustment to AssertSameExtensionFlags in this patch to make it print cpu names, useful when the TargetParser unit tests are run through lit to distinguish which cpu is failing. Differential Revision: https://reviews.llvm.org/D142087 -
Guilherme Valarini authored
The entries inside a "target data end" is processed in three steps: 1. Query internal data maps for the entries and dispatch any necessary device-side operations (i.e., data retrieval); 2. Synchronize the such operations; 3. Update the host-side pointers and remove any entry which reference counter reached zero. Such steps may be executed by multiple threads which may even operate on the same entries. The current implementation (D121058) tries to synchronize these threads by tracking the "owner" for the deletion of each entry using their thread ID. Unfortunately it may failed to do so because of the following reasons: 1. The owner is always assigned at the first step only if the reference count is 0 when the map is queried. This does not work when such owner thread is faster than a previous one that is also processing the same entry on another "target data end", leading to user-after-free problems. 2. The entry is only added for post-processing (step 3) if its reference count was 0 at query time (step 1). This does not allow for threads to exchange responsibility for the deletion, leading again to user-after-free problems. 3. An entry may appear multiple times in the arguments array of a "target data end", which may lead to deleting the entry prematurely, leading, again, to user-after-free problems. This patch addresses these problems by tracking all the threads that are using an entry at "target data end" region through a counter, ensuring only the last one deletes it when needed. It also ensures that all entries that are successfully found inside the data maps in step 1 are also processed in step 3, regardless if their reference count was zeroed or not at query time. This ensures the deletion ownership may be passed to any thread that is using such entry. Reviewed By: ye-luo Differential Revision: https://reviews.llvm.org/D132676 -
Ben Mudd authored
DexExpectStepOrder uses the line to expect a debugger step from the actual line of the command in the Dexter source file. Now Dexter scripts have mainly moved to thier own script files instead of the actual source, there should be a option to override this behaviour to choose your own debugger step location. Reviewed By: Orlando Differential Revision: https://reviews.llvm.org/D142099
-
Nikita Popov authored
And regenerate test checks to pick up new names.
-
Nikita Popov authored
I made a typo here, this was supposed to be !align rather than !aligned. But then !align can only be applied to loads, not calls (where one would use the return attribute instead). And freeze can't be pushed through loads anyway, so there's no way to test this case (same as !nonnull).
-
Yitzhak Mandelbaum authored
Also adds uses of the new printing in analysis inner loop. Differential Revision: https://reviews.llvm.org/D141716
-
Michal Paszkowski authored
This change makes the AsmPrinter emit OpExecutionMode ContractionOff when both opencl.enable.FP_CONTRACT and spirv.ExecutionMode metadata are not present. Differential Revision: https://reviews.llvm.org/D141734
-
Nikita Popov authored
These are currently being miscompiled, see PR59888.
-
Alex Brachet authored
Add {e,r}flags, {g,f}s.base registers so they can be referenced in cfi directives,. They are not otherwise useable in any instructions, but can be implicitly pushed to the stack like with pushf for {e,r}flags. Differential Revision: https://reviews.llvm.org/D141879 -
Nico Weber authored
-
Nico Weber authored
-
Haojian Wu authored
A prototype of using include-cleaner library in clangd: - (re)implement clangd's "unused include" warnings with the library - the new implementation is hidden under a flag `Config::UnusedIncludesPolicy::Experiment` Differential Revision: https://reviews.llvm.org/D140875
-
Christian Ulmann authored
This commit purges direct accesses to MD_prof metadata and replaces them with the accessors provided from the utility file wherever possible. This commit can be seen as the first step towards switching the branch weights to 64 bits. See post here: https://discourse.llvm.org/t/extend-md-prof-branch-weights-metadata-from-32-to-64-bits/67492 Reviewed By: davidxl, paulkirth Differential Revision: https://reviews.llvm.org/D141393
-
Amaury Séchet authored
Fix a regression in D141883 Depends on D141883 Reviewed By: lebedev.ri Differential Revision: https://reviews.llvm.org/D141884
-
Haojian Wu authored
Support building UsingType for elaborated type specifiers: ``` namespace ns { class Foo {}; } using ns::Foo; // The TypeLoc of `Foo` below should be a ElaboratedTypeLoc with an // inner UsingTypeLoc rather than the underlying `CXXRecordTypeLoc` class Foo foo; ``` Differential Revision: https://reviews.llvm.org/D141280 -
Jean Perier authored
Adds support for: - referencing a whole allocatable/pointer symbol - passing allocatable/pointer in a call This required update in HLFIRTools.cpp helpers so that the raw address, extents, lower bounds, and type parameters of a fir.box/fir.class can be extracted. This is required because in hlfir lowering, dereferencing a pointer/alloc is only doing the fir.load fir.box part, and the helpers have to be able to reason about that fir.box without the help of a "fir::FortranVariableOpInterface". Missing: - referencing part of allocatable/pointer (will need to update Designator lowering to dereference the pointer/alloc). Same for whole allocatable and pointer components. - allocate/deallocate/pointer assignment statements. - Whole allocatable assignment. - Lower inquires. Differential Revision: https://reviews.llvm.org/D142043
-
serge-sans-paille authored
When used to find an exact match, some extra context can be used to totally cut some computations. This saves 1% of the instruction count when pre processing sqlite3.c through valgrind --tool=callgrind ./bin/clang -E sqlite3.c -o/dev/null Differential Revision: https://reviews.llvm.org/D142026
-
Groverkss authored
This patch adds support for divisions in the union of two PWMAFunction. This is now possible because of previous patches, which made divisions explicitly stored in MultiAffineFunction (MAF). This patch also refactors the previous implementation, moving the implementation for obtaining a set of points where a MAF is lexicographically "better" than the other to MAF. Reviewed By: arjunp Differential Revision: https://reviews.llvm.org/D138118
-
Michal Paszkowski authored
Differential Revision: https://reviews.llvm.org/D142061
-
Timm Bäder authored
This reverts commit 5b54cf1a. This breaks a builder: https://lab.llvm.org/buildbot/#/builders/5/builds/30854
-
Timm Bäder authored
This reverts commit 490e8214. This breaks a builder: https://lab.llvm.org/buildbot/#/builders/214/builds/5415
-
Timm Bäder authored
-
Amaury Séchet authored
This transofrm loses information that can be useful for other transforms. Reviewed By: lebedev.ri Differential Revision: https://reviews.llvm.org/D141883
-
Evan Smal authored
This patch improves the diagnostic message "initializer-string for char array is too long" by specifying an expected array length and by indicating that the initializer string implicitly includes the null terminator. Fixes #58829 Differential Revision: https://reviews.llvm.org/D141283
-
Florian Hahn authored
This patch adds a new VPBlockShallowTraversalWrapper struct to provide graph traits specialization that do not traverse through VPRegionBlocks. This matches the behavior of the existing traits for plain VPBlockBase and is a step before moving the graph traits for VPBlockBase to traverse through VPRegionBlocks to enable cross region support in VPDominatorTree. Depends on D140511. Reviewed By: Ayal Differential Revision: https://reviews.llvm.org/D140512
-
Mariusz Sikora authored
Introducing feature predicate VFmacF64Inst for targets which supports v_fmac_f64 instructions. Differential Revision: https://reviews.llvm.org/D142017
-
Timm Bäder authored
This reverts commit fddf6418. Apparently this also breaks some builders: /usr/bin/ld: EvalEmitter.cpp:(.text._ZN5clang6interp11EvalEmitter7emitShlENS0_8PrimTypeES2_RKNS0_10SourceInfoE+0x1f54): undefined reference to `bool clang::interp::CheckShift<clang::interp::Integral<16u, true> >(clang::interp::InterpState&, clang::interp::CodePtr, clang::interp::Integral<16u, true> const&, unsigned int)' /usr/bin/ld: EvalEmitter.cpp:(.text._ZN5clang6interp11EvalEmitter7emitShlENS0_8PrimTypeES2_RKNS0_10SourceInfoE+0x1fd4): undefined reference to `bool clang::interp::CheckShift<clang::interp::Integral<32u, true> >(clang::interp::InterpState&, clang::interp::CodePtr, clang::interp::Integral<32u, true> const&, unsigned int)' /usr/bin/ld: EvalEmitter.cpp:(.text._ZN5clang6interp11EvalEmitter7emitShlENS0_8PrimTypeES2_RKNS0_10SourceInfoE+0x2058): undefined reference to `bool clang::interp::CheckShift<clang::interp::Integral<32u, true> >(clang::interp::InterpState&, clang::interp::CodePtr, clang::interp::Integral<32u, true> const&, unsigned int)' (etc)
-
Dominik Adamski authored
simd nontemporal construct is represented as a list of variables which have low locality accross simd iterations Added verifier of nontemporal clause. MLIR tests were updated to test correctness of MLIR definition of nontemporal clause. Differential Revision: https://reviews.llvm.org/D140553 Reviewed By: kiranchandramohan
-
Jordan Rupprecht authored
-
Timm Bäder authored
This reverts commit 9ee0d749.
-
Timm Bäder authored
Implement mul, div, rem, etc. compound assign operators. Differential Revision: https://reviews.llvm.org/D137071
-
Timm Bäder authored
Just like we do with all the other Check* functions.
-
Timm Bäder authored
-
Timm Bäder authored
This way we can check for this flag in the new interpreter as well.
-
Quentin Colombet authored
The `lower_vectors` operation of the transform dialect takes a lot of arguments to build. In order to make C++ code easier to work with when using this instruction, introduce a new structure, named `LowerVectorsOptions`, that aggregates all the options that are used to build this instruction. This allows to use patterns like: ``` LowerVectorsOptions opts; opts.setOptZ(...) .setOptY(...)...; builder.create<LowerVectorsOp>(target, opts); ``` Instead of having to pass all N options directly to the builder and set them in the right order. NFC Differential Revision: https://reviews.llvm.org/D141923
-
Christian Ulmann authored
This commit introduces LLVM's `MemoryEffects` attribute and replaces the deprecated usage of `llvm.readnone` in the LLVM dialect. The absence of the attribute on a `LLVMFuncOp` implies that it might access all kinds of memory. This semantic corresponds to `llvm::Function`'s behaviour. Depends on D142002 Differential Revision: https://reviews.llvm.org/D142013
-
Alex Zinenko authored
-
Alex Zinenko authored
Simplify the handling of silenceable failures in the transform dialect. Previously, the logic of `TransformEachOpTrait` required that `applyToEach` returned a list of null pointers when a silenceable failure was emitted. This was not done consistently and also crept into ops without this trait although they did not require it. Handle this case earlier in the interpreter and homogeneously associated preivously unset transform dialect values (both handles and parameters) with empty lists of the matching kind. Ignore the results of `applyToEach` for the targets for which it produced a silenceable failure. As a result, one never needs to set results to lists containing nulls. Furthermore, the objects associated with transform dialect values must never be null. Depends On D140980 Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D141305
-
Alex Zinenko authored
Use the recently introduced transform dialect parameter mechanism to perform controllable multi-size tiling with sizes computed at the transformation time rather than at runtime. This requires to generalize tile and split structured transform operations to work with any transform dialect handle types, which is desirable in itself to avoid unchecked overuse of PDL OperationType. Reviewed By: shabalin Differential Revision: https://reviews.llvm.org/D140980
-
Nikita Popov authored
Similarly to what backref() does, add an "easy path" to slow() that can handle some non-branching cases, in particular simple character matches. This has the dual effect of reducing the number of characters we need to match, and the number of states in the NFA. This reduces FileCheck runtime on vloxseg.c from 17s to 12s on my machine.
-