- Jul 11, 2023
-
-
Ties Stuij authored
Currently in LowerConstantFP, when we compile for execute-only (XO) we don't check what architecture we're compiling for (v6m=< or >v6m). We shouldn't get here for v6m, so put in an assert. Reviewed By: simonwallis2, dmgreen Differential Revision: https://reviews.llvm.org/D154506
-
Simon Pilgrim authored
Extend coverage for lowering wide vector types during type legalization to allow us to use PACKSS/PACKUS patterns instead of dropping down to shuffle lowering. First step towards avoiding premature folds of TRUNCATE to PACKSS/PACKUS nodes as described on Issue #63710 - which causes a large number of regressions on D152928 - we will next need to tweak the TRUNCATE widening in ReplaceNodeResults Differential Revision: https://reviews.llvm.org/D154592
-
Lorenzo Chelini authored
If we deal with statically known tensors and tiles and a given tile perfectly divides a given dimension, we can omit the padding attribute. As a bonus point, we can now run pack and unpack propagation (currently, we bail out during propagation if we have the padding attribute). Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D154607
-
pvanhout authored
Depends on D153757 NOTE: This would land iff D153757 (RFC) lands too. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153861
-
pvanhout authored
Only a few minor test changes needed because I removed the "helper" suffix from the combiner name, as it's not really a helper anymore but more like the implementation itself. Depends on D153757 NOTE: This would land iff D153757 (RFC) lands too. Reviewed By: aemerson Differential Revision: https://reviews.llvm.org/D153850
-
pvanhout authored
Use the new matchtable-based combiner backend for all AMDGPU combiners. This drop-in from the user's perspective; there are no test changes, the new combiner behaves exactly like the old one. Depends on D153757 NOTE: This would land iff D153757 (RFC) lands too. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153758
-
Timm Bäder authored
-
Matthias Springer authored
Remove patterns that fold tensor subset ops into vector transfer ops from the vector dialect. These patterns already exist in the tensor dialect. Differential Revision: https://reviews.llvm.org/D154932
-
Jay Foad authored
-
Jim Lin authored
-
Chuanqi Xu authored
Close https://github.com/llvm/llvm-project/issues/63544. Background: We landed std modules in libcxx recently but we haven't landed the corresponding in-tree tests. According to @Mordante, there are only 1% libcxx tests failing with std modules. And the major blocking issue is the lambda expression in the require clauses. The root cause of the issue is that previously we never consider any lambda expression as the same. Per [temp.over.link]p5: > Two lambda-expressions are never considered equivalent. I thought this is an oversight at first but @rsmith explains that in the wording, the program is as if there is only a single definition, and a single lambda-expression. So we don't need worry about this in the spec. The explanation makes sense. But it didn't reflect to the implementation directly. Here is a cycle in the implementation. If we want to merge two definitions, we need to make sure its implementation are the same. But according to the explanation above, we need to judge if two lambda-expression are the same by looking at its parent definitions. So here is the problem. To solve the problem, I think we have to profile the lambda expressions actually to get the accurate information. But we can't do this universally. So in this patch I tried to modify the interface of `Stmt::Profile` and only profile the lambda expression during the process of merging the constraint expressions. Differential Revision: https://reviews.llvm.org/D153957
-
Nabeel Omer authored
Fixes #63692. In reference to volatile memory accesses, the langref says: > the backend should never split or merge target-legal volatile load/store instructions. Differential Revision: https://reviews.llvm.org/D154609
-
Maksim Panchenko authored
Add missing EOL in a warning message. Reviewed By: rafauler Differential Revision: https://reviews.llvm.org/D154895
-
Chuanqi Xu authored
This comes from https://reviews.llvm.org/D153003 By @rsmith, the test case is valid since: > Per [temp.type]/1.4 (http://eel.is/c++draft/temp.type#1.4), > >> Two template-ids are the same if [...] their corresponding template >> template-arguments refer to the same template. > so B<A> and B<NS::A> are the same type. The stricter "same sequence of > tokens" rule doesn't apply here, because using-declarations are not > definitions. > we should either (preferably) be including only the syntactic form of > the base specifier (because local syntax is what the ODR covers), or > the canonical type (which should be the same for both > using-declarations). Here we adopt the second suggested solutions. Reviewed By: cor3ntin, v.g.vassilev Differential Revision: https://reviews.llvm.org/D154324
-
Zi Xuan Wu (Zeson) authored
[RISCV] Don't fold RISCVISD::VMV_V_X_VL series node and scalar load to vector load when scalar load is update load We try to fold RISCVISD::VMV_V_X_VL series node + scalar load -> vector load. But if scalar load is indexed load (load update form), it's not profitable to fold because load update node can't be removed after fold. Differential Revision: https://reviews.llvm.org/D152222
-
pvanhout authored
Adds a new backend to power the GISel Combiners using the InstructionSelector's match tables. This does not depend on any of the data structures created for the current combiner and is intended to replace it entirely. See the RFC for more details: https://discourse.llvm.org/t/rfc-matchtable-based-globalisel-combiners/71457/6 Note: this would replace D141135. Reviewed By: aemerson, arsenm Differential Revision: https://reviews.llvm.org/D153757
-
pvanhout authored
Move all of the reusable logic out of `GlobalISelEmitter.cpp` into a `GlobalISelMatchTableExecutorEmitter` class so the future combiner backend can use it as well. Depends on D153755 Reviewed By: aemerson Differential Revision: https://reviews.llvm.org/D153756
-
pvanhout authored
Makes `InstructionSelector.h`/`InstructionSelectorImpl.h` generic so the match tables can also be used for the combiner. Some notes: - Coverage was made an optional parameter of `executeMatchTable`, combines won't use it for now. - `GIPFP_` -> `GICXXPred_` so it's more generic. Those are just C++ predicates and aren't PatFrag-specific. - Pass the MatcherState directly to testMIPredicate_MI, the combiner will need it. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153755
-
Petr Hosek authored
On Linux crt is typically use in combination with builtins. In the Clang driver the use of builtins and crt is controlled by the --rtlib option. Both builtins and crt also have similar build requirements where they need to be built before any other runtimes and must avoid dependencies. We also want builtins and crt these to be buildable separately from the rest of compiler-rt for bootstrapping purposes. Given how simple crt is, rather than maintaining a separate directory with its own separate build setup, it's more efficient to just move crt into builtins. We still use separate CMake option to control whether to built crt same as before. This is an alternative to D89492 and D136664. Differential Revision: https://reviews.llvm.org/D153989
-
Job Noorman authored
BOLT used `ToolOutputFile::keep` to make sure the intermediary object file was written to disk for debugging purposes when `--keep-tmp` was passed. However, since and intermediary `buffer_ostream` was used to stream to, and this class only writes to its output stream in its destructor, the object file was lost whenever its destructor wouldn't run. This could happen, for example, if there is a crash while linking. This patch makes sure the object file is written to disk immediately after we're done creating it. This is very useful while debugging JITLink crashes. This patch also gets rid of creating a temporary file when `--keep-tmp` is not passed by streaming the object file directly to a `SmallString`. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D154826
-
Tobias Gysi authored
A distinct attribute associates a referenced attribute with a unique identifier. Every call to its create function allocates a new distinct attribute instance. The address of the attribute instance temporarily serves as its unique identifier. Similar to the names of SSA values, the final unique identifiers are generated during pretty printing. Examples: #distinct = distinct[0]<42.0 : f32> #distinct1 = distinct[1]<42.0 : f32> #distinct2 = distinct[2]<array<i32: 10, 42>> This mechanism is meant to generate attributes with a unique identifier, which can be used to mark groups of operations that share a common properties such as if they are aliasing. The design of the distinct attribute ensures minimal memory footprint per distinct attribute since it only contains a reference to another attribute. All distinct attributes are stored outside of the storage uniquer in a thread local store that is part of the context. It uses one bump pointer allocator per thread to e...
-
Haojian Wu authored
Similar to the https://reviews.llvm.org/D86048 (it only sets the bit for C++ code), we propagate the contains-errors bit for C-code path. Fixes https://github.com/llvm/llvm-project/issues/50236 Fixes https://github.com/llvm/llvm-project/issues/50243 Fixes https://github.com/llvm/llvm-project/issues/48636 Fixes https://github.com/llvm/llvm-project/issues/50320 Differential Revision: https://reviews.llvm.org/D154861
-
Balazs Benics authored
When we construct a `NonParamVarRegion`, we canonicalize the decl to always use the same entity for consistency. At the moment that is the canonical decl - which is the first decl in the redecl chain. However, this can cause problems with tentative declarations and extern declarations if we declare an array with unknown bounds. Consider this C example: https://godbolt.org/z/Kdvr11EqY ```lang=C typedef typeof(sizeof(int)) size_t; size_t clang_analyzer_getExtent(const void *p); void clang_analyzer_dump(size_t n); extern const unsigned char extern_redecl[]; const unsigned char extern_redecl[] = { 1,2,3,4 }; const unsigned char tentative_redecl[]; const unsigned char tentative_redecl[] = { 1,2,3,4 }; const unsigned char direct_decl[] = { 1,2,3,4 }; void test_redeclaration_extent(void) { clang_analyzer_dump(clang_analyzer_getExtent(direct_decl)); // 4 clang_analyzer_dump(clang_analyzer_getExtent(extern_redecl)); // shoul...
-
Yeting Kuo authored
The patch uses a way similiar to vp.load/store and consider the mask popcount as the effetive vector length. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D151713
-
Martin Braenne authored
Reviewed By: sammccall Differential Revision: https://reviews.llvm.org/D154834
-
Manish Kausik H authored
If modified attribute list is invalid, reverting the change is a low-cost maintainence solution as compared to examples like [this](https://github.com/llvm/llvm-project/blob/main/llvm/tools/bugpoint/CrashDebugger.cpp#L368). This will ensure that the ListReducer maintains the sanctity of any new attribute dependencies added in the future/already present. Reviewed By: modocache Differential Revision: https://reviews.llvm.org/D154348
-
wangpc authored
The code to resolve class/multiclass arguments are similar, we extract them to `resolveArguments`s to simplify code. Reviewed By: tra, reames Differential Revision: https://reviews.llvm.org/D154065
-
wangpc authored
For `CSR_Interrupt`, we can generate the register list via a single `sequence`. For `CSR_XLEN_F32_Interrupt` and `CSR_XLEN_F64_Interrupt`, I don't see the reason why we need to keep the order the same as how we used to allocate registers (and we have changed the order in D146488), so I fold them into one `sequence`. There are some *.ll changes because of the order change. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D154837
-
Piyou Chen authored
Reviewed By: asb Differential Revision: https://reviews.llvm.org/D154690
-
Caroline Tice authored
In two calls to ReadMemory in DWARFExpression.cpp, the buffer size passed to ReadMemory is not checked and can be bigger than the actual size of the buffer. This caused a buffer overflow bug, which we found through Address Sanitizer. This patch fixes the problem by checking the address size when it is first read out of the DWARF, and setting an error and returning immediatley if the size is invalid. This is the second attempt to fix this issue; I reverted the first one, as it was not quite correct. Differential Revision: https://reviews.llvm.org/D154907
-
Brad Smith authored
Have ToolChain::IsIntegratedAssemblerDefault default to true. Almost all of the ToolChains are using IAS nowadays. There are a few exceptions like XCore, some NaCl archs, and NVPTX/XCore in Generic_GCC::IsIntegratedAssemblerDefault. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D154902
-
Rahul Kayaith authored
Since op `Attribute`s are automatically downcasted on access, these mappings aren't necessary anymore. Instead we just always generate the getters/setters for attributes even if there isn't a `PythonAttr` mapping. depends on D154462 Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D154468
-
Rahul Kayaith authored
Update remaining `PyAttribute`-returning APIs to return `MlirAttribute` instead, so that they go through the downcasting mechanism. Reviewed By: makslevental Differential Revision: https://reviews.llvm.org/D154462
-
Matt Arsenault authored
We need to replace the other uses of the call chain with the new load chain. Fixes not preserving the return def with unused x86_fp80 results. Regression reported here: https://reviews.llvm.org/rGb15bf305ca3e9ce63aaef7247d32fb3a75174531#1224999
-
Ashay Rane authored
Older versions of clang (for example, v12) throw an error when compiling CStringChecker.cpp that the initializers for `SourceArgExpr`, `DestinationArgExpr`, and `SizeArgExpr` are missing braces around initialization of subobject. Newer clang versions don't throw this error. This patch adds the initialization braces to satisfy clang. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D154871
-
Peiming Liu authored
Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D154912
-
Wang Rui authored
This revision explicitly specifies the machine instruction properties instead of relying on guesswork. This is because guessing instruction properties has proven to be inaccurate, such as the machine LICM not working: ``` void func(char *a, char *b) { int i; for (i = 0; i != 72526; i++) a[i] = b[i]; } ``` Guessing instruction properties: ``` func: # @func move $a2, $zero .LBB0_1: # =>This Inner Loop Header: Depth=1 ldx.b $a3, $a1, $a2 stx.b $a3, $a0, $a2 addi.d $a2, $a2, 1 lu12i.w $a3, 17 ori $a3, $a3, 2894 bne $a2, $a3, .LBB0_1 ret .Lfunc_end0: ``` Explicitly specify instruction properties: ``` func: # @func lu12i.w $a2, 17 ori $a2, $a2, 2894 move $a3, $zero .LBB0_1: # =>This Inner Loop Header: Depth=1 ldx.b $a4, $a1, $a3 stx.b $a4, $a0, $a3 addi.d $a3, $a3, 1 bne $a3, $a2, .LBB0_1 ret .Lfunc_end0: ``` Reviewed By: SixWeining, xen0n Differential Revision: https://reviews.llvm.org/D154192 -
Kai Sasaki authored
As TOSA does not support the tensor with zero dimensions, we can check the zero value for the static shape input. Ideally, we should be able to check the tensor shape more broadly, such as using `CPred` in the TOSA type definition. But based on [[ https://discourse.llvm.org/t/where-can-we-put-the-shared-verification-among-multiple-dialect-ops/71806 | the discussion here ]] It makes input type verification complicated and hard to maintain and still only applies to the case the input is statically shaped. Therefore, in this change, we have put the zero dimension check in the verification of each op, which would be flexible and maintainable. See: https://github.com/llvm/llvm-project/issues/63212 Reviewed By: eric-k256 Differential Revision: https://reviews.llvm.org/D154569
-
Teresa Johnson authored
Previously the MemProf profile was expected to be in the same profile file as a normal PGO profile, passed via the usual -fprofile-use= option, and was matched in the same pass. To simplify profile preparation, since the raw MemProf profile requires the binary for symbolization and may be simpler to index separately from the raw PGO profile, and also to enable providing a MemProf profile for a SamplePGO build, separate out the MemProf feedback option and matching pass. This patch adds the -fmemory-profile-use=${file} option, and the provided file is passed down to LLVM and ultimately used in a new MemProfUsePass which performs the matching of just the memory profile contents of that file. Note that a single profile file containing both normal PGO and MemProf profile data is still supported, and the relevant profile data is matched by the appropriate matching pass(es) based on which option(s) the profile is provided with (the same profile file can be supplied to both feedback options). Differential Revision: https://reviews.llvm.org/D154856