- Oct 23, 2023
-
-
Bjorn Pettersson authored
This is a follow up to commit 8511ade5 which removed legacy PM support in LowerExpectIntrinsicPass. That commit made some of the includes redundant.
-
Utkarsh Saxena authored
`S.getScopeForContext` determins the **active** scope associated with the given `declContext`. This fails to find the matching `operator!=` if candidate `operator==` was found via ADL since that scope is not active. Instead, just directly lookup using the namespace decl of `operator==` Fixes #68901
-
Aaron Ballman authored
-
Aaron Ballman authored
Post-commit feedback (https://reviews.llvm.org/D156565#4654773) found that the changes in 84a3aadf caused us to diagnose use of VLAs in C89 mode by default which was an unintended change. This adds -Wvla-cxx-extension as a warning group and adds the C++- specific warnings to it while leaving the C warnings under -Wvla-extension. -Wvla-cxx-extension is then added to -Wall.
-
Igor Kirillov authored
This patch allows the generation of SVE code with masks that mimic Neon.
-
Benjamin Maxwell authored
This allows folding extracts from `vector.create_mask` ops that have a known value. Currently, there's no fold for this, but you get the same effect from the unrolling in LowerVectorMask (part of -convert-vector-to-llvm), then folds after that. However, for a future patch, this simplification needs to be done before lowering to LLVM, hence the need for this fold. E.g.: ``` %0 = vector.create_mask %c1, %dimA, %dimB : vector<1x[4]x[4]xi1> %1 = vector.extract %mask[0] : vector<[4]x[4]xi1> ``` -> ``` %0 = vector.create_mask %dimA, %dimB : vector<[4]x[4]xi1> ```
-
Ilya Biryukov authored
Make it a strict weak order. Fixes #64121. Current implementation uses the definition of ordering from the C++ Standard. The definition provides only a partial order and cannot be used in sorting algorithms. The debug builds of libc++ are capable of detecting that problem and this failure was found when building Clang with libc++ and those extra checks enabled, see #64121. The new ordering is a strict weak order and still pushes most interesting functions to the start of the list. In some cases, it leads to better results, e.g. ``` struct Foo { operator int(); operator const char*(); }; void test() { Foo() - Foo(); } ``` Now produces a list with two most relevant builtin operators at the top, i.e. `operator-(int, int)` and `operator-(const char*, const char*)`. Previously `operator-(const char*, const char*)` was the first element, but `operator-(int, int)` was only the 13th element in the output. This is a consequence of `stable_sort` now being able to compare those two candidates, which are indistinguishable in the semantic partial order despite being two local minimums in their respective comparable subsets. However, new implementation does not take into account some aspects of C++ semantics, e.g. which function template is more specialized. This can also lead to worse ordering sometimes. Reviewed By: #clang-language-wg, aaron.ballman Differential Revision: https://reviews.llvm.org/D159351 -
Benjamin Kramer authored
c312f025 worked around a bug here, but that has since been fixed in 47747da6
-
Kerry McLaughlin authored
visitCallInst already looks for fixed width vector extracts where number of elements in the source and destination types are equal. This patch modifies the function to also identify scalable extracts which can be removed.
-
ZhangYin authored
-
Boian Petkantchin authored
scf.forall.parallel_insert_slice -> tensor.parallel_insert_slice add -> linalg.add map -> linalg.map matmul -> linalg.matmul
-
Sarthak Gupta authored
Fixes https://github.com/llvm/llvm-project/issues/67760 `levelCheckRank` ensures that the tensors for tosa operations are not unranked During tosa validation in `levelCheckRank`, we were trying to get the rank of a tensor without checking if it is ranked or unranked, which leads to an `assert` error. I see two ways to fix this: - Only check `type.getRank() > tosa_level.MAX_RANK` if the tensor is ranked, and then proceed as usual. (like `if (type.hasRank() && type.getRank() > tosa_level.MAX_RANK)` , OR - Throw an error for unranked tensors as result.
-
Mariya Podchishchaeva authored
Skip anonymous members when rebuilding `DesignatedInitExpr` since designated inits for them are meant to be created during `InitListChecker::CheckDesignatedInitializer` routine. Fixes https://github.com/llvm/llvm-project/issues/65143
-
Hans Wennborg authored
-
jeanPerier authored
Intrinsic analysis in semantics reorder the actual arguments so that they match the dummy order. This was not done for MIN/MAX because they are special: these are the only intrinsics with a variadic number of arguments. This caused bugs in lowering that only check the optionality of actual arguments from the third position (since A1 and A2 are mandatory). Update semantics to place A1/A2 first. This also allow removing some checks that were specific to MIN/MAX. There is no point in sorting/placing the rest of the arguments which would be tedious and tricky because of the variadic aspect.
-
Job Noorman authored
In #67707, the minimum function alignment on RISC-V was set to 4. When RVC (compressed instructions) is enabled, the minimum alignment can be reduced to 2. This patch implements this by delegating the choice of minimum alignment to a new `MCPlusBuilder::getMinFunctionAlignment` function. This way, the target-dependent code in `BinaryFunction` is minimized.
-
Nikita Popov authored
Use BatchAA with EarliestEscapeInfo instead of callCapturesBefore() in MemDepAnalysis. The advantage of this is that it will also take not-captured-before information into account for non-calls (see test_store_before_capture for a representative example), and that this is a cached analysis. The disadvantage is that EII is slightly less precise than full CapturedBefore analysis. In practice the impact is positive, with gvn.NumGVNLoad going from 22022 to 22808 on test-suite. The impact to compile-time is also positive, mainly in the ThinLTO configuration.
-
jeanPerier authored
https://reviews.llvm.org/D143819 turned mismatched DATA substring from a crash into a warning. However, the resulting DATA was incorrect when there were subsequent DATA values after the mismatch because the DATA to init conversion stopped. This change let the DATA to init continue. I added a LengthMismatch tag instead of using SizeMismatch because the other situation where SizeMismatch is returned seem like bug situations to me (the DATA value is dropped and the offset is not advanced), so I did not want to continue DATA processing in these cases.
-
Aviad Cohen authored
-
Job Noorman authored
We used to hard-code target features for RISC-V. However, most features (with the exception of relax) are stored in the object file. This patch extracts those features to ensure BOLT's output doesn't use any features not present in the input file.
-
martinboehme authored
It's only used in its own unit tests.
-
Owen Pan authored
-
Adrian Kuegel authored
-
Adrian Kuegel authored
-
Craig Topper authored
RISCVGenDAGISel.inc can call this before it checks the node type. Ensure the type is scalar before wasting time to do the more computationally expensive checks. This also avoids an assertion if we hit a VMV_X_S instruction which doesn't have a VL operand which vectorPseudoHasAllNBitUsers expects.
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Wang Pengcheng authored
MultiClass argument is not used any more since aa843266. Besides, for maintainability, we should put the implementation of qualifying name in one place (that is `QualifyName` function), so `Scoper` is removed and we use `IsMC` to indicate that we are in a multiclass.
-
Wang Pengcheng authored
To match GCC's behaviors. Fixes #67596
-
Chuanqi Xu authored
[C++20] [Modules] [Driver] Don't enable -fdelayed-template-parsing by default on windows with C++20 (#69431) There are already 3 issues about the broken state of -fdelayed-template-parsing and C++20 modules: - https://github.com/llvm/llvm-project/issues/61068 - https://github.com/llvm/llvm-project/issues/64810 - https://github.com/llvm/llvm-project/issues/65027 The problem is more complex than I thought. I am not sure how to fix it properly now. Given the complexities and -fdelayed-template-parsing is actually an extension to support old MS codes, I think it may make sense to not enable the -fdelayed-template-parsing option by default with C++20 modules to give more user friendly experience. Users who still want -fdelayed-template-parsing can specify it explicitly. Also according to https://learn.microsoft.com/en-us/cpp/build/reference/permissive-standards-conformance?view=msvc-170, MSVC actually defaults to -fno-delayed-template-parsing (/Zc:twoPhase- with MSVC CLI) if using C++20. So we match the behavior with MSVC here to not enable -fdelayed-template-parsing by default after C++20.
-
Johannes Doerfert authored
-
Matthias Springer authored
Add a new attribute `bufferization.manual_deallocation` that can be attached to allocation and deallocation ops. Buffers that are allocated with this attribute are assigned an ownership of "false". Such buffers can be deallocated manually (e.g., with `memref.dealloc`) if the deallocation op also has the attribute set. Previously, the ownership-based buffer deallocation pass used to reject IR with existing deallocation ops. This is no longer the case if such ops have this new attribute. This change is useful for the sparse compiler, which currently deallocates the sparse tensor buffers by itself.
-
Hongtao Yu authored
Tweaking warnings more to avoid flooding user log.
-
Fangrui Song authored
When linking an executable with a slightly larger executable, ld.lld --call-graph-profile-sort=cdsort can be very slow (see #68638). ``` 4.6% 20.7Mi .text.hot 3.5% 15.9Mi .text 3.4% 15.2Mi .text.unknown ``` Add cl option `cdsort-max-chain-size`, which is similar to `ext-tsp-max-chain-size`, and set it to 128, to improve performance. In `ld.lld @response.txt --threads=4 --call-graph-profile-sort=cdsort --time-trace" builds, the "Total Sort sections" time is measured as follows: * -mllvm -cdsort-max-chain-size=64: 1.321813 * -mllvm -cdsort-max-chain-size=128: 2.030425 * -mllvm -cdsort-max-chain-size=256: 2.927684 * -mllvm -cdsort-max-chain-size=512: 5.493106 * unlimited: 9 minutes The rest part takes 6.8s.
-
Fangrui Song authored
Using the legacy PM for the optimization pipeline was deprecated in 13.0.0. Many legacy passes have been removed in 2022.
-
cmtice authored
lldb/test/Shell/Breakpoint/breakpoint-command.test adds a python command, to be executed when a breakpoint hits, that writes out a number. It then runs, hits the breakpoint and checks that the number is present exactly once. The problem is that on some systems the test can be run in a filepath that happens to contain the number (e.g. auto-generated directory names). The number is then detected multiple times and the test fails. This patch fixes the issue by using a string instead, particularly a string with spaces, which is very unlikely to be auto-generated by any system.
-
Kazu Hirata authored
Identified with misc-include-cleaner.
-
Justin Fargnoli authored
[mlir][DeadCodeAnalysis] Don't Require `RegionBranchTerminatorOpInterface` in `visitRegionTerminator()` (#69043) Fix for a crash reported in #64975. The crash occurs in the cast located [here](https://github.com/llvm/llvm-project/blob/ece5dd101c7e4dc2fd23428abd312f75fd3d3eaf/mlir/lib/Analysis/DataFlow/DeadCodeAnalysis.cpp#L262) because `llvm.unreachable` doesn't implement `RegionBranchTerminatorOpInterface`. The crash is caused by `DeadCodeAnalysis` assuming that `isa<RegionBranchOpInterface>(op->getParentOp())` implies `isa<RegionBranchTerminatorOpInterface>(op)` in `DeadCodeAnalysis::visit()`. This patch tried to fix this by enabling the analysis to proceed regardless of whether `op` is a `RegionBranchTerminatorOpInterface`.
-
Jonas Hahnfeld authored
The pass only requires that it can determine a uniquely identified target at some offsets. Multiple relocations at the same offset are fine otherwise and will be required when adding exception handling support for RISC-V.
-