- May 06, 2022
-
-
Ben Shi authored
Reviewed By: aykevl, dylanmckay Differential Revision: https://reviews.llvm.org/D125077
-
Max Kazantsev authored
Because this instruction is a noop, we can simply go through it in search of the base.
-
Nico Weber authored
needed by check-hwasan as of 4af9392e
-
Max Kazantsev authored
-
Yitzhak Mandelbaum authored
Adjusts the comment to specify that the output vector's size matches the number of CFG blocks. Differential Revision: https://reviews.llvm.org/D125091
-
Nikita Popov authored
https://reviews.llvm.org/D124075 causes MLIR to no longer build when using make rather than ninja, due to a tablegen-generated header being used before it is created. It seems that this is related to the use of LLVM_ENABLE_OBJLIB when using add_tablgen with a non-Ninja/Xcode generator. In that case an intermediate objlib target is generated. This patch fixes the issue by a) declaring dependencies in add_tablegen for mlir-pdll and b) making sure those dependencies are added to the objlib target. Differential Revision: https://reviews.llvm.org/D125010
-
Aaron Ballman authored
-
Nikita Popov authored
To make it either to extend to the case where the other operand is not a constant.
-
Simon Pilgrim authored
Only pre-SSE41 targets double-pump the fp comparison ops
-
Fraser Cormack authored
D113035 enhanced the matching of bitwise selects from vector types. This change unfortunately introduced crashes as it tries to cast scalable vector types to integers. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D124997
-
Simon Pilgrim authored
Based off the script from D103695 - Jaguar, Bulldozer, Silvermont (et al) and Haswell all have slow BLENDV ops, so adjust the worse case cost values
-
Kiran Chandramohan authored
The OpenMP worksharing loop operation in the dialect is a proper loop operation and not a container of a loop. So we have to lower the parse-tree OpenMP loop construct and the do-loop inside the construct to a omp.wsloop operation and there should not be a fir.do_loop inside it. This is achieved by skipping fir.do_loop creation and calling genFIR for the nested evaluations in the lowering of the do construct. Note: Handling of more clauses, parallel do, storage of loop index variable etc will come in separate patches. Part of the upstreaming effort to move LLVM Flang from fir-dev branch of https://github.com/flang-compiler/f18-llvm-project to the LLVM Project. Reviewed By: peixin Differential Revision: https://reviews.llvm.org/D125024 Co-authored-by:
Sourabh Singh Tomar <SourabhSingh.Tomar@amd.com> Co-authored-by:
Shraiysh Vaishay <Shraiysh.Vaishay@amd.com>
-
Fraser Cormack authored
This test starts failing with the changes in D125021.
-
Nikolas Klauser authored
Reviewed By: var-const, #libc Spies: libcxx-commits, mgorny Differential Revision: https://reviews.llvm.org/D124440
-
Simon Pilgrim authored
Based off the script from D103695, we now mainly use BLENDV or OR(AND,ANDN) to select scalar float/double ops
-
Sam McCall authored
The mechanism behind "check-all" is recording params of add_lit_testsuite() calls in global variables LLVM_LIT_*, and then creating an extra suite with their union at the end. This avoids composing the check-* targets directly, which doesn't work well. We generalize this by allowing multiple families of variables LLVM_{name}_LIT_*: umbrella_lit_testsuite_begin(check-foo) ... test suites here will be added to LLVM_FOO_LIT_* variables ... umbrella_lit_testsuite_end(check-foo) (This also moves some implementation muck out of {llvm,clang}/CMakeLists.txt This patch also changes check-clang-tools to use be an umbrella test target, which means the clangd and clang-pseudo tests are included in it, along with the the other testsuites that already are (like check-clang-extra-clang-tidy). Differential Revision: https://reviews.llvm.org/D121838 -
Simon Pilgrim authored
Based off the script from D103695, on AVX1, Jaguar/Bulldozer both have low throughput for ymm select patterns (BLENDV + OR(AND,ANDN))), and even on AVX2 Haswell still struggles with BLENDV ops
-
Simon Pilgrim authored
-
Sam McCall authored
This includes only the taken branch of conditional sections. The API allows for producing a stream for a particular PP branch, which will be used later for the secondary GLR parses of not-taken branches. Differential Revision: https://reviews.llvm.org/D123243
-
Balazs Benics authored
It seems like multiple users are affected by a crash introduced by this commit, thus I'm reverting it for the time being. Read more about the found reproducers at Phabricator. Differential Revision: https://reviews.llvm.org/D124658 This reverts commit f0d6cb4a.
-
David Spickett authored
[libcxx] Reject month 0 in get_date/__get_month This fixes #47663. Months in dates should be >= 1 and <= 12. We parse up to two digits then minus one, because we want to store this as "months since January" (0-11). However we didn't check that the result of that was not -1. For example if you had (MM/DD/YYYY) 00/21/2022. Added tests for: * Failing if month is 0 * Failing if month is 13 * Allowing a leading zero in month e.g. "01" Note that libc++ and libstdc++ return different values on parsing failure, and MSVC STL returns end of stream instead. Handle the first two by checking for defines, MSVC STL expects these tests to fail for other reasons already: https://github.com/microsoft/STL/blob/main/tests/libcxx/expected_results.txt#L372 so not handling that case here. Reviewed By: #libc, Mordante Differential Revision: https://reviews.llvm.org/D124175
-
John Brawn authored
On targets without unistd.h or sys/wait.h (such as bare metal targets) any test that uses check_assertion.h will fail, so add REQUIRES: has-unix-headers to them and autodetect whether we have these headers or not. These tests currently have unsupported on windows, but that's exactly because windows doesn't have these headers so we can remove the specific check for windows. Differential Revision: https://reviews.llvm.org/D124623
-
David Green authored
If the mask is made up of elements that form a mask in the higher type we can convert shuffle(bitcast into the bitcast type, simplifying the instruction sequence. A v4i32 2,3,0,1 for example can be treated as a 1,0 v2i64 shuffle. This helps clean up some of the AArch64 concat load combines, along with helping simplify a number of other tests. The PowerPC combine for v16i8 splat vector loads needed some fixes to keep it working for v16i8 vectors. This improves the handling of v2i64 shuffles to match too, hopefully improving them in general. Differential Revision: https://reviews.llvm.org/D123801
-
Simon Pilgrim authored
We were only testing basic vector types
-
Matthias Springer authored
Ops that are created during the bufferization were not analyzed (when run with One-Shot Bufferize), and users should instead create memref ops directly. Futhermore, this fixes an issue where an op was erased (and put on the `erasedOps` list), but subsequently a new tensor op was created at the same memory location. This op was then not bufferized. Disallowing the creation of new tensor ops simplifies the bufferization and fixes such issues. Differential Revision: https://reviews.llvm.org/D125017
-
wangpc authored
Enable default outlining when the function has the minsize attribute. `addr-label.ll` crashed after enabling this, so a barrier is added before instruction selection as a workaround. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D122213
-
David Spickett authored
"A-f" -> "A-F"
-
Kiran Chandramohan authored
The FIR `do_loop` is designed as a structured operation with a single block inside it. Presence of unstructured constructs like jumps, exits inside the loop will cause the loop to be marked as unstructured. These loops are lowered using the `control-flow` dialect branch operations. Fortran semantics do not allow the loop variable to be modified inside the loop. To prevent accidental modification, the iteration of the loop is modeled by two variables, trip-count and loop-variable. -> The trip-count and loop-variable are initialized in the pre-header. The trip-count is set as (end-start+step)/step where end, start and step have the usual meanings. The loop-variable is initialized to start. -> The header block contains a conditional branch instruction which selects between branching to the body of the loop or the exit block depending on the value of the trip-count. -> Inside the body, the trip-count is decremented and the loop-variable incremented by the step value. Finally it branches to the header of the loop. Part of the upstreaming effort to move LLVM Flang from fir-dev branch of https://github.com/flang-compiler/f18-llvm-project to the LLVM Project. Reviewed By: awarzynski Differential Revision: https://reviews.llvm.org/D124837 Co-authored-by:
Val Donaldson <vdonaldson@nvidia.com> Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com> Co-authored-by:
Jean Perier <jperier@nvidia.com> Co-authored-by:
Peter Klausler <pklausler@nvidia.com>
-
Florian Hahn authored
After D97756, collectHomogenousInstGraphLoopInvariants may collect conditions for both logical ANDs and logical ORs in case the root is a select that matches both logical AND & OR. This means the function won't return invariant values of either AND/OR chains, but both. This can result in incorrect transformations. See llvm/test/Transforms/SimpleLoopUnswitch/trivial-unswitch-logical-and-or.ll. Without the patch, Alive2 rejects the modified tests with: Source and target don't have the same return domain. Note that this also applies to the test case added in D97756 (@test_partial_condition_unswitch_or_select). We can't unswitch on %cond6, because the graph leading to it contains and AND and an OR. This only fixes trivial unswitching for now, but a similar problem likely exists with non-trivial unswitching. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D124526 -
Andrzej Warzynski authored
This patch adds support for `-save-temps` in `flang-new`, Flang's compiler driver. The semantics of this option are inherited from Clang. The file extension for temporary Fortran preprocessed files is set to `i`. This is identical to what Clang uses for C (or C++) preprocessed files. I have tried researching what other compilers do here, but I couldn't find any definitive answers. One GFortran thread [1] suggests that indeed it is not clear what the right approach should be. Normally, various phases in Clang/Flang are combined. The `-save-temps` option works by forcing the compiler to run every phase separately. As there is no integrated assembler driver in Flang, user will have to use `-save-temps` together with `-fno-integrated-as`. Otherwise, an invocation to the integrated assembler would be generated generated, which is going to fail (i.e. something equivalent to `clang -cc1as` from Clang). There are no specific plans for implementing an integrated assembler for Flang for now. One possible solution would be to share it entirely with Clang. Note that on Windows you will get the following error when using `-fno-integrated-as`: ```bash flang-new: error: there is no external assembler that can be used on this platform ``` Unfortunately, I don't have access to a Windows machine to investigate this. Instead, I marked the tests in this patch as unsupported on Windows. [1] https://gcc.gnu.org/bugzilla//show_bug.cgi?id=81615 Differential Revision: https://reviews.llvm.org/D124669
-
Matthias Springer authored
Buffers with undefined contents (e.g., the result of an init_tensor) are no longer copied. Differential Revision: https://reviews.llvm.org/D125015
-
Matthias Springer authored
This follows the same implementation strategy as scf::ForOp and common functionality is extracted into helper functions. This implementation works well in cases where each yielded value (from either body/condition region) is equivalent to the corresponding bbArg of the parent block. In that case, each OpResult of the loop may be aliasing with the corresponding OpOperand of the loop (and with no other OpOperand). In the absence of said equivalence relationship, new buffer copies must be inserted, so that the aliasing OpOperand/OpResult contract of scf::WhileOp is honored. In essence, by yielding a newly allocated buffer, we can enforce the specified may-alias relationship. (Newly allocated buffers cannot alias with any OpOperands of the loop.) Differential Revision: https://reviews.llvm.org/D124929
-
Luo, Yuanke authored
-
Diana Picus authored
This seems to be the consensus in https://github.com/flang-compiler/f18-llvm-project/issues/1316 The patch adds ExternalNameConversion to the default FIR CodeGen pass pipeline, right before the FIRtoLLVM pass. It also adds a flag to optionally disable it, and sets it in `tco`. In other words, `flang-new` and `flang-new -fc1` will both run the pass by default, whereas `tco` will not, so none of the tests need to be updated. Differential Revision: https://reviews.llvm.org/D121171
-
Sam McCall authored
As confirmation, running this locally found 2 crashes: - trivial: crashes on file with no tokens - lexer: hits an assertion failure on bytes: 0x5c,0xa,0x5c,0x1,0x65,0x5c,0xa Differential Revision: https://reviews.llvm.org/D125037
-
Marco Elver authored
Factor our InstrumentationIRBuilder and share it between ThreadSanitizer and SanitizerCoverage. Simplify its usage at the same time (use function of passed Instruction or BasicBlock). This class may be used in other instrumentation passes in future. NFCI. Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D125038
-
David Green authored
This patch adds a combine to attempt to reduce the costs of certain select-shuffle patterns. The form of code it attempts to detect is: %x = shuffle ... %y = shuffle ... %a = binop %x, %y %b = binop %x, %y shuffle %a, %b, selectmask A classic select-mask will pick items from each lane of a or b. These do not always have a great lowering on many architectures. This patch attempts to pack a and b into the lower elements, creating a differently ordered shuffle for reconstructing the orignal which may be better than the select mask. This can be better for performance, especially if less elements of a and b need to be computed and the input shuffles are cheaper. Because select-masks are just one form of shuffle, we generalize to any mask. So long as the backend has decent costmodel for the shuffles, this can generally improve things when they come up. For more basic cost models the folds do not appear to be profitable, not getting past the cost checks. Differential Revision: https://reviews.llvm.org/D123911
-
Martin Storsjö authored
Adding a mingw based config is easy in the current CI environment (where we can just choose the different target by calling `i686-w64-mingw32-clang`), while adding a clang-cl based config would require setting up different environment variables pointing to the i386 library directory. Just adding one config (DLL) instead of exhaustively testing both (DLL and static) as very few tests would differ in practice, to keep the CI load reasonable. Differential Revision: https://reviews.llvm.org/D124991
-
Sam McCall authored
BumpPtrAllocator::Allocate() is marked __attribute__((returns_nonnull)) when the compiler supports it, which makes it UB to return null. When there have been no allocations yet, the current slab is [nullptr, nullptr). A zero-sized allocation fits in this range, and so Allocate(0, 1) returns null. There's no explicit docs whether Allocate(0) is valid. I think we have to assume that it is: - the implementation tries to support it (e.g. >= tests instead of >) - malloc(0) is allowed - requiring each callsite to do a check is bug-prone - I found real LLVM code that makes zero-sized allocations Differential Revision: https://reviews.llvm.org/D125040
-
Sam McCall authored
It turns out clang::expandUCNs only works on tokens that contain valid UCNs and no other random escapes, and clang only uses it on raw_identifiers. Currently we can hit an assertion by creating tokens with stray non-valid-UCN backslashes in them. Fortunately, expanding UCNs in raw_identifiers is actually all we need. Most tokens (keywords, punctuation) can't have them. UCNs in literals can be treated as escape sequences like \n even this isn't the standard's interpretation. This more or less matches how clang works. (See https://isocpp.org/files/papers/P2194R0.pdf which points out that the standard's description of how UCNs work is misaligned with real implementations) Differential Revision: https://reviews.llvm.org/D125049
-