- Apr 30, 2021
-
-
Roman Lebedev authored
The profitability check is: we don't want to create more than a single PHI per instruction sunk. We need to create the PHI unless we'll sink all of it's would-be incoming values. But there is a caveat there. This profitability check doesn't converge on the first iteration! If we first decide that we want to sink 10 instructions, but then determine that 5'th one is unprofitable to sink, that may result in us not sinking some instructions that resulted in determining that some other instruction we've determined to be profitable to sink becoming unprofitable. So we need to iterate until we converge, as in determine that all leftover instructions are profitable to sink. But, the direct approach of just re-iterating seems dumb, because in the worst case we'd find that the last instruction is unprofitable, which would result in revisiting instructions many many times. Instead, i think we can get away with just two passes - forward and backward. However then it isn't obvious what is the most performant way to update InstructionsToSink.
-
Sam Clegg authored
Unlike the existing `--export` option this will not causes errors or warnings if the specified symbol is not defined. See: https://github.com/emscripten-core/emscripten/issues/13736 Differential Revision: https://reviews.llvm.org/D99887
-
Mark de Wever authored
While working on D70631, Microsoft's unit tests discovered an issue. Our `std::to_chars` implementation for bases != 10 uses the range `[first,last)` as temporary buffer. This violates the contract for to_chars: [charconv.to.chars]/1 http://eel.is/c++draft/charconv#to.chars-1 `to_chars_result to_chars(char* first, char* last, see below value, int base = 10);` "If the member ec of the return value is such that the value is equal to the value of a value-initialized errc, the conversion was successful and the member ptr is the one-past-the-end pointer of the characters written." Our implementation modifies the range `[member ptr, last)`, which causes Microsoft's test to fail. Their test verifies the buffer `[member ptr, last)` is unchanged. (The test is only done when the conversion is successful.) While looking at the code I noticed the performance for bases != 10 also is suboptimal. This is tracked in D97705. This patch fixes the issue and adds a benchmark. ...
-
Petr Hosek authored
We overrite CXX_FLAGS to enable relative vtables, but doing so overwrites generic Fuchsia CXX_FLAGS leading to a build failure on Windows. Differential Revision: https://reviews.llvm.org/D101551
-
Raphael Isemann authored
Right now to get the 'NSSet *` pointer value we first derefence it and then take the address of the result. Beside being inefficient this potentially can cause an infinite recursion if the `pointer` value we get is a pointer of a type that the TypeSystem can't derefence. If the pointer is for example some form of `void *` that the dynamic type resolution can't resolve to an actual type, then the `Derefence` call goes back to asking the formatters how to reference it. If the NSSet formatter then checks if it's an NSSet variation under the hood then we just end infinitely often recursion. In practice this seems to happen with some form of Builtin.RawPointer we get from a NSDictionary in Swift. FWIW, no other formatter is doing the same deref->addressOf as here and there doesn't seem to be any specific reason to do so in the git history (it's just part of the initial formatter commit) Reviewed By: JDevlieghere Differential Revision: https://reviews.llvm.org/D101537
-
LLVM GN Syncbot authored
-
Benjamin Kramer authored
This reverts commit 3b8ec86f. Revert "[X86] Refine AMX fast register allocation" This reverts commit c3f95e91. This pass breaks using LLVM in a multi-threaded environment by introducing global state.
-
Vitaly Buka authored
This reverts commit 7ad4dee3.
-
Sanjay Patel authored
-
Martin Storsjö authored
As the libcxx tests link with -nostdlib, libraries that normally are added by default by the compiler driver has to be added manually. The "oldnames" library is automatically added when driving linking with clang-cl. When linking with the plain clang driver, as the libcxx tests do, the clang driver does the same but only since Clang 12.0). But when linking with -nostdlib, like the libcxx tests do, the driver defaults aren't added at all, and we need to specify the defaults manually. This allows removing a TODO from the Windows CI setup; it turns out that upgrading to Clang 12.0 didn't help here as expected, sorry about that mixup. Differential Revision: https://reviews.llvm.org/D101434
-
Vitaly Buka authored
Attribute guaranties safe static initialization of globals. Differential Revision: https://reviews.llvm.org/D101514
-
Craig Topper authored
[RISCV] Teach DAG combine to fold (and (select_cc lhs, rhs, cc, -1, c), x) -> (select_cc lhs, rhs, cc, x, (and, x, c)) Similar for or/xor with 0 in place of -1. This is the canonical form produced by InstCombine for something like `c ? x & y : x;` Since we have to use control flow to expand select we'll usually end up with a mv in basic block. By folding this we may be able to pull the and/or/xor into the block instead and avoid a mv instruction. The code here is based on code from ARM that uses this to create predicated instructions. I'm doing it on SELECT_CC so it happens late, but we could do it on select earlier which is what ARM does. I'm not sure if we lose any combine opportunities if we do it earlier. I left out add and sub because this can separate sext.w from the add/sub. It also made a conditional i64 addition/subtraction on RV32 worse. I guess both of those would be fixed by doing this earlier on select. The select-binop-identity.ll test has not been commited yet, but I made the diff show the changes to it. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D101485
-
Craig Topper authored
-
Alexey Bataev authored
Added an extra analysis for better choosing of shuffle kind in getShuffleCost functions for better cost estimation if mask was provided. Differential Revision: https://reviews.llvm.org/D100865
-
Fangrui Song authored
-
Fangrui Song authored
Add tests which can catch the issue in 0ce723cb (If any function needs CFISection::EH, the module should use CFISection::EH). Reviewed By: echristo Differential Revision: https://reviews.llvm.org/D101339
-
Sanjay Patel authored
-
Sanjay Patel authored
We should handle other cases (undef/poison), so reduce the duplication of repeated switches.
-
Sanjay Patel authored
-
- Apr 29, 2021
-
-
Anirudh Prasad authored
- Currently, the "." (Dot) character, when not identifying an Identifier or a Constant, refers to the current PC (Program Counter) - However, in z/OS, for the HLASM dialect, it strictly accepts only the "*" as the current PC (Support for this will be put up in a follow-up patch) - The changes in this patch allow individual platforms to choose whether they would like to use the "." (Dot) character as a marker for the current PC or not. - It is achieved by introducing a new field in MCAsmInfo.h called `DotIsPC` (similar to `DollarIsPC`) Reviewed By: abhina.sreeskantharajan Differential Revision: https://reviews.llvm.org/D100975
-
Fangrui Song authored
GNU ld -r can create .rela.eh_frame with unordered r_offset values. (With LLD, we can craft such a case by reordering sections in .eh_frame.) This is currently unsupported and will trigger `assert(pieces[i].inputOff <= off ...` in `OffsetGetter::get` (the content is corrupted in a -DLLVM_ENABLE_ASSERTIONS=off build). This patch supports this case. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D101116
-
Craig Topper authored
This replaces D98479. This allows type legalization to form SPLAT_VECTOR_PARTS so we don't lose the splattedness when the scalar type is split. I'm handling SPLAT_VECTOR_PARTS for fixed vectors separately so we can continue using non-VL nodes for scalable vectors. I limited to RV32+vXi64 because DAGCombiner::visitBUILD_VECTOR likes to form SPLAT_VECTOR before seeing if it can replace the BUILD_VECTOR with other operations. Especially interesting is a splat BUILD_VECTOR of the extract_vector_elt which can become a splat shuffle, but won't if we form SPLAT_VECTOR first. We either need to reorder visitBUILD_VECTOR or add visitSPLAT_VECTOR. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D100803
-
Craig Topper authored
This seems like a reasonable upper bound on VL. WG discussions for the V spec would probably allow us to use 2^16 as an upper bound on VLEN, but this is good enough for now. This allows us to remove sext and zext if user happens to assign the size_t result into an int and then uses it as a VL intrinsic argument which is size_t. Reviewed By: frasercrmck, rogfer01, arcbbb Differential Revision: https://reviews.llvm.org/D101472
-
Sander de Smalen authored
Temporarily reverting this patch due to some unexpected issue found by one of the PPC buildbots. This reverts commit 584e9b6e.
-
Jay Foad authored
-
Chirag Khandelwal authored
This patch is child of D89671, contains the clang implementation to use the OpenMP IRBuilder's section construct. Co-author: @anchu-rajendran Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D91054
-
David Zarzycki authored
-
Anastasia Stulova authored
Differential Revision: https://reviews.llvm.org/D101092
-
Chirag Khandelwal authored
This patch adds section support in the OpenMP IRBuilder module, along with a test for the same. Reviewed By: fghanim Differential Revision: https://reviews.llvm.org/D89671
-
Anastasia Stulova authored
This extension is primarily targeting SPIR-V compilations flow as the IR translation is the same between 1.x and 2.x atomics. Differential Revision: https://reviews.llvm.org/D101089
-
Florian Hahn authored
As suggested in D99294, this adds a getVPSingleValue helper to use for recipes that are guaranteed to define a single value. This replaces uses of getVPValue() which used to default to I = 0.
-
Arnamoy Bhattacharyya authored
-
Alex Zinenko authored
-
Bradley Smith authored
At the intrinsic layer the sve.insr operation takes a scalar. When this scalar is an integer we are forcing a data transition between GPRs and ZPRs that is potentially costly. Often the integer scalar is the result of a vector extract, when performing a reduction for example. In such cases we should keep all data within the ZPRs. Co-authored-by:
Paul Walker <paul.walker@arm.com> Differential Revision: https://reviews.llvm.org/D101169
-
Bradley Smith authored
By converting the SVE intrinsic to a normal LLVM insertelement we give the code generator a better chance to remove transitions between GPRs and VPRs Co-authored-by:
Paul Walker <paul.walker@arm.com> Depends on D101302 Differential Revision: https://reviews.llvm.org/D101167
-
Bradley Smith authored
As part of this the ptrue coalescing done in SVEIntrinsicOpts has been modified to not introduce redundant converts, since the convert removal will no longer run after that optimisation to clean up. Differential Revision: https://reviews.llvm.org/D101302
-
Alex Zinenko authored
This enables to express more complex parallel loops in the affine framework, for example, in cases of tiling by sizes not dividing loop trip counts perfectly or inner wavefront parallelism, among others. One can't use affine.max/min and supply values to the nested loop bounds since the results of such affine.max/min operations aren't valid symbols. Making them valid symbols isn't an option since they would introduce selection trees into memref subscript arithmetic as an unintended and undesired consequence. Also add support for converting such loops to SCF. Drop some API that isn't used in the core repo from AffineParallelOp since its semantics becomes ambiguous in presence of max/min bounds. Loop normalization is currently unavailable for such loops. Depends On D101171 Reviewed By: bondhugula Differential Revision: https://reviews.llvm.org/D101172
-
Alex Zinenko authored
Introduce a basic support for parallelizing affine loops with reductions expressed using iteration arguments. Affine parallelism detector now has a flag to assume such reductions are parallel. The transformation handles a subset of parallel reductions that are can be expressed using affine.parallel: integer/float addition and multiplication. This requires to detect the reduction operation since affine.parallel only supports a fixed set of reduction operators. Reviewed By: chelini, kumasento, bondhugula Differential Revision: https://reviews.llvm.org/D101171
-
Lorenzo Chelini authored
-
Nathan Sidwell authored
This libstc++ hack isn't ready for removal. Updating the comment to note what I found. While I have not proven Ville's __is_throw_swappable patch made this go away, that patch did remove the use of noexcept(noexcept(swap(....))). I'm not sure when gcc grew deferred noexcept parsing. Differential Revision: https://reviews.llvm.org/D101441
-