- Sep 20, 2021
-
-
Vladislav Vinogradov authored
The discussion on forum: https://llvm.discourse.group/t/bug-in-partial-dialect-conversion/4115 The `applyPartialConversion` didn't handle the operations, that were marked as illegal inside dynamic legality callback. Instead of reporting error, if such operation was not converted to legal set, the method just added it to `unconvertedSet` in the same way as unknown operations. This patch fixes that and handle dynamically illegal operations as well. The patch includes 2 fixes for existing passes: * `tensor-bufferize` - explicitly mark `std.return` as legal. * `convert-parallel-loops-to-gpu` - ugly fix with marking visited operations to avoid recursive legality checks. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D108505
-
Vladislav Vinogradov authored
Reviewed By: lattner, ftynse Differential Revision: https://reviews.llvm.org/D109223
-
Pavel Labath authored
These have been here since r215992, guarding the calls to HostInfo, but their purpose unclear -- HostInfoLinux provides these functions and they work fine.
-
Siva Chandra Reddy authored
-
Siva Chandra Reddy authored
-
Max Kazantsev authored
All transforms of IndVars have prerequisite requirement of LCSSA and LoopSimplify form and rely on it. Added test that shows that this actually stands.
-
Max Kazantsev authored
This reverts commit 6fec6552. The patch was reverted on incorrect claim that this patch may break LCSSA form when the loop is not in a simplify form. All IndVars' transform insure that the loop is in simplify and LCSSA form, so if it wasn't broken before this transform, it will also not be broken after it.
-
Siva Chandra Reddy authored
Reviewed By: michaelrj Differential Revision: https://reviews.llvm.org/D109952
-
Max Kazantsev authored
There is a piece of logic that uses the fact that signed and unsigned versions of the same predicate are equivalent when both values are non-negative. It's also true when both of them are negative. Differential Revision: https://reviews.llvm.org/D109957 Reviewed By: nikic
-
David Blaikie authored
-
David Blaikie authored
This should probably be rendered as "std::nullptr_t" but for now clang uses the unqualified name (which is ambiguous with possible user defined name in the global namespace), so match that here.
-
David Blaikie authored
-
David Blaikie authored
-
Simon Pilgrim authored
Helps with the regression noted on D109065 - don't truncate a broadcast source if the source has multiple uses.
-
Kazu Hirata authored
-
Chris Jackson authored
The scev-based salvaging for LSR can sometimes produce unnecessarily verbose expressions. This patch adds logic to detect when the value to be recovered and the induction variable differ by only a constant offset. Then, the expression to derive the current iteration count can be omitted from the dbg.value in favour of the offset. Reviewed by: aprantl Differential Revision: https://reviews.llvm.org/D109044
-
David Blaikie authored
llvm-dwarfdump: Pretty printing types including a space between const and parenthesized references/pointers to arrays
-
Craig Topper authored
Unlike psadbw, mpsadbw is not commutable because of how it operates on blocks. We already marked as not commutable for MachineIR, but had it commutable for the tablegened isel patterns. Fixes PR51908.
-
Craig Topper authored
-
David Blaikie authored
-
David Blaikie authored
DWARFDie: Improve type printing for function and array types - with qualifiers (cv/reference) and pointers to them
-
Simon Pilgrim authored
Both ports are required in most cases. Update the uops counts + port usage based off the most recent llvm-exegesis captures (PR36895) and what Intel AoM / Agner / InstLatX64 reports as well. Noticed while trying to improve fp costs for vectorization via the D103695 helper script.
-
Simon Pilgrim authored
We can combine unary shuffles into either of SHUFPS's inputs and adjust the shuffle mask accordingly. Unlike general shuffle combining, we can be more aggressive and handle multiuse cases as we're not going to accidentally create additional shuffles.
-
David Blaikie authored
Move most type tests to a pre-generated assembly file to make it easier to add more weird cases without having to hand craft more DWARF. Move the novel array types that aren't reachable via clang-generated DWARF to a separate file for easy maintenance.
-
- Sep 19, 2021
-
-
Simon Pilgrim authored
Based off a mixture of llvm-exegesis captures (PR36895) and Intel AoM / Agner / InstLatX64 reports.
-
Roman Lebedev authored
[X86][TLI] SimplifyDemandedVectorEltsForTargetNode(): don't break apart broadcasts from which not just the 0'th elt is demanded Apparently this has no test coverage before D108382, but D108382 itself shows a few regressions that this fixes. It doesn't seem worthwhile breaking apart broadcasts, assuming we want the broadcasted value to be preset in several elements, not just the 0'th one. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D108411
-
Roman Lebedev authored
[X86] lowerShuffleAsDecomposedShuffleMerge(): if both inputs are broadcastable/identities, canonicalize broadcasts as such Split off from D108253. Broadcast is simpler than any other shuffle we might produce to do what we want to do here, so prefer it. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D108382
-
Roman Lebedev authored
-
Roman Lebedev authored
[X86] combineX86ShufflesRecursively(): call SimplifyMultipleUseDemandedVectorElts() on after finishing recursing This was suggested in https://reviews.llvm.org/D108382#inline-1039018, and it avoids regressions in that patch. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D109065
-
Sanjay Patel authored
If we transform these, we have to propagate no-wrap/undef carefully.
-
David Green authored
These were apparently missing, having no pattern that could convert a VGETLANEu of a v4f16 to an i32. Added bf16 whilst here, following the same code.
-
xndcn authored
1. Add missing indent in CondBranchOp 2. Remove indent in block label Differential Revision: https://reviews.llvm.org/D109805
-
Simon Pilgrim authored
Both ports are required, for reg and mem variants - we can also use the WriteFComX class directly and remove the unnecessary InstRW overrides. Matches what Intel AoM / Agner / InstLatX64 report as well.
-
Sylvestre Ledru authored
-
Ben Shi authored
Optimize (add (shl x, c0), (shl y, c1)) -> (SLLI (SH*ADD x, y), c1), if c0-c1 == 1/2/3. Reviewed By: craig.topper, luismarques Differential Revision: https://reviews.llvm.org/D108916 -
David Blaikie authored
-
David Blaikie authored
-
Roman Lebedev authored
Should avoid some regressions in D109065 Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D109989
-
Nikita Popov authored
Missed this one in 80110aaf. This is another test mixing up alias scopes and alias scope lists.
-
Nikita Popov authored
Mostly this fixes cases where !noalias or !alias.scope were passed a scope rather than a scope list. In some cases I opted to drop the metadata entirely instead, because it is not really relevant to the test.
-