- Nov 27, 2020
-
-
Felipe de Azevedo Piovezan authored
Many pages have had their titles renamed over time, causing broken links to spread throughout the documentation. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D92093
-
Georgii Rymar authored
This patch starts emitting the `EShNum` key, when the `e_shnum = 0` and the section header table exists. `e_shnum` might be 0, when the the number of entries in the section header table is larger than or equal to SHN_LORESERVE (0xff00). In this case the real number of entries in the section header table is held in the `sh_size` member of the initial entry in section header table. Currently, obj2yaml crashes when an object has `e_shoff != 0` and the `sh_size` member of the initial entry in section header table is `0`. This patch fixes it. Differential revision: https://reviews.llvm.org/D92098
-
Marek Kurdej authored
-
Georgii Rymar authored
The following line asserts when `sh_addralign > MAX_UINT32 && (uint32_t)sh_addralign == 0`: ``` ExpectedOffset = alignTo(ExpectedOffset, SecHdr.sh_addralign ? SecHdr.sh_addralign : 1); ``` it happens because `sh_addralign` is truncated to 32-bit value, but `alignTo` doesn't accept `Align == 0`. We should change `1` to `1uLL`. Differential revision: https://reviews.llvm.org/D92163 -
David Green authored
-
Sjoerd Meijer authored
This adds LLVM_DEBUG messages to dump the (intermediate) tree cost calculations, which is useful to trace and see how the final cost is calculated.
-
Simon Pilgrim authored
[DAG] Legalize umin(x,y) -> sub(x,usubsat(x,y)) and umax(x,y) -> add(x,usubsat(y,x)) iff usubsat is legal If usubsat() is legal, this is likely to result in smaller codegen expansion than the default cmp+select codegen expansion. Allows us to move the x86-specific lowering to the generic expansion code. Differential Revision: https://reviews.llvm.org/D92183
-
Simon Pilgrim authored
Rename prefix from X32 to X86 as we typically use X32 for gnux32 triples
-
Cheng Wang authored
Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D90381
-
Jay Foad authored
As a bonus this makes it (IMO) obvious that the iterator is not invalidated, so remove the comment explaining that.
-
Jay Foad authored
The Direction parameter to AnalysisResolver::getAnalysisIfAvailable has never been documented or used for anything.
-
Raphael Isemann authored
The test case isn't using the AST matchers for all checks as there doesn't seem to be support for matching NonTypeTemplateParmDecl default arguments. Otherwise this is simply importing the default arguments. Reviewed By: martong Differential Revision: https://reviews.llvm.org/D92106
-
Richard Sandiford authored
This patch adds tests for things that happened to be fixed by previous patches, but that should continue working if we do decide to treat sizeless types as incomplete types. Differential Revision: https://reviews.llvm.org/D79584
-
Cullen Rhodes authored
Folding a select of vector constants that include undef elements only applies to fixed vectors, but there's no earlier check the type is not scalable so it crashes for scalable vectors. This adds a check so this optimization is only attempted for fixed vectors. Reviewed By: sdesmalen Differential Revision: https://reviews.llvm.org/D92046
-
Roman Lebedev authored
This was orginally committed in 2245fb8a. but was immediately reverted in f3abd549 because of a PHI handling issue. Original commit message: 1. It doesn't make sense to enforce that the bonus instruction is only used once in it's basic block. What matters is whether those user instructions fit within our budget, sure, but that is another question. 2. It doesn't make sense to enforce that said bonus instructions are only used within their basic block. Perhaps the branch condition isn't using the value computed by said bonus instruction, and said bonus instruction is simply being calculated to be used in successors? So iff we can clone bonus instructions, to lift these restrictions, we just need to carefully update their external uses to use the new cloned instructions. Notably, this transform (even without this change) appears to be poison-unsafe as per alive2, but is otherwise (including the patch) legal. We don't introduce any new PHI nodes, but only "move" the instructions around, i'm not really seeing much potential for extra cost modelling for the transform, especially since now we allow at most one such bonus instruction by default. This causes the fold to fire +11.4% more (13216 -> 14725) as of vanilla llvm test-suite + RawSpeed. The motivational pattern is IEEE-754-2008 Binary16->Binary32 extension code: https://github.com/darktable-org/rawspeed/blob/ca57d77fb2ba81f21fc712cfac26e54f46406473/src/librawspeed/common/FloatingPoint.h#L115-L120 ^ that should be a switch, but it is not now: https://godbolt.org/z/bvja5v That being said, even thought this seemed like this would fix it: https://godbolt.org/z/xGq3TM apparently that fold is happening somewhere else afterall, so something else also has a similar 'artificial' restriction.
-
Roman Lebedev authored
This is the problematic pattern i didn't think of, that lead to revert of 2245fb8a in f3abd549.
-
Frederik Gossen authored
Overcome the assumption that parallel loops are only nested in other parallel loops. Differential Revision: https://reviews.llvm.org/D92188
-
Christian Sigg authored
The ops are very similar to the std variants, but support async GPU execution. gpu.alloc does not currently support an alignment attribute, and the new ops do not have canonicalizers/folders like their std siblings do. Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D91698
-
Wang, Pengfei authored
Scalar intrinsics x86.avx512*_cvt* have an extra rounding mode operand. We can directly ignore it to reuse the SSE/AVX math. This fix the bug https://bugs.llvm.org/show_bug.cgi?id=48298. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D92206
-
Nicolas Vasilache authored
This adds LLVM triple propagation and updates the test that did not check it properly. Differential Revision: https://reviews.llvm.org/D92182
-
Cheng Wang authored
-
Markus Lavin authored
This reverts commit 06758c6a. Bug: https://bugs.llvm.org/show_bug.cgi?id=48166 Additional discussion in: https://reviews.llvm.org/D91711
-
Georgii Rymar authored
This is related to MIPS. Currently we might report an error and exit, though there is no problem to report a warning and try to continue dumping an object. The code uses `MipsGOTParser<ELFT> Parser`, which is isolated in this method. Differential revision: https://reviews.llvm.org/D92090
-
Craig Topper authored
[RISCV] Don't remove (and X, 0xffffffff) from inputs when matching RISCVISD::DIVUW/REMUW to 64-bit DIVU/REMU. These patterns are using zexti32 which matches either assertzexti32 or (and X, 0xffffffff). But if we match (and X, 0xffffffff) it will remove the AND and the inputs may no longer have the zero bits needed to guarantee the result has enough zeros. This commit changes the patterns to only match assertzexti32. I'm not sure how to test the broken case since the DIVUW/REMUW nodes are created during type legalization, but type legalization won't create an (and X, 0xfffffffff) directly on the inputs. I've also changed the zexti32 on the root of the pattern to just checking for AND. We were previously also matching assertzexti32, but I doubt that pattern would ever occur.
-
Max Kazantsev authored
-
Kazu Hirata authored
-
Max Kazantsev authored
When widening an IndVar that has LCSSA Phi users outside the loop, we can safely widen it as usual and then truncate the result outside the loop without hurting the performance. Differential Revision: https://reviews.llvm.org/D91593 Reviewed By: skatkov
-
Kirill Bobyrev authored
This was originally a part of D71880 but is separated for simplicity and ease of reviewing. Fixes: https://github.com/clangd/clangd/issues/582 Reviewed By: hokein Differential Revision: https://reviews.llvm.org/D91952
-
QingShan Zhang authored
For now, we will hardcode the result as 0.0 if the input is denormal or 0. That will have the impact the precision. As the fsqrt added belong to the cold path of the cmp+branch, it won't impact the performance for normal inputs for PowerPC, but improve the precision if the input is denormal. Reviewed By: Spatel Differential Revision: https://reviews.llvm.org/D80974
-
Kazu Hirata authored
-
Juneyoung Lee authored
This patch adds a description about the newly added poison constant to LangRef. Differential Revision: https://reviews.llvm.org/D92162
-
Craig Topper authored
-
Sam McCall authored
We need a real target now, and it was only being created if grpc was built from source or imported from homebrew. Differential Revision: https://reviews.llvm.org/D92107
-
Nikita Popov authored
Add a flag that disables caching when computing aliasing results potentially based on a phi-phi NoAlias assumption. We'll still insert cache entries temporarily to catch infinite recursion, but will drop them afterwards, so they won't persist in BatchAA. Differential Revision: https://reviews.llvm.org/D91936
-
Arthur Eubanks authored
Also clean it up a bit.
-
Louis Dionne authored
python3-distutils is required to use `import distutils.spawn`, which is required by the ABI list targets.
-
Roman Lebedev authored
Many bots are unhappy, at the very least missed a few codegen tests, and possibly this has a logic hole inducing a miscompile (will be really awesome to have ready reproducer..) Need to investigate. This reverts commit 2245fb8a.
-
Mariusz Ceier authored
Without this change mesa fails while looking for llvm components like amdgpu, engine or native: Run-time dependency LLVM (modules: amdgpu(missing), bitreader, bitwriter, core, engine(missing), executionengine, instcombine, ipo, mcdisassembler, mcjit, native(missing), scalaropts, transformutils, coroutines) Looking for a fallback subproject for the dependency llvm (modules: bitwriter, engine, mcdisassembler, mcjit, core, executionengine, scalaropts, transformutils, instcombine, amdgpu, native, bitreader, ipo) This change adds component groups (like all-targets, engine, native, amdgpu) to the "all" component. Differential Revision: https://reviews.llvm.org/D92158 -
Roman Lebedev authored
1. It doesn't make sense to enforce that the bonus instruction is only used once in it's basic block. What matters is whether those user instructions fit within our budget, sure, but that is another question. 2. It doesn't make sense to enforce that said bonus instructions are only used within their basic block. Perhaps the branch condition isn't using the value computed by said bonus instruction, and said bonus instruction is simply being calculated to be used in successors? So iff we can clone bonus instructions, to lift these restrictions, we just need to carefully update their external uses to use the new cloned instructions. Notably, this transform (even without this change) appears to be poison-unsafe as per alive2, but is otherwise (including the patch) legal. We don't introduce any new PHI nodes, but only "move" the instructions around, i'm not really seeing much potential for extra cost modelling for the transform, especially since now we allow at most one such bonus instruction by default. This causes the fold to fire +11.4% more (13216 -> 14725) as of vanilla llvm test-suite + RawSpeed. The motivational pattern is IEEE-754-2008 Binary16->Binary32 extension code: https://github.com/darktable-org/rawspeed/blob/ca57d77fb2ba81f21fc712cfac26e54f46406473/src/librawspeed/common/FloatingPoint.h#L115-L120 ^ that should be a switch, but it is not now: https://godbolt.org/z/bvja5v That being said, even thought this seemed like this would fix it: https://godbolt.org/z/xGq3TM apparently that fold is happening somewhere else afterall, so something else also has a similar 'artificial' restriction.
-
Roman Lebedev authored
[NFC][SimplifyCFG] Add test coverage for FoldBranchToCommonDest xform with live-out bonus instuctions The uses of the bonus instructions should not be preventing the transformation.
-