- Oct 26, 2020
-
-
Craig Topper authored
-
Sanjay Patel authored
This is a modified 2nd try of 22d10b8a (reverted by 1c837169 because it managed to expose an existing crashing bug that should be fixed by 74ffc823 ). Original commit message: This is similar in spirit to 01ea93d8 (memcpy) except that here the underlying caller assumptions were created for vectorizer use (throughput) rather than other passes. That meant targets could have an enormous throughput cost with no corresponding size, latency, or blended cost increase. The ARM costs show a small difference between throughput and size because there's an underlying difference in cmp/sel costs that is also predicated on cost-kind. Paraphrasing from the previous commits: This may not make sense for some callers, but at least now the costs will be consistently wrong instead of mysteriously wrong. Targets should provide better overrides if the current modeling is not accurate.
-
Sanjay Patel authored
I'm not sure if/how this ever worked, but it must not be tested currently because the basic tests added here were crashing as noted in the post-review comments for 1c837169 (which reverted another cost-model fix in 22d10b8a).
-
Nikita Popov authored
Same change as 0dda6333, but for mul expressions. We want to first fold any constant operans and then strengthen the nowrap flags, as we can compute more precise flags at that point.
-
Aaron Puchert authored
The constructor of Project asserts that the contained ValueDecl is not null, use that in the ThreadSafetyAnalyzer. In the case of LiteralPtr it's the other way around. Also dyn_cast<> is sufficient if we know something isn't null.
-
Aaron Puchert authored
Instead of just mutex members we also consider mutex globals. Unsurprisingly they are always in scope. Now the paper [1] says that > The scope of a class member is assumed to be its enclosing class, > while the scope of a global variable is the translation unit in > which it is defined. But I don't think we should limit this to TUs where a definition is available - a declaration is enough to acquire the mutex, and if a mutex is really limited in scope to a translation unit, it should probably be only declared there. The previous attempt in 9dcc82f3 was causing false positives because I wrongly assumed that LiteralPtrs were always globals, which they are not. This should be fixed now. [1] https://static.googleusercontent.com/media/research.google.com/en/us/pubs/archive/42958.pdf Fixes PR46354. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D84604
-
Nikita Popov authored
Establish parity with the handling of add expressions, by always constant folding mul expression operands before checking the depth limit (this is a non-recursive simplification). The code was already unconditionally constant folding the case where all operands were constants, but was not folding multiple constant operands together if there were also non-constant operands. This requires picking out a different demonstration for depth-based folding differences in the limit-depth.ll test.
-
Nikita Popov authored
Separate out the code handling constant folding into a separate block, that is independent of other folds that need a constant first operand. Also make some minor adjustments to make the constant folding look nearly identical to the same code in getAddExpr(). The only reason this change is not strictly NFC is that the C1*(C2+V) fold is moved below the constant folding, which means that it now also applies to C1*C2*(C3+V), as it should.
-
Nikita Popov authored
We should first try to constant fold the add expression and only strengthen nowrap flags afterwards. This allows us to determine stronger flags if e.g. only two operands are left after constant folding (and thus "guaranteed no wrap region" code applies) or the resulting operands are non-negative and thus nsw->nuw strengthening applies.
-
Nikita Popov authored
Also run the test case through -instnamer.
-
- Oct 25, 2020
-
-
Sanjay Patel authored
This extends D78430 to solve cases like: https://llvm.org/PR47858 There are still missed opportunities shown in the tests, and as noted in the earlier patches, we have related functionality in InstCombine, so we may want to extend other folds in a similar way. A semi-random sampling of test diff proofs in this patch: https://rise4fun.com/Alive/sS4C
-
Sanjay Patel authored
One variant of this is shown in: https://llvm.org/PR47858
-
Melanie Blower authored
Correct LIT test failure detected on buildbot after mibintc committed rG2e204e23: [clang] Enable support for #pragma STDC FENV_ACCESS D87528
-
Florian Hahn authored
This patch adds an additional set of tests that can be vectorized efficiently on AArch64, using CMxx & BFI.
-
Simon Pilgrim authored
-
Melanie Blower authored
Reviewers: rjmccall, rsmith, sepavloff Differential Revision: https://reviews.llvm.org/D87528
-
Simon Pilgrim authored
I'm not certain InstCombinerImpl::matchBSwapOrBitReverse needs to filter the or(op0(),op1()) ops - there are just too many cases that recognizeBSwapOrBitReverseIdiom/collectBitParts handle now (and quickly).
-
Simon Pilgrim authored
Currently InstCombinerImpl::matchBSwapOrBitReverse won't match starting from funnel shifts.
-
Richard Smith authored
-
Craig Topper authored
It's required to be a constant and can never be in a register so make it explicit.
-
Martin Storsjö authored
This reverts commit 22d10b8a. This broke compilation e.g. like this: $ cat synth.c *a; float *b; c() { for (;;) { float d = -*b * *a++; d -= *--b * *a++; d -= *--b * *a; d -= *--b * *a; e(d); } } $ clang -target x86_64-linux-gnu -c -O2 -ffast-math synth.c clang: ../include/llvm/Support/Casting.h:104: static bool llvm::isa_impl _cl<To, const From*>::doit(const From*) [with To = llvm::PointerType; Fr om = llvm::Type]: Assertion `Val && "isa<> used on a null pointer"' fail ed.
-
Teresa Johnson authored
Disable the part of this test that started failing only on the llvm-avr-linux bot after 5c20d7db. Unfortunately, "XFAIL: avr" does not work. Still in the process of trying to figure out how to debug.
-
Richard Smith authored
to disallowed objects or have non-constant destruction.
-
Nathan Ridge authored
TestWorkspace allows easily writing tests involving multiple files that can have inclusion relationships between them. BackgroundIndexTest.RelationsMultiFile is refactored to use TestWorkspace, and moved to FileIndexTest as it no longer depends on BackgroundIndex. Differential Revision: https://reviews.llvm.org/D89297
-
Arthur Eubanks authored
-
Fangrui Song authored
While MC did not produce R_X86_64_GOTPCRELX for test/binop instructions (movl/adcl/addl/andl/...) before the previous commit, this code path has been exercised by -fno-integrated-as for GNU as since 2016: -no-pie relaxing may incorrectly access loc[-3] and produce a corrupted instruction. Simply handle test/binop R_X86_64_GOTPCRELX like R_X86_64_GOTPCREL.
-
Fangrui Song authored
[X86] Produce R_X86_64_GOTPCRELX for test/binop instructions (MOV32rm/TEST32rm/...) when -Wa,-mrelax-relocations=yes is enabled We have been producing R_X86_64_REX_GOTPCRELX (MOV64rm/TEST64rm/...) and R_X86_64_GOTPCRELX for CALL64m/JMP64m without the REX prefix since 2016 (to be consistent with GNU as), but not for MOV32rm/TEST32rm/...
-
Drew Fisher authored
While some platforms call `AsanThread::Init()` from the context of the thread being started, others (like Fuchsia) call `AsanThread::Init()` from the context of the thread spawning a child. Since `AsyncSignalSafeLazyInitFakeStack` writes to a thread-local, we need to avoid calling it from the spawning thread on Fuchsia. Skipping the call here on Fuchsia is fine; it'll get called from the new thread lazily on first attempted access. Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D89607
-
Drew Fisher authored
When enabling stack use-after-free detection, we discovered that we read the thread ID on the main thread while it is still set to 2^24-1. This patch moves our call to AsanThread::Init() out of CreateAsanThread, so that we can call SetCurrentThread first on the main thread. Reviewed By: mcgrathr Differential Revision: https://reviews.llvm.org/D89606
-
Fangrui Song authored
-
Nico Weber authored
This reverts commit fa66bcf4. Seems to break tests, see https://reviews.llvm.org/D89827#2351930
-
Sanjay Patel authored
This is similar in spirit to 01ea93d8 (memcpy) except that here the underlying caller assumptions were created for vectorizer use (throughput) rather than other passes. That meant targets could have an enormous throughput cost with no corresponding size, latency, or blended cost increase. The ARM costs show a small difference between throughput and size because there's an underlying difference in cmp/sel costs that is also predicated on cost-kind. Paraphrasing from the previous commits: This may not make sense for some callers, but at least now the costs will be consistently wrong instead of mysteriously wrong. Targets should provide better overrides if the current modeling is not accurate.
-
Benjamin Kramer authored
No scheduling, no autodetection.
-
Benjamin Kramer authored
No scheduling, no autodetection. Just enough so -march=znver3 works.
-
Benjamin Kramer authored
-
dfukalov authored
1. Throughput and codesize costs estimations was separated and updated. 2. Updated fdiv cost estimation for different cases. 3. Added scalarization processing for types that are treated as !isSimple() to improve codesize estimation in getArithmeticInstrCost() and getArithmeticInstrCost(). The code was borrowed from TCK_RecipThroughput path of base implementation. Next step is unify scalarization part in base class that is currently works for TCK_RecipThroughput path only. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D89973
-
David Green authored
-
Andrzej Warzynski authored
Without this change LIT tests for Flang fail with: ``` TypeError: append() takes exactly one argument (2 given) ```
-
- Oct 24, 2020
-
-
Stefan Gränitz authored
Root cause of the test failure was fixed with: [JITLink][ELF] PCRel32GOTLoad edge offset can be smaller three This reverts commit 10b1a61b.
-
Stefan Gränitz authored
Offset is 2 for MOVL instruction in test ELF_x86-64_common. This should fix the test failures. Differential Revision: https://reviews.llvm.org/D89795
-