- Jun 15, 2020
-
-
Matt Arsenault authored
-
Matt Arsenault authored
I tried to use an IR inline asm test, but that doesn't work since the inline asm handling asserts without an MVT to use.
-
Simon Pilgrim authored
This would help reduce XMM->GPR traffic for some reduction cases.
-
- Jun 14, 2020
-
-
Qiu Chaofan authored
This patch adds handling of constrained FP intrinsics about round, truncate and extend for PowerPC target, with necessary tests. Reviewed By: steven.zhang Differential Revision: https://reviews.llvm.org/D64193
-
Qiu Chaofan authored
On PowerPC, we have vnmsubfp Altivec instruction for fnmsub operation on v4f32 type. Default pattern for this instruction never works since we don't have legal fneg for v4f32 when VSX disabled. Reviewed By: steven.zhang Differential Revision: https://reviews.llvm.org/D80617
-
Qiu Chaofan authored
Current implementation of division estimation isn't correct for some cases like 1.0/0.0 (result is nan, not expected inf). And this change exposes a potential infinite loop: we use isConstOrConstSplatFP in combineRepeatedFPDivisors to look up if the divisor is some constant. But it doesn't work after legalized on some platforms. This patch restricts the method to act before LegalDAG. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D80542
-
Sanjay Patel authored
As noted in D80236 - the early-cse pass was included here before: D75145 / rG71a31688 But it got moved outside of the "extra" option there, then it got dropped while adjusting -vector-combine: rG6438ea45 rG57bb4787 So this is restoring the behavior and adding a test to prevent accidental changes again. I don't see an equivalent option for the new pass manager.
-
Joachim Protze authored
This patch allows to specify a prefix (default:empty) to be included into print-out written by callback.h. Also adding a cmake target to find the header file from other tests. Reviewed by: jdoerfert Differential Revision: https://reviews.llvm.org/D76008
-
Joachim Protze authored
Reviewed by: AndreyChurbanov Differential Revision: https://reviews.llvm.org/D81497
-
Nikita Popov authored
This class uses a mix of different indentation levels, normalize it.
-
Nikita Popov authored
When LVI is performing assume intersections, it also checks for llvm.experimental.guard intrinsics. To avoid unnecessary block scans, it first checks whether this intrinsic is declared in the module at all. I've noticed that we end up spending quite a lot of time looking up that function again and again... Avoid this by only looking it up once when LazyValueInfo is constructed. This of course assumes that we don't introduce new guard intrinsics (which is the case for all existing uses of LVI -- and even if it weren't, it would not introduce miscompiles, just potentially lose optimization power.) Differential Revision: https://reviews.llvm.org/D81796
-
David Green authored
This adds additional cast cpst tests useful for MVE, notably around half types.
-
Sanjay Patel authored
(a[0] + a[1] + a[2] + a[3]) - (b[0] + b[1] + b[2] +b[3]) --> (a[0] - b[0]) + (a[1] - b[1]) + (a[2] - b[2]) + (a[3] - b[3]) This should be the last step in solving PR43953: https://bugs.llvm.org/show_bug.cgi?id=43953 We started emitting reduction intrinsics with: D80867/ rGe50059f6 So it's a relatively easy pattern match now to re-order those ops. Also, I have not seen any complaints for the switch to intrinsics yet, so I'll propose to remove the "experimental" tag from the intrinsics soon. Differential Revision: https://reviews.llvm.org/D81491
-
Sanjay Patel authored
This is a hacky, but low-risk fix to avoid the infinite loop in PR46271: https://bugs.llvm.org/show_bug.cgi?id=46271 As discussed there, the problem is that FoldOpIntoSelect() can get into a conflict with a transform that wants to pull a 'not' op through min/max via SimplifyDemandedVectorElts(). We need to relax our matching of min/max to include undefined elements in vector constants to avoid that. Alternatively, we could improve or cripple the demanded elements analysis, but that could create even more problems. The likely better, safer alternative will be to create min/max intrinsics, so we can remove all of the hacks related to min/max matching in instcombine. Differential Revision: https://reviews.llvm.org/D81698
-
Simon Pilgrim authored
Even without PTEST, we can still efficiently perform an OR reduction as PMOVMSKB(PCMPEQB(X,0)) == 0, avoiding xmm->gpr extractions.
-
Uday Bondhugula authored
Add a few more commonly used ops and missing keywords.
-
njames93 authored
-
Simon Pilgrim authored
Ensure codegen is still reasonable - ideally we'd make use of MOVMSK for this.
-
Xing GUO authored
[NFC] mv llvm/test/tools/obj2yaml/macho-DWARF-debug-ranges.yaml llvm/test/ObjectYAML/MachO/DWARF-debug_ranges.yaml
-
Xing GUO authored
This patch adds a new field `bool Is64bit` in `DWARFYAML::Data` to indicate the address size of target. It's helpful for inferring the `AddrSize` in some DWARF sections. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D81709
-
Fangrui Song authored
[IteratedDominanceFrontier] Decrease number of SmallPtrSet::insert and delete unneeded SmallVector::clear Also, fix the argument name to be consistent with the declaration.
-
Craig Topper authored
[X86] Add mayLoad flag to FARCALL*m/FARJMP memory instrutions. Add 'm' to the end of FARJMP64/FARCALL64 instruction names. We never codegen them so this doesn't matter in practice. But sometimes someone comes along and tries to use these flags for something else. LIke the Load Value Inject inline assembly handling.
-
Craig Topper authored
Previously, the X86AsmParser would issue a warning whenever a ret instruction is encountered. This patch changes the behavior to automatically transform each ret instruction in an inline assembly stream into: shlq $0, (%rsp) lfence ret which is secure, according to https://software.intel.com/security-software-guidance/insights/deep-dive-load-value-injection#specialinstructions. Patch by Scott Constable with some minor changes by Craig Topper.
-
Craig Topper authored
If the input to the bitcast is a sign bit test, it makes sense to directly use vpmovmskb or vmovmskps/pd. This removes the need to copy the sign bits to a k-register and then to a GPR. Fixes PR46200. Differential Revision: https://reviews.llvm.org/D81327
-
Craig Topper authored
[X86] Move -x86-use-vzeroupper command line flag into runOnMachineFunction for the pass itself rather than the pass pipeline construction This pass has no dependencies on other passes so conditionally including it in the pipeline doens't do much. Just move it the pass itself to keep it isolated.
-
Roman Lebedev authored
-
Vladimir Vereschaka authored
This reverts commit 3ea9450b. The commit fails the remote library tests on the toolchain builders: http://lab.llvm.org:8011/builders/llvm-clang-win-x-armv7l http://lab.llvm.org:8011/builders/llvm-clang-win-x-aarch64
-
Florian Hahn authored
isOverwrite expects the later location as first argument and the earlier result later. The adjusted call is intended to check whether CC overwrites DefLoc.
-
Craig Topper authored
A lot of what EVEX->VEX does is equivalent to what the prioritization in the assembly parser does. When an AVX mnemonic is used without any EVEX features or XMM16-31, the parser will pick the VEX encoding. Since codegen doesn't go through the parser, we should also use VEX instructions when we can so that the code coming out of integrated assembler matches what you'd get from outputing an assembly listing and parsing it. The pass early outs if AVX isn't enabled and uses TSFlags to check for EVEX instructions before doing the more costly table lookups. Hopefully that's enough to keep this from impacting -O0 compile times.
-
Craig Topper authored
relocImm was a complexPattern that handled both ConstantSDNode and X86Wrapper. But it was only applied selectively because using it would cause patterns to be not importable into FastISel or GlobalISel. So it only got applied to flag setting instructions, stores, RMW arithmetic instructions, and rotates. Most of the test changes are a result of making patterns available to GlobalISel or FastISel. The absolute-cmp.ll change is due to this fixing a pattern ordering issue to make an absolute symbol match to an 8-bit immediate before trying a 32-bit immediate. I tried to use PatFrags to reduce the repetition, but I was getting errors from TableGen.
-
- Jun 13, 2020
-
-
Amanieu d'Antras authored
Summary: Bugzilla: https://bugs.llvm.org/show_bug.cgi?id=46060 I've also added the Extra_IsConvergent flag which was missing from FastISel. Reviewers: echristo Reviewed By: echristo Subscribers: hiraditya, llvm-commits Tags: #llvm Differential Revision: https://reviews.llvm.org/D80759
-
Xing GUO authored
-
Bruno Ricci authored
This saves sizeof(void *) bytes per LambdaExpr. Review-after-commit since this is a straightforward change similar to the work done on other nodes. NFC.
-
mydeveloperday authored
Summary: This patch fixes bug #44192 When clang-format is run with option AllowShortBlocksOnASingleLine, it is expected to either succeed in putting the short block with its control statement on a single line or fail and leave the block as is. When brace wrapping after control statement is activated, if the block + the control statement length is superior to column limit but the block alone is not, clang-format puts the block in two lines: one for the control statement and one for the block. This patch removes this unexpected behaviour. Current unittests are updated to check for this behaviour. Patch By: Bouska Reviewed By: MyDeveloperDay Differential Revision: https://reviews.llvm.org/D71512
-
Bruno Ricci authored
This test illustrate the bug fixed in D81787.
-
Bruno Ricci authored
...as done. This is a NAD which has always been implemented correctly.
-
Bruno Ricci authored
...lambda-expression) as done. They have been allowed since at least clang 3.3.
-
Xing GUO authored
-