- Oct 10, 2020
-
-
Krzysztof Parzyszek authored
The selection of HVX shuffles can produce more nodes in the DAG, which need special handling, or otherwise they would be left unselected by the main selection code. Make the handling of such nodes more general.
-
Christian Sigg authored
See https://llvm.discourse.group/t/rfc-new-dialect-for-modelling-asynchronous-execution-at-a-higher-level/1345 Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D88954
-
Nicolas Vasilache authored
This revision belongs to a series of patches that reduce reliance of Linalg transformations on templated rewrite and conversion patterns. Instead, this uses a MatchAnyTag pattern for the vast majority of cases and dispatches internally. Differential revision: https://reviews.llvm.org/D89133
-
Nicolas Vasilache authored
This revision belongs to a series of patches that reduce reliance of Linalg transformations on templated rewrite and conversion patterns. Instead, this uses a MatchAnyTag pattern for the vast majority of cases and dispatches internally. Differential Revision: https://reviews.llvm.org/D89133
-
Nicolas Vasilache authored
This reverts commit 0a34492f. This change turned out to be very intrusive wrt some internal projects. Reverting until this can be sorted out.
-
Arthur Eubanks authored
This pass hasn't been touched in a long time and isn't used in tree.
-
Vy Nguyen authored
Make use of the newly added thread-properties API (available since 31). Differential Revision: https://reviews.llvm.org/D85927 -
Mircea Trofin authored
-
Stella Laurenzo authored
-
Stella Laurenzo authored
* Isolates the visibility controlled parts of its implementation to a detail namespace. * Applies a struct level visibility attribute which applies to the static local within the get() functions. * The prior version was not emitting a symbol for the static local "instance" fields when the user TU was compiled with -fvisibility=hidden. Differential Revision: https://reviews.llvm.org/D89153
-
Scott Linder authored
There doesn't seem to be a direct test of this, and I'm planning to make future changes which will affect it. I'm not particularly familiar with the blocks extension, so suggestions for better tests are welcome. Differential Revision: https://reviews.llvm.org/D88754
-
Craig Topper authored
[X86] When expanding LCMPXCHG16B_NO_RBX in EmitInstrWithCustomInserter, directly copy address operands instead of going through X86AddressMode. I suspect getAddressFromInstr and addFullAddress are not handling all addresses cases properly based on a report from MaskRay. So just copy the operands directly. This should be more efficient anyway.
-
Craig Topper authored
The expansion code creates a copy to RBX before the real LCMPXCHG16B. It's possible this copy uses a register that is also used by the real LCMPXCHG16B. If we set the kill flag on the use in the copy, then we'll fail the machine verifier on the use on the LCMPXCHG16B. Differential Revision: https://reviews.llvm.org/D89151
-
Nikita Popov authored
-
Louis Dionne authored
-
Louis Dionne authored
To make it clearer this is about whether the library supports the debug mode at all, not whether the debug mode is enabled. Per comment by Nico Weber on IRC.
-
Louis Dionne authored
It greatly increases readability because defining the methods out-of-line involves a ton of boilerplate template declarations.
-
Arthur Eubanks authored
Or else on optnone functions we get the following during instruction selection: fatal error: error in backend: Cannot select: intrinsic %llvm.preserve.struct.access.index Currently the -O0 pipeline doesn't properly run passes registered via TargetMachine::registerPassBuilderCallbacks(), so don't add that RUN line yet. That will be fixed after this. Reviewed By: yonghong-song Differential Revision: https://reviews.llvm.org/D89083
-
Simon Pilgrim authored
Based on offline discussions regarding D89139 and D88783 - we want to make sure targets aren't doing anything particularly dumb Tests copied from aarch64 which has a mixture of general, legalization and special case tests
-
Jonas Devlieghere authored
Buildbot got upgraded and now the (LLDB) builders have different URLs.
-
Giorgis Georgakoudis authored
There are cases that generated OpenMP code consists of multiple, consecutive OpenMP parallel regions, either due to high-level programming models, such as RAJA, Kokkos, lowering to OpenMP code, or simply because the programmer parallelized code this way. This optimization merges consecutive parallel OpenMP regions to: (1) reduce the runtime overhead of re-activating a team of threads; (2) enlarge the scope for other OpenMP optimizations, e.g., runtime call deduplication and synchronization elimination. This implementation defensively merges parallel regions, only when they are within the same BB and any in-between instructions are safe to execute in parallel. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D83635
-
Louis Dionne authored
Due to the need to support compilers that implement builtin operator new/delete but not their align_val_t overloaded versions, there was a lot of complexity. By assuming that a compiler that supports the builtin new/delete operators also supports their align_val_t overloads, the code can be simplified quite a bit. Differential Revision: https://reviews.llvm.org/D88301
-
Louis Dionne authored
Currently, Clang looks for libc++ headers alongside the installation directory of Clang, and it also adds a search path for headers in the -isysroot. This is problematic if headers are found in both the toolchain and in the sysroot, since #include_next will end up finding the libc++ headers in the sysroot instead of the intended system headers. This patch changes the logic such that if the toolchain contains libc++ headers, no C++ header paths are added in the sysroot. However, if the toolchain does *not* contain libc++ headers, the sysroot is searched as usual. This should not be a breaking change, since any code that previously relied on some libc++ headers being found in the sysroot suffered from the #include_next issue described above, which renders any libc++ header basically useless. Differential Revision: https://reviews.llvm.org/D89001
-
Louis Dionne authored
We don't support any compiler that doesn't support variadics and rvalue references in C++03 mode, so these workarounds can be dropped. There's still *a lot* of cruft related to these workarounds, but I try to tackle a bit of it here and there.
-
Arthur Eubanks authored
In the NPM, a pass cannot depend on another non-analysis pass. So pin the test that tests that -lowerswitch is run automatically to legacy PM. Reviewed By: sameerds Differential Revision: https://reviews.llvm.org/D89051
-
Arthur Eubanks authored
Reviewed By: fhahn Differential Revision: https://reviews.llvm.org/D89058
-
Jay Foad authored
Following on from D88890, this makes the newly added patterns conditional on NoFP32Denormals. mad/mac f32 instructions always flush denormals regardless of the MODE register setting, and I believe the legacy variants do the same. Differential Revision: https://reviews.llvm.org/D89123
-
Tres Popp authored
Without this PatternRewriting infrastructure does not know of modifications and cannot properly legalize nor rollback changes. Differential Revision: https://reviews.llvm.org/D89129
-
- Oct 09, 2020
-
-
Simon Pilgrim authored
[InstCombine] Support lshr(trunc(lshr(x,c1)), c2) -> trunc(lshr(lshr(x,c1),c2)) uniform vector tests FoldShiftByConstant is hardcoded for scalar/uniform outer shift amounts atm so that needs to be fixed first to support non-uniform cases
-
Simon Pilgrim authored
-
Eugene Zhulenev authored
Async execute operation can take async arguments as dependencies. Change `async.execute` custom parser/printer format to use `%value as %unwrapped: !async.value<!type>` sytax. Reviewed By: mehdi_amini, herhut Differential Revision: https://reviews.llvm.org/D88601
-
Andrzej Warzynski authored
Reverts one breaking change introduced in https://reviews.llvm.org/D88846. Differential Revision: https://reviews.llvm.org/D89111
-
David Green authored
There were some existing tests that were not super useful. New ones are added for testing MVE specific patterns.
-
Scott Linder authored
Reformat to avoid unrelated changes in diff of future patch. Committed as obvious.
-
Simon Pilgrim authored
Note: we already fold srem to undef if any denominator vector element is undef.
-
Krzysztof Parzyszek authored
-
Sanjay Patel authored
There might be a better way to specify the pre-conditions, but this is hopefully clearer than the way it was written: https://rise4fun.com/Alive/Jhk3 Pre: C2 < 0 && isShiftedMask(C2) && (C1 == C1 & C2) %a = and %x, C2 %r = add %a, C1 => %a2 = add %x, C1 %r = and %a2, C2
-
Simon Pilgrim authored
[InstCombine] Add tests for X shift (A srem B) -> X shift (A and B-1) pow2 nonuniform constant vectors
-
Anastasia Stulova authored
Extended -cl-std/std flag with CL3.0 and added predefined version macros. Patch by Anton Zabaznov (azabaznov)! Tags: #clang Differential Revision: https://reviews.llvm.org/D88300
-
Louis Dionne authored
To make sure we don't store a mutable object (which could be modified by outside code without us noticing) as the cache key, we pickle the cache key to get a byte stream. If two keys are unequal, we know for sure they will not have the same pickling. And if they are equal, there's a large chance they will have the same pickling. If they don't, we might end up not reusing a cached entry when we could have, but at least the behavior we'll have is semantically correct.
-