- May 06, 2021
-
-
thomasraoux authored
-
thomasraoux authored
Differential Revision: https://reviews.llvm.org/D101955
-
Nemanja Ivanovic authored
This was reverted in 3761b9a2 just as I was about to commit the fix. This patch inlcudes the necessary fix.
-
Austin Kerbow authored
NFC.
-
David Spickett authored
Since https://reviews.llvm.org/D101849 this test has been failing on bots that only enable either Arm or AArch64 targets. See: https://lab.llvm.org/buildbot/#/builders/107/builds/7601 Temporarily requires X86 for this test while the difference is figured out.
-
Louis Dionne authored
This is a rough reapplication of the change that fixed std::to_address to avoid relying on element_type (da456167). It is somewhat different because the fix to avoid breaking Clang (which caused it to be reverted in 347f69c5) was a bit more involved. Differential Revision: https://reviews.llvm.org/D101638
-
Victor Huang authored
- Add branch absolute reloction R_RBA, R_TLS relocation for the variable offset for the tlsgd model and R_TLSM for the region handle for the tlsgd model - Properly set the relocation fixed values for R_TLS and R_TLSM - Emit the TCEntry with the variant kind in the XCOFFStreamer Reviewed by: sfertile, nemanjai, DiggerLin Differential Revision: https://reviews.llvm.org/D100214
-
Nico Weber authored
This reverts commit ed87f512. Breaks check-clang, see e.g. https://lab.llvm.org/buildbot/#/builders/139/builds/3818
-
Raphael Isemann authored
-
Jay Foad authored
-
Nemanja Ivanovic authored
This adds additional support for XL compatibility. There are a number of functions in altivec.h that produce a single instruction (or a very short sequence) for Power8 but can be done on Power7 without scalarization. XL provides these implementations. This patch adds the following overloads for doubleword vectors: vec_add vec_cmpeq vec_cmpgt vec_cmpge vec_cmplt vec_cmple vec_sl vec_sr vec_sra
-
Paul C. Anagnostopoulos authored
Differential Revision: https://reviews.llvm.org/D101766
-
Anastasia Stulova authored
This patch simplifies the parser and makes the language semantics consistent. There is no extension pragma requirement in the spec for the subgroup functions in enqueue kernel or pipes and all other builtin functions are available without the pragama. Differential Revision: https://reviews.llvm.org/D100984
-
Jon Chesterfield authored
[amdgpu-arch] Fix rpath to run from build dir Prior to this, amdgpu-arch has RUNPATH set to $ORIGIN/../lib which works for some installs, but not from the build directory where clang executes the tool from when running tests. This cmake option adds the location of the rocr runtime to the RUNPATH (note, it amends RUNPATH here, despite the cmake option referring to RPATH) to create a binary that runs from build or install location. Before: RUNPATH [$ORIGIN/../lib] After: RUNPATH [$ORIGIN/../lib:$HOME/llvm-install/lib] Credit to Greg for knowing this trick and pointing to examples of it in use for the aomp build scripts. Reviewed By: pdhaliwal Differential Revision: https://reviews.llvm.org/D101926
-
Carl Ritson authored
Instruction test for inactive kill/demote needs to be based on actual opcode not whether instruction would be lowered to demote. Reviewed By: piotr Differential Revision: https://reviews.llvm.org/D101966
-
Malhar Jajoo authored
Reverting commit since it causes failure (10462). This reverts commit b856f4a2.
-
Benjamin Kramer authored
-
David Green authored
The loop vectorizer will currently assume a large trip count when calculating which of several vectorization factors are more profitable. That is often not a terrible assumption to make as small trip count loops will usually have been fully unrolled. There are cases however where we will try to vectorize them, and especially when folding the tail by masking can incorrectly choose to vectorize loops that are not beneficial, due to the folded tail rounding the iteration count up for the vectorized loop. The motivating example here has a trip count of 5, so either performs 5 scalar iterations or 2 vector iterations (with VF=4). At a high enough trip count the vectorization becomes profitable, but the rounding up to 2 vector iterations vs only 5 scalar makes it unprofitable. This adds an alternative cost calculation when we know the max trip count and are folding tail by masking, rounding the iteration count up to the correct number for the vector width. We still do not account for anything like setup cost or the mixture of vector and scalar loops, but this is at least an improvement in a few cases that we have had reported. Differential Revision: https://reviews.llvm.org/D101726
-
Ben Dunbobbin authored
This is a slight improvement to the help text, as I was slightly surprised when strip-all did more than remove the symbol table. Currently, we match gold's help text for strip-all and strip-debug. I think that the GNU documentation for these options is not particularly clear. However, I have opted to make only a minor change here and keep the help text similar to gold's as these are mature options that are well understood. ld.bfd (https://sourceware.org/binutils/docs/ld/Options.html) has a similar implication although it defines strip-debug as a subset of strip-all. However, felt that noting that strip-all implies strip-debug is better; because, with the ld.bfd approach you have to read both the --strip-debug and the --strip-all help text to understand the behaviour of --strip-all (and the --strip-all help text doesn't indicate that he --strip-debug help text is related). Differential Revision: https://reviews.llvm.org/D101890
-
Christian Sigg authored
Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D101757
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
Jonas Paulsson authored
In order to use __builtin_frame_address(0) with packed stack and no backchain, the address of where the backchain would have been written is returned (like GCC). This address may either contain a saved register or be unused. Review: Ulrich Weigand Differential Revision: https://reviews.llvm.org/D101897
-
Kerry McLaughlin authored
Adds support for scalable vectorization of loops containing first-order recurrences, e.g: ``` for(int i = 0; i < n; i++) b[i] = a[i] + a[i - 1] ``` This patch changes fixFirstOrderRecurrence for scalable vectors to take vscale into account when inserting into and extracting from the last lane of a vector. CreateVectorSplice has been added to construct a vector for the recurrence, which returns a splice intrinsic for scalable types. For fixed-width the behaviour remains unchanged as CreateVectorSplice will return a shufflevector instead. The tests included here are the same as test/Transform/LoopVectorize/first-order-recurrence.ll Reviewed By: david-arm, fhahn Differential Revision: https://reviews.llvm.org/D101076
-
Eliza Velasquez authored
Reviewed By: curdeius Differential Revision: https://reviews.llvm.org/D101862
-
Eliza Velasquez authored
This fixes two errors: Previously, clang-format was splitting up type identifiers from the nullable ?. This changes this behavior so that the type name sticks with the operator. Additionally, nullable operators attached to return types in interface functions were not parsed correctly. Digging deeper, it looks like interface bodies were being parsed differently than classes and structs, causing MustBeDeclaration to be incorrect for interface members. They now share the same logic. One other change is reintroducing the CSharpNullable type independent of JsTypeOptionalQuestion. Despite having a similar semantic purpose, their actual syntax differs quite a bit. Reviewed By: MyDeveloperDay, curdeius Differential Revision: https://reviews.llvm.org/D101860
-
Eliza Velasquez authored
This adds support for the null-coalescing assignment and null-forgiving operators. https://docs.microsoft.com/en-us/dotnet/csharp/language-reference/operators/null-coalescing-operator https://docs.microsoft.com/en-us/dotnet/csharp/language-reference/operators/null-forgiving Reviewed By: krasimir, curdeius Differential Revision: https://reviews.llvm.org/D101702
-
Jay Foad authored
First clean up the strange API of tryConstantFoldOp where it took an immediate operand value, but no indication of which operand it was the value for. Second clean up the loop that calls tryConstantFoldOp so that it does not have to restart from the beginning every time it folds an instruction. This is NFCI but there are some minor changes caused by the order in which things are folded. Differential Revision: https://reviews.llvm.org/D100031
-
Andrzej Warzynski authored
`%f18` was originally introduced to represent the old Flang driver, `f18`. With the introduction of the new driver, `flang-new`, we have been switching to `%flang` (compiler driver) and `%flang_fc1` (frontend driver) as more generic alternatives. As most tests have been portend to use the new LIT variables instead of `%f18`, this is good time to remove it from lit.cfg.py. There's only one test left that requires the old driver to run. It's updated with: ``` ! REQUIRES: old-flang-driver ``` This way we preserve its semantics while reducing the number of variables in LIT configuration. Differential Revision: https://reviews.llvm.org/D101281
-
Malhar Jajoo authored
This patch converts llvm.memcpy intrinsic into Tail Predicated Hardware loops for a target that supports the Arm M-profile Vector Extension (MVE). From an implementation point of view, the patch - adds an ARM specific SDAG Node (to which the llvm.memcpy intrinsic is lowered to, during first phase of ISel) - adds a corresponding TableGen entry to generate a pseudo instruction, with a custom inserter, on matching the above node. - Adds a custom inserter function that expands the pseudo instruction into MIR suitable to be (by later passes) into a WLSTP loop. Note: A cli option is used to control the conversion of memcpy to TP loop and this option is currently disabled by default. It may be enabled in the future after further downstream testing. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D99723
-
James Henderson authored
Previously, if the search_env argument was specified, and the tool was found at that location, the path was not reported, unlike other situations when this function was called. Adding the reporting makes the function consistent. Reviewed by: thopre Differential Revision: https://reviews.llvm.org/D101896
-
Tim Renouf authored
Fix up my recent commit rG1128311a to use std::make_unique instead of std::unique_ptr(new), as requested by David Blaikie. Differential Revision: https://reviews.llvm.org/D101822
-
Guillaume Chatelet authored
Differential Revision: https://reviews.llvm.org/D101910
-
Guillaume Chatelet authored
Differential Revision: https://reviews.llvm.org/D101909
-
Guillaume Chatelet authored
Differential Revision: https://reviews.llvm.org/D101907
-
Guillaume Chatelet authored
Differential Revision: https://reviews.llvm.org/D101906
-
Guillaume Chatelet authored
Differential Revision: https://reviews.llvm.org/D101905
-
Johannes Doerfert authored
This patch fixes various issues with our prior `declare target` handling and extends it to support `omp begin declare target` as well. This started with PR49649 in mind, trying to provide a way for users to avoid the "ref" global use introduced for globals with internal linkage. From there it went down the rabbit hole, e.g., all variables, even `nohost` ones, were emitted into the device code so it was impossible to determine if "ref" was needed late in the game (based on the name only). To make it really useful, `begin declare target` was needed as it can carry the `device_type`. Not emitting variables eagerly had a ripple effect. Finally, the precedence of the (explicit) declare target list items needed to be taken into account, that meant we cannot just look for any declare target attribute to make a decision. This caused the handling of functions to require fixup as well. I tried to clean up things while I was at it, e.g., we should not "parse declarations and defintions" as part of OpenMP parsing, this will always break at some point. Instead, we keep track what region we are in and act on definitions and declarations instead, this is what we do for declare variant and other begin/end directives already. Highlights: - new diagnosis for restrictions specificed in the standard, - delayed emission of globals not mentioned in an explicit list of a declare target, - omission of `nohost` globals on the host and `host` globals on the device, - no explicit parsing of declarations in-between `omp [begin] declare variant` and the corresponding end anymore, regular parsing instead, - precedence for explicit mentions in `declare target` lists over implicit mentions in the declaration-definition-seq, and - `omp allocate` declarations will now replace an earlier emitted global, if necessary. --- Notes: The patch is larger than I hoped but it turns out that most changes do on their own lead to "inconsistent states", which seem less desirable overall. After working through this I feel the standard should remove the explicit declare target forms as the delayed emission is horrible. That said, while we delay things anyway, it seems to me we check too often for the current status even though that is often not sufficient to act upon. There seems to be a lot of duplication that can probably be trimmed down. Eagerly emitting some things seems pretty weak as an argument to keep so much logic around. --- Reviewed By: ABataev Differential Revision: https://reviews.llvm.org/D101030 -
Johannes Doerfert authored
A user reported an assertion (below) but without a reproducer. I failed to create a test myself but from the assertion one can derive the problem. I set the DefaultMapperId location now to make sure this doesn't cause trouble. ``` clang-13: .../DeclTemplate.h:1940: void clang::ClassTemplateSpecializationDecl::setPointOfInstantiation(clang::SourceLocation): Assertion `Loc.isValid() && "point of instantiation must be valid!"' failed. ``` Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D100621
-
Johannes Doerfert authored
We do provide `operator delete(void*)` in `<new>` but it should be available by default. This is mostly boilerplate to test it and the unconditional include of `<new>` in the header we always in include on the device. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D100620
-