- May 17, 2022
-
-
Ken Matsui authored
This patch is itended to avoid suggesting typoed directives in `.S` files to support the cases of `#` directives treated as comments or various pseudo-ops. The feature is implemented in https://reviews.llvm.org/D124726. Fixes: https://reviews.llvm.org/D124726#3516346. Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D125727
-
wren romano authored
This enables the compiler to perform devirtualization. And benchmarks indicate devirtualization can sometimes give considerable speedup. Depends On D122061 Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D125428
-
wren romano authored
Fixes: https://github.com/llvm/llvm-project/issues/51652 Depends On D122060 Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D122061
-
River Riddle authored
In the overwhelmingly majority of cases only one dialect is generated at a time anyways, and this restriction more easily catches user error when multiple dialects might be generated. We hit this semi-recently with the PDL dialect, and circt+other downstream users are also actively hitting this as well. Differential Revision: https://reviews.llvm.org/D125651
-
Jason Molenda authored
This patch addresses two perf issues when we find a dSYM on macOS after calling into the DebugSymbols framework. First, when we have a local (probably stripped) binaary, we find the dSYM and we may be told about the location of the symbol rich binary (probably unstripped) which may be on a remote filesystem. We don't need the unstripped binary, use the local binary we already have. Second, after we've found the path to the dSYM, save that in the Module so we don't call into DebugSymbols a second time later on to rediscover it. If the user has a DBGShellCommands set, we need to exec that process twice, serially, which can add up. Differential Revision: https://reviews.llvm.org/D125616 rdar://84576917
-
Joseph Huber authored
The Clang driver additional stages to build a complete offloading program for applications using CUDA or OpenMP offloading. This normally requires either a source file input or a valid object file to be handled. This would cause problems when trying to compile an assembly or LLVM IR file through clang with flags that would enable offloading. This patch simply adds a check to prevent the offloading toolchain from being used if we don't have a valid source file. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D125705
-
Joseph Huber authored
The OpenMP device offloading library is a bitcode library and thus only expect to build and linked with the same version of clang that was used to create it. This somewhat copmlicates the building process as we require the Clang that was just built to be used to create the library. This is either done with a two-step build, where OpenMP is built with the Clang that was just installed, or through the `-DLLLVM_ENABLE_RUNTIMES=openmp` option. This has always been the case, but recent changes have caused this to make it difficult to build the rest of OpenMP. This patchs adds a check to not build the OpenMP device runtime if the current compiler is not Clang with the same version as the LLVM installation. This should allow users to build OpenMP as a project using any compiler without it erroring out due to the bitcode library, but if users require it they will need to use the above methods to compile it. Reviewed By: jdoerfert, tianshilei1992, ...
-
Alex Zinenko authored
Op registration mechanism does not allow for ops with the same name to be re-registered. This is okay to avoid name conflicts and debug double-registration, but may be problematic for dialect extensions that may get registered several times (unlike dialects that are deduplicated in the registry). When registering ops through the Transform dialect extension mechanism, check first if the ops are already registered and only complain in the case of repeated registration with the same name but different TypeID. Differential Revision: https://reviews.llvm.org/D125554
-
Jonas Devlieghere authored
Avoid a OverflowError (an underflow really) when the pc is zero. This can happen for "unknown frames" where the crashlog generator reports a zero pc. We could omit them altogether, but if they're part of the crashlog it seems fair to display them in lldb as well. rdar://92686666 Differential revision: https://reviews.llvm.org/D125716
-
Sanjay Patel authored
This reverts commit 3794cc0e. This change is suspected of causing bots to hang at stage 2 compiles, so reverting to confirm and investigate.
-
Martin Storsjö authored
This allows sharing opcodes between prolog and epilog even when there is more than one epilog. I didn't make any handcrafted special MC level testcases for this (yet at least), but it does seem to have the expected effect on two existing CodeGen level testcases. Differential Revision: https://reviews.llvm.org/D125619
-
Martin Storsjö authored
[MC] [Win64EH] Try writing an ARM64 "packed epilog" even if the epilog doesn't share opcodes with the prolog The "packed epilog" form only implies that the epilog is located exactly at the end of the function (so the location of the epilog is implicit from the epilog opcodes), but it doesn't have to share opcodes with the prolog - as long as the total number of opcode bytes and the offset to the epilog fit within the bitfields. This avoids writing a 4 byte epilog scope in many cases. (I haven't measured how much this shrinks actual xdata sections in practice though.) Differential Revision: https://reviews.llvm.org/D125536
-
Martin Storsjö authored
In f8b0a7af in 2016, this parameter was generalized on the caller side (previously passing STI.isTargetMachO(), now passing STI.splitFramePushPop()). Rename the parameter on the receiver side to match the generalization. Differential Revision: https://reviews.llvm.org/D125681
-
Eli Friedman authored
The spinlock requires that lock-free operations are available; otherwise, the implementation just calls itself. As discussed in D120026. Differential Revision: https://reviews.llvm.org/D123080
-
Paul Walker authored
NEON does not have native support for xor reductions. However, when reducing predicate vectors the operation is synonymous with an add reduction that is supported. Differential Revision: https://reviews.llvm.org/D125605
-
Ellis Hoag authored
When using counter relocations, two instructions are emitted to compute the address of the counter variable. ``` %BiasAdd = add i64 ptrtoint <__profc_>, <__llvm_profile_counter_bias> %Addr = inttoptr i64 %BiasAdd to i64* ``` When promoting a counter, these instructions might not be available in the block, so we need to copy these instructions. This fixes https://github.com/llvm/llvm-project/issues/55125 Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D125710
-
Matthias Springer authored
Return immediately when an op bufferization patterns fails. Differential Revision: https://reviews.llvm.org/D125087
-
Mogball authored
An attribute without a type builder followed by a colon in an assembly format is potentially ambiguous because the parser will read ahead to parse the colon-type and pass this as the type argument to the attribute's constructor. However, the previous verifier that checks for this ambiguity erroneously produces an error in the case of ``` let assemblyFormat = "( `(` $attr `)` )? `:`"; ``` This patch fixes the bug by implementing a checker that correctly handles all edge cases, including very strange assembly formats like: ``` let assemblyFormat = "( `(` $attr ) : (`>`)? attr-dict (`>` $a^) : (`<`)? `:`"; ``` Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125445
-
River Riddle authored
This was carry over from LLVM IR where the alias definition can be ambiguous, but MLIR type aliases have no such problems. Having the `type` keyword is superfluous and doesn't add anything. This commit drops it, which also nicely aligns with the syntax for attribute aliases (which doesn't have a keyword). Differential Revision: https://reviews.llvm.org/D125501
-
Mogball authored
This patch adds a topological sort utility and pass. A topological sort reorders the operations in a block without SSA dominance such that, as much as possible, users of values come after their producers. The utility function sorts topologically the operation range in a given block with an optional user-provided callback that can be used to virtually break cycles. The toposort pass itself recursively sorts graph regions under the target op. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D125063
-
zhijian authored
Summary: Use object::Archive::create so that the returned archive object has a dynamic type of either Archive or BigArchive. Reviewers: James Henderson,Fangrui Song Differential Revision: https://reviews.llvm.org/D124940
-
Mogball authored
The attribute self type parameter is currently treated like any other attribute parameter in the assembly format. The self type parameter should be handled by the operation parser and printer and play no role in the generated parsers and printers of attributes. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125724
-
Aart Bik authored
This is the first implementation of complex (f64 and f32) support in the sparse compiler, with complex add/mul as first operations. Note that various features are still TBD, such as other ops, and reading in complex values from file. Also, note that the std::complex<float> had a bit of an ABI issue when passed as single argument. It is still TBD if better solutions are possible. Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D125596
-
Yang Keao authored
In D123677, @YangKeao provided an implementation of `DOTGraphTraits{Viewer,Printer}` in the new pass manager. This commit migrates the `DomPrinter` and `DomViewer` to the new pass manager. Reviewed By: Meinersbur Differential Revision: https://reviews.llvm.org/D124904 -
Paul Walker authored
During early gather/scatter enablement two different approaches were taken to represent scaled indices: * A Scale operand whereby byte_offsets = Index * Scale * An IndexType whereby byte_offsets = Index * sizeof(MemVT.ElementType) Having multiple representations is bad as shown by this patch which fixes instances where the two are out of sync. The dedicated scale operand is more flexible and pervasive so this patch removes the UNSCALED values from IndexType. This means all indices are scaled but the scale can be one, hence unscaled. SDNodes now use the scale operand to answer the "isScaledIndex" question. I toyed with the idea of keeping the UNSCALED enums and helper functions but because they will have no uses and force SDNodes to validate the set of supported values I figured it's best to remove them. We can re-add them if there's a real need. For similar reasons I've kept the IndexType enum when a bool could be used as I think being explicitly looks better. Depends On D123347 Differential Revision: https://reviews.llvm.org/D123381
-
Louis Dionne authored
As mentionned in D97044, it is fine if users include <atomic> and then include <stdatomic.h> -- we don't need to error out for that case. Differential Revision: https://reviews.llvm.org/D125579
-
Louis Dionne authored
I think this notion of libc++abi's version was relevant a long time ago on Apple platforms when we were using a Xcode project to build the library. As part of moving Apple's build to CMake, D59489 made it possible to specify the "ABI version" of libc++abi in use. However, it's not possible to build libc++abi with that old ABI anymore and we don't need the ability to link against that version from libc++ anymore. Hence, we can clean this up and stop falsely pretending that libc++abi has more than one ABI version. Differential Revision: https://reviews.llvm.org/D125687
-
John Paul Adrian Glaubitz authored
Linux on SPARC uses fstatat64 instead. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D125572
-
Utkarsh Saxena authored
Reduces time spent in findRef by 66%. Differential Revision: https://reviews.llvm.org/D125675
-
Robert Suderman authored
We were custom counting per bit for the clz instruction. Math dialect now has an intrinsic to do this in one instruction. Migrated to this instruction and fixed a minor bug math-to-llvm for the intrinsic. Reviewed By: mravishankar Differential Revision: https://reviews.llvm.org/D125592
-
zhijian authored
Summary: llvm-ar can not read empty big archive correctly. it output error as error: unable to load 'empty.a': truncated or malformed archive (characters in size field in archive member header are not all decimal numbers: '<bigaf>' Reviewers: James Henderson Differential Revision: https://reviews.llvm.org/D124017
-
Stanislav Mekhanoshin authored
This reverts ffbee7ac, see also bug 37653 which it was fixing. The bug claims this is an undocumented feature which actually works. In the reality it is documented as not working for a good reason. It likely does something, but it is useless anyway. These instructions write into the LDS. The LDS address is: M0 + inst_offset + (TIDinWave * 4). For a store wider than a DWORD neighboring lanes will overwrite each other. Differential Revision: https://reviews.llvm.org/D125409
-
Utkarsh Saxena authored
Differential Revision: https://reviews.llvm.org/D125682
-
Alex Bradbury authored
The logic around IsCanonical previously used getAsString and compared to "1". Just using getValueAsBit is simpler.
-
Jakub Kuderski authored
In this instance, the trailing return type does not improve readability as it repeats what is returned in the same line. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D125697
-
Fangrui Song authored
The former is a bit misleading and a user may try installing zlib which will not help.
-
Fangrui Song authored
-
Rahman Lavaee authored
[llvm-objdump] Let --symbolize-operands symbolize basic block addresses based on the SHT_LLVM_BB_ADDR_MAP section. `--symbolize-operands` already symbolizes branch targets based on the disassembly. When the object file is created with `-fbasic-block-sections=labels` (ELF-only) it will include a SHT_LLVM_BB_ADDR_MAP section which maps basic blocks to their addresses. In such case `llvm-objdump` can annotate the disassembly based on labels inferred on this section. In contrast to the current labels, SHT_LLVM_BB_ADDR_MAP-based labels are created for every machine basic block including empty blocks and those which are not branched into (fallthrough blocks). The old logic is still executed even when the SHT_LLVM_BB_ADDR_MAP section is present to handle functions which have not been received an entry in this section. Reviewed By: jhenderson, MaskRay Differential Revision: https://reviews.llvm.org/D124560
-
Adrian Prantl authored
-
David Green authored
There have been some patterns in the AArch64 backend to optimize code of the form: ldrsh w8, [x0] scvtf s0, w8 to: ldr h0, [x0] sshll v0.4s, v0.4h, #0 scvtf s0, s0 The idea is to remove the GRP->FPR move, but in reality is making code larger and slower (or the same) on all the cpus I tried. This patch adds the UseAlternateSExtLoadCVTF32 predicate similar to nearby related pattern. Differential Revision: https://reviews.llvm.org/D125470
-