- Oct 05, 2022
-
-
Sanjay Patel authored
-
Fangrui Song authored
The output is similar to objdump --no-addresses since binutils 2.35. Depends on D135039 Close #58088 Differential Revision: https://reviews.llvm.org/D135040
-
Fangrui Song authored
It seems to make sense to omit offsets when --no-leading-addr is specified. The output is now closer to objdump -dr --no-addresses (non-wide output). Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D135039
-
Nathaniel McVicar authored
This restores the fix from D134925 to make MSVC and clang happy. Reviewed By: stella.stamenova Differential Revision: https://reviews.llvm.org/D135126
-
Mark de Wever authored
This adds support for the new code points in the Extended Grapheme Cluster algorithm. The algorithm itself has remained unchanged. The width estimation still follows the rules of the Standard. @cor3ntin filed LWG3780 format's width estimation is too approximate and not forward compatible to improve the estimate. Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D134106
-
Daniel Rodríguez Troitiño authored
This is a split of D134250. Supports for parsing and dumping the LC_DATA_IN_CODE contents (as binary data). This allows more complete testing of llvm-objdump in D133974. Reviewed By: Higuoxing Differential Revision: https://reviews.llvm.org/D134569
-
Craig Topper authored
There are few changes mixed in here. -Try to reuse the destination register from ADDI instead of always creating a virtual register. This way we lean on the register scavenger in fewer case. -Explicitly reuse the primary virtual register when possible. There's still a case where both getVLENFactoredAmount and handling large fixed offsets can both create a secondary virtual register. -Combine similar BuildMI calls by manipulating the Register variables. There are still a couple early outs for ADDI, but overall I tried to arrange the code into steps. Reviewed By: reames Differential Revision: https://reviews.llvm.org/D135009
-
Craig Topper authored
The old code took two different paths based on whether there is a scalable offset, but these two paths had some code in common. The main difference between the two code paths was whether we needed to create a GPR or not for the ADDI that gets created for RVVSpill. If we had a scalable offset, the same GPR was used as the destination for adding the scalable offset and the ADDI. To manage this, we now cache the scratch register and reuse it if it has already been created. This is a pre-patch for D135009. Reviewed By: reames, frasercrmck Differential Revision: https://reviews.llvm.org/D135092
-
Nicolas Vasilache authored
Differential Revision: https://reviews.llvm.org/D135152
-
Daniel Rodríguez Troitiño authored
The `dumpExportEntry` was dumping everything using signed LEB128, but the format seems to use unsigned LEB128. This can be cross-checked with the implementation in MachOObjectFile.cpp, the implementation in LLD's ExportTrie.cpp, and the implementation in macho2yaml.cpp, which all use ULEB128 functions.. The difference is only apparent when encoding some values with specific bit patterns (bit active in the 7th, 14th, ... bits of the binary). The encoding was not always creating problems in the resulting binaries because if the extra byte was part of the padding, the result of decoding it as ULEB128 is the same as decoding as SLEB128, however, the code of MachOObjectFile.cpp (used by llvm-objdump) checks the buffer decoding position against the reported length, which triggered an error. Modified a test that used an address with this pattern (0x3FA0, the 14th bit is active), to show that a round trip still produces the same results, and added a check using llvm-objdump to use their extra checks to verify this implementation. Reviewed By: pete Differential Revision: https://reviews.llvm.org/D134563
-
Nicolas Vasilache authored
Differential Revision: https://reviews.llvm.org/D135135
-
Aaron Ballman authored
Post-commit feedback observed that returning the TypeOfKind from the type instead of a Boolean cleans up code using that interface.
-
Tom Honermann authored
Previously, a lambda expression in a dependent context with a default argument containing an immediately invoked lambda expression would produce a closure class object that, if invoked such that the default argument was used, resulted in a compiler crash or one of the following assertion failures during code generation. The failures occurred regardless of whether the lambda expressions were dependent. clang/lib/CodeGen/CGCall.cpp: Assertion `(isGenericMethod || Ty->isVariablyModifiedType() || Ty.getNonReferenceType()->isObjCRetainableType() || getContext() .getCanonicalType(Ty.getNonReferenceType()) .getTypePtr() == getContext().getCanonicalType((*Arg)->getType()).getTypePtr()) && "type mismatch in call argument!"' failed. clang/lib/AST/Decl.cpp: Assertion `!Init->isValueDependent()' failed. Default arguments in declarations in local context are instantiated along with their enclosing function or variable template (since such declarations can't be explicitly specialized). Previously, such instantiations were performed at the same time that their associated parameters were instantiated. However, that approach fails in cases like the following in which the context for the inner lambda is the outer lambda, but construction of the outer lambda is dependent on the parameters of the inner lambda. This change resolves this dependency by delyaing instantiation of default arguments in local contexts until after construction of the enclosing context. template <typename T> auto f() { return [](T = []{ return T{}; }()) { return 0; }; } Refactoring included with this change results in the same code now being used to instantiate default arguments that appear in local context and those that are only instantiated when used at a call site; previously, such code was duplicated and out of sync. Fixes https://github.com/llvm/llvm-project/issues/49178 Reviewed By: erichkeane Differential Revision: https://reviews.llvm.org/D133500
-
- Oct 04, 2022
-
-
Sanjay Patel authored
https://alive2.llvm.org/ce/z/oShzr3 This was noted as a missing fold in D134876 (with additional examples based on issue #58046). I'm assuming that fmul with a zero operand is rare enough that the use of ValueTracking will not noticeably increase compile-time. This adjusts a PowerPC codegen test that was added with D88388 because it would get folded away and no longer provide coverage for the bug fix.
-
Alexey Bataev authored
In the canonical form of the shuffle the poison/undef operand is the second operand, the patch tries to emit canonical form for partial vectorization of the buildvector sequence. Also, this patch starts emitting freeze instruction for shuffles with undef indices if the second shuffle operan is undef, not poison. It is an initial step to D93818, where undef mask element are treated as returning poison value. Differential Revision: https://reviews.llvm.org/D134377
-
Pierre van Houtryve authored
The bitmask used to extract the bits assumed 16 bit elements and wasn't taking the size of the elements into account. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D135156
-
Sanjay Patel authored
-
Sanjay Patel authored
This is a modification of the earlier attempt from: 7b7940f9 For fma callers, we only want to swap a 0.0 or 1.0 constant.
-
Pierre van Houtryve authored
Make it illegal, remove InstructionSelector logic for it Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D134967
-
Jay Foad authored
Fix a crash in the FMA combine added by D132837 and amended by D134810. In cases where the newly created node could be folded, the combiner would fail this assertion: llc: DAGCombiner.cpp:268: void (anonymous namespace)::DAGCombiner::AddToWorklist(llvm::SDNode *): Assertion `N->getOpcode() != ISD::DELETED_NODE && "Deleted Node added to Worklist"' failed. Differential Revision: https://reviews.llvm.org/D135150
-
Amara Emerson authored
-
Dominik Adamski authored
If 'order(concurrent)' clause is specified, then the iterations of SIMD loop can be executed concurrently. This patch adds support for LLVM IR codegen via OMPIRBuilder for SIMD loop with 'order(concurrent)' clause. The functionality added to OMPIRBuilder is similar to the functionality implemented in 'CodeGenFunction::EmitOMPSimdInit'. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D134046 Signed-off-by:
Dominik Adamski <dominik.adamski@amd.com>
-
Florian Hahn authored
-
Pierrick Bouvier authored
- Detect VS devcmd error (missing VS) - Detect missing python install - Show commands executed - Removed pause (blocking CI usage) Differential revision: https://reviews.llvm.org/D135138
-
Nikita Popov authored
Reapply with a fix for the case where an operand simplified back to the original phi: We need to map this case to the new phi node. ----- foldOpIntoPhi() currently only folds operations into the phi if all but one operands constant-fold. The two exceptions to this are freeze and select, where we allow more general simplification. This patch makes foldOpIntoPhi() generally simplification based and removes all the instruction-specific logic. We just try to simplify the instruction for each operand, and for the (potentially) one non-simplified operand, we move it into the new block with adjusted operands. This fixes https://github.com/llvm/llvm-project/issues/57448, which was my original motivation for the change.
-
Alex Richardson authored
This currently does not make much of a difference (only one tests is affected), but it is helpful e.g. for the out-of-tree CHERI target where Builder.CreateMemCpy() can add attributes other than parameter alignment. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D135075
-
Nikita Popov authored
Degenerate case for D134954.
-
Louis Dionne authored
This is a breaking change. If you were passing one of those three runtimes in LLVM_ENABLE_PROJECTS, you need to start passing them in LLVM_ENABLE_RUNTIMES instead. The runtimes in LLVM_ENABLE_RUNTIMES will start being built using the "bootstrapping build" instead, which means that they will be built using the just-built Clang. This is usually what you wanted anyway. If you were using LLVM_ENABLE_PROJECTS=all with the explicit goal of building these three runtimes, you can now use LLVM_ENABLE_RUNTIMES=all and these runtimes will be built using the bootstrapping build. NOTE: This is a re-application of 887b8bd7 which had been reverted in 6b03a4fe because it broke the Sphinx documentation publishers. The Sphinx documentation publishers have now been moved to using the runtimes build, so this should not be an issue anymore. Differential Revision: https://reviews.llvm.org/D132480
-
Denys Shabalin authored
Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D135139
-
Uday Bondhugula authored
Simplify affine expressions and maps while exploiting simple range and step info of any IVs that are operands. This simplification is local, O(1) and practically useful in several scenarios. Accesses with floordiv's and mod's where the LHS is non-negative and bounded or is a known multiple of a constant can often be simplified. This is implemented as a canonicalization for all affine ops in a generic way: all affine.load/store, vector_load/store, affine.apply, affine.min/max, etc. ops. Eg: For tiled loop nests accessing buffers this way: affine.for %i = 0 to 1024 step 32 { affine.for %ii = 0 to 32 { affine.load [(%i + %ii) floordiv 32, (%i + %ii) mod 32] } } // Note that %i is a multiple of 32 and %ii < 32, hence: (%i + %ii) floordiv 32 is the same as %i floordiv 32 (%i + %ii) mod 32 is the same as %ii mod 32. The simplification leads to simpler index/subscript arithmetic for multi-dimensional arrays and also in turn enables detection of spatial locality (for vectorization for eg.), temporal locality or loop invariance for hoisting or scalar replacement. Differential Revision: https://reviews.llvm.org/D135085 -
Thomas Symalla authored
-
Adrian Kuegel authored
loop variable is copied but only used as const reference.
-
Alex Zinenko authored
Relax the restriction in the transform dialect interpreter utilities that expected a payload IR op to be assocaited with at most one transform IR handle value. This was useful during the initial bootstrapping to avoid use-after-free error equivalents when a payload IR op could be erased through one of the handles associated with it and then accessed through another. It was, however, possible to erase an ancestor of the payload IR operation in question. The expensive-checks mode of interpretation is able to detect both cases and has proven sufficiently robust in debugging use-after-free errors. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D134964
-
Guray Ozen authored
This revision adds GPU transform dialect. It also introduce a prefix such as "transform.gpu" for all ops related to this dialect. MLIR already had two GPU transform op in linalg. This revision moves these ops into GPUTransformOps. The Ops are as follows: `transform.structured.map_nested_foreach_thread_to_gpu_blocks` -> `transform.gpu.map_foreach_to_blocks` This op selects the outermost (toplevel) foreach_thread and parallelize across GPU blocks. It can also generate `gpu_launch`. `transform.structured.map_nested_foreach_thread_to_gpu_threads` -> `transform.gpu.map_nested_foreach_to_threads` This op parallelizes nested foreach_thread that are inside `gpu_launch` across GPU threads. It doesn't add new functionality, but there are some minor refactoring of the code. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D134800
-
Bjorn Pettersson authored
The helpers in BuildLibCalls normally expect that the Value arguments already have the correct type (matching the lib call signature). And exception has been emitFPutC which casted the Char argument to 'int' using CreateIntCast. This patch moves the cast to the caller instead of doing it inside emitFPutC. I think it makes sense to make the BuildLibCall API:s a bit more consistent this way, despite the need to handle the int cast in two different places now. Differential Revision: https://reviews.llvm.org/D135066
-
Bjorn Pettersson authored
Stop assuming that an 'int' is 32 bits in helpers that emit libcalls to lib functions that had 'int' in the signature. For most targets this is NFC. For a target with 16 bit 'int' type this could help out detecting if trying to emit a libcall with incorrect signature. Similarly we now derive the type mapping to 'size_t' by asking TLI about the size of 'size_t'. This should be NFC (at least for in-tree targets) since getSizeTSize(), in TLI, is deriving the size in the same way as DataLayout::getIntPtrType(). Differential Revision: https://reviews.llvm.org/D135065
-
Bjorn Pettersson authored
Lots of BuildLibCalls helpers are using Builder::getInt32Ty to get a type matching an 'int', and DataLayout::getIntPtrType to get a type matching 'size_t'. The former is not true for all targets, since and 'int' isn't always 32 bits. And the latter is a bit weird as well as the definition of DataLayout::getIntPtrType isn't clearly mapping it to 'size_t'. This patch is not aiming at solving any such problems. It is merely highlighting when a libcall is expecting to use 'int' and 'size_t' by naming the types as IntTy and SizeTTy when preparing the type signatures for the emitted libcalls. Differential Revision: https://reviews.llvm.org/D135064
-
Florian Hahn authored
Use LoopAccessInfoManager directly instead of various GetLAA lambdas. Depends on D134608. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D134609
-
Amara Emerson authored
This reverts commit dcd02a52. We should instead use the generic combine.
-
Nikita Popov authored
Handle constant expressions by falling through to the general operator-based code. In particular, this adds support for bitcast and GEP expressions.
-