- Apr 13, 2023
-
-
Lawrence Benson authored
Adds a DAG combine checks for vector comparisons followed by a bitcast to a scalar value. Previously, this resulted in an expand. Now, this is done with a constant number of instructions that take one bit per vector value (via an AND mask) and perfom a horizontal add to get a single value. This is especially useful for Clang's __builtin_convertvector() to a bool vector. Issue: https://github.com/llvm/llvm-project/issues/59829 Differential Revision: https://reviews.llvm.org/D145301
-
Simon Pilgrim authored
Based off D148215, when expanding a min/max reduction we should be creating min/max intrinsics directly instead of relying on instcombine to fold them back together. This patch handles integer min/max cases. Hopefully we can add floating point support soon (at least for fastmath/nnan cases) - but we're missing some of the plumbing to pass the correct FMF to the intrinsic at the moment. Differential Revision: https://reviews.llvm.org/D148221
-
Caroline Tice authored
This test previously relied on just segfaulting or not. This commit adds a CHECK statement to the test. Differential Revision: https://reviews.llvm.org/D148151
-
Chris Jones authored
`str` must be valid UTF-8, which is not guaranteed for C++ strings. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D147818
-
Nicole Rabjohn authored
These test cases were fixed with AIX 73TL1, and are currently passing on AIX machines with that fix. This fix has also been backported to the 7.2 service line. These were tested on a machine with AIX 7.2 TL 5 SP4 installed. Differential Revision: https://reviews.llvm.org/D148040
-
Zequan Wu authored
Fix a crash when an expression/statement can have valid start location but invalid end location in some situations. For example: https://github.com/llvm/llvm-project/blob/llvmorg-16.0.1/clang/lib/Sema/SemaExprCXX.cpp#L1536 This confuses `CounterCoverageMappingBuilder` when popping a region from region stack as if the end location is a macro or include location. Reviewed By: hans, aaron.ballman Differential Revision: https://reviews.llvm.org/D147073
-
Utkarsh Saxena authored
Using FileEntry for retrieving filenames give the name used for the last access. FileEntryRef gives the first access name. Last access name is suspected to change with unrelated changes in clang or the underlying filesystem. First access name gives more stability to the name and makes it easier to track. Differential Revision: https://reviews.llvm.org/D148213
-
Vlad Serebrennikov authored
We've been using https://wg21.link for C++ DR status page, but it forwards non-resolved issues to EDG wiki, which is not useful for general public. This patch replace it with https://cplusplus.github.io/CWG/issues/ .
-
Job Noorman authored
The following tests fail when enabling UBSan due to an unaligned memory load: > runtime error: load of misaligned address 0x620000000643 for type > 'const uint32_t' (aka 'const unsigned int'), which requires 4 byte > alignment BOLT :: AArch64/asm-func-debug.test BOLT :: AArch64/update-debug-reloc.test BOLT :: X86/asm-func-debug.test BOLT :: X86/dwarf5-df-dualcu.test BOLT :: X86/dwarf5-df-mono-dualcu.test BOLT :: X86/dwarf5-ftypes-dwp-input-dwo-output.test BOLT :: X86/dwarf5-locaddrx.test BOLT :: X86/dwarf5-split-dwarf4-monolithic.test BOLT :: X86/inlined-function-mixed.test BOLT :: non-empty-debug-line.test This patch fixes this by using read32le for the load. Reviewed By: ayermolo Differential Revision: https://reviews.llvm.org/D148217
-
Nikita Popov authored
-
Archibald Elliott authored
We were still seeing occasional crashes with inline assembly blocks using fp16/bf16 after my previous patches: - https://reviews.llvm.org/rGff4027d152d0 - https://reviews.llvm.org/rG7d15212b8c0c - https://reviews.llvm.org/rG20b2d11896d9 It turns out: - The original two commits were wrong, and we should have always been choosing the SPR register class, not the HPR register class, so that LLVM's SelectionDAGBuilder correctly did the right splits/joins. - The `splitValueIntoRegisterParts`/`joinRegisterPartsIntoValue` changes from rG20b2d118 are still correct, even though they sometimes result in inefficient codegen of casts between fp16/bf16 and i32/f32 (which is visible in these tests). This patch fixes crashes in `getCopyToParts` and when trying to select `(bf16 (bitconvert (fp16 ...)))` dags when Neon is enabled. This patch also adds support for passing fp16/bf16 values using the 'x' constraint that is LLVM-specific. This should broadly match how we pass with 't' and 'w', but with a different set of valid S registers. Differential Revision: https://reviews.llvm.org/D147715
-
Timm Bäder authored
-
Timm Bäder authored
just because we're being told to evaluate it twice. This sometimes happens when a variable is evaluated again during codegen. Differential Revision: https://reviews.llvm.org/D147535
-
Jun Zhang authored
Signed-off-by:Jun Zhang <jun@junz.org>
-
Timm Bäder authored
Otherwise, we run into an assertion when trying to use the current variable scope while creating temporaries for constructor initializers. Differential Revision: https://reviews.llvm.org/D147534
-
David Spickett authored
This change uses the information from target.xml sent by the GDB stub to produce C types that we can use to print register fields. lldb-server *does not* produce this information yet. This will only work with GDB stubs that do. gdbserver or qemu are 2 I know of. Testing is added that uses a mocked lldb-server. ``` (lldb) register read cpsr x0 fpcr fpsr x1 cpsr = 0x60001000 = (N = 0, Z = 1, C = 1, V = 0, TCO = 0, DIT = 0, UAO = 0, PAN = 0, SS = 0, IL = 0, SSBS = 1, BTYPE = 0, D = 0, A = 0, I = 0, F = 0, nRW = 0, EL = 0, SP = 0) ``` Only "register read" will display fields, and only when we are not printing a register block. For example, cpsr is a 32 bit register. Using the target's scratch type system we construct a type: ``` struct __attribute__((__packed__)) cpsr { uint32_t N : 1; uint32_t Z : 1; ... uint32_t EL : 2; uint32_t SP : 1; }; ``` If this register had unallocated bits in it, those would have been filled in by RegisterFlags as anonymous fields. A new option "SetChildPrintingDecider" is added so we can disable printing those. Important things about this type: * It is packed so that sizeof(struct cpsr) == sizeof(the real register). (this will hold for all flags types we create) * Each field has the same storage type, which is the same as the type of the raw register value. This prevents fields being spilt over into more storage units, as is allowed by most ABIs. * Each bitfield size matches that of its register field. * The most significant field is first. The last point is required because the most significant bit (MSB) being on the left/top of a print out matches what you'd expect to see in an architecture manual. In addition, having lldb print a different field order on big/little endian hosts is not acceptable. As a consequence, if the target is little endian we have to reverse the order of the fields in the value. The value of each field remains the same. For example 0b01 doesn't become 0b10, it just shifts up or down. This is needed because clang's type system assumes that for a struct like the one above, the least significant bit (LSB) will be first for a little endian target. We need the MSB to be first. Finally, if lldb's host is a different endian to the target we have to byte swap the host endian value to match the endian of the target's typesystem. | Host Endian | Target Endian | Field Order Swap | Byte Order Swap | |-------------|---------------|------------------|-----------------| | Little | Little | Yes | No | | Big | Little | Yes | Yes | | Little | Big | No | Yes | | Big | Big | No | No | Testing was done as follows: * Little -> Little * LE AArch64 native debug. * Big -> Little * s390x lldb running under QEMU, connected to LE AArch64 target. * Little -> Big * LE AArch64 lldb connected to QEMU's GDB stub, which is running an s390x program. * Big -> Big * s390x lldb running under QEMU, connected to another QEMU's GDB stub, which is running an s390x program. As we are not allowed to link core code to plugins directly, I have added a new plugin RegisterTypeBuilder. There is one implementation of this, RegisterTypeBuilderClang, which uses TypeSystemClang to build the CompilerType from the register fields. Reviewed By: jasonmolenda Differential Revision: https://reviews.llvm.org/D145580 -
Alex Zinenko authored
-
Timm Bäder authored
Introduced in 09effa70
-
Archibald Elliott authored
I landed this test with a typo, the callsites all show `fp16_inner` returning `half`, so the declaration should too.
-
Timm Bäder authored
Differential Revision: https://reviews.llvm.org/D146819
-
Archibald Elliott authored
The main change here is to add a `widenScalarToNextPow2` before the `clampScalar` so that non-power-of-two sizes between 32 and 64 are turned into s64 count trailing zeroes. However, if you make the legalisation rules depend on TypeIdx 0 (the output), then you still get crashes for the s65 testcase, which I solved by instead flipping the rules around to be about TypeIdx 1 (the input), with a `scalarSameSizeAs` at the end to tie index 0 to index 1. This, incidentally, is how things are written for `G_CTLZ`. Differential Revision: https://reviews.llvm.org/D147602
-
Alex Zinenko authored
Add a set of transform operations into the "structured" extension of the Transform dialect that allow one to select transformation targets more specifically than the currently available matching. In particular, add the mechanism for identifying the producers of operands (input and init in destination-passing style) and users of results, as well as mechanisms for reasoning about the shape of the iteration space. Additionally, add several transform operations to manipulate parameters that could be useful to implement more advanced selectors. Specifically, new operations let one produce and compare parameter values to implement shape-driven transformations. New operations are placed in separate files to decrease compilation time. Some relayering of the extension is necessary to avoid repeated generation of enums. Depends on D148013 Depends on D148014 Depends on D148015 Reviewed By: chelini Differential Revision: https://reviews.llvm.org/D148017
-
David Spickett authored
This teaches ProcessGDBRemote to look for "flags" nodes in the target XML that tell you what fields a register has. https://sourceware.org/gdb/onlinedocs/gdb/Target-Description-Format.html It will check for various invalid inputs like: * Flags nodes with 0 fields in them. * Start or end being > the register size. * Fields that overlap. * Required properties not being present (e.g. no name). * Flag sets being redefined. If anything untoward is found, we'll just drop the field or the flag set altogether. Register fields are a "nice to have" so LLDB shouldn't be crashing because of them, instead just log anything we throw away. So the user can fix their XML/file a bug with their vendor. Once that is done it will sort the fields and pass them to the RegisterFields class I added previously. There is no way to see these fields yet, so tests for this code will come later when the formatting code is added. The fields are stored in a map of unique pointers on the ProcessGDBRemote class. It will give out raw pointers on the assumption that the GDB process lives longer than the users of those pointers do. Which means RegisterInfo is still a trivial struct but we are properly destroying the fields when the GDB process ends. We can't store the fields directly in the map because adding new items may cause its storage to be reallocated, which would invalidate pointers we've already given out. Reviewed By: jasonmolenda, JDevlieghere Differential Revision: https://reviews.llvm.org/D145574
-
Stefan Gränitz authored
ObjectFileELF::ApplyRelocations() considered all 32-bit input objects to be i386 and didn't provide good error messages for AArch32 objects. Please find an example in https://github.com/llvm/llvm-project/issues/61948 While we are here, let' improve the situation for unsupported architectures as well. I think we should report the error here too and not silently fail (or crash with assertions enabled). Reviewed By: SixWeining Differential Revision: https://reviews.llvm.org/D147627
-
Job Noorman authored
The following test fails when enabling UBSan due to a left shift of a negative value: > runtime error: left shift of negative value -2 BOLT :: AArch64/ext-island-ref.s This patch fixes this by using a multiplication instead of a shift. Reviewed By: yota9 Differential Revision: https://reviews.llvm.org/D148218
-
Nicolas Vasilache authored
Previously, hoisting through an iter_arg would mistakenly yield the unpadded value and cast it to the padded value. This was incorrect and resulted in out-of-bounds accesses. The correct formulation is to yield the padded value and extract a smaller dynamic slice out of it. Differential Revision: https://reviews.llvm.org/D148173
-
David Spickett authored
These are used to store new state added by the Scalable Matrix Extension which is documented in https://developer.arm.com/documentation/ddi0616/aa/. The values match those defined by Linux, see: https://github.com/torvalds/linux/blob/e62252bc55b6d4eddc6c2bdbf95a448180d6a08d/include/uapi/linux/elf.h#L435 The ZT register(s) are added by SME2 which is not yet publicly documented but has support in LLVM and Linux already. Also added descriptions for SVE and PAC_MASK notes since those were missing. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D148126
-
Jorge Pinto Sousa authored
"static assertion failed" pointed to the static_assert token and then underlined the static assertion expression: <source>:3:1: error: static assertion failed static_assert(false); ^ ~~~~~ 1 error generated. See Godbolt: https://godbolt.org/z/r38booz59 Now it points to and highlights the assertion expression. Fixes https://github.com/llvm/llvm-project/issues/61951 Differential Revision: https://reviews.llvm.org/D147745 -
Timm Bäder authored
-
Timm Bäder authored
-
Kadir Cetinkaya authored
Instantiation pattern is null for incomplete template types and using specializaiton decl results in not seeing re-declarations. Differential Revision: https://reviews.llvm.org/D148158
-
David Spickett authored
This structure is supposed to be trivial, so we cannot simply do "= nullptr;" on the new member. Doing that means you are non trivial, regardless of whether you emulate the previously implied constructor somehow. The next option is to update every use of brace initialisation. Given that this is some hundreds of lines, this change just adds a dummy pointer that is set to nullptr. Subsequent changes will actually use that to point to register flags information. Note: This change is not clang-format-ted because it changes a bunch of areas that are not themselves formatted. It would just add noise. Reviewed By: jasonmolenda, JDevlieghere Differential Revision: https://reviews.llvm.org/D145568
-
Simon Pilgrim authored
[TTI][X86] getMinMaxCost - use existing float minnum/maxnum intrinsic cost values instead of maintaining a duplicate cost table Without fastmath (nnan) flags, minnum/maxnum must perform isnan handling as well as fmin/fmax - meaning the costs are notably higher, this is correctly handled in getIntrinsicInstrCost but was missing from the getMinMaxCost cost tables (which assumed fastmath). Followup to 63c38953 which handled the integer cases
-
David Green authored
This adjusts some of the tests to use the architecture features directly as opposed to -mcpu=cortex-m33 names.
-
Alex Zinenko authored
Conversions to the LLVM dialect have an option to use the "bare pointer" calling convention that converts memref types differently than the default convention. It has crept into the conversion of operations that are not related to calls but do require multiresult-to-struct packing. Use a similar mechanism for the latter without using the calling convention. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D148086
-
LLVM GN Syncbot authored
-
David Spickett authored
This models the "flags" node from GDB's target XML: https://sourceware.org/gdb/onlinedocs/gdb/Target-Description-Format.html This node is used to describe the fields of registers like cpsr on AArch64. RegisterFlags is a class that contains a list of register fields. These fields will be extracted from the XML sent by the remote. We assume that there is at least one field, that the fields are sorted in descending order and do not overlap. That will be enforced by the XML processor (the GDB client code in our case). The fields may not cover the whole register. To account for this RegisterFields will add anonymous padding fields so that sizeof(all fields) == sizeof(register). This will save a lot of hasssle later. Reviewed By: jasonmolenda, JDevlieghere Differential Revision: https://reviews.llvm.org/D145566
-
Quentin Colombet authored
This patch recognizes when tensor.pack/unpack operations are simple tensor.pad/unpad (a.k.a. tensor.extract_slice) and lowers them in a simpler sequence of instruction. For pack, instead of doing: ``` pad expand_shape transpose ``` we do ``` pad insert_slice ``` For unpack, instead of doing: ``` transpose collapse_shape extract_slice ``` we do ``` extract_slice ``` Note: returning nullptr for the transform dialect is fine. The related handles are just ignored by the following transformation. Differential Revision: https://reviews.llvm.org/D148159
-
Nicolas Vasilache authored
These old patterns are not in use in either MLIR or downstream projects except for one test. Additionally this is redundant with logic in the tensor.pad tiling implementation. Drop SplitPaddingPatterns to reduce entropy. Differential Revision: https://reviews.llvm.org/D148207
-
Simon Pilgrim authored
[TTI] Remove unnecessary default CostKind args from getExtendedReductionCost/getMulAccReductionCost wrappers. NFC. We should only ever call these from the TargetTransformInfo interface, which provides all the args.
-