- Jun 29, 2022
-
-
Mehdi Amini authored
-
Pengxuan Zheng authored
It's not clear what Microsoft's LIB.exe actually does based on the official description of the flag (link below). We can probably ignore it for now. https://docs.microsoft.com/en-us/cpp/build/reference/managing-a-library?view=msvc-170 Reviewed By: thieta Differential Revision: https://reviews.llvm.org/D128458
-
Aart Bik authored
Enforce the assumption made on tensor buffers explicitly. When in-place, reuse the buffer, but fill with all zeroes for the non-update case, since the kernel assumes all elements are written to. When not in-place, zero out the new buffer when materializing or when no-updates occur. Copy the original tensor value when updates occur. This prepares migrating to the new bufferization strategy, where these assumptions must be made explicit. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D128691
-
Craig Topper authored
I believe we already checked that the destination of the first CMOV is only used by the second CMOV so I don't think there is any reason we need the PHI to write the register that was used by the first CMOV. We can directly use the second CMOV destination and avoid the copy. This may be a left over from when the cascaded select handling was part of the main algorithm before it was refactored in D35685. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D128124
-
Mitch Phillips authored
Looks like with https://reviews.llvm.org/D127911, Windows emits more globals with mangled names into the IR. Relax the tests in order to allow these mangled names.
-
Valentin Clement authored
This patch fixes a couple of issues with the lowering of user defined assignment. This patch is part of the upstreaming effort from fir-dev branch. Reviewed By: klausler Differential Revision: https://reviews.llvm.org/D128730 Co-authored-by:
Jean Perier <jperier@nvidia.com> Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com>
-
- Jun 28, 2022
-
-
Lei Zhang authored
Reviewed By: hanchung Differential Revision: https://reviews.llvm.org/D128692
-
Arjun P authored
Also updated the tests, which were asserting the wrong behaviour. Reviewed By: Groverkss Differential Revision: https://reviews.llvm.org/D128735
-
Michał Górny authored
Sponsored by: The FreeBSD Foundation
-
Rahman Lavaee authored
[Propeller] Encode address offsets of basic blocks relative to the end of the previous basic blocks. This is a resurrection of D106421 with the change that it keeps backward-compatibility. This means decoding the previous version of `LLVM_BB_ADDR_MAP` will work. This is required as the profile mapping tool is not released with LLVM (AutoFDO). As suggested by @jhenderson we rename the original section type value to `SHT_LLVM_BB_ADDR_MAP_V0` and assign a new value to the `SHT_LLVM_BB_ADDR_MAP` section type. The new encoding adds a version byte to each function entry to specify the encoding version for that function. This patch also adds a feature byte to be used with more flexibility in the future. An use-case example for the feature field is encoding multi-section functions more concisely using a different format. Conceptually, the new encoding emits basic block offsets and sizes as label differences between each two consecutive basic block begin and end label. When decoding, offsets must be aggregated along with basic block sizes to calculate the final offsets of basic blocks relative to the function address. This encoding uses smaller values compared to the existing one (offsets relative to function symbol). Smaller values tend to occupy fewer bytes in ULEB128 encoding. As a result, we get about 17% total reduction in the size of the bb-address-map section (from about 11MB to 9MB for the clang PGO binary). The extra two bytes (version and feature fields) incur a small 3% size overhead to the `LLVM_BB_ADDR_MAP` section size. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D121346
-
Sam McCall authored
-
Aaron Ballman authored
This mostly finishes the DRs for C89, though there are still a few outliers which remain. It also corrects some of the statuses of DRs where it's not clear if it was fully resolved by the committee or not. As a drive-by, it also adds -fsyntax-only to the tests which are verifying diagnostic results. This was previously missed by accident.
-
Michał Górny authored
Sponsored by: The FreeBSD Foundation
-
Egor Zhdan authored
This is already possible for e.g. `cstring_literals`, but the entry for zerofill was unnamed. rdar://90336380 Differential Revision: https://reviews.llvm.org/D128654
-
Sam McCall authored
-
Sam McCall authored
Treat captures as a uniform list, rather than default-captures being special snowflakes that may only appear at the start. This accepts a larger set of (incorrect) code, and simplifies error-handling by making this fit into the usual homogeneous-list pattern. Differential Revision: https://reviews.llvm.org/D128708
-
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D128259
-
Sam McCall authored
This isn't allowed by the standard grammar but is allowed in C, and clang/GCC permit it as an extension. It avoids the need to determine which type of list we have in error-recovery. While here, also support array index designators `{ [4]=1 }` which are also legal in C, and common extensions in C++. Differential Revision: https://reviews.llvm.org/D128687 -
Joe Nash authored
Differential Revision: https://reviews.llvm.org/D128527
-
Valentin Clement authored
Fix bugs relating to support for characters of different kinds. Lowering was creating bad FIR and MLIR that crashed in conversion to LLVM IR. This patch is part of the upstreaming effort from fir-dev branch. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D128723 Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com>
-
Mehdi Amini authored
This attribute is similar to DenseElementsAttr but does not support splat. As such it has a much simpler API and does not need any smart iterator: it exposes direct ArrayRef access. A new syntax is introduced so that the generic printing/parsing looks like: [:i64 1, -2, 3] This attribute beings like an ArrayAttr but has a `:` token after the opening square brace to introduce the element type (supported are I8, I16, I32, I64, F32, F64) and the comma separated list for the data. This is particularly convenient for attributes intended to be small, like those referring to shapes. For example a `transpose` operation with a `dims` attribute could be defined as such: let arguments = (ins AnyTensor:$input, DenseI64ArrayAttr:$dims); let assemblyFormat = "$input `dims` `=` $dims attr-dict : type($input)"; And printed this way (the element type is elided in this case): transpose %input dims = [0, 2, 1] : tensor<2x3x4xf32> The C++ API for dims would just directly return an ArrayRef<int64> RFC: https://discourse.llvm.org/t/rfc-introduce-a-new-dense-array-attribute/63279 Recommit with a custom DenseArrayBaseAttrStorage class to ensure over-alignment of the storage to the largest type. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D123774
-
Valentin Clement authored
For the rapid triage push, just add a TODO for the degenerate POINTER assignment case. The LHD ought to be a variable of type !fir.box, but it is currently returning a shadow variable for the raw data pointer. More investigation is needed there. Make sure that conversions are applied in FORALL degenerate contexts. This patch is part of the upstreaming effort from fir-dev branch. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D128724 Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com>
-
Valentin Clement authored
Add lowering tests left behind during the upstreaming. This patch is part of the upstreaming effort from fir-dev branch. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D128721 Co-authored-by:
Jean Perier <jperier@nvidia.com> Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com>
-
Vladislav Khmelevsky authored
The gold linker veneers are written between functions without symbols, so we to handle it specially in BOLT. Vladislav Khmelevsky, Advanced Software Technology Lab, Huawei Differential Revision: https://reviews.llvm.org/D128082
-
Nikita Popov authored
Migrate extractelement, insertelement and shufflevector to use the FoldXYZ rather than CreateXYZ APIs. This is probably NFC in practice, because the places using InstSimplifyFolder probably aren't using vector operations.
-
Mehdi Amini authored
This reverts commit 508eb41d. UBSAN indicates some pointer mis-alignment I need to investigate
-
Yi Kong authored
PERF_COUNT_SW_DUMMY is introduced in Linux 3.12. Differential Revision: https://reviews.llvm.org/D128707
-
Pavel Samolysov authored
It makes sense to handle byval promotion in the same way as non-byval but also allowing `store` instructions. However, these should use the same checks as the `load` instructions do, i.e. be part of the `ArgsToPromote` collection. For these instructions, the check for interfering modifications can be disabled, though. The promotion algorithm itself has been modified a lot: all the accesses (i.e. loads and stores) are rewritten to the emitted `alloca` instructions. To optimize these new `alloca`s out, the `PromoteMemToReg` function from `Transforms/Utils/PromoteMemoryToRegister.cpp` file is invoked after promotion. In order to let the `PromoteMemToReg` promote as many `alloca`s as it is possible, there should be no `GEP`s from the `alloca`s. To eliminate the `GEP`s, its own `alloca` is generated for every argument part because a single `alloca` for the whole argument (that significantly simplifies the code of the pass though) unfortunately cannot be used. The idea comes from the following discussion: https://reviews.llvm.org/D124514#3479676 Differential Revision: https://reviews.llvm.org/D125485
-
Mehdi Amini authored
This attribute is similar to DenseElementsAttr but does not support splat. As such it has a much simpler API and does not need any smart iterator: it exposes direct ArrayRef access. A new syntax is introduced so that the generic printing/parsing looks like: [:i64 1, -2, 3] This attribute beings like an ArrayAttr but has a `:` token after the opening square brace to introduce the element type (supported are I8, I16, I32, I64, F32, F64) and the comma separated list for the data. This is particularly convenient for attributes intended to be small, like those referring to shapes. For example a `transpose` operation with a `dims` attribute could be defined as such: let arguments = (ins AnyTensor:$input, DenseI64ArrayAttr:$dims); let assemblyFormat = "$input `dims` `=` $dims attr-dict : type($input)"; And printed this way (the element type is elided in this case): transpose %input dims = [0, 2, 1] : tensor<2x3x4xf32> The C++ API for dims would just directly return an ArrayRef<int64> RFC: https://discourse.llvm.org/t/rfc-introduce-a-new-dense-array-attribute/63279 Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D123774
-
Ting Wang authored
opportunities There are straight forward splat load opportunities blocked by getNormalLoadInput(), since those cases involve consecutive bitcasts. Improve by looking through bitcasts. Reviewed By: nemanjai Differential Revision: https://reviews.llvm.org/D128703
-
Alex Bradbury authored
Implements the ratified RISC-V Base Cache Management Operation ISA Extension: Zicbop, as described in https://github.com/riscv/riscv-CMOs/blob/master/specifications/cmobase-v1.0.pdf. This is implemented in a separate patch to Zicbom and Zicboz due to it requiring a new ASM operand type to be defined. Differential Revision: https://reviews.llvm.org/D117433
-
Alex Bradbury authored
Implements the ratified RISC-V Base Cache Management Operation ISA Extensions: Zicbom and Zicboz, as described in https://github.com/riscv/riscv-CMOs/blob/master/specifications/cmobase-v1.0.pdf. Zicbop is implemented in a separate patch due to it requiring a new ASM operand type to be defined. As discussed in the relevant issue in the upstream spec https://github.com/riscv/riscv-CMOs/issues/47, the cbo.* instructions use the format (rs1) or 0(rs1) for their operand, similar to the AMOs. Differential Revision: https://reviews.llvm.org/D117432
-
Nikita Popov authored
Hopefully fixes clang-ppc64-aix. Apparently std::function can't be instantiated with a forward declared type in some environments.
-
Mehdi Amini authored
-
Mehdi Amini authored
-
Tim Northover authored
Before, we were trying to sign extend half -> float, and asserted in getNode.
-
Ting Wang authored
Reviewed By: shchenz Differential Revision: https://reviews.llvm.org/D128718
-
Matthias Springer authored
This was previous implemented as part of the BufferizableOpInterface of ForEachThreadOp. Moving the implementation to ParallelInsertSliceOp to be consistent with the remaining ops and to have a nice example op that can serve as a blueprint for other ops. Differential Revision: https://reviews.llvm.org/D128666
-
Guillaume Chatelet authored
-
LLVM GN Syncbot authored
-