- Mar 09, 2023
-
-
Michael Kruse authored
Polly-ACC is unmaintained and since it has never been ported to the NPM pipeline, since D136621 it is not even accessible anymore without manually specifying the passes on the `opt` command line. Since there is no plan to put it to a maintainable state, remove it from Polly. Reviewed By: grosser Differential Revision: https://reviews.llvm.org/D142580
-
wren romano authored
This helps to reduce the confusion from using `unsigned` everywhere. Depends On D145606 Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D145611
-
Pavel Iliin authored
Differential Revision: https://reviews.llvm.org/D145538
-
wren romano authored
The copy assignment is already implicitly deleted, but making it explicit helps clean up compilation error messages. Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D145606
-
Kirill Stoimenov authored
Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D145615
-
Siva Chandra authored
-
Chia-hung Duan authored
Shuffle the regions' base address so that the layout of all regions is less predictable. Reviewed By: cferris, cryptoad Differential Revision: https://reviews.llvm.org/D145407
-
Michael Kruse authored
The linker-flags.f90 test checks for the linker command line. The `-target` indicates cross-compiling, the toolchain executables themselves are still running on the native platform. If it is Windows, the driver will try to fully resolve the path to `ld` which may include an `.exe` suffix. In my case, it resolves to the MinGW installation (`"C:\\tools\\msys64\\usr\\bin\\ld.exe"`) found in `PATH`. The GNU ld that comes with the MSYS2 distribution does not support `elf64lppc` or MacOS emulation modes (`acosx_version_min`), but the test also does not require executing the linker. Reviewed By: awarzynski Differential Revision: https://reviews.llvm.org/D144592
-
Devajith Valaparambil Sreeramaswamy authored
Decompose conv_2d -> conv_1d. This MR follows a similar approach to https://reviews.llvm.org/D112928. This patch adds support to convert conv_2D operation with either unit height or unit width to conv_1D operation. This is useful when 2D convolution is tiled to have a single dimension for either height or width and then can be vectorized once it is decomposed into 1D convolution. This patch https://reviews.llvm.org/D145160 adds vector support for linalg.conv_1d operation and thereby allowing us to vectorize linalg.conv_2d operation after proper tiling. This missing feature is reported here: https://discourse.llvm.org/t/vectorization-of-convolution-op/60458. Reviewed By: hanchung Differential Revision: https://reviews.llvm.org/D145162
-
Anna Thomas authored
-
Devajith Valaparambil Sreeramaswamy authored
This MR add vectorization support for linalg.conv_1D operation. Reviewed By: nicolasvasilache, hanchung, dcaballe, vmurali Differential Revision: https://reviews.llvm.org/D145160
-
Augusto Noronha authored
Differential Revision: https://reviews.llvm.org/D145612
-
Craig Topper authored
Makes it harder to write an inexact constant that gets parsed as a valid constant.
-
Vitaly Buka authored
This is max acceptable value with pow of 2 for DefaultSizeClassMap, the same as for ASAN. Reviewed By: kstoimenov Differential Revision: https://reviews.llvm.org/D145536
-
Stanislav Mekhanoshin authored
This turns an idempotent atomic operation into an atomic load. Fixes: SWDEV-385135 Differential Revision: https://reviews.llvm.org/D144759
-
Min-Yih Hsu authored
This patch adds support for 'm', 'Q', and 'U' memory constraints. Differential Revision: https://reviews.llvm.org/D143529
-
Min-Yih Hsu authored
In order to support inline asm with memory constraints, AsmPrinter::PrintAsmMemOperand needs to be implemented, which has lots of overlaps with MCInstPrinter especially on the format of complex addressing modes. This patch factors out the common printing logics from MCInstPrinter into a separate class inherited by both AsmPrinter and MCInstPrinter, in which the derived classes only need to provide primitives like printOperand and printDisp. This change is basically NFC. See D143529 for changes on AsmPrinter. Differential Revision: https://reviews.llvm.org/D143528
-
Alexey Bataev authored
-
V Donaldson authored
-
Chia-hung Duan authored
Given the memory group, we are unlikely to need a huge page map to record entire region. This CL reduces the size of default page map buffer from 2048 to 512 and increase the number of static buffers to 2. Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D144754
-
Alexey Bataev authored
The indeces of the dependent loops are properly ordered, just start from 1, so need just subtract 1 to get correct loop index. Differential Revision: https://reviews.llvm.org/D145514
-
Nikolas Klauser authored
Reviewed By: ldionne, #libc, #libc_abi Spies: #libc_vendors, smeenai, libcxx-commits Differential Revision: https://reviews.llvm.org/D145320
-
Valentin Clement authored
When a derived-type as no component, its elem_len will be set to zero when emboxed. Update the function to let empty derived-type pointer/target succeed the test. Example extracted from gfortran test pointer_init_8 ``` module m type :: c end type c type, extends(c) :: d end type d type(c), target :: x end module use m class(c), pointer :: px => x if (.not. associated(px, x)) STOP 1 end ``` Reviewed By: klausler Differential Revision: https://reviews.llvm.org/D145604
-
Mark de Wever authored
During the implementation of P2286 a second Unicode decoder was added. The original decoder was only used for the width estimation. Changing an ill-formed Unicode sequence to the replacement character, works properly for this use case. For P2286 an ill-formed Unicode sequence needs to be formatted as a sequence of code units. The exact wording in the Standard as a bit unclear and there was odd example in the WP. This made it hard to use the same decoder. SG16 determined the odd example in the WP was a bug and this has been fixed in the WP. This made it possible to combine the two decoders. The P2286 decoder kept track of the size of the ill-formed sequence. However this was not needed since the output algorithm needs to keep track of size of a well-formed and an ill-formed sequence. So this feature has been removed. The error status remains since it's needed for P2286, the grapheme clustering can ignore this unneeded value. (In general, grapheme clustering is only has specified behaviour for Unicode. When the string is in a non-Unicode encoding there are no requirements. Ill-formed Unicode is a non-Unicode encoding. Still libc++ does a best effort estimation.) There UTF-8 decoder accepted several ill-formed sequences: - Values in the surrogate range U+D800..U+DFFF. - Values encoded in more code units than required, for example 0+0020 in theory can be encoded using 1, 2, 3, or 4 were accepted. This is not allowed by the Unicode Standard. - Values larger than U+10FFFF were not always rejected. Reviewed By: #libc, ldionne, tahonermann, Mordante Differential Revision: https://reviews.llvm.org/D144346
-
Simon Pilgrim authored
Check we write to the entire memory span of the inlined memset Simplifies future update_llc_test_checks regenerations
-
Arthur Eubanks authored
Funnel fetching and building LLVM instructions into GettingStarted. Modernize the build steps a little. Remove comments saying CMAKE_BUILD_TYPE defaults to Debug as that's not true anymore (must explicitly pass it). Reviewed By: MaskRay, hans Differential Revision: https://reviews.llvm.org/D145413
-
Craig Topper authored
Integers are ambiguous as to whether it's an index or an FP value without a decimal. Looks like maybe AArch64 equivalent treates integers in hex as index and any other integer as a FP value without a decimal. We need to work with the RVI community to decide what we should do.
-
Renaud-K authored
Differential revision: https://reviews.llvm.org/D145602
-
Mikhail R. Gadelha authored
This patch now enables full build. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D145594
-
Thomas Raoux authored
Fix bug when pipelining while interleaving stages. Re-do the logic to only consider cloned operands when updating the use-def chain. Differential Revision: https://reviews.llvm.org/D145598
-
Siva Chandra authored
Its test is currently failing of real riscv64 hardware.
-
Han Zhu authored
Second try at A-Wadhwani's https://reviews.llvm.org/D132096, which was reverted. The original patch had three issues: * https://reviews.llvm.org/D134032, which bjope kindly fixed. That patch is merged into this one. * [GHI #57796](https://github.com/llvm/llvm-project/issues/57796). Fixed and added a test. * [GHI #57821](https://github.com/llvm/llvm-project/issues/57821). I believe this is an undefined behavior which is not the fault of the original patch. Please see the issue for more details. Original diff summary: This patch adds additional vector types to be considered when doing promotion in SROA, based on the types of the store and load slices. This provides more promotion opportunities, by potentially using an optimal "intermediate" vector type. For example, the following code would currently not be promoted to a vector, since `__m128i` is a `<2 x i64>` vector. ``` __m128i packfoo0(int a, int b, int c, int d) { int r[4] = {a, b, c, d}; __m128i rm; std::memcpy(&rm, r, sizeof(rm)); return rm; } ``` ``` packfoo0(int, int, int, int): mov dword ptr [rsp - 24], edi mov dword ptr [rsp - 20], esi mov dword ptr [rsp - 16], edx mov dword ptr [rsp - 12], ecx movaps xmm0, xmmword ptr [rsp - 24] ret ``` By also considering the types of the elements, we could find that the `<4 x i32>` type would be valid for promotion, hence removing the memory accesses for this function. In other words, we can explore other new vector types, with the same size but different element types based on the load and store instructions from the Slices, which can provide us more promotion opportunities. Additionally, the step for removing duplicate elements from the `CandidateTys` vector was not using an equality comparator, which has been fixed. Differential Revision: https://reviews.llvm.org/D143225
-
Aaron Ballman authored
This adds test coverage for N2607, which makes arrays and their elements identically qualified. Clang already implements much of the functionality from this paper, but is still missing some support. It also adds some details to the C status page so users have this information as well.
-
Peiming Liu authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D145603
-
Ganesh Gopalasubramanian authored
-
Mikhail R. Gadelha authored
Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D145593
-
Dave Lee authored
The `v` (`frame variable`) command can directly access ivars/fields of `this` or `self`. Such as `v field`, instead of `v this->field`. This change relaxes the criteria for finding `this`/`self` variables. There are cases where a `this`/`self` variable does exist, but up to now the `v` command has not made use of it. The user would have to explicitly run `v this->field` or `self->_ivar` to access ivars. This change allows such cases to also work (without explicitly dereferencing `this`/`self`). A very common example in Objective-C (and Swift) is weakly capturing `self`: ``` __weak Type *weakSelf = self; void (^block)(void) = ^{ Type *self = weakSelf; // Re-establish strong reference. // `v _ivar` should work just as well as `v self->_ivar`. }; ``` In this case, `self` exists but `v` would not have used it. With this change, the fact that a variable named `self` exists is enough for it to be used. Differential Revision: https://reviews.llvm.org/D145276 -
Mikhail R. Gadelha authored
Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D145592
-
Adam Paszke authored
Right now the bindings assume that all DenseElementsAttrs correspond to tensor values, making it impossible to create vector-typed constants. I didn't want to change the API significantly, so I opted for reusing the current signature of `.get`. Its `type` argument now accepts both element types (in which case `shape` and `signless` can be specified too), or a shaped type, which specifies the full type of the created attr (`shape` cannot be specified in that case). Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D145053
-
Florian Hahn authored
This patch adds the predicate as additional operand to VPReplicateRecipe during initial construction. The predicated recipes are later moved into replicate regions. This simplifies constructions and some VPlan transformations, like fixed-order recurrence handling. It also improves codegen in some cases (e.g. for in-loop reductions), because the recipes remain in the same block. Reviewed By: Ayal Differential Revision: https://reviews.llvm.org/D143865
-