- Apr 18, 2024
-
-
Alexey Bataev authored
We can try to vectorize long store sequences, if short ones were unsuccessful because of the non-profitable vectorization. It should not increase compile time significantly (stores are sorted already, complexity is n x log n), but vectorize extra code. Metric: size..text Program size..text results results0 diff test-suite :: External/SPEC/CINT2006/400.perlbench/400.perlbench.test 1088012.00 1088236.00 0.0% test-suite :: SingleSource/UnitTests/matrix-types-spec.test 480396.00 480476.00 0.0% test-suite :: External/SPEC/CINT2017rate/525.x264_r/525.x264_r.test 664613.00 664661.00 0.0% test-suite :: External/SPEC/CINT2017speed/625.x264_s/625.x264_s.test 664613.00 664661.00 0.0% test-suite :: External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test 2041105.00 2040961.00 -0.0% test-suite :: MultiSource/Applications/JM/lencod/lencod.test 836563.00 836387.00 -0.0% test-suite :: MultiSource/Benchmarks/7zip/7zip-benchmark.test 1035100.00 1032140.00 -0.3% In all benchmarks extra code gets vectorized Reviewers: RKSimon Reviewed By: RKSimon Pull Request: https://github.com/llvm/llvm-project/pull/88563 -
Abdul Raheem authored
Signed-off: Abdul Raheem Beigh <abdulraheembeigh@gmail.com>
-
Jorge Gorbe Moya authored
-
Michael Maitland authored
-
Jorge Gorbe Moya authored
-
fabrizio-indirli authored
-
Nick Desaulniers authored
Implement: - pthread_condattr_destroy - pthread_condattr_getclock - pthread_condattr_getpshared - pthread_condattr_init - pthread_condattr_setclock - pthread_condattr_setpshared Fixes: #88581
-
Alexandros Lamprineas authored
As explained in https://github.com/ARM-software/acle/pull/315 we are deprecating features which aren't adding any value. These are: sha1, pmull, dit, dgh, ebf16, sve-bf16, sve-ebf16, sve-i8mm, sve2-pmull128, memtag2, memtag3, ssbs2, bti, ls64_v, ls64_accdata
-
- Apr 17, 2024
-
-
Valentin Clement (バレンタイン クレメン) authored
Replace the runtime call to `AllocatableAllocate` for CUDA device variable to the newly added `fir.cuda_allocate` operation.
-
Valentin Clement (バレンタイン クレメン) authored
Add MemRead effect on the box operand as the descriptor might be read when performing the allocation of the data. Also update the expected type of the box operand to be a reference. Check in the verifier that this is a reference to a box or class type. This addresses the comment made post commit on #88586
-
Alexander Richardson authored
The code in this file dates back to 2012 when Clang's support for atomic builtins was still quite limited. The bugs referenced in the comment at the top of the file have long been fixed and using the compiler builtins directly should now generate slightly better code. Additionally, this allows using the atomic builtin header for platforms where the __sync_builtins are lacking (e.g. Arm Morello). This change does not introduce any code generation changes for __tsan_read*/__tsan_write* or __tsan_func_{entry,exit} on x86, which indicates the previously noted compiler issues have been fixed. We also have to touch the non-clang codepaths here since the only way we can make this work easily is by making the memory_order enum match the compiler-provided macros, so we have to update the debug checks that assumed the enum was always a bitflag. The one downside of this change is that 32-bit MIPS now definitely requires libatomic (but that may already have been needed for RMW ops). Reviewed By: dvyukov Pull Request: https://github.com/llvm/llvm-project/pull/84439 -
Krystian Stasiowski authored
Clang currently allows the following: ``` auto x = requires (this int) { true; }; ``` This patch addresses that. -
Robin Caloudis authored
Provide C23 `fetestexceptflag` function according to 7.6.4.6 in the latest [revision of the C standard](https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3096.pdf) from 2023-04-02. Closes https://github.com/llvm/llvm-project/issues/87565.
-
Dinar Temirbulatov authored
This reverts commit 4e85e1ff
-
Vlad Serebrennikov authored
This patch converts the enum into scoped enum, and moves it into its own header for the time being. It's definition is needed in `Sema.h`, and is going to be needed in upcoming `SemaObjC.h`. `Lookup.h` can't hold it, because it includes `Sema.h`.
-
Mark de Wever authored
The formatting of years has been done manually since the results of %Y outside the "typical" range may produce unexpected values. The same applies to %F which is identical to %Y-%m-%d. None of these conversion specifiers is affected by the locale used. So it's trivial to manually handle this case. This removes several platform specific ifdefs from the tests.
-
Florian Hahn authored
Since ne After a separate recipe has been introduced for wide loads in a9bafe91, we can directly check for load recipes in the early bail-out and remove the redundant bail out for stores.
-
Simon Pilgrim authored
These confuse the update_test_checks.py script when run by DOS cmd.exe
-
Simon Pilgrim authored
Inspired by the recent patches by @shamithoke - we have real scheduler model numbers for GFNI instructions now, allowing us to calculate an upper bounds costs table instead of performing it analytically.
-
Pavel Labath authored
-
yronglin authored
Signed-off-by:yronglin <yronglin777@gmail.com>
-
Rajveer Singh Bharadwaj authored
Resolves #88328
-
Guray Ozen authored
This PR adds NVGPU dialects' TensorMapDescriptorType in the py bindings. This is a follow-up issue from [this PR](https://github.com/llvm/llvm-project/pull/87153#discussion_r1546193095)
-
Jay Foad authored
Use OtherPredicates to avoid interfering with other uses of SubtargetPredicate for GFX12.
-
Florian Hahn authored
Factor out logic to collect all users recursively to be re-used in https://github.com/llvm/llvm-project/pull/87816.
-
Aaron Ballman authored
This paper is about type compatibility rules that changed in C99, but this is only applicable across translation units and so there's nothing for us to test. The specific change was that C89 allowed different tag types (e.g., struct and union) to be compatible and C99 tightened that restriction. This is a case where the user gets whatever they get if they link two TUs with incompatible tag types.
-
Luke Lau authored
If an instruction between MI and NextMI uses VL or VTYPE we demand the respective fields so as to not clobber them at their uses. But we don't consider if something might modify VL or VTYPE, and will happily coalesce two vsetvlis when we need to preserve them. This fixes this by skipping to the next vsetvli. Demanding the fields isn't enough, as we need to preserve the VL and VTYPE values even if no fields are demanded. In practice this doesn't happen, presumably due to there not being any instructions that write to VL or VTYPE without reading them. But I noticed this whilst working on a separate patch and split it out.
-
Zaara Syeda authored
This patch adds the pseudo op ADDItocL for 32-bit large code-model support for toc-data.
-
Luke Lau authored
-
Krzysztof Parzyszek authored
Re-enable an old assertion in `decreaseSetPressure`.
-
Oleksandr "Alex" Zinenko authored
This functionality is available in C++, make it available in Python directly to operate on transform modules.
-
DianQK authored
-
Fabian Ritter authored
The link to the Heterogeneous-race-free Memory Models ASPLOS'14 paper by Hower et al. pointed to a bogus website, probably because the domain ownership has changed. This patch updates it to a version hosted on research.cs.wisc.edu.
-
Kevin P. Neal authored
Correct missing cases in a switch that result in @llvm.vp.fma.v4f32 getting lowered to a constrained fma intrinsic. Vector predicated lowering to contrained intrinsics is not supported currently, and there's no consensus on the path forward. We certainly shouldn't be introducing constrained intrinsics into a function that isn't strictfp. Problem found with D146845.
-
Pavel Labath authored
After 281d7160, llvm generates 32-bit relocations, which overflow when we load these objects into high memory. Interestingly, setting the code model to "large" does not help here (perhaps it is the default?). I'm not completely sure that this is the right thing to do, but it doesn't seem to cause any ill effects. I'll follow up with the author of that patch about the expected behavior here.
-
Quentin Dian authored
[TailDuplicator] Add maximum predecessors and successors to consider tail duplicating blocks (#78582) Fixes #78578. Duplicating a BB which has both multiple predecessors and successors will result in a complex CFG and also may cause huge amount of PHI nodes. See https://github.com/llvm/llvm-project/issues/78578#issuecomment-1962363580 for a detailed description of the limit.
-
Luke Lau authored
-
Oleksandr "Alex" Zinenko authored
Greedy rewrite driver has options to control the number of rewrites applies. Expose those via the corresponding transform op.
-
Simon Pilgrim authored
-
Louis Dionne authored
Also add tests for those, and add a few missing requirements to testing iterators in the test suite.
-