- May 03, 2021
-
-
Nikita Popov authored
We currently can't determine any exit counts here, because there is no "controlling exit".
-
- May 02, 2021
-
-
William S. Moses authored
1) Canonicalize IndexCast(SExt(x)) => IndexCast(x) 2) Provide constant folds of sign_extend and truncate Differential Revision: https://reviews.llvm.org/D101714
-
Mark de Wever authored
- Use the proper review for 'Fix integral conformance'. - Mark 'Fix integral conformance' as completed. - Move some tasks to in progress.
-
Juneyoung Lee authored
This is a patch that folds select of select to salvage some optimizations after select -> and/or folding is disabled. ``` select (select a, true, b), c, false -> select a, c, false select c, (select a, true, b), false -> select c, a, false if c implies that b is false (isImpliedCondition). ``` https://alive2.llvm.org/ce/z/ANatjt, https://alive2.llvm.org/ce/z/rv8zTB ``` sel (sel c, a, false), true, (sel !c, b, false) -> sel c, a, b sel (sel !c, a, false), true, (sel c, b, false) -> sel c, b, a ``` https://alive2.llvm.org/ce/z/U2kp-t, https://alive2.llvm.org/ce/z/bc88EE See D101191 Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D101375
-
Juneyoung Lee authored
-
William S. Moses authored
Differential Revision: https://reviews.llvm.org/D101712
-
Christopher Di Bella authored
C++20 revised the definition of what it means to be an iterator. While all _Cpp17InputIterators_ satisfy `std::input_iterator`, the reverse isn't true. D100271 introduces a new test adaptor to accommodate this new definition (`cpp20_input_iterator`). In order to help readers immediately distinguish which input iterator adaptor is _Cpp17InputIterator_, the current `input_iterator` adaptor has been prefixed with `cpp17_`. Differential Revision: https://reviews.llvm.org/D101242
-
Juneyoung Lee authored
-
Arthur Eubanks authored
To reduce dependence on pointee types for opaque pointers. Reviewed By: dblaikie Differential Revision: https://reviews.llvm.org/D101706
-
Juneyoung Lee authored
This is an NFC that reruns update_test_checks.py on the tests that are going to be updated in D101191.
-
Juneyoung Lee authored
This is a patch that adds ctpop intrinsics to propagatesPoison. Splitted from D101191
-
eopXD authored
Added canonicalization for vector_load and vector_store. An existing pattern SimplifyAffineOp can be reused to compose maps that supplies result into them. Added AffineVectorStoreOp and AffineVectorLoadOp into static_assert of SimplifyAffineOp to allow operation to use it. This fixes the bug filed: https://bugs.llvm.org/show_bug.cgi?id=50058 Reviewed By: bondhugula Differential Revision: https://reviews.llvm.org/D101691
-
Juneyoung Lee authored
This update supports the following transformation: ``` select(extract(mul_with_overflow(a, _), _), (a == 0), false) => and(extract(mul_with_overflow(a, _), _), (a == 0)) ``` which is correct because if `a` was poison the select's condition was also poison. This update is splitted from D101423.
-
LLVM GN Syncbot authored
-
Juneyoung Lee authored
As discussed in D101191, this patch adds a poison-safe folding of overflow bit check: ``` %Op0 = icmp ne i4 %X, 0 %Agg = call { i4, i1 } @llvm.[us]mul.with.overflow.i4(i4 %X, i4 %Y) %Op1 = extractvalue { i4, i1 } %Agg, 1 %ret = select i1 %Op0, i1 %Op1, i1 false => %Y.fr = freeze %Y %Agg = call { i4, i1 } @llvm.[us]mul.with.overflow.i4(i4 %X, i4 %Y.fr) %Op1 = extractvalue { i4, i1 } %Agg, 1 %ret = %Op1 ``` https://alive2.llvm.org/ce/z/zgPUGT https://alive2.llvm.org/ce/z/h2gZ_6 Note that there are cases where inserting freeze is not necessary: e.g. %Y is `noundef`. In this case, LLVM is already good because `%ret` is already successfully folded into `and`, triggering the pre-existing optimization in InstSimplify: https://godbolt.org/z/v6qena15K Differential Revision: https://reviews.llvm.org/D101423 -
Juneyoung Lee authored
-
Yaxun (Sam) Liu authored
Choose optimized device lib bitcode by fp options for performance. Reviewed by: Artem Belevich, Fangrui Song Differential Revision: https://reviews.llvm.org/D101654
-
Fangrui Song authored
-
Harald van Dijk authored
X32 uses 32-bit ELF object files with 32-bit alignment, so the .note.gnu.property section needs to be emitted as it is for X86. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D101689
-
Nikita Popov authored
If V & Mask != 0, we know that at least one of the bits in Mask must be set, so the value must be >= the lowest bit in Mask.
-
Nikita Popov authored
-
Michał Górny authored
Commit 88a5b35d changed the API of RegisterInfoPOSIX_arm64 and effectively broke the FreeBSD plugin. Update it to work with the new API. Differential Revision: https://reviews.llvm.org/D101521
-
Craig Topper authored
-
Chris Lattner authored
Don't get RegionKindInterface if we won't use it. Noticed by inspection.
-
Roman Lebedev authored
Introduce basic schedule model for AMD Zen 3 CPU's, a.k.a `znver3`. This is fully built from scratch, from llvm-mca measurements and documented reference materials. Nothing was copied from `znver2`/`znver1`. I believe this is in a reasonable state of completion for inclusion, probably better than D52779 `bdver2` was :) Namely: * uops are pretty spot-on (at least what llvm-mca can measure) {F16422596} * latency is also pretty spot-on (at least what llvm-mca can measure) {F16422601} * throughput is within reason {F16422607} I haven't run much benchmarks with this, however RawSpeed benchmarks says this is beneficial: {F16603978} {F16604029} I'll call out the obvious problems there: * i didn't really bother with X87 instructions * i didn't really bother with obviously-microcoded/system instructions * There are large discrepancy in throughput for `mr` and `rm` instructions. I'm not really sure if it's a modelling defect that needs to be fixed, or it's a defect of measurments. * Pipe distributions are probably bad :) I can't do much here until AMD allows that to be fixed by documenting the appropriate counters and updating libpfm That being said, as @RKSimon notes: >>! In D94395#2647381, @RKSimon wrote: > I'll mention again that all the znver* models appear to be very inaccurate wrt SIMD/FPU instructions <...> so how much worse this could possibly be?! Things that aren't there: * Various tunings: zero idioms, etc. That is follow-ups. Differential Revision: https://reviews.llvm.org/D94395 -
Pratyush Das authored
The code example: ``` constexpr const char kEta[] = "Eta"; template <const char*, typename T> class Column {}; using quick = Column<kEta,double>; void lookup() { quick c1; c1.ls(); } ``` emits error: no member named 'ls' in 'Column<&kEta, double>'. The patch fixes the printed type name by not printing the ampersand for array types. Differential Revision: https://reviews.llvm.org/D36368 -
Chris Lattner authored
-
- May 01, 2021
-
-
Nikita Popov authored
This seems to be a leftover from when the BackedgeTakenInfo stored multiple exit counts with manual memory management. At some point this was switchted to a simple vector, and there should be no need to micro-manage the clearing anymore. We can simply drop the loop from the map and the the destructor do its job.
-
Nikita Popov authored
This is checked again directly below this condition.
-
LemonBoy authored
Apply the same logic used to check if CMPXCHG nodes should be expanded at -O0: the register allocator may end up spilling some register in between the atomic load/store pairs, breaking the atomicity and possibly stalling the execution. Fixes PR48017 Reviewed By: efriedman Differential Revision: https://reviews.llvm.org/D101163
-
Nikita Popov authored
-
LemonBoy authored
Pre-requisite for D101163, the `NOLSE-0O` case shows registers being spilled inside the rmw loop. Use two separate prefixes for the `LSE-O0` case as some outputs differ only by a comment that update_llc_test_checks.py ignores but lit does not, causing the test to fail unexpectedly when run.
-
Arthur O'Dwyer authored
This reverts another of the macros just added in D101613, because it turns out that the <optional> and <filesystem> headers use the identifier __opt.
-
Arthur O'Dwyer authored
This reverts one of the macros just added in D101613, because it turns out that the <utility> header actually uses the identifiers __x, __y, __z. We probably *shouldn't* use __z if it's reserved on Windows; but since it's not causing us any active problem even on Windows, I think this is the safest way to unbreak the test.
-
Yaxun (Sam) Liu authored
AMDGPU backend need to know whether floating point opcodes that support exception flag gathering quiet and propagate signaling NaN inputs per IEEE754-2008, which is conveyed by a function attribute "amdgpu-ieee". "amdgpu-ieee"="false" turns this off. Without this function attribute backend assumes it is on for compute functions. -mamdgpu-ieee and -mno-amdgpu-ieee are added to Clang to control this function attribute. By default it is on. -mno-amdgpu-ieee requires -fno-honor-nans or equivalent. Reviewed by: Matt Arsenault Differential Revision: https://reviews.llvm.org/D77013
-
LemonBoy authored
Pre-requisite for D101163, the NOLSE-0O case shows registers being spilled inside the rmw loop.
-
Nikita Popov authored
or-ne is the conjugated pattern for and-eq.
-
Vitaly Buka authored
-
Vitaly Buka authored
Attribute guaranties safe static initialization of globals. Reviewed By: hctim Differential Revision: https://reviews.llvm.org/D101514
-