- May 02, 2021
-
-
Juneyoung Lee authored
-
Arthur Eubanks authored
To reduce dependence on pointee types for opaque pointers. Reviewed By: dblaikie Differential Revision: https://reviews.llvm.org/D101706
-
Juneyoung Lee authored
This is an NFC that reruns update_test_checks.py on the tests that are going to be updated in D101191.
-
Juneyoung Lee authored
This is a patch that adds ctpop intrinsics to propagatesPoison. Splitted from D101191
-
eopXD authored
Added canonicalization for vector_load and vector_store. An existing pattern SimplifyAffineOp can be reused to compose maps that supplies result into them. Added AffineVectorStoreOp and AffineVectorLoadOp into static_assert of SimplifyAffineOp to allow operation to use it. This fixes the bug filed: https://bugs.llvm.org/show_bug.cgi?id=50058 Reviewed By: bondhugula Differential Revision: https://reviews.llvm.org/D101691
-
Juneyoung Lee authored
This update supports the following transformation: ``` select(extract(mul_with_overflow(a, _), _), (a == 0), false) => and(extract(mul_with_overflow(a, _), _), (a == 0)) ``` which is correct because if `a` was poison the select's condition was also poison. This update is splitted from D101423.
-
LLVM GN Syncbot authored
-
Juneyoung Lee authored
As discussed in D101191, this patch adds a poison-safe folding of overflow bit check: ``` %Op0 = icmp ne i4 %X, 0 %Agg = call { i4, i1 } @llvm.[us]mul.with.overflow.i4(i4 %X, i4 %Y) %Op1 = extractvalue { i4, i1 } %Agg, 1 %ret = select i1 %Op0, i1 %Op1, i1 false => %Y.fr = freeze %Y %Agg = call { i4, i1 } @llvm.[us]mul.with.overflow.i4(i4 %X, i4 %Y.fr) %Op1 = extractvalue { i4, i1 } %Agg, 1 %ret = %Op1 ``` https://alive2.llvm.org/ce/z/zgPUGT https://alive2.llvm.org/ce/z/h2gZ_6 Note that there are cases where inserting freeze is not necessary: e.g. %Y is `noundef`. In this case, LLVM is already good because `%ret` is already successfully folded into `and`, triggering the pre-existing optimization in InstSimplify: https://godbolt.org/z/v6qena15K Differential Revision: https://reviews.llvm.org/D101423 -
Juneyoung Lee authored
-
Yaxun (Sam) Liu authored
Choose optimized device lib bitcode by fp options for performance. Reviewed by: Artem Belevich, Fangrui Song Differential Revision: https://reviews.llvm.org/D101654
-
Fangrui Song authored
-
Harald van Dijk authored
X32 uses 32-bit ELF object files with 32-bit alignment, so the .note.gnu.property section needs to be emitted as it is for X86. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D101689
-
Nikita Popov authored
If V & Mask != 0, we know that at least one of the bits in Mask must be set, so the value must be >= the lowest bit in Mask.
-
Nikita Popov authored
-
Michał Górny authored
Commit 88a5b35d changed the API of RegisterInfoPOSIX_arm64 and effectively broke the FreeBSD plugin. Update it to work with the new API. Differential Revision: https://reviews.llvm.org/D101521
-
Craig Topper authored
-
Chris Lattner authored
Don't get RegionKindInterface if we won't use it. Noticed by inspection.
-
Roman Lebedev authored
Introduce basic schedule model for AMD Zen 3 CPU's, a.k.a `znver3`. This is fully built from scratch, from llvm-mca measurements and documented reference materials. Nothing was copied from `znver2`/`znver1`. I believe this is in a reasonable state of completion for inclusion, probably better than D52779 `bdver2` was :) Namely: * uops are pretty spot-on (at least what llvm-mca can measure) {F16422596} * latency is also pretty spot-on (at least what llvm-mca can measure) {F16422601} * throughput is within reason {F16422607} I haven't run much benchmarks with this, however RawSpeed benchmarks says this is beneficial: {F16603978} {F16604029} I'll call out the obvious problems there: * i didn't really bother with X87 instructions * i didn't really bother with obviously-microcoded/system instructions * There are large discrepancy in throughput for `mr` and `rm` instructions. I'm not really sure if it's a modelling defect that needs to be fixed, or it's a defect of measurments. * Pipe distributions are probably bad :) I can't do much here until AMD allows that to be fixed by documenting the appropriate counters and updating libpfm That being said, as @RKSimon notes: >>! In D94395#2647381, @RKSimon wrote: > I'll mention again that all the znver* models appear to be very inaccurate wrt SIMD/FPU instructions <...> so how much worse this could possibly be?! Things that aren't there: * Various tunings: zero idioms, etc. That is follow-ups. Differential Revision: https://reviews.llvm.org/D94395 -
Pratyush Das authored
The code example: ``` constexpr const char kEta[] = "Eta"; template <const char*, typename T> class Column {}; using quick = Column<kEta,double>; void lookup() { quick c1; c1.ls(); } ``` emits error: no member named 'ls' in 'Column<&kEta, double>'. The patch fixes the printed type name by not printing the ampersand for array types. Differential Revision: https://reviews.llvm.org/D36368 -
Chris Lattner authored
-
- May 01, 2021
-
-
Nikita Popov authored
This seems to be a leftover from when the BackedgeTakenInfo stored multiple exit counts with manual memory management. At some point this was switchted to a simple vector, and there should be no need to micro-manage the clearing anymore. We can simply drop the loop from the map and the the destructor do its job.
-
Nikita Popov authored
This is checked again directly below this condition.
-
LemonBoy authored
Apply the same logic used to check if CMPXCHG nodes should be expanded at -O0: the register allocator may end up spilling some register in between the atomic load/store pairs, breaking the atomicity and possibly stalling the execution. Fixes PR48017 Reviewed By: efriedman Differential Revision: https://reviews.llvm.org/D101163
-
Nikita Popov authored
-
LemonBoy authored
Pre-requisite for D101163, the `NOLSE-0O` case shows registers being spilled inside the rmw loop. Use two separate prefixes for the `LSE-O0` case as some outputs differ only by a comment that update_llc_test_checks.py ignores but lit does not, causing the test to fail unexpectedly when run.
-
Arthur O'Dwyer authored
This reverts another of the macros just added in D101613, because it turns out that the <optional> and <filesystem> headers use the identifier __opt.
-
Arthur O'Dwyer authored
This reverts one of the macros just added in D101613, because it turns out that the <utility> header actually uses the identifiers __x, __y, __z. We probably *shouldn't* use __z if it's reserved on Windows; but since it's not causing us any active problem even on Windows, I think this is the safest way to unbreak the test.
-
Yaxun (Sam) Liu authored
AMDGPU backend need to know whether floating point opcodes that support exception flag gathering quiet and propagate signaling NaN inputs per IEEE754-2008, which is conveyed by a function attribute "amdgpu-ieee". "amdgpu-ieee"="false" turns this off. Without this function attribute backend assumes it is on for compute functions. -mamdgpu-ieee and -mno-amdgpu-ieee are added to Clang to control this function attribute. By default it is on. -mno-amdgpu-ieee requires -fno-honor-nans or equivalent. Reviewed by: Matt Arsenault Differential Revision: https://reviews.llvm.org/D77013
-
LemonBoy authored
Pre-requisite for D101163, the NOLSE-0O case shows registers being spilled inside the rmw loop.
-
Nikita Popov authored
or-ne is the conjugated pattern for and-eq.
-
Vitaly Buka authored
-
Vitaly Buka authored
Attribute guaranties safe static initialization of globals. Reviewed By: hctim Differential Revision: https://reviews.llvm.org/D101514
-
Martin Storsjö authored
This reverts a224bf8e and fixes the underlying issue. The underlying issue is simply that MSVC headers contains a define like "#define __in", where __in is one macro in the MSVC Source Code Annotation Language, defined in sal.h Just use a different variable name than "__in" __indirectly_readable_impl, and add "__in" to nasty_macros.h just like the existing __out. (Also adding a couple more potentially conflicting ones.) Differential Revision: https://reviews.llvm.org/D101613
-
Nathan James authored
-
Martin Storsjö authored
If libc++ is built as a DLL, calls to operator new within the DLL aren't overridden if a user provides their own operator in calling code. Therefore, the alloc counter doesn't pick up on allocations done within std::string, so skip that check if running on windows. (Technically, we could keep the checks if running on windows when not built as a DLL, but trying to keep the conditionals simple.) Differential Revision: https://reviews.llvm.org/D100219
-
Nathan Chancellor authored
Revert "Re-reapply "[DebugInfo] Use variadic debug values to salvage BinOps and GEP instrs with non-const operands"" This reverts commit 791930d7, as per https://llvm.org/docs/DeveloperPolicy.html#patch-reversion-policy. I observed breakage with the Linux kernel, as reported at https://reviews.llvm.org/D91722#2724321 Fixes exist at https://reviews.llvm.org/D101523 https://reviews.llvm.org/D101540 but they have not landed so to unbreak the tree for the weekend, revert this commit. Commit b11e4c99 ("Revert "[DebugInfo] Drop DBG_VALUE_LISTs with an excessive number of debug operands"") only reverted one follow-up fix, not the original patch that broke the kernel. e
-
Arthur O'Dwyer authored
A span has no idea what container (if any) "owns" its iterators, nor under what circumstances they might become invalidated. However, continue to use `__wrap_iter<T*>` instead of raw `T*` outside of debug mode, because we've been shipping `std::span` since Clang 7 and ldionne doesn't want to break ABI. (Namely, the mangling of functions taking `span::iterator` as a parameter.) Permit using raw `T*` there, but only under an ABI macro: `_LIBCPP_ABI_SPAN_POINTER_ITERATORS`. Differential Revision: https://reviews.llvm.org/D101003
-
Aart Bik authored
(1) migrates the encoding from TensorDialect into the new SparseTensorDialect (2) replaces dictionary-based storage and builders with struct-like data Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D101669
-
Alex Lorenz authored
when passing -platform_version to the linker The use of a valid SDK version is preferred over an empty SDK version (0.0.0) as the system's runtime might expect the linked binary to contain a valid SDK version in order for the binary to work correctly rdar://66795188
-