- Jun 17, 2022
-
-
Alexey Bataev authored
If the root scalar is mapped to to the smallest bit width, the vector is truncated and the types between original buildvector and extracted value mismatched. For extract, we emit sext/zext instructions, for shuffles we can reuse oringal vector instead of the truncated one. Differential Revision: https://reviews.llvm.org/D127974
-
Mark de Wever authored
-
Philip Reames authored
Apparently the parser/verifier is more lax than it should be. The typo'd names should have been rejected.
-
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D127955
-
Jay Foad authored
This includes: - New llvm.amdgcn.image.msaa.load.* intrinsics - NSA changes, because MIMG-NSA is now limited to 3 dwords - Split CD forms of IMAGE_SAMPLE instructions out into separate test files since they are no longer supported in GFX11 Differential Revision: https://reviews.llvm.org/D127837
-
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D127671
-
David Green authored
-
Arthur Eubanks authored
Reviewed By: #opaque-pointers, rnk, nikic Differential Revision: https://reviews.llvm.org/D126309
-
Peter Klausler authored
There's a few (3) cases where Fortran allows two distinct symbols to have the same name in the same scope. Module file output copes with only two of them. The third involves a separate module procedure that isn't separate: both the procedure and its declared interface appear in the same (sub)module. Fix to ensure that the interface is included in the module file output, so that the module file reader doesn't suffer a bogus error about a "separate module procedure without an interface". Differential Revision: https://reviews.llvm.org/D127784
-
Craig Topper authored
This removes one of the uses of ForceTailUndisturbed.
-
Peter Klausler authored
Previous one was returning a bogus error status about a bad WAIT statement ID number. Differential Revision: https://reviews.llvm.org/D127979
-
Michael Jones authored
Previously, any line buffered write of size 0 would cause an error. The variable used to track the index of the last newline started at the size of the write - 1, which underflowed. Now it's handled properly, and a test has been added to prevent regressions. Reviewed By: sivachandra, lntue Differential Revision: https://reviews.llvm.org/D127914
-
Michael Jones authored
The hex converter handles the %x and %X conversions. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D126082
-
Thomas Raoux authored
contraction op can have mixed type, add support for this case to the pattern lowering contraction op to outerproduct. Differential Revision: https://reviews.llvm.org/D127926
-
Arthur Eubanks authored
Reviewed By: pcc Differential Revision: https://reviews.llvm.org/D127876
-
Thomas Raoux authored
Support the case where convolution does float extension of the inputs. Differential Revision: https://reviews.llvm.org/D127925
-
Philip Reames authored
-
Alex Brachet authored
-
Mircea Trofin authored
This requires DominatorTree be updated, which we do in the ml inliner case, but not in the default case, and the cost of doing so is noticeable to compile time for the latter[1]. So the patch only affects the ML inliner. [1] https://llvm-compile-time-tracker.com/compare.php?from=9fc0aa45e3312944431ba7e1ca0cec99c613992b&to=7af461b1ce0d9138211ef5f883f35d5b9ddf47be&stat=wall-time Differential Revision: https://reviews.llvm.org/D127899
-
Craig Topper authored
We may eventually need tail undisturbed patterns, but we will need a policy operand on the ISD node to communicate it.
-
Adrian Tong authored
This helps handling a case where the BUILD_VECTOR has i16 element type and i32 constant operands t2: v8i16 = setcc t8, t17, setult:ch t3: v8i16 = BUILD_VECTOR Constant:i32<1>, ... t4: v8i16 = and t2, t3 t5: v8i16 = add t8, t4 This can be turned into t5: v8i16 = sub t8, t2, and allows us to remove t3 and t4 from the DAG. Differential Revision: https://reviews.llvm.org/D127354 -
Nikolas Klauser authored
This reverts commit 147f74b6.
-
- Jun 16, 2022
-
-
Philip Reames authored
This change just moves some code around, and extracts out a helper function expected to be useful when reusing the demanded field logic in the forward dataflow.
-
Ahsan Saghir authored
This patch fixes the load and store quadword instructions on PowerPC to use correct offset and base address. Reviewed By: #powerpc, nemanjai, lkail Differential Revision: https://reviews.llvm.org/D126807
-
Philip Reames authored
-
Alex Brachet authored
Fuchsia's dynamic linker does not and will never support IFUNC's. Differential revision: https://reviews.llvm.org/D127933
-
Craig Topper authored
If the merge operand isn't undef we need to be using tail undisturbed. Turns out all of our uses of riscv_slidedown_vl use undef so this doesn't affect any tests.
-
Corentin Jabot authored
LLVM had 2 methods to convert a number to an hexa string, this remove one of them. Differential Revision: https://reviews.llvm.org/D127958
-
Kevin P. Neal authored
-
Philip Reames authored
-
Philip Reames authored
The motivating case, and the only one actually enabled by this patch, is a load or store followed by another op with the same SEW/LMUL ratio. As an example, consider: define void @test1(ptr %in, ptr %out) { entry: %0 = load <8 x i16>, ptr %in, align 2 %1 = sext <8 x i16> %0 to <8 x i32> store <8 x i32> %1, ptr %out, align 4 ret void } Without this patch, we get: vsetivli zero, 8, e16, mf4, ta, mu vle16.v v8, (a0) vsetvli zero, zero, e32, mf2, ta, mu vsext.vf2 v9, v8 vse32.v v9, (a1) ret Whereas with the patch we get: vsetivli zero, 8, e32, mf2, ta, mu vle16.v v8, (a0) vsext.vf2 v9, v8 vse32.v v9, (a1) ret We have rewritten the first vsetvli and thus removed the second one. As is strongly hinted by the code structure and todos, I am planning on communing this with all (or most all?) of the cases from isCompatible used in the forward data flow. This will be done in a series of following changes - some NFC reworks, and some reviewed optimization extensions. Differential Revision: https://reviews.llvm.org/D127780 -
Lei Zhang authored
The previous approach does not work as the Adreno driver is clever at optimizing away the selection. So now check two inputs together. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D127930
-
David Sherwood authored
-
Samuel Eubanks authored
TurnSwitchRangeIntoICmp crashes when given a switch with a default destination of unreachable Addresses issue #53208 https://github.com/llvm/llvm-project/issues/53208 Differential revision: https://reviews.llvm.org/D127712
-
Florian Hahn authored
Instead of using the underlying instruction and VF to get the type, use the type of the incoming value. This removes an unnecessary dependence on the underlying instruction and enables using the recipe without an underlying instruction.
-
Alexey Bataev authored
Currently scatter vectorize nodes can be emitted only for GEPs with constant indices. But we can also emit such nodes for GEPs with the same ptr and non-constant vectorizable/gathered indices, if profitable. Patch adds support for such nodes and tries to improve handling of GEPs with non-const indeces for such nodes. Metric: SLP.NumVectorInstructions Program SLP.NumVectorInstructions results results0 diff test-suite :: External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s.test 5243.00 5240.00 -0.1% test-suite :: External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r.test 5243.00 5240.00 -0.1% test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 27550.00 27507.00 -0.2% test-suite :: External/SPEC/CFP2006/453.povray/453.povray.test 5395.00 5380.00 -0.3% test-suite :: External/SPEC/CFP2017rate/511.povray_r/511.povray_r.test 5389.00 5374.00 -0.3% test-suite :: External/SPEC/CINT2017rate/520.omnetpp_r/520.omnetpp_r.test 961.00 958.00 -0.3% test-suite :: External/SPEC/CINT2017speed/620.omnetpp_s/620.omnetpp_s.test 961.00 958.00 -0.3% test-suite :: External/SPEC/CFP2006/447.dealII/447.dealII.test 5664.00 5643.00 -0.4% test-suite :: External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test 13202.00 13127.00 -0.6% test-suite :: External/SPEC/CINT2006/445.gobmk/445.gobmk.test 212.00 207.00 -2.4% test-suite :: MultiSource/Benchmarks/7zip/7zip-benchmark.test 890.00 850.00 -4.5% test-suite :: External/SPEC/CINT2006/464.h264ref/464.h264ref.test 1695.00 1581.00 -6.7% test-suite :: MultiSource/Applications/JM/lencod/lencod.test 2338.00 2140.00 -8.5% test-suite :: SingleSource/UnitTests/matrix-types-spec.test 63.00 55.00 -12.7% test-suite :: SingleSource/Benchmarks/Adobe-C++/loop_unroll.test 468.00 356.00 -23.9% Geomean difference -0.3% All numbers show increased number of generated vector instructions. Diff: SingleSource/Benchmarks/Adobe-C++/loop_unroll - better without LTO, but need an extra analysis with LTO (with LTO compiler generates masked_gather, while before regular loads were emitted because of extra data, availbale at LTO time). SingleSource/UnitTests/matrix-types-spec - more vector code. MultiSource/Applications/JM/lencod/lencod - same. External/SPEC/CINT2006/464.h264ref/464.h264ref - same. MultiSource/Benchmarks/7zip/7zip-benchmark - same. External/SPEC/CINT2006/445.gobmk/445.gobmk - no changes. External/SPEC/CFP2017rate/510.parest_r/510.parest_r - more vector code. External/SPEC/CFP2006/447.dealII/447.dealII - same External/SPEC/CINT2017speed/620.omnetpp_s/620.omnetpp_s - same External/SPEC/CINT2017rate/520.omnetpp_r/520.omnetpp - same External/SPEC/CFP2017rate/511.povray_r/511.povray - same External/SPEC/CFP2006/453.povray/453.povray - same External/SPEC/CFP2017rate/526.blender_r/526.blender_r - same External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r - same External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s - same Differential Revision: https://reviews.llvm.org/D127219 -
Nikolas Klauser authored
Reviewed By: ldionne, #libc, EricWF Spies: EricWF, libcxx-commits, mgrang Differential Revision: https://reviews.llvm.org/D127669
-
Matheus Izvekov authored
Signed-off-by:
Matheus Izvekov <mizvekov@gmail.com> Reviewed By: dyung Differential Revision: https://reviews.llvm.org/D127943
-
Jay Foad authored
-
Dmitry Preobrazhensky authored
Differential Revision: https://reviews.llvm.org/D127847
-