- Aug 04, 2023
-
-
Zhongyunde authored
this PR tries to match the following pattern, seperate from D156881 ``` %vscale = call i64 @llvm.vscale.i64() %shift = shl nuw nsw i64 %vscale, 11 ``` Now, we only check the shl recursively when the OrZero is true. Reviewed By: goldstein.w.n Differential Revision: https://reviews.llvm.org/D157062 -
Alex Langford authored
Differential Revision: https://reviews.llvm.org/D156934
-
Alexey Bataev authored
Use O(nlogn) instead of O(N2) (N <= 32) sorting approach and do not try to revectorize all possible combinations of stores, if they definitely cannot be combined because of mem/data dependencies. Compile time (O3 + lto, skylake_avx512): External/SPEC/CINT2006/483.xalancbmk/483.xalancbmk.test 117.15 120.11 2.5% External/SPEC/CINT2017speed/623.xalancbmk_s/623.xalancbmk_s.test 203.67 207.42 1.8% External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 232.43 235.01 1.1% External/SPEC/CINT2017rate/523.xalancbmk_r/523.xalancbmk_r.test 205.49 207.25 0.9% External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test 310.46 306.23 -1.4% Link time (O3+lto, skylake_avx512): External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 1383.69 1475.94 6.7% Other changes are too small, cannot rely on them. size..text Program size..text results results0 diff test-suite :: SingleSource/Regression/C/Regression-C-sumarray.test 392.00 1439.00 267.1% test-suite :: MultiSource/Applications/JM/ldecod/ldecod.test 394258.00 394818.00 0.1% test-suite :: MultiSource/Applications/JM/lencod/lencod.test 846355.00 847075.00 0.1% test-suite :: External/SPEC/CINT2006/464.h264ref/464.h264ref.test 782816.00 783360.00 0.1% test-suite :: External/SPEC/CFP2017rate/508.namd_r/508.namd_r.test 779667.00 779923.00 0.0% test-suite :: MultiSource/Benchmarks/mafft/pairlocalalign.test 224398.00 224446.00 0.0% test-suite :: MultiSource/Applications/oggenc/oggenc.test 185019.00 185035.00 0.0% test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 12487610.00 12488010.00 0.0% test-suite :: MultiSource/Benchmarks/7zip/7zip-benchmark.test 1051772.00 1051804.00 0.0% test-suite :: MultiSource/Applications/SPASS/SPASS.test 529586.00 529602.00 0.0% test-suite :: External/SPEC/CINT2006/400.perlbench/400.perlbench.test 1084684.00 1084716.00 0.0% test-suite :: MultiSource/Benchmarks/tramp3d-v4/tramp3d-v4.test 1014245.00 1014261.00 0.0% test-suite :: MultiSource/Benchmarks/MallocBench/espresso/espresso.test 223494.00 223478.00 -0.0% test-suite :: External/SPEC/CINT2017speed/625.x264_s/625.x264_s.test 660843.00 660795.00 -0.0% test-suite :: External/SPEC/CINT2017rate/525.x264_r/525.x264_r.test 660843.00 660795.00 -0.0% test-suite :: MultiSource/Applications/ClamAV/clamscan.test 568824.00 568760.00 -0.0% espresso - 2 more stores vectorized x264 - small number of changes in 3-4 functions, generated a bit more vector stores (2 4x zeroinitializer stores + some other small variations). clamscan - emitted 32xi8 store instead of several scalar stores + several 4x-8x stores. Differential Revision: https://reviews.llvm.org/D155246 -
Alex Voicu authored
`__dynamic_cast` relies on `type_info`, which its signature assumed to be in the generic / default address space. This patch corrects the oversight (we know that `type_info` resides in the GlobalVar address space) and adds an associated test. Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D155870
-
Chia-hung Duan authored
`atomic_compare_exchange` was using `_strong` and `memory_order_acquire` by default. This is not aligned with general use, for example, in C++, the default is `memory_order_seq_cst`. To reduce the ambiguity, make the version and ordering explicitly. Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D156952
-
Jan Svoboda authored
This will make it possible to accept the spelling as `StringLiteral` in D157029 and avoid some unnecessary allocations in a later patch. Reviewed By: benlangmuir Differential Revision: https://reviews.llvm.org/D157035
-
Zhen Wang authored
-
Bjorn Pettersson authored
Differential Revision: https://reviews.llvm.org/D157016
-
Bjorn Pettersson authored
Differential Revision: https://reviews.llvm.org/D156911
-
Florian Hahn authored
The last dependency of code defined in LoopVectorize.cpp has been removed a while ago. Move VPTransformState::get() to VPlan.cpp where other members are also defined.
-
Andrzej Warzynski authored
Differential Revision: https://reviews.llvm.org/D156823
-
Andrew Lenharth authored
OpInterface inheritance will duplicate base interfaces, causing compilation failure. Unique the set of base interfaces. Reviewed By: rriddle, jdd Differential Revision: https://reviews.llvm.org/D156964
-
Tom Yang authored
Summary: In cases where the PC has no function name, lldb-vscode crashes. `lldb::SBFrame::GetDisplayFunctionName()` returns a `nullptr`, and when we attempt to construct an `std::string`, it raises an exception. Test plan: This can be observed with creating a test file (credit to @clayborg for the example): ``` int main() { typedef void (*FooCallback)(); FooCallback foo_callback = (FooCallback)0; foo_callback(); // Crash at zero! return 0; } ``` and attempting to debug the file through VSCode. I add a test case in D156732 which should cover this. Differential Revision: https://reviews.llvm.org/D156970 -
Emilio Cota authored
-
Krishna-13-cyber authored
Differential revision: https://reviews.llvm.org/D156877
-
Kiran Chandramohan authored
The copyprivate clause is not yet implemented. Provide a TODO error message when this clause is seen. Reviewed By: NimishMishra Differential Revision: https://reviews.llvm.org/D155596
-
Craig Topper authored
Tweak the immediate on two vror.vi test cases to use a uimm6 immediate that would have failed before D156974 when we were looking for a simm6 immediate.
-
Volodymyr Sapsai authored
-
Jan Svoboda authored
The `Twine::str()` function currently always allocates heap memory via `std::string`. However, some instances of `Twine` don't need an intermediate buffer at all, and the rest can attempt to print into a stack buffer first. This is intentionally not making use of `Twine::isSingleStringLiteral()` from D157010 to skip saving the string in the bump-pointer allocator, since the `StringSaver` documentation suggests that MUST happen for every given string. Reviewed By: benlangmuir Differential Revision: https://reviews.llvm.org/D157015
-
Rainer Orth authored
Since GCC 11, the bundled Solaris/SPARC GCC uses the `sparcv8plus` subdirectory for 32-bit objects, just like upstream GCC. Before that, it used `32` instead from a local patch. Since `clang` doesn't know about that `sparcv8plus` subdirectory, it wouldn't properly use GCC 11+ installations. The new `solaris-sparc-gcc-search.test` testcase wasn't run initially (like the existing `crash-report-null.test`) because the `.test` suffix wasn't handled. Tested on `sparcv9-sun-solaris2.11`, `amd64-pc-solaris2.11`, and `x86_64-pc-linux-gnu`. Differential Revision: https://reviews.llvm.org/D157013
-
Jan Svoboda authored
The `const char *` storage backing StringLiteral has static lifetime. Making `Twine` aware of that allows us to avoid allocating heap memory in some contexts (e.g. avoid passing it to `StringSaver::save()` in a follow-up Clang patch). Reviewed By: benlangmuir Differential Revision: https://reviews.llvm.org/D157010
-
Alexander Yermolovich authored
Fixed a bug where when Skelton CU had DW_AT_ranges, it the output CU DW_AT_ranges offset was relative, and not absolute. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D156958
-
Alexander Yermolovich authored
Now that we have new DWARF Rewriter we can remove DW_AT_low_pc when converting DW_AT_low_pc/DW_AT_high_pc to DW_AT_ranges. Which closer follows DWARF spec. Leaving CU DW_AT_low_pc in place. Reading the spec I think it's needed. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D156957
-
Valentin Clement authored
Lower link clause with data entry operation. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D156913
-
Nikolas Klauser authored
This is to avoid clang-tidy complaining all over the tests that the naming is wrong.
-
Nikolas Klauser authored
-
Jay Foad authored
Also rename the flag from supportsBackwardScavenger to eliminateFrameIndicesBackwards to reflect what it actually does. X86 is the only target still using forwards frame index elimination. This will not block removing support for forwards register scavenging, because X86 does not use the register scavenger. Differential Revision: https://reviews.llvm.org/D156983
-
- Aug 03, 2023
-
-
Benjamin Maxwell authored
This regressed in D154458 due to the added tracking of used variable names that now also has to be cleared alongside the counter. Reviewed By: rafaelubalmw, c-rhodes, awarzynski Differential Revision: https://reviews.llvm.org/D156547
-
Nikolas Klauser authored
Reviewed By: #libc, Mordante, ldionne Spies: Mordante, libcxx-commits Differential Revision: https://reviews.llvm.org/D155382
-
Joseph Huber authored
Summary: Older gcc can't figure out the copy elision and needs an explicit move.
-
Ingo Müller authored
This patch adds a mix-in class for the only transform op of the tensor dialect that can benefit from one: the MakeLoopIndependentOp. It adds an overload that makes providing the return type optional. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D156918
-
Podchishchaeva, Mariya authored
A bunch of classes in APValue free resources in the destructor but don't have user-written copy c'tor or assignment operator, so copying them using default ones can cause double free. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D156975
-
Podchishchaeva, Mariya authored
ToolInvocation frees resources in the destructor but doesn't have user-written copy c'tor or assignment operator, so copying it using default ones can cause double free. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D156896
-
Craig Topper authored
Part of this test file was stolen from D156895. We should merge them when committing. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D156926
-
Craig Topper authored
Don't access leaf 7 subleaf 1 unless subleaf 0 says it is supported via EAX. Intel documentation says invalid subleaves return 0. We had been relying on that behavior instead of checking the max sublef number. It appears that some Sandy Bridge CPUs return at least the subleaf 0 EDX value for subleaf 1. Best guess is that this is a bug in a microcode patch since all of the bits we're seeing set in EDX were introduced after Sandy Bridge was originally released. This is causing avxvnniint16 to be incorrectly enabled with -march=native on these CPUs. Reviewed By: pengfei, anna Differential Revision: https://reviews.llvm.org/D156963
-
David Spickett authored
This makes anlysing test failures much more easy. For SendJSON this is simple, just use llvm::format instead. For GetNextObject/ReadJSON it's a bit more tricky. * Print the "Content-Length:" line in ReadJSON, but not the json. * Back in GetNextObject, if the JSON doesn't parse, it'll be printed as a normal string along with an error message. * If we didn't error before, we have a JSON value so we pretty print it. * Finally, if it isn't an object we'll log an error for that, not including the JSON. Before: ``` <-- Content-Length: 81 {"command":"disconnect","request_seq":5,"seq":0,"success":true,"type":"response"} ``` After: ``` <-- Content-Length: 81 { "command": "disconnect", "request_seq": 5, "seq": 0, "success": true, "type": "response" } ``` There appear to be some responses that include strings that are themselves JSON, and this won't pretty print those but I think it's still worth doing. Reviewed By: wallace Differential Revision: https://reviews.llvm.org/D156979 -
David Spickett authored
Reviewed By: wallace Differential Revision: https://reviews.llvm.org/D156977
-
pvanhout authored
Previous heuristics had a big flaw: they only looked at single PHI at a time, and didn't take into account the whole "chain". The concept of "chain" is important because if we only break a chain partially, we risk forcing regalloc to reserve twice as many registers for that vector. We also risk adding a lot of copies that shouldn't be there and can inhibit backend optimizations. The solution I found is to consider the whole "PHI chain" when looking at PHI. That is, we recursively look at the PHI's incoming value & users for other PHIs, then make a decision about the chain as a whole. The currrent threshold requires that at least `ceil(chain size * (2/3))` PHIs have at least one interesting incoming value. In simple terms, two-thirds (rounded up) of the PHIs should be breakable. This seems to work well. A lower threshold such as 50% is too aggressive because chains can often have 7 or 9 PHIs, and breaking 3+ or 4+ PHIs in those case often causes performance issue. Fixes SWDEV-409648, SWDEV-398393, SWDEV-413487 Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D156414
-
Joseph Huber authored
This feature was supposed to allow you to trace execution inside of Libomptarget. However, this never really worked properly. The printing was always reoganized, only worked for single threads, and pretty much only told you a handful of things about a runtime library that's an implementation detail to all users. Despite this, it contributed about 40% of the total filesize of the deviceRTL. This patch simply removes this functionalit which I think was past due. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D157001
-