- Jun 23, 2023
-
-
Valentin Clement authored
Lower multi-dimensional arrays reduction for add and mul operator. Depends on D153448 Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D153455
-
Valentin Clement authored
Lower 1d array reduction for add and mul operator. Multi-dimensional arrays and other operator will follow. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D153448
-
Manna, Soumi authored
This patch uses castAs instead of getAs which will assert if the type doesn't match and adds nullptr check if needed. Also this patch improves the codes and passes I.getData() instead of doing a lookup in dumpVarDefinitionName() since we're iterating over the same map in LocalVariableMap::dumpContex(). Reviewed By: aaron.ballman, aaronpuchert Differential Revision: https://reviews.llvm.org/D153033
-
Vitaly Buka authored
-
Manna, Soumi authored
This patch adds missing assignment operator to the class which has user-defined copy constructor. Reviewed By: tahonermann, aaronpuchert Differential Revision: https://reviews.llvm.org/D150931
-
Vitaly Buka authored
Sanitizers allocate shadow and memory as MAP_NORESERVE. User memory can stay this way and do not increase RSS as long as we don't store there. The shadow unpoisoning also can avoid RSS increase for zeroed pages. However as soon we poison the shadow, we need the page in RSS. To avoid unnececary RSS increase we should not poison memory just before unpoisoning them. Depends on D153497. Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D153500
-
Fangrui Song authored
When the MCAssembler is non-null and the MCAsmLayout is null, we can fold A-B when * A and B are in the same fragment, or * A's fragment suceeds B's fragment, and they are not separated by non-data fragments (D69411) This patch allows folding when A's fragment precedes B's fragment so that `9997b - . == 0` below can be evaluated as true: ``` nop .arch_extension sec 9997:nop // old behavior: error: expected absolute expression .if 9997b - . == 0 .endif ``` Add a case to llvm/test/MC/ARM/directive-if-subtraction.s. Note: for MCAsmStreamer, we cannot evaluate `.if . - 9997b == 0` at parse time due to MCAsmStreamer::getAssemblerPtr returning nullptr (D45164). Some Darwin tests check that this folding does not work. Add `.p2align 2` to block some label difference folding or adjust the tests. Reviewed By: nickdesaulniers Differential Revision: https://reviews.llvm.org/D153096
-
Florian Hahn authored
Test cases for #62565.
-
Joseph Huber authored
These headers are currently broken when included from the offloading languages like OpenMP, OpenCL, CUDA, and HIP. Turn this logic off so we can compile these languages when the GPU libc is installed. I am currently trying to remedy this and have made an RFC for it in libc, see https://discourse.llvm.org/t/rfc-implementing-gpu-headers-in-the-llvm-c-library/71523. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D153578
-
Fangrui Song authored
-
Manna, Soumi authored
Reviewed By: erichkeane, steakhal, tahonermann, shafik Differential Revision: https://reviews.llvm.org/D150744
-
Sam McCall authored
This appears to be just an accidental copy rather than move from a scratch variable. As well as doing redundant work, these copies introduce extra SAT variables which make debugging harder (each Enviroment has a unique FC token). Example flow condition before: ``` (B0:1 = V15) (B1:1 = V8) (B2:1 = V10) (B3:1 = (V4 & (!V7 => V6))) (V10 = (B3:1 & !V7)) (V12 = B1:1) (V13 = B2:1) (V15 = (V12 | V13)) (V3 = V2) (V4 = V3) (V8 = (B3:1 & !!V7)) B0:1 V2 ``` after: ``` (B0:1 = (V9 | V10)) (B1:1 = (B3:1 & !!V6)) (B2:1 = (B3:1 & !V6)) (B3:1 = (V3 & (!V6 => V5))) (V10 = B2:1) (V3 = V2) (V9 = B1:1) B0:1 V2 ``` (with labelling from D153488) There are also some more copies that can be avoided here (when multiple blocks without terminating statements are joined), but they're less trivial, so I'll put those in another patch. Differential Revision: https://reviews.llvm.org/D153491
-
Joe Nash authored
The _e64_dpp suffix can be added to an instruction to force the AsmParser to encode it as VOP3 with DPP if possible on GFX11+. This has been the behavior since GFX11 was introduced; this patch only updates the documentation. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153564
-
Vitaly Buka authored
ComplexDeinterleavingPass.cpp:1849:3: error: default label in switch which covers all enumeration values This reverts commit 116953b8.
-
Craig Topper authored
As the extension list continues to grow it probably makes sense to use a binary search rather than linear search. Sorting the strings will make this possible. This also avoids any question about where to add new strings in the tables. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D153170
-
Vitaly Buka authored
For the secondary allocation we don't need poison and fill memory if we skip quarantine, and we don't need to poison after quarantine. In both cases the secondary allocator will unmap memory and unpoison the shadow from get_allocator().Deallocate(). Depends on D153496. Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D153497
-
Matt Arsenault authored
The select-of-different-exp pattern appears in the device libraries. I haven't seen the select-of-values case.
-
Matt Arsenault authored
-
Paul Robinson authored
-
Florian Hahn authored
getExpr is missing a check to make sure the result is invertible. This can lead to incorrect results, so return nullptr in those cases like in other places in IVUsers. Fixes #62660. Reviewed By: qcolombet Differential Revision: https://reviews.llvm.org/D153202
-
Yann Girsberger authored
There is a gap between running opt -Oz and running opt -passes="OZ_PASSES" where OZ_PASSES is taken from running opt -Oz -print-pipeline-passes. One of the reasons causing this is that -Oz uses non-default setting for LoopRotate but LoopRotate does not expose its settings when printing the pipeline. This commit fixes this by exposing LoopRotates parameters. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D153437
-
Aiden Grossman authored
Revert "[llvm-exegesis] Introduce Subprocess Executor Mode" This reverts commit 5e9173c4. This reverts commit 4d618b52. Reverting the PID commit as it is currently breaking MinGW builds and the way I'm checking for the presence of pid_t needs to be fixed and I need to do some testing. The subprocess executor mode patch is a dependent patch so also needs to be reverted and also needs some work as it is currently failing tests where libpfm is installed and the kernel version is less than 5.6.
-
Kamlesh Kumar authored
Fix build failure on windows system with msvc toolchain Reviewed By: ellis Differential Revision: https://reviews.llvm.org/D153318
-
Kazuki Sakamoto authored
D152759 introduced the Android .zip so file support, but it only considered POSIX path. The code also runs on Windows, so the path could be Windows path. Support both patterns on Windows. Differential Revision: https://reviews.llvm.org/D153390
-
Vitaly Buka authored
-
Vitaly Buka authored
-
Vitaly Buka authored
Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D153496
-
Michael Maitland authored
Since the scheduling resources for reductions and ordered reductions now account for LMUL and SEW, we can modify the Latency and ResourceCycles for these resoruces. * Most reductions take a total of approx `vl*SEW/DLEN + 5*(4 + log2(DLEN/SEW))` cycles. * Ordered floating-point reductions take a total of approx `5*vl` cycles. This commit re-commits 208fc34c. It was failing because it used wrong version of SchedSEWSet. Differential Revision: https://reviews.llvm.org/D153474
-
Zahira Ammarguellat authored
Differential Revision: https://reviews.llvm.org/D146148
-
Michael Maitland authored
This reverts commit 208fc34c. Reverting because build failure.
-
Michael Maitland authored
Since the scheduling resources for reductions and ordered reductions now account for LMUL and SEW, we can modify the Latency and ResourceCycles for these resoruces. * Most reductions take a total of approx `vl*SEW/DLEN + 5*(4 + log2(DLEN/SEW))` cycles. * Ordered floating-point reductions take a total of approx `5*vl` cycles. Differential Revision: https://reviews.llvm.org/D153474
-
Michael Maitland authored
* Unit-stride loads and stores can operate at the full bandwidth of the memory pipe. The memory pipe is DLEN bits wide. * Strided loads and stores operate at one element per cycle and should be scheduled accordingly. * Indexed loads and stores operate at one element per cycle, and they stall the machine until all addresses have been generated, so they cannot be scheduled. * Unit stride seg2 load is number of DLEN parts * seg3-8 are one segment per cycle, unless the segment is larger than DLEN in which each segment takes multiple cycles. Differential Revision: https://reviews.llvm.org/D153475
-
Jon Chesterfield authored
Also moves the wait-until-inbox-changes test into a shared method. Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D153573
-
Vitaly Buka authored
Almost NFC, as blocks over max quarantine size will trigger immediate drain anyway. In followup patches we can optimize passthrough case. Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D153495
-
Fangrui Song authored
The `__DATA,xray_instr_map` section has label differences like `.quad Lxray_sled_0-Ltmp0` that is represented as a pair of UNSIGNED and SUBTRACTOR relocations. LLVM integrated assembler attempts to rewrite A-B into A-B'+offset where B' can be included in the symbol table. B' is called an atom and should be a non-temporary symbol in the same section. However, since `xray_instr_map` does not define a non-temporary symbol, the SUBTRACTOR relocation will have no associated symbol, and its `r_extern` value will be 0. Therefore, we will see linker errors like: error: SUBTRACTOR relocation must be extern at offset 0 of __DATA,xray_instr_map in a.o To fix this issue, we need to define a non-temporary symbol in the section. We can accomplish this by renaming `Lxray_sleds_start0` to `lxray_sleds_start0` ("L" to "l"). `lxray_sleds_start0` serves as the atom for this dead-strippable subsection. With the `S_ATTR_LIVE_SUPPORT` attribute, `ld -dead_strip` will retain subsections that reference live functions. Special thanks to Oleksii Lozovskyi for reporting the issue and providing initial analysis. Differential Revision: https://reviews.llvm.org/D153239 -
Sindhu Chittireddy authored
Replace getAs with castAs and add assert if needed. Differential revision: https://reviews.llvm.org/D153236
-
Igor Kirillov authored
Adds the capability to recognize SelectInst that appear in the IR. These instructions are generated during scalable vectorization for reduction and when the code contains conditions inside the loop body or when "-prefer-predicate-over-epilogue=predicate-dont-vectorize" is set. Differential Revision: https://reviews.llvm.org/D152558
-
Jun Zhang authored
Signed-off-by:
Jun Zhang <jun@junz.org> Differential Revision: https://reviews.llvm.org/D153572
-
Jon Chesterfield authored
This makes the interface less error prone. The acquire was previously forgotten. Release is currently missing if recv() is the last operation made before close. Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D153571
-