- Jun 19, 2023
-
-
Yevgeny Rouban authored
Relax condition on runtime trip count unrolling loops with 1 non-latch exit that leads to a deop block. There are cases when the deopt blocks are common exits for different loops. LoopSimplify pass splits such edges to the common deopting blocks to make sure that all exit nodes of the loop only have predecessors that are inside of the loop (See simplifyOneLoop()). This breaks the current condition for unrolling. This patch allows the split transitive blocks that still lead to the deopting blocks. Differential Revision: https://reviews.llvm.org/D152639
-
Chuanqi Xu authored
Close https://github.com/llvm/llvm-project/issues/61940. The root cause is that clang will generate vtable as strong symbol now even if the corresponding class is defined in other module units. After I check the wording in Itanium ABI, I find this is not inconsistent. Itanium ABI 5.2.3 (https://itanium-cxx-abi.github.io/cxx-abi/abi.html#vague-vtable) says: > The virtual table for a class is emitted in the same object containing > the definition of its key function, i.e. the first non-pure virtual > function that is not inline at the point of class definition. So the current behavior is incorrect. This patch tries to address this. Also I think we need to do a similar change for MSVC ABI. But I don't find the formal wording. So I don't address this in this patch. Reviewed By: rjmccall, iains, dblaikie Differential Revision: https://reviews.llvm.org/D150023
-
Fangrui Song authored
Add the `S_ATTR_LIVE_SUPPORT` attribute to the sections so that `ld -dead_strip` will retain subsections that reference live functions, once we we add linker private "l" symbols as atoms.
-
Jianjian GUAN authored
Since we use match shl (v, splat 1) to vadd, we could also expand to widening add. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D153112
-
Fangrui Song authored
Fixes: c26c5e47 (essentially a no-op) The newly created MCDataFragment should inherit Atom (see MCMachOStreamer::finishImpl). To the best of my knowledge, this change cannot be tested at present, but this is important to ensure MCExpr.cpp:AttemptToFoldSymbolOffsetDifference gives the same result in case we evaluate the expression again with a MCAsmLayout. In the following case, ``` .section __DATA,xray_instr_map lxray_sleds_start1: .space 16 Lxray_sleds_end1: .section __DATA,xray_fn_idx .quad (Lxray_sleds_end1-lxray_sleds_start1)>>4 // can be folded without a MCAsmLayout ``` When we have a MCAsmLayout, without this change, evaluating (Lxray_sleds_end1-lxray_sleds_start1)>>4 again will fail due to `FA->getAtom() == nullptr && FB.getAtom() != nullptr` in MachObjectWriter::isSymbolRefDifferenceFullyResolvedImpl, called by AttemptToFoldSymbolOffsetDifference.
-
Fangrui Song authored
When the MCAssembler is non-null and the MCAsmLayout is null, we can fold A-B in these additional cases: * when A is a pending label (will be reassigned to a real fragment in flushPendingLabels()) * A and B are separated by a MCFillFragment with a constant size
-
Fangrui Song authored
If FA == FB, we can use SA.getOffset() - SB.getOffset() even if FA is not a MCDataFragment, as the only case this can be problematic (different offsets for a variable-size fragment) is invalid/unreachable. If FA != FB, the `if (FI->getKind() != MCFragment::FT_Data)` check below can bail out correctly. This change will help Mach-O fold more expressions. For ELF this is NFC, unless evaluateFixup has a bug that would evaluate an expression differently.
-
Fangrui Song authored
The newly created MCDataFragment should inherit Atom (see MCMachOStreamer::finishImpl). I cannot think of a case to test the behavior, but this is one step towards folding the Mach-O label difference below and making Mach-O more similar to ELF. ``` .section __DATA,xray_instr_map lxray_sleds_start1: .space 16 Lxray_sleds_end1: .section __DATA,xray_fn_idx .quad (Lxray_sleds_end1-lxray_sleds_start1)>>4 // error: expected relocatable expression // Mach-O ```
-
Alfred Persson Forsberg authored
Differential Revision: https://reviews.llvm.org/D153231
-
Fangrui Song authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Kazu Hirata authored
-
Hristo Hristov authored
[libc++][spaceship][NFC] P1612R2: Mark remove `operator!=` from "Ranges Library" items as "Complete" Marked already implemented parts of P1612R2 as "Complete": - `ranges::iota_view::iterator` https://reviews.llvm.org/D110774 - `iota_view::sentinel` https://reviews.llvm.org/D107396 - `filter_view::iterator` https://reviews.llvm.org/D109086 - `filter_view::sentinel` https://reviews.llvm.org/D109086 - `ranges::transform_view::iterator` https://reviews.llvm.org/D110774 - `transform_view::sentinel` https://reviews.llvm.org/D103056 - `take_view::sentinel` https://reviews.llvm.org/D123600 - `join_view::iterator` https://reviews.llvm.org/D107671 - `join_view::sentinel ` https://reviews.llvm.org/D107671 - `split_view::outer_iterator` https://reviews.llvm.org/D142063 - `split_view::inner_iterator` https://reviews.llvm.org/D142063 Note these operators were added and removed in C++20. Reviewed By: Mordante, #libc Differential Revision: https://reviews.llvm.org/D152721
-
Pranav Kant authored
-
Uday Bondhugula authored
Provide the bare pointer memref lowering option on gpu-to-nvvm pass. This is needed whenever we lower memrefs on the host function side and the kernel calls on the host-side (gpu-to-llvm) with the bare ptr convention. The GPU module side of the lowering should also "align" and use the bare pointer convention. Reviewed By: krzysz00 Differential Revision: https://reviews.llvm.org/D152480
-
Matt Arsenault authored
Missed these in 43fd46fd
-
Serge Pavlov authored
-
Simon Pilgrim authored
This function was lifted from fast-isel, and still referred to the Instruction::SRem/URrem opcodes, instead of the G_SREM/G_UREM opcodes. But it turns out these aren't necessary at all as only the G_SREM/G_UREM codepaths will use the AH register for DivRemResultReg anyhow.
-
Simon Pilgrim authored
-
- Jun 18, 2023
-
-
Serge Pavlov authored
A new builtin function __builtin_isfpclass is added. It is called as: __builtin_isfpclass(<floating point value>, <test>) and returns an integer value, which is non-zero if the floating point argument falls into one of the classes specified by the second argument, and zero otherwise. The set of classes is an integer value, where each value class is represented by a bit. There are ten data classes, as defined by the IEEE-754 standard, they are represented by bits: 0x0001 (__FPCLASS_SNAN) - Signaling NaN 0x0002 (__FPCLASS_QNAN) - Quiet NaN 0x0004 (__FPCLASS_NEGINF) - Negative infinity 0x0008 (__FPCLASS_NEGNORMAL) - Negative normal 0x0010 (__FPCLASS_NEGSUBNORMAL) - Negative subnormal 0x0020 (__FPCLASS_NEGZERO) - Negative zero 0x0040 (__FPCLASS_POSZERO) - Positive zero 0x0080 (__FPCLASS_POSSUBNORMAL) - Positive subnormal 0x0100 (__FPCLASS_POSNORMAL) - Positive normal 0x0200 (__FPCLASS_POSINF) - Positive infinity They have corresponding builtin macros to facilitate using the builtin function: if (__builtin_isfpclass(x, __FPCLASS_NEGZERO | __FPCLASS_POSZERO) { // x is any zero. } The data class encoding is identical to that used in llvm.is.fpclass function. Differential Revision: https://reviews.llvm.org/D152351 -
Yingwei Zheng authored
Fixes issue https://github.com/llvm/llvm-project/issues/63365 Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D153194
-
luxufan authored
-
Ivan Butygin authored
Add pass to uplift from arith mulf + addf ops to math.fma if fastmath flags allow it. Differential Revision: https://reviews.llvm.org/D152633
-
Simon Pilgrim authored
-
Simon Pilgrim authored
-
Paul Walker authored
Consider: add(pg, a, mul_u(pg, b, c)) Although the multiply's inactive lanes are undefined, they don't contribute to the final result. The overall result of the inactive lanes come from "a" and thus the above is another form of mla rather than mla_u.
-
Paul Walker authored
-
Paul Walker authored
-
AMS21 authored
We now display a simple note if the reason is that the used class does not support move semantics. This fixes llvm#62550 Reviewed By: PiotrZSL Differential Revision: https://reviews.llvm.org/D153220
-
AMS21 authored
For a declaration the `FunctionDecl` begin location does not include the template parameter lists, but for some reason if you have a separate definitions to the declaration the begin location does include them. With this patch we now correctly handle that case. This fixes llvm#62746 Reviewed By: PiotrZSL Differential Revision: https://reviews.llvm.org/D153218
-
LLVM GN Syncbot authored
-
AMS21 authored
As discussed in the https://reviews.llvm.org/D148697 review. Reviewed By: PiotrZSL Differential Revision: https://reviews.llvm.org/D153198
-
NAKAMURA Takumi authored
-
Youngsuk Kim authored
* Add `Address::withElementType()` as a replacement for `CGBuilderTy::CreateElementBitCast`. * Partial progress towards replacing `CreateElementBitCast`, as it no longer does what its name suggests. Either replace its uses with `Address::withElementType()`, or remove them if no longer needed. * Remove unused parameter 'Name' of `CreateElementBitCast` Reviewed By: barannikov88, nikic Differential Revision: https://reviews.llvm.org/D153196
-
-
-
Krzysztof Parzyszek authored
-
LLVM GN Syncbot authored
-
Nico Weber authored
-