- Apr 11, 2022
-
-
LLVM GN Syncbot authored
-
Guoxiong Li authored
Fixes https://llvm.org/PR54521 Differential Revision: https://reviews.llvm.org/D123452
-
Momchil Velikov authored
This pass inserts the necessary CFI instructions to compensate for the inconsistency of the call-frame information caused by linear (non-CGA aware) nature of the unwind tables. Unlike the `CFIInstrInserer` pass, this one almost always emits only `.cfi_remember_state`/`.cfi_restore_state`, which results in smaller unwind tables and also transparently handles custom unwind info extensions like CFA offset adjustement and save locations of SVE registers. This pass takes advantage of the constraints taht LLVM imposes on the placement of save/restore points (cf. `ShrinkWrap.cpp`): * there is a single basic block, containing the function prologue * possibly multiple epilogue blocks, where each epilogue block is complete and self-contained, i.e. CSR restore instructions (and the corresponding CFI instructions are not split across two or more blocks. * prologue and epilogue blocks are outside of any loops Thus, during execution, at the beginning and at the end of each basic block the function can be in one of two states: - "has a call frame", if the function has executed the prologue, or has not executed any epilogue - "does not have a call frame", if the function has not executed the prologue, or has executed an epilogue These properties can be computed for each basic block by a single RPO traversal. From the point of view of the unwind tables, the "has/does not have call frame" state at beginning of each block is determined by the state at the end of the previous block, in layout order. Where these states differ, we insert compensating CFI instructions, which come in two flavours: - CFI instructions, which reset the unwind table state to the initial one. This is done by a target specific hook and is expected to be trivial to implement, for example it could be: ``` .cfi_def_cfa <sp>, 0 .cfi_same_value <rN> .cfi_same_value <rN-1> ... ``` where `<rN>` are the callee-saved registers. - CFI instructions, which reset the unwind table state to the one created by the function prologue. These are the sequence: ``` .cfi_restore_state .cfi_remember_state ``` In this case we also insert a `.cfi_remember_state` after the last CFI instruction in the function prologue. Reviewed By: MaskRay, danielkiss, chill Differential Revision: https://reviews.llvm.org/D114545 -
Christian Sigg authored
Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D123490
-
Marius Brehler authored
Replaces `!emitc.opaque` types used to express pointers with `!emitc.ptr` types.
-
Sanjay Patel authored
fshl (or X, Y), X, C ==/!= 0 --> or (shl Y, C), X ==/!= 0 fshl X, (or X, Y), C ==/!= 0 --> or (srl Y, BW-C), X ==/!= 0 This is similar to an existing setcc-of-rotate fold, but the matching requires more checks for the more general funnel op: https://alive2.llvm.org/ce/z/Ab2jDd We are effectively decomposing the funnel shift into logical shifts, reassociating, and removing a shift. This should get us the final improvements for x86-64 that were originally shown in D111530 ( https://github.com/llvm/llvm-project/issues/49541 ); x86-32 still shows some SHLD/SHRD, so the pattern is not matching there yet. Differential Revision: https://reviews.llvm.org/D122919
-
Florian Hahn authored
-
Tim Northover authored
It was on ToT when I pushed and committed unintentionally.
-
Tim Northover authored
-
Tim Northover authored
Asynchronous exception support for the prologue means that there can be multiple .cfi_def_cfa_offset instructions in a single function, which tripped up an assertion in the compact unwind generator. In reality the compact unwind format is far too restrictive to represent asynchronous frames so if we ever wanted that on Darwin we'd fall back to DWARF (possibly keeping compact unwind around for synchronous users). So the compact format should continue to represent the synchronous situation, and the assertion can be removed.
-
Tim Northover authored
arm64_32 guarantees the high 32 bits of pointer parameters are passed as 0, and this is modelled in the IR by inserting an AssertZExt after the CopyFromReg. The function deciding whether registers that need to be preserved actually are wasn't expecting this so it banned perfectly legitimate tail calls.
-
Nikita Popov authored
Does not run on x86, so I missed this before. The test currently has typed pointer check lines.
-
Simon Pilgrim authored
As discussed on Issue #37628, we can flip a min/max node if we're subtracting from the sum of the node's operands Alive2: https://alive2.llvm.org/ce/z/W_KXfy Differential Revision: https://reviews.llvm.org/D123399
-
Iain Sandoe authored
This adds in testcases reflecting the remaining example in section 10.2 of the C++20 standard. Differential Revision: https://reviews.llvm.org/D122124
-
gysit authored
Rewrite tensor::ExtractSliceOp(vector::TransferWriteOp) to vector::TransferWriteOp(tensor::ExtractSliceOp) if the full slice is overwritten and inserted into another tensor. After this rewrite, the operations bufferize in-place since all of them work on the same %iter_arg slice. For example: ```mlir %0 = vector.transfer_write %vec, %init_tensor[%c0, %c0] : vector<8x16xf32>, tensor<8x16xf32> %1 = tensor.extract_slice %0[0, 0] [%sz0, %sz1] [1, 1] : tensor<8x16xf32> to tensor<?x?xf32> %r = tensor.insert_slice %1 into %iter_arg[%iv0, %iv1] [%sz0, %sz1] [1, 1] : tensor<?x?xf32> into tensor<27x37xf32> ``` folds to ```mlir %0 = tensor.extract_slice %iter_arg[%iv0, %iv1] [%sz0, %sz1] [1, 1] : tensor<27x37xf32> to tensor<?x?xf32> %1 = vector.transfer_write %vec, %0[%c0, %c0] : vector<8x16xf32>, tensor<?x?xf32> %r = tensor.insert_slice %1 into %iter_arg[%iv0, %iv1] [%sz0, %sz1] [1, 1] : tensor<?x?xf32> into tensor<27x37xf32> Reviewed By: nicolasvasilache, hanchung Differential Revision: https://reviews.llvm.org/D123190 -
Sven van Haastregt authored
Align guards of these builtins with opencl-c.h.
-
Simon Pilgrim authored
znver1/2 models were incorrectly modelling these as single uop instructions, instead of the microcoded nightmares they really are. Now matches AMD SoG, Agner and instlatx64 numbers. Fixes #54811
-
Nikita Popov authored
We need to make sure that the stored type matches the return type.
-
gysit authored
Clarify the in_bounds attribute is specified for the vector dimensions. Reviewed By: nicolasvasilache Differential Revision: https://reviews.llvm.org/D123188
-
Jean Perier authored
-
Haojian Wu authored
There is a TemplateName::getTemplateDecl which does the same work.
-
Mats Petersson authored
Most Fortran compilers appear to return the process time for calls to CPU_TIME, where the flang implementation prior to this change was returning the time used by the current thread. This would cause incorrect time being reported when for example OpenMP is used to share work across multiple CPUs. This patch changes the order so the selection of "what time to return" so that if there is a process time to report, that is the reported value, and only if that is not available, the thread time is considerd instead. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D123416
-
Simon Pilgrim authored
Revert rG88ff6f70 "[X86] Extend vselect(cond, pshufb(x), pshufb(y)) -> or(pshufb(x), pshufb(y)) to include inner or(pshufb(x), pshufb(y)) chains" Reverting while I investigate reports of internal test regressions/failures
-
Nikita Popov authored
All users of NewPM=false for the (legacy) ThinLTOCodeGenerator have been removed, so we can remove this functionality entirely.
-
Kiran Chandramohan authored
Privatisation creates local copies of variables in the OpenMP region. Two functions `createHostAssociateVarClone` and `copyHostAssociateVar` are added to create a clone of the variable for basic privatisation and to copy the contents for first-privatisation. Note: Tests for more data-types will be added when the fir.do_loop is upstreamed. This is part of the upstreaming effort from the fir-dev branch in [1]. [1] https://github.com/flang-compiler/f18-llvm-project Reviewed By: peixin, NimishMishra Differential Revision: https://reviews.llvm.org/D122595 Co-authored-by:
Jean Perier <jperier@nvidia.com> Co-authored-by:
Eric Schweitz <eschweitz@nvidia.com> Co-authored-by:
Peter Klausler <pklausler@nvidia.com> Co-authored-by:
Valentin Clement <clementval@gmail.com> Co-authored-by:
Sourabh Singh Tomar <SourabhSingh.Tomar@amd.com> Co-authored-by:
Nimish Mishra <neelam.nimish@gmail.com> Co-authored-by:
Peixin-Qiao <qiaopeixin@huawei.com>
-
Nikita Popov authored
Enable opaque pointers by default in clang, which can be disabled either via cc1 option -no-opaque-pointers or cmake flag -DCLANG_ENABLE_OPAQUE_POINTERS=OFF. See https://llvm.org/docs/OpaquePointers.html for context. Differential Revision: https://reviews.llvm.org/D123300
-
Nikita Popov authored
-
Iain Sandoe authored
This addresses a post commit review comment by removing an unused and empty 'else' (replaced with a comment).
-
Simon Pilgrim authored
This doesn't hit if we use pshufb intrinsics directly due to a change in lowering order
-
Nikita Popov authored
This removes support for the legacy pass manager in llvm-lto and llvm-lto2. In this case I've dropped the use-new-pm option entirely, as I don't think this is considered part of the public interface. This also makes -debug-pass-manager work with llvm-lto, because that was needed to migrate some tests to NewPM. Differential Revision: https://reviews.llvm.org/D123376
-
jeanPerier authored
Handle dynamic optional argument in GET_COMMAND_ARGUMENT and GET_ENVIRONMENT_VARIABLE (previously compiled but caused segfaults). The previous code handled static presence/absence aspects, but not when an absent dummy optional was passed to one of the optional intrinsic arguments. Simplify the runtime call lowering to simply lower the runtime call without dealing with optionality there. This keeps the optional handling logic in IntrinsicCall.cpp. Note that the new code will generate some extra "if (not null addr )/then/else" when the actual arguments are always there at runtime. That makes the implementation a lot simpler/safer, and I think it is OK for now (I do not expect these runtime function to be called in hot loop nests). Differential Revision: https://reviews.llvm.org/D123388
-
Jean Perier authored
Add a check that CheckUnitNumberInRangeImpl is not needlessly instantiated. Differential Revision: https://reviews.llvm.org/D123285
-
Alexander Shaposhnikov authored
-
Alexander Shaposhnikov authored
This diff splits fuse-literals feature and enables fuse-adrp-add by default, in particular, it adjusts instruction scheduling to place ADRP+ADD pairs together. This also enables the linker to apply the relaxations described in https://github.com/ARM-software/abi-aa/commit/d2ca58c54b8e955cfef25c71822f837ae0439d73. Differential revision: https://reviews.llvm.org/D120104 Test plan: make check-all
-
Patryk Wychowaniec authored
This commit contains a refactoring that merges AVRRelaxMemOperations into AVRExpandPseudoInsts, so that we have a single place in code that expands the STDWPtrQRr opcode. Seizing the day, I've also fixed a couple of potential bugs with our previous implementation (e.g. when the destination register was killed, the previous implementation would try to .addDef() that killed register, crashing LLVM in the process - that's fixed now, as proved by the test). Reviewed By: benshi001 Differential Revision: https://reviews.llvm.org/D122533
-
LiaoChunyu authored
Scalable vectors llvm.experimental.stepvector intrinsic will crash due to an invalid cost when run the code through the loopunroll. Reviewed By: kito-cheng Differential Revision: https://reviews.llvm.org/D122782
-
Yaxun (Sam) Liu authored
kernels in anonymous name space needs to have unique name to avoid duplicate symbols. Fixes: https://github.com/llvm/llvm-project/issues/54560 Reviewed by: Artem Belevich Differential Revision: https://reviews.llvm.org/D123353
-
Nico Weber authored
No behavior change.
-
Sheng authored
`cast` will assert instead of returning null pointer.
-
Florian Hahn authored
Add a simpler test for D114487/D108699.
-