- Jun 24, 2023
-
-
Matt Arsenault authored
-
Zahira Ammarguellat authored
a '#pragma clang fp eval_method, it can lead to ABI breakage. See https://godbolt.org/z/56zG4Wo91 This patch prevents this. Differential Revision: https://reviews.llvm.org/D153590
-
LLVM GN Syncbot authored
-
Jonas Devlieghere authored
Addresses Jason's post-commit feedback in D153644.
-
Chia-hung Duan authored
Ensure the thread that refills freelist will get the Batch without contending the lock in SizeClassAllocator64. Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D152419
-
Chia-hung Duan authored
Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D152420
-
Joseph Huber authored
The patch in D152592 changed the logic for this. We could never check if we were on the GPU as this was before the variable was defined so I moved it later. Secondly, we cannot use the `LLVM_BINARY_DIR` here, and I do not know if that works in general. The problem is that it will isntall the headers under a normal path outside of the `LLVM_ENABLE_RUNTIMES` build. I don't know if that's correct for the other targets, but for the GPU I need to set it back to the CMAKE_BINARY_DIR so it works. Reviewed By: phosek Differential Revision: https://reviews.llvm.org/D153637
-
Paul Kirth authored
There seems to be a problem on arm buildbots. Reverting until I can investigate. https://lab.llvm.org/buildbot#builders/245/builds/10184 This reverts commit a67208e1 and dependent commit e54a3112.
-
Sami Tolvanen authored
With `-fsanitize=kcfi` (Kernel Control-Flow Integrity), Clang emits "kcfi" operand bundles to indirect call instructions. Similarly to the target-specific lowering added in D119296, implement KCFI operand bundle lowering for RISC-V. This patch disables the generic KCFI pass for RISC-V in Clang, and adds the KCFI machine function pass in `RISCVPassConfig::addPreSched` to emit target-specific `KCFI_CHECK` pseudo instructions before calls that have KCFI operand bundles. The machine function pass also bundles the instructions to ensure we emit the checks immediately before the calls, which is not possible with the generic pass. `KCFI_CHECK` instructions are lowered in `RISCVAsmPrinter` to a contiguous code sequence that traps if the expected hash in the operand bundle doesn't match the hash before the target function address. This patch emits an `ebreak` instruction for error handling to match the Linux kernel's `BUG()` implementation. Just like for X86, we also emit trap locations to a `.kcfi_traps` section to support error handling, as we cannot embed additional information to the trap instruction itself. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D148385
-
LLVM GN Syncbot authored
-
Benjamin Kramer authored
-
Jonas Devlieghere authored
When specifying the C-string format for dumping memory, we treat unprintable characters as signed. Whether a character is signed or not is implementation defined, but all printable characters are signed. Therefore it's fair to assume that unprintable characters are unsigned. Before this patch, "\xcf\xfa\xed\xfe\f" would be printed as "\xffffffcf\xfffffffa\xffffffed\xfffffffe\f". Now we correctly print the original string. rdar://111126134 Differential revision: https://reviews.llvm.org/D153644
-
Artem Belevich authored
This avoids unnecessary vector splitting that was needed for vectorized store instruction. Differential Revision: https://reviews.llvm.org/D152593
-
Paul Kirth authored
Fat LTO objects contain both LTO compatible IR, as well as generated object code. This allows users to defer the choice of whether to use LTO or not to link-time. This is a feature available in GCC for some time, and makes the existing -ffat-lto-objects flag functional in the same way as GCC's. Within LLVM, we add a new EmbedBitcodePass that serializes the module to the object file, and expose a new pass pipeline for compiling fat objects. The new pipeline initially clones the module and runs the selected (Thin)LTOPrelink pipeline, after which it will serialize the module into a `.llvm.lto` section of an ELF file. When compiling for (Thin)LTO, this normally the point at which the compiler would emit a object file containing the bitcode and metadata. After that point we compile the original module using the PerModuleDefaultPipeline used for non-LTO compilation. We generate standard object files at the end of this pipeline, which contain machine code and the new `.llvm.lto` section containing bitcode. Since the two pipelines operate on different copies of the module, we can be sure that the bitcode in the `.llvm.lto` section and object code in `.text` are congruent with the existing output produced by the default and LTO pipelines. Original RFC: https://discourse.llvm.org/t/rfc-ffat-lto-objects-support/63977 Reviewed By: tejohnson, MaskRay, nikic Differential Revision: https://reviews.llvm.org/D146776
-
Artem Belevich authored
Fixes https://github.com/llvm/llvm-project/issues/63436 Improves lowering of extending FP vector loads. We were previously splitting them unnecessarily. Differential Revision: https://reviews.llvm.org/D153477
-
Alex Langford authored
The use of ConstString in StructuredDataPlugin is unneccessary as fast comparisons are not neeeded for StructuredDataPlugins. Differential Revision: https://reviews.llvm.org/D153482
-
Valentin Clement authored
acc.reduction operation is used as data entry operation for the reduction operands. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D153367
-
Hongtao Yu authored
Llvm-profgen internally uses the llvm libraries and the MCDesc interface to do disassembling and symblization and it never checks against target-specific instruction operators. This makes it quite transparent to targets and a first attempt for an aarch64 binary just works. Therefore I'm removing the unnecessary triple check to unblock for new targets. Reviewed By: wenlei Differential Revision: https://reviews.llvm.org/D153449
-
Fangrui Song authored
GNU addr2line exits immediately if -e (default to a.out) specifies a file that cannot be open or a directory. llvm-addr2line used to wait for input on if the input file cannot be open and addresses are not specified in command line. Replace the D147652 checkFileExists with getOrCreateModuleInfo to avoid a separate `sys::fs::status` operation. Reviewed By: sepavloff Differential Revision: https://reviews.llvm.org/D153595
-
Wenlei He authored
This change allows sinking defs from loop preheader with PHI-use into loop body. Loop sink can now see through PHI-use and select incoming blocks of value being used as candidate sink destination. It makes loop sink more effective so more LICM can be undone if proven unprofitable with profile info. It addresses the motivating case in D87551, without resorting to profile guided LICM which breaks canonicalization. This is the 2nd attempt after D152772.
-
Valentin Clement authored
Add support for RecordType in getTypeAsString Depends on D153461 Reviewed By: razvanlupusoru, jeanPerier Differential Revision: https://reviews.llvm.org/D153467
-
Fangrui Song authored
`__xray_customevent` and `__xray_typedevent` are built-in functions in Clang. With -fxray-instrument, they are lowered to intrinsics llvm.xray.customevent and llvm.xray.typedevent, respectively. These intrinsics are then lowered to TargetOpcode::{PATCHABLE_EVENT_CALL,PATCHABLE_TYPED_EVENT_CALL}. The target is responsible for generating a code sequence that calls either `__xray_CustomEvent` (with 2 arguments) or `__xray_TypedEvent` (with 3 arguments). Before patching, the code sequence is prefixed by a branch instruction that skips the rest of the code sequence. After patching (compiler-rt/lib/xray/xray_AArch64.cpp), the branch instruction becomes a NOP and the function call will take effects. This patch implements the lowering process for {PATCHABLE_EVENT_CALL,PATCHABLE_TYPED_EVENT_CALL} and implements the runtime. ``` // Lowering of PATCHABLE_EVENT_CALL .Lxray_sled_N: b #24 stp x0, x1, [sp, #-16]! x0 = reg of op0 x1 = reg of op1 bl __xray_CustomEvent ldrp x0, x1, [sp], #16 ``` As a result, two updated tests in compiler-rt/test/xray/TestCases/Posix/ now pass on AArch64. Reviewed By: peter.smith Differential Revision: https://reviews.llvm.org/D153320 -
Piotr Zegar authored
Add missing documentation for DelimiterStem and ReplaceShorterLiterals options. Fixes #54662 Reviewed By: Eugene.Zelenko Differential Revision: https://reviews.llvm.org/D153639
-
- Jun 23, 2023
-
-
Ties Stuij authored
Temporarily disabling the execute-only tests. We recently added codegen for armv6-m, which is still in heavy development (D152795). Disabling the tests while we're figuring out what's going on is probably the least disruptive option, as a patch dependent on it also already landed.
-
Simon Pilgrim authored
If we have an excessive number of stores in a single chain then the candidate WideVT may exceed the maximum width of an EVT integer type (and will assert) - but since mergeTruncStores doesn't support anything wider than a i64 store we should just early-out if we've collected more than stores than that. Fixes #63306
-
Nikita Popov authored
-
Nikita Popov authored
-
Nikita Popov authored
-
eopXD authored
The template was created in D151396 but was not aware of the change in D153067. This commit adds the operand and keep similar templates aligned. Reviewed By: reames, craig.topper Differential Revision: https://reviews.llvm.org/D153506
-
Emilia Kond authored
Previously, using ColumnLimit: 0 with extended inline asm with the BreakBeforeInlineASMColon: OnlyMultiline option (the default style), the formatter would act as if in Always mode, meaning a line break was added before every colon in an extended inline assembly block. This patch respects the already existing line breaks, and doesn't add any new ones, if in ColumnLimit 0 mode. Behaviour with Always stays as expected, with a break before every colon regardless of any existing line breaks. Behaviour with Never was broken before, and remains broken with this patch, it is just never respected in ColumnLimit 0 mode. Fixes https://github.com/llvm/llvm-project/issues/62754 Reviewed By: HazardyKnusperkeks, owenpan Differential Revision: https://reviews.llvm.org/D150848
-
Nikita Popov authored
Use InstCombine's insertion helper for the created extracts, so they become part of the worklist and will be revisited.
-
David Green authored
This call to reassociateReduction is used by both fminnum/fmaxnum and fminimum/fmaximum. In adding support for fminimum/fmaximum we appear to be fixing the use of an incorrect reduction type, which should have only applied to minnum/maxnum. I also believe that it doesn't need nsz and reassoc to perform the reassociation. For float min/max it should always be valid. Differential Revision: https://reviews.llvm.org/D153247
-
Nikita Popov authored
The inserted instructions can usually be simplified. Make sure this happens in the same InstCombine iteration by adding them to the worklist. We happen to get some better optimization in two cases, but this is just a lucky accident. https://github.com/llvm/llvm-project/issues/63472 tracks implementing a fold for that case. This doesn't track all inserted instructions yet, for that we would also have to include those created by ObjectSizeOffsetEvaluator.
-
Takuya Shimizu authored
Makes some comments conform to bugprone-argument-comment (https://clang.llvm.org/extra/clang-tidy/checks/bugprone/argument-comment.html)
-
Alex Bradbury authored
My MC layer support patches missed adding these to RISCVUsage. Also update the link to the most recent spec PDF (including the recently committed encoding fix for vfwmaccbf16.
-
Alex Bradbury authored
For the same reasons as D151284, this requires custom lowering of the truncate libcall on hard float ABIs (the normal libcall code path is used on soft ABIs). The extend operation is implemented by a shift just as in the standard legalisation, but needs to be custom lowered because i32 isn't a legal type on RV64. This patch aims to make the minimal changes that result in correct codegen for the bfloat.ll tests. Differential Revision: https://reviews.llvm.org/D151663
-
Matt Arsenault authored
-
Matt Arsenault authored
-
Matt Arsenault authored
Find and replace on the new log tests (plus <3 x half> which was missing). Apparently exp10 never worked.
-
Alex Bradbury authored
The encoding matched the one given in the bf16 extension specification PDF, but per https://github.com/riscv/riscv-bfloat16/issues/45 it seems this encoding was not the one that is intended and was incorrectly modified due to an issue in the PDF generation process. This patch corrects the opcode to 111011 from 100011. The correct encoding is shown in the new spec PDF <https://github.com/riscv/riscv-bfloat16/releases/tag/20230614>. Differential Revision: https://reviews.llvm.org/D152894
-