- Jun 20, 2023
-
-
Joseph Huber authored
The GPU vendors currently provide bitcode files for their device runtime. These files need to be handled specially as they are not built to be linked in with a standard `llvm-link` call or through LTO linking. This patch adds an alternative to use the existing clang handling of these libraries that does the necessary magic to make this work. We do this by causing the LTO backend to emit bitcode before running the backend. We then pass this through to clang which uses the existing support which has been fixed to support this by D152391. The backend will then be run with the merged module. This patch adds the `--builtin-bitcode=<triple>=file.bc` to specify a single file, or just `--clang-backend` to let the toolchain handle its defaults (currently nothing for NVPTX and the ROCm device libs for AMDGPU). This may have a performance impact due to running the optimizations again, we could potentially disable optimizations in LTO and only do the linking if this is an issue. This should allow us to resolve issues when relying on the `linker-wrapper` to do a late linking that may depend on vendor libraries. Depends on D152391 Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D152442
-
Joseph Huber authored
Clang provides the `-mlink-bitcode-file` and `-mlink-builtin-bitcode` options to insert LLVM-IR into the current TU. These are usefuly primarily for including LLVM-IR files that require special handling to be correct and cannot be linked normally, such as GPU vendor libraries like `libdevice.10.bc`. Currently these options can only be used if the source input goes through the AST consumer path. This patch makes the changes necessary to also support this when the input is LLVM-IR. This will allow the following operation: ``` clang in.bc -Xclang -mlink-builtin-bitcode -Xclang libdevice.10.bc ``` Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D152391
-
Matthias Springer authored
Add extra error checking to prevent passes from being run on unsupported ops through the pass manager infrastructure. Differential Revision: https://reviews.llvm.org/D153144
-
Haojian Wu authored
This fixes a false positive where a ParamVarDecl happend to be the same name of some C standard symbol and has a global namespace. ``` using A = int(int time); // we suggest <ctime> for the `int time`. ``` Differential Revision: https://reviews.llvm.org/D153330
-
Vladislav Dzhidzhoev authored
Revert "Reland "[DebugMetadata][DwarfDebug] Support function-local types in lexical block scopes (4/7)" (2)" This reverts commit cb9ac705. It causes an assert in clang: virtual void llvm::DwarfDebug::endFunctionImpl(const llvm::MachineFunction*): Assertion `LScopes.getAbstractScopesList().size() == NumAbstractSubprograms && "getOrCreateAbstractScope() inserted an abstract subprogram scope"' failed. https://bugs.chromium.org/p/chromium/issues/detail?id=1456288#c2
-
Alexey Lapshin authored
This patch is a followup for D153162. It cures one more place where indexed address was incorrectly read. It also moves handling of indexed address into DWARFUnit. Differential Revision: https://reviews.llvm.org/D153297
-
Alex Zinenko authored
LLVM build system separates between `add_llvm_example_library` and `add_llvm_library`, which is presumably used to package examples separately from the regular library. Introduce a similar approach to building example libraries in MLIR and use it for the transform dialect tutorial. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D153265
-
Harvin Iriawan authored
- Update the Cortex-A510 mcpu target to use A510 scheduling info instead of A55. Values taken are based on the A510 software optimisation guide https://developer.arm.com/documentation/PJDOC-466751330-536816/latest - Make latency of most integer ops to 1. CPU uarch is able to resolve most integer ops in 1 cycle Differential Revision: https://reviews.llvm.org/D152688
-
ManuelJBrito authored
Drop alignment to allow test to run in different platforms. Differential Revision: https://reviews.llvm.org/D152547
-
Chuanqi Xu authored
Try to address part of https://github.com/llvm/llvm-project/issues/61900. It is not completely addressed since the original reproducer is not fixed due to the final suspend point is optimized out in its special case. But that is a relatively independent issue.
-
Ivan Kosarev authored
They cause failures on the llvm-clang-x86_64-expensive-checks-debian buildbot. This partially reverts D153269 [AMDGPU][GFX11] Add test coverage for FMA instructions.
-
Jan Svoboda authored
This is a follow-up to D151938 that should fix GCC's -Wcast-qual warning.
-
Francesco Petrogalli authored
Differential Revision: https://reviews.llvm.org/D153325
-
Ivan Kosarev authored
Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153269
-
Francesco Petrogalli authored
The option `-misched-detail-resource-booking` prints the following information every time the method `SchedBoundary::getNextResourceCycle` is invoked: 1. counters of the resources that have already been booked; 2. the values returned by `getNextResourceCycle`, which is the next available cycle in which a resource can be booked. The method is useful to debug low-level checks inside the machine scheduler that make decisions based on the values returned by `getNextResourceCycle`. Reviewed By: andreadb Differential Revision: https://reviews.llvm.org/D153116
-
Francesco Petrogalli authored
Reverting because of https://lab.llvm.org/buildbot#builders/75/builds/32485: llvm-project/llvm/lib/CodeGen/MachineScheduler.cpp:2374:7: error: use of undeclared identifier 'MischedDetailResourceBooking' if (MischedDetailResourceBooking) This reverts commit fc06262c.
-
Francesco Petrogalli authored
The option `-misched-detail-resource-booking` prints the following information every time the method `SchedBoundary::getNextResourceCycle` is invoked: 1. counters of the resources that have already been booked; 2. the values returned by `getNextResourceCycle`, which is the next available cycle in which a resource can be booked. The method is useful to debug low-level checks inside the machine scheduler that make decisions based on the values returned by `getNextResourceCycle`. Reviewed By: andreadb Differential Revision: https://reviews.llvm.org/D153116
-
Matthias Springer authored
All `apply` functions now have a `TransformRewriter &` parameter. This rewriter should be used to modify the IR. It has a `TrackingListener` attached and updates the internal handle-payload mappings based on rewrites. Implementations no longer need to create their own `TrackingListener` and `IRRewriter`. Error checking is integrated into `applyTransform`. Tracking listener errors are reported only for ops with the `ReportTrackingListenerFailuresOpTrait` trait attached, allowing for a gradual migration. Furthermore, errors can be silenced with an op attribute. Additional API will be added to `TransformRewriter` in subsequent revisions. This revision just adds an "empty" `TransformRewriter` class and updates all `apply` implementations. Differential Revision: https://reviews.llvm.org/D152427
-
Nikita Popov authored
If there are no post-inc loops, normalization is a no-op. Don't bother rewriting the SCEV in that case.
-
Diana Picus authored
Co-authored-by:
Nicolai Hähnle <nicolai.haehnle@amd.com> Differential Revision: https://reviews.llvm.org/D151997
-
Diana Picus authored
Add a section to AMDGPUUsage.rst about calling conventions and list the ones from the CallingConv enum. Full descriptions can come later (help appreciated). Differential Revision: https://reviews.llvm.org/D151996
-
serge-sans-paille authored
This makes the new `order` subcommand part of the help. As a side effect, also make llvm::map_range compatible with plain arrays. Differential Revision: https://reviews.llvm.org/D153303
-
Michael Buch authored
Currently we emit `DW_AT_deleted` for `deleted` special-member functions (i.e., ctors/dtors). However, in C++ one can mark any member function as deleted. This patch expands the set of member functions for which we emit `DW_AT_deleted`. The DWARFv5 spec section 5.7.8 says: ``` <non-normative> In C++, a member function may be declared as deleted. This prevents the compiler from generating a default implementation of a special member function such as a constructor or destructor, and can affect overload resolution when used on other member functions. </non-normative> If the member function entry has been declared as deleted, then that entry has a DW_AT_deleted attribute. ``` Thus this change is conforming. Differential Revision: https://reviews.llvm.org/D153282
-
Nuno Lopes authored
-
Ben Shi authored
Try to break a multiplication with a specific immediate to an/a addition/subtraction of left shifts. Reviewed By: zixuan-wu Differential Revision: https://reviews.llvm.org/D153106
-
Ben Shi authored
Reviewed By: zixuan-wu Differential Revision: https://reviews.llvm.org/D153105
-
Diana Picus authored
This reverts commit aa7b127c. ...because I really ought to install sphinx.
-
Martin Braenne authored
This reverts commit dfbcee28. This was causing unit tests to fail on Gentoo, see comments on https://reviews.llvm.org/D152696.
-
Diana Picus authored
-
Diana Picus authored
Add a section to AMDGPUUsage.rst about calling conventions and list the ones from the CallingConv enum. Full descriptions can come later (help appreciated). Differential Revision: https://reviews.llvm.org/D151996
-
Nathan Ridge authored
Fixes https://github.com/clangd/clangd/issues/1666 Differential Revision: https://reviews.llvm.org/D153251
-
Haojian Wu authored
The binary tool only works on working source code, if the source code is not compilable, don't perform any analysis and edits. Differential Revision: https://reviews.llvm.org/D153271
-
Matthias Springer authored
This transform op runs a pass on the target op. Differential Revision: https://reviews.llvm.org/D153143
-
Kazu Hirata authored
-
Fangrui Song authored
Optimize (cmp+beq => cbz), duduplicate code (SAVE_REGISTERS/RESTORE_REGISTERS), improve portability (use ASM_SYMBOL to be compatible with Mach-O), and fix style issues. Also, port D37965 (x86 tail call) to __xray_FunctionTailExit.
-
Jaroslav Sevcik authored
-
Jaroslav Sevcik authored
This removes dependence on the libc abort function.
-
Zhongyunde authored
For simd vector selects, use cmeq + bsl for v2f32/v4f32/v2f64, so their cost are cheep. Fix https://github.com/llvm/llvm-project/issues/63082 Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D152523
-
Bing1 Yu authored
%tile = call x86_amx @llvm.x86.tileloadd64.internal(i16 8, i16 32, i8* %src_ptr, i64 64) %vec = call <256 x i8> @llvm.x86.cast.tile.to.vector.v256i8(x86_amx...%tile) store <256 x i8> %vec, <256 x i8>* %dst_ptr, align 256 => %tile = call x86_amx @llvm.x86.tileloadd64.internal(i16 8, i16 32, i8* %src_ptr, i64 64) %stride = sext i16 32 to i64 call void @llvm.x86.tilestored64.internal(i16 8, i16 32, i8* %dst_ptr, i64 32, x86_amx %tile) Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D153002
-
Fangrui Song authored
The intrinsic has a smaller integer type than the parameter type of builtin-function/API. Fix this similar to commit 3fa3cb40.
-