- Jul 26, 2023
-
-
Mark de Wever authored
As mentioned in D155024 libc++ now has release notes for LLVM 17 and 18 to make cherry-picking easier.
-
Johannes Doerfert authored
This reverts commit 0d126830 and reapplies ef9ec4bb with an extension to fix the Flang build. Differential Revision: https://reviews.llvm.org/D156184
-
Kevin P. Neal authored
Correct AMDGPU strictfp tests to follow the rules documented in the LangRef: https://llvm.org/docs/LangRef.html#constrained-floating-point-intrinsics Mostly these tests just needed the strictfp attribute on function definitions. I've also removed the strictfp attribute from uses of the constrained intrinsics because it comes by default since D154991, but I only did this in tests I was changing anyway. I also removed attributes added to declare lines of intrinsics. The attributes of intrinsics cannot be changed in a test so I eliminated attempts to do so. Test changes verified with D146845.
-
Quinn Dawkins authored
Adds `apply_patterns.gpu.unroll_vectors_subgroup_mma` which allows specifying a native MMA shape of `m`, `n`, and `k` to unroll to, greedily unrolling the inner most dimension of contractions and other vector operations based on expected usage. Differential Revision: https://reviews.llvm.org/D156079
-
Jan Sjodin authored
This patch ensures that all outlined functions parameters are i64 or ptr when compiling for a target device, which is what the OpenMP runtime expects. The values are then cast to the correct type inside the kernel. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D155628
-
Paul Robinson authored
Differential revision: https://reviews.llvm.org/D156248
-
Thomas Köppe authored
Currently, objcopy cannot set the new flag SHF_X86_64_LARGE. This change introduces the named flag "large" which translates to that section flag. An "invalid argument" error is produced if a user attempts to set the flag on an architecture other than X86_64. Reviewed By: jhenderson, MaskRay Differential Revision: https://reviews.llvm.org/D153262
-
Arthur Eubanks authored
-
Craig Topper authored
This does change the InstrFormat from I to R. It is closer to R, but the 2 source register fields are used for immediates. Thankfully the R and I format values in TSFlags aren't used for anything.
-
Aaron Ballman authored
Spotted during post-commit review of 4b15bb9a
-
Cyndy Ishida authored
InterfaceFile is the in-memory representation for tbd files. Add APIs to merge, extract, remove, and inline reexported libraries. Reviewed By: zixuw Differential Revision: https://reviews.llvm.org/D153398
-
Sam McCall authored
Most headers that we discover are likely to be fairly portable across at least clang/gcc, which is why we get away with adding them to clangd's search path. This is not the case for the compiler builtin headers, which are shipped with the parser and rely on parser builtins/features. We're better off *not* using the the scanned builtin headers, and hoping our own intrinsics provide the interface the code is relying on. Fixes https://github.com/clangd/clangd/issues/1695 Differential Revision: https://reviews.llvm.org/D156044
-
Craig Topper authored
If we are selecting between two setccs that need to be legalized with xor, the select will be legalized first. Detect this pattern so we can pull the xor through to expose it to additional optimizations. We could generalize this to other operations, but those normally get handled in DAG combine before select legalization. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D156159
-
Craig Topper authored
Instead of checking for setcc, look for any 0/1 value. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D156153
-
Craig Topper authored
If we're selecting the result of two setccs that have been legalized by introducing an xor with 1, we can pull the xor with 1 through the select to enable more optimizations. We could generalize this to other binary operators with identical conditions, but those are usually caught before we legalize the select. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D156144
-
Yaxun (Sam) Liu authored
Rename -fcuda-approx-transcendentals as -fgpu-approx-transcendentals and pass it to both device and host clang -cc1. Fix its interaction with -ffast-math to allow -fno-gpu-approx-transcendentals to override the implicit -fcuda-approx-transcendentals due to -ffast-math. Rename the predefined macro to be __CLANG_GPU_APPROX_TRANSCENDENTALS__. Emit the macro for both device and host compilation. Reviewed by: Artem Belevich, Fangrui Song Differential Revision: https://reviews.llvm.org/D154797
-
- Jul 25, 2023
-
-
Craig Topper authored
This avoids creating an unnecessary register pressure set. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D156196
-
Oleg Shyshkov authored
-
Weining Lu authored
This reverts commit 92c06114.
-
Jakub Kuderski authored
1-element vectors are not valid in SPIR-V and fail `Bitcast` op verification. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D156207
-
Andrei Homescu authored
Back-ported from https://r.android.com/2591905. Reviewed By: Chia-hungDuan Differential Revision: https://reviews.llvm.org/D155144
-
Marco Nelissen authored
Back-ported from https://r.android.com/2591905. Reviewed By: Chia-hungDuan Differential Revision: https://reviews.llvm.org/D152888
-
Chia-hung Duan authored
BatchClassId will go through pushBatchClassBlocks Reviewed By: cferris Differential Revision: https://reviews.llvm.org/D156148
-
Aaron Ballman authored
This addresses the issue found by: https://lab.llvm.org/buildbot/#/builders/92/builds/47877
-
Mark de Wever authored
This test fails on clang-18.
-
Nikita Popov authored
When commuting the operands, don't create a constant expression for undesirable binops. Only invoke the constant folding function in that case.
-
Jacek Caban authored
This is a preparation for ARM64EC/ARM64X binaries, which may contain both ARM64 and x86_64 code in the same file. llvm-objdump already has partial support for mixing disassemblers for ARM thumb mode support. However, for ARM64EC we can't share MCContext, MCInstrAnalysis and PrettyPrinter instances. This patch provides additional abstraction which makes adding mixed code support later in the series easier. Reviewed By: jhenderson, MaskRay Differential Revision: https://reviews.llvm.org/D149093
-
Yaxun (Sam) Liu authored
CUDA allows replacing standard math functions with less accurate native math functions when -use_fast_math is specified (https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#intrinsic-functions). Cuda-clang does this when -fcuda-approx-transcendentals or -ffast-math is specified. HIP-clang currently passes option -fcuda-approx-transcendentals to clang -cc1 and predefines __CLANG_CUDA_APPROX_TRANSCENDENTALS__ but does not replace standard math functions with native math functions. This patch implements this in a similar approach as cuda-clang. Reviewed by: Brian Sumner, Matt Arsenault Differential Revision: https://reviews.llvm.org/D154790
-
Haojian Wu authored
C99 is the the earliest C version provided by the language opt. Differential Revision: https://reviews.llvm.org/D155816
-
Teresa Johnson authored
Guard FoldBranchToCommonDest in SimplifyCFG with the SpeculateBlocks flag as it can also speculate instructions. This was split out of D155997. Differential Revision: https://reviews.llvm.org/D156194
-
Yaxun (Sam) Liu authored
start with example usage, predefined macros, and path setting. Reviewed by: Brian Sumner, Siu Chi Chan, Matt Arsenault, Ronan Keryell Differential Revision: https://reviews.llvm.org/D154123
-
Yaxun (Sam) Liu authored
Recover the checking for the default language standard for C++. Reviewed by: Douglas Yung, Paul Robinson
-
Nikita Popov authored
This reapplies the change for and, but also marks or as undesirable at the same time. Only handling one of them can cause infinite combine loops due to the asymmetric handling. ----- In preparation for removing support for and/or expressions, mark them as undesirable. As such, we will no longer implicitly create such expressions, but they still exist.
-
Weining Lu authored
As described in [1][2], `-mtune=` is used to select the type of target microarchitecture, defaults to the value of `-march`. The set of possible values should be a superset of `-march` values. Currently possible values of `-march=` and `-mtune=` are `native`, `loongarch64` and `la464`. D136146 has supported `-march={loongarch64,la464}` and this patch adds support for `-march=native` and `-mtune=`. A new ProcessorModel called `loongarch64` is defined in LoongArch.td to support `-mtune=loongarch64`. `llvm::sys::getHostCPUName()` returns `generic` on unknown or future LoongArch CPUs, e.g. the not yet added `la664`, leading to `llvm::LoongArch::isValidArchName()` failing to parse the arch name. In this case, use `loongarch64` as the default arch name for 64-bit CPUs. And these two preprocessor macros are defined: - __loongarch_arch - __loongarch_tune [1]: https://github.com/loongson/LoongArch-Documentation/blob/2023.04.20/docs/LoongArch-toolchain-conventions-EN.adoc [2]: https://github.com/loongson/la-softdev-convention/blob/v0.1/la-softdev-convention.adoc Differential Revision: https://reviews.llvm.org/D155824 -
Joseph Huber authored
Summary: This is supposed to be enabled to say that we want correct sqrt by default.
-
Matt Arsenault authored
-
Matt Arsenault authored
Remove test hack that was accidentally pushed.
-
LLVM GN Syncbot authored
-
Michael Halkenhaeuser authored
[OpenMP] [OMPT] [7/8] Invoke tool-supplied callbacks before and after target launch and data transfer operations Implemented RAII objects, initialized at target entry points, that invoke tool-supplied callbacks. Updated status of target callbacks as implemented. Depends on D127365 Patch from John Mellor-Crummey <johnmc@rice.edu> With contributions from: Dhruva Chakrabarti <Dhruva.Chakrabarti@amd.com> Jan-Patrick Lehr <janpatrick.lehr@amd.com> Reviewed By: jdoerfert, dhruvachak, jplehr Differential Revision: https://reviews.llvm.org/D127367
-
Matt Arsenault authored
-