- Mar 07, 2021
-
-
Fangrui Song authored
The hackery is due to glibc clock_gettime crashing from preinit_array (D40679). 32-bit musl architectures do not define `__NR_clock_gettime` so the code causes a compile error. Tested on Alpine Linux x86-64 (musl) and FreeBSD x86-64. Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D96925
-
William S. Moses authored
Enable Attributor's heap-to-stack to lower unbounded allocations given a max size of -1 Differential Revision: https://reviews.llvm.org/D97873
-
Philip Reames authored
-
Philip Reames authored
GVN basically doesn't handle phi nodes at all. This is for a reason - we can't value number their inputs since the predecessor blocks have probably not been visited yet. However, it also creates a significant pass ordering problem. As it stands, instcombine and simplifycfg ends up implementing CSE of phi nodes. This means that for any series of CSE opportunities intermixed with phi nodes, we end up having to alternate instcombine/simplifycfg and gvn to make progress. This patch handles the simplest case by simply preprocessing the phi instructions in a block, and CSEing them if they are syntactically identical. This turns out to be powerful enough to handle many cases in a single invocation of GVN since blocks which use the cse'd phi results are visited after the block containing the phi. If there's a CSE opportunity in one the phi predecessors required to recognize the phi CSE opportunity, that will require a second iteration on the function. (Still within a single run of gvn though.) Compile time wise, this could go either way. On one hand, we're potentially causing GVN to iterate over the function more. On the other, we're cutting down on iterations between two passes and potentially shrinking the IR aggressively. So, a bit unclear what to expect. Note that this does still rely on instcombine to canonicalize block order of the phis, but that's a one time transformation independent of the values incoming to the phi. Differential Revision: https://reviews.llvm.org/D98080
-
Martin Storsjö authored
Differential Revision: https://reviews.llvm.org/D98107
-
Philip Reames authored
As a pragmatic tradeoff, the ease of updating the tests outweighs the slightly easier to understand test conditions. Where revevant, debug output was converted to comments to help human understanding.
-
Philip Reames authored
-
Philip Reames authored
I'd originally intended to build on this for another purpose and have decided not to, but at a minimum, the stronger asserts are useful.
-
Philip Reames authored
-
Vy Nguyen authored
Differential Revision: https://reviews.llvm.org/D98115
-
- Mar 06, 2021
-
-
Nikita Popov authored
Separate out some conditions with early exits, to make it easier to support additional cases.
-
Nikita Popov authored
Examples of things we mostly don't handle.
-
KareemErgawy-TomTom authored
To unify the naming scheme across all ops in the SPIR-V dialect, we are moving from spv.camelCase to spv.CamelCase everywhere. For ops that don't have a SPIR-V spec counterpart, we use spv.mlir.snake_case. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D98014
-
Yaxun (Sam) Liu authored
Spack is a package management tool extensively used by HPC community. As ROCm packages are built by Spack by HPC community, we need to teach clang driver to detect ROCm installation built by Spack. Reviewed by: Artem Belevich Differential Revision: https://reviews.llvm.org/D97340
-
Lei Zhang authored
Normally tensors will be stored in buffers before converting to SPIR-V, given that is how a large amount of data is sent to the GPU. However, SPIR-V supports converting from tensors directly too. This is for the cases where the tensor just contains a small amount of elements and it makes sense to directly inline them as a small data array in the shader. To handle this, internally the conversion might create new local variables. SPIR-V consumers in GPU drivers may or may not optimize that away. So this has implications over register pressure. Therefore, a threshold is used to control when the patterns should kick in. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D98052
-
Alexey Lapshin authored
That review is extracted from D69372. It fixes https://bugs.llvm.org/show_bug.cgi?id=42219 bug. For the noimplicitfloat mode, the compiler mustn't generate floating-point code if it was not asked directly to do so. This rule does not work with variable function arguments currently. Though compiler correctly guards block of code, which copies xmm vararg parameters with a check for %al, it does not protect spills for xmm registers. Thus, such spills are generated in non-protected areas and could break code, which does not expect floating-point data. The problem happens in -O0 optimization mode. With this optimization level there is used FastRegisterAllocator, which spills virtual registers at basic block boundaries. Register Allocator does not protect spills with additional control-flow modifications. Thus to resolve that problem, it is suggested to not copy incoming physical registers into virtual registers. Instead, store incoming physical xmm registers into the memory from scratch. Differential Revision: https://reviews.llvm.org/D80163
-
Nikita Popov authored
There seems to be an impedance mismatch between what the type system considers an aggregate (structs and arrays) and what constants consider an aggregate (structs, arrays and vectors). Rather than adjusting the type check, simply drop it entirely, as getAggregateElement() is well-defined for non-aggregates: It simply returns null in that case.
-
Nikita Popov authored
-
David Zarzycki authored
This partially reverts commit e1173c87 until we find out why libcxx tests are failing under runtimes build.
-
Juneyoung Lee authored
-
Nikita Popov authored
Instead of handling a number of special cases for selects, handle this generally when inferring ranges from conditions. We already infer ranges from `x + C pred C2` to `x`, so doing the same for `x pred C2` to `x + C` is straightforward.
-
Nikita Popov authored
These are the same as the existing tests, but using different predicates that are not handled by the current code.
-
Raul Tambre authored
It moved the logic for CMake target arguments into llvm_ExternalProject_Add(). No handling was added for CMAKE_CROSSCOMPILING, which has a separate set of compiler_args. This broke crosscompiling, as now the runtimes builds defaulted to the compiler's default. I've also added passing of CMAKE_ASM_COMPILER, which was missing before although we were passing the triple for it. Reviewed By: zero9178 Differential Revision: https://reviews.llvm.org/D97855
-
Nikita Popov authored
Instead of by pointer. This allows us to use offsets that are not materialized in the IR.
-
Nikita Popov authored
These tests didn't test the pattern they were supposed to, because %a instead of %add was used in the select, which turned this into a normal min/max). Noticed this when commenting out the clamp handling code did not result in any test failures...
-
Jay Foad authored
This reverts commit e58d68fc. This reinstates commit fc28f600 with a fix to initialize HasShaderCyclesRegister. See https://reviews.llvm.org/D97928.
-
Aleksandr Platonov authored
Without this patch the file list of the preamble index contains URIs, but other indexes file lists contain file paths. This makes `indexedFiles()` always returns `IndexContents::None` for the preamble index, because current implementation expects file paths inside the file list of the index. This patch fixes this problem and also helps to avoid a lot of URI to path conversions during indexes merge. Reviewed By: kadircet Differential Revision: https://reviews.llvm.org/D97535
-
Martin Storsjö authored
If cross testing (and manually specifying a LIBCXX_TARGET_INFO in the cmake configuration, as the default is to match the build platform), we want the accessors for querying the target platform, is_windows, is_darwin, to return the right value depending on which target info class is used, not based on what platform is running the build and driving the tests. When LIBCXX_TARGET_INFO isn't defined, the right target info class is chosen automatically based on the platform one is running on, so this shouldn't make any practical difference for such setups. Differential Revision: https://reviews.llvm.org/D98045
-
Martin Storsjö authored
For MinGW targets, we distinguish between an explicitly shared unwinder library (requested via -shared-libgcc), an explicitly static one (requested via -static-libgcc or -static) and the default case (which just passes -lunwind to the linker, which will pick either shared or static depending on what's available, with the normal linker logic). This makes the implicit default case (as added in D79995) actually work as it was intended, when using the g++ driver (which is the main usecase for libunwind as far as I know). Differential Revision: https://reviews.llvm.org/D98023
-
Martin Storsjö authored
CLANG_DEFAULT_RTLIB had a typo, and libunwind isn't a valid option for it. This keeps the actual behaviour from before, defaulting to none if using compiler-rt as rtlib. Differential Revision: https://reviews.llvm.org/D98022
-
Fangrui Song authored
BFD_RELOC_NONE is useful for ld --gc-sections: it provides a generic way indicating a dependency between two sections.
-
Fangrui Song authored
BFD_RELOC_NONE is useful for ld --gc-sections: it provides a generic way indicating a dependency between two sections.
-
Fangrui Song authored
-
Mehdi Amini authored
This is fixing the missing title and menu entry on the MLIR website.
-
Fangrui Song authored
BFD_RELOC_NONE is useful for ld --gc-sections: it provides a generic way indicating a dependency between two sections.
-
Fangrui Song authored
BFD_RELOC_NONE is useful for ld --gc-sections: it provides a generic way indicating a dependency between two sections.
-
Fangrui Song authored
The names are unfortunate, but BFD_RELOC_NONE provides a generic way indicating a dependency between two sections, which is useful for ld --gc-sections. See https://sourceware.org/bugzilla/show_bug.cgi?id=27530
-
Vitaly Buka authored
ABORTING message is inconsistent across sanitizers. Another followup for D98089
-
Christopher Di Bella authored
Implements parts of: - P0898R3 Standard Library Concepts - P1754 Rename concepts to standard_case for C++20, while we still can Depends on D96742 Differential Revision: https://reviews.llvm.org/D97162 -
Jianzhou Zhao authored
-