- Oct 06, 2022
-
-
Christopher Bate authored
The ConvertVectorToGpu pass implementation contained a small private support library for performing various calculations during conversion between `vector` and `nvgpu.mma.sync` and `nvgpu.ldmatrix` operations. The support library is moved under `Dialect/NVGPU/Utils` because the functions have wider utility. Some documentation comments are added or improved. Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D135303
-
Peiming Liu authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D135337
-
wren romano authored
This is a followup to D135004, to correct one of the tests that didn't get caught by the buildbot. Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D135336
-
Ilia Diachkov authored
The patch introduces reading the attributes of kernel arguments both from function-attached and module-level metadata, during kernel arguments lowering. Two tests are added to show the improvement. Differential Revision: https://reviews.llvm.org/D135106 Co-authored-by:
Aleksandr Bezzubikov <zuban32s@gmail.com> Co-authored-by:
Michal Paszkowski <michal.paszkowski@outlook.com> Co-authored-by:
Andrey Tretyakov <andrey.tretyakov@mail.com> Co-authored-by:
Konrad Trifunovic <konrad.trifunovic@intel.com>
-
wren romano authored
This differential adjusts the numeric values for DimLevelType values: using the low-order two bits for recording the "No" and "Nu" properties, and the high-order bits for the formats per se. (The choice of encoding may seem a bit peculiar, since the bits are mapped to negative properties rather than positive properties. But this was done in order to preserve the collation order of DimLevelType values. If we don't care about collation order, then we may prefer to flip the semantics of the property bits, so that they're less surprising to readers.) Using distinguished bits for the properties and formats enables faster implementation for the predicates detecting those properties/formats, which matters because this is in the runtime library itself (rather than on the codegen side of things). This differential pushes through the changes to the enum values, and optimizes the basic predicates. However it does not optimize all the places where we check compound predicates (e.g., "is compressed or singleton"), to help reduce rebasing conflict with D134933. Those optimizations will be done after this differential and D134933 are landed. Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D135004
-
Carl Ritson authored
Invariant loads can always be sunk. Reviewed By: foad, arsenm Differential Revision: https://reviews.llvm.org/D135133
-
Murali Vijayaraghavan authored
vector.extract, if the result is a scalar) only if all reduction dimensions are of size 1. Differential Revision: https://reviews.llvm.org/D135333
-
Daniel Rodríguez Troitiño authored
Fix typo in error message. No other changes Differential Revision: https://reviews.llvm.org/D135318
-
Yaxun (Sam) Liu authored
In HIP a library is usually compiled with default target ID e.g. gfx906 so that it can be used in all GPU configurations. The bitcode is saved in bundled bitcode with gfx906 in entry ID. In runtime compilation, a HIP program is compiled with a target ID matching the GPU configuration, e.g. gfx906:xnack-. This program needs to link with a library bundled bitcode with target ID gfx906. For example: clang --offload-arch=gfx906 -o lib.o lib.hip clang --offload-arch=gfx906:xnack- program.hip lib.o This common use case requires that clang-offlod-bundler to be able to extract entry with compatible target ID, e.g. extracting an gfx906 entry when requesting gfx906:xnack-. Currently clang-offload-bundler only allow extracting entry with exact match of target ID. This patch relaxes that so that it can extract entries with compatible target ID. Reviewed by: Artem Belevich, Saiyedul Islam Differential Revision: https://reviews.llvm.org/D134546
-
Dominic Chen authored
This reverts commit 5470b1fc.
-
Rob Suderman authored
If the scaling factor is by 1 with no offset or border, then the resize is a no-op. Reviewed By: dcaballe Differential Revision: https://reviews.llvm.org/D135329
-
Dominic Chen authored
Differential Revision: https://reviews.llvm.org/D134917
-
wren romano authored
Handle more cases of singleton DLT including direct sparse2sparse conversion. (Followup to D134096) Depends On D134926 Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D134933
-
Ben Langmuir authored
When dep-scanning, canonicalize the module map path as much as we can. This avoids unnecessarily needing to build multiple versions of a module due to symlinks or case-insensitive file paths. Despite the name `tryGetRealPathName`, the previous implementation did not actually return the realpath most of the time, and indeed it would be incorrect to do so since the realpath could be outside the module directory, which would have broken finding headers relative to the module. Instead, use a canonicalization that is specific to the needs of modulemap files (canonicalize the directory separately from the filename). Differential Revision: https://reviews.llvm.org/D134923
-
Nathaniel McVicar authored
Restores the fix from D134925 for MSVC without breaking cpu runner. Differential Revision: https://reviews.llvm.org/D135304
-
Sami Tolvanen authored
Specify the correct size for the KCFI_CHECK pseudo instruction, which is lowered into six 4-byte instructions in AArch64AsmPrinter::LowerKCFI_CHECK. Link: https://github.com/ClangBuiltLinux/linux/issues/1730
-
Matthias Braun authored
Avoid crash in `reduceOperandsOneDeltaPass` function for operands with vector of pointer type. While on it add a `reduce-operands-ptr.ll` test in the spirit of the existing `reduce-operands-int.ll`/`reduce-operands-fp.ll` tests. Differential Revision: https://reviews.llvm.org/D135307
-
Aart Bik authored
Makes individual testing and debugging easier. Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D135319
-
Murali Vijayaraghavan authored
This reverts commit c16f3260. There's a bug in the commit creates a scalar result with `ShapeCastOp`. Reverting till that fix is done.
-
Quentin Colombet authored
Forgot another SDValue check and a boolean initialization.
-
Valentin Clement authored
It is useful for couple of test suite like NAG to keep failing with a TODO until the polymorphic entities is implemented all the way done to codegen. This pass adds a flag to LoweringOptions for experimental development. This flag is off by default and can be enable in `bbc` with `-polymorphic-type`. Options can be added in the driver and tco when needed. Reviewed By: PeteSteinfeld Differential Revision: https://reviews.llvm.org/D135283
-
Sam McCall authored
These are 8 bytes and we don't care about the actual ordering, so use integer compare. The array generated code has some extra byte swaps (clang), calls memcmp (gcc) or inlines a big chain of comparisons (MSVC): https://godbolt.org/z/e79r6jM6K
-
Quentin Colombet authored
Explicitly check an SDValue with the invalid SDValue. UBSan reports: runtime error: load of value 36, which is not a valid value for type 'bool' https://lab.llvm.org/buildbot/#/builders/85/builds/11231
-
Quentin Colombet authored
This patch allows the combines that fold extensions in binary operations to have more than one use. The approach here is pretty conservative: if all the users of an extension can fold the extension, then the folding is done, otherwise we don't fold. This is the first step towards avoiding the one-use limitation. As a result, we make a decision to fold/don't fold for a web of instructions. An instruction is part of the web of instructions as soon as it consumes an extension that needs to be folded for all its users. Because of how SDISel works a web of instructions can be visited over and over. More precisely, if the folding happens, it happens for the whole web and that's the end of it, but if the folding fails, the whole web may be revisited when another member of the web is visited. To avoid a compile time explosion in pathological cases, we bail out earlier for webs that are bigger than a given threshold (arbitrarily set at 18 for now.) This size can be changed using `--riscv-lower-ext-max-web-size=<maxWebSize>`. At the current time, I didn't see a better scheme for that. Assuming we want to stick with doing that in SDISel. Differential Revision: https://reviews.llvm.org/D133739
-
David Green authored
This adds some missing tablegen patterns to handle trn1/trn2/zip1/zip2/uzp1/uzp2, similar to the Arm handling in 5e1a9d31, but via tablegen patterns for the AArch64 backend.
-
oToToT authored
As PR56606 stated, the current implementation of CityHash in libc++ would drop some bits unintentionally. Cast the 32bit int to the 64bit int to avoid this happened. Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D134124
-
David Blaikie authored
(& remove PSPTargetInfo because it's unused - it had the wrong ctor in it anyway, so wouldn't've been able to be instantiated - must've happened due to bitrot over the years)
-
Mark de Wever authored
This concept is introduced in P2286, but was implemented in libc++ before. This implementation was used in the library internally. This implementation lacked the resolution of LWG3636. The original formatter had a non-const member function that wasn't trivial to make a const member. The recent parser improvements made this member a const member in preparation of LWG3636. Note LWG3636 isn't voted in. Its status is Ready. P2286's concept has been written as-if LWG3636 is accepted and refers to that LWG issue. Updates some tests make format a const member function and removes a tests that's mainly a duplicate of the formattable concept test. Implements - LWG3636 formatter<T>::format should be const-qualified Implements parts of - P2286R8 Formatting Ranges Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D134110
-
TatWai Chong authored
Attribute stride and shift are removed, and has new scale and border. Signed-off-by:
TatWai Chong <tatwai.chong@arm.com> Change-Id: I6cdbeb3978f5ee540bc6cf59eb7c273eb0131430 Reviewed By: rsuderman Differential Revision: https://reviews.llvm.org/D131629
-
Ben Langmuir authored
Update SourceManager::ContentCache::OrigEntry to keep the original FileEntryRef, and use that to enable ModuleMap::getModuleMapFile* to return the original FileEntryRef. This change should be NFC for most users of SourceManager::ContentCache, but it could affect behaviour for users of getNameAsRequested such as in compileModuleImpl. I have not found a way to detect that difference without additional functional changes, other than incidental cases like changes from / to \ on Windows so there is no new test. Differential Revision: https://reviews.llvm.org/D135220
-
Filipp Zhinkin authored
Baseline tests for D135302.
-
River Riddle authored
The current splicing behavior dates back to when all blocks had terminators, so we would "helpfully" splice before the terminator. This doesn't make sense anymore, and leads to somewhat unexpected results when parsing multiple pieces of IR into the same block. Differential Revision: https://reviews.llvm.org/D135096
-
Zixu Wang authored
ExtractAPI doesn't care about locations of anonymous TagDecls. Set the printing policy to exclude that from anonymous decl names. Differential Revision: https://reviews.llvm.org/D135295
-
Ivan Butygin authored
Motivation: we have lowering pipeline based on upstream gpu and spirv dialects and and we are using host shared gpu memory to transfer data between host and device. Add `host_shared` flag to `gpu.alloc` to distinguish between shared and device-only gpu memory allocations. Differential Revision: https://reviews.llvm.org/D133533
-
Slava Zakharin authored
nullptr matches both against ::mlir::UnitAttr and ::mlir::TypeRange, so the following two candidates fit: static void mlir::omp::OrderedRegionOp::build(::mlir::OpBuilder &odsBuilder, ::mlir::OperationState &odsState, /*optional*/::mlir::UnitAttr simd) static void mlir::omp::OrderedRegionOp::build(::mlir::OpBuilder &odsBuilder, ::mlir::OperationState &odsState, ::mlir::TypeRange resultTypes, /*optional*/bool simd = false) -
Argyrios Kyrtzidis authored
In the context of caching clang invocations it is important to emit diagnostics in deterministic order; the same clang invocation should result in the same diagnostic output. rdar://100336989 Differential Revision: https://reviews.llvm.org/D135118
-
Joseph Huber authored
A previous patch merged the static and bitcode versions of the deviceRTL. We previously used the static library's separate compilation to set a special flag that prevented `IsSPMDMode` from being put in the used list and preventing it from being optimized out. When they were merged we could no longer do this separate compilation that allowed users of LTO to get more optimal code. This patch rearranges the code. The `IsSPMDMode` global is now transitively used by its inclusion in the changed `__keep_alive` function. This allows us to then manually delete the `__keep_alive` function from the module when building the static library via `llvm-extract`. The result is that the bitcode library correctly will maintain the needed shared state, while the static library will be able to internalize it and optimize it out. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D135280
-
Joseph Huber authored
We use protected visibility for almost everything with offloading. This is because it provides us with the ability to read things from the host without the expectation that it will be preempted by a shared library load, bugs related to this have happened when offloading to the host. This patch just makes the `exec_mode` global generated for each plugin have protected visibility. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D135285
-
Sam McCall authored
The simplest way to ensure full canonicalization is to canonicalize recursively in most cases. This fixes an assertion failure and presumably correctness bugs. It does show up that D132797's index-based virtual method renames doesn't handle templates well (the AST behavior is different and IMO better). We could choose to disable in this case or change the index behavior, but this patch doesn't do either. Differential Revision: https://reviews.llvm.org/D133415
-
Jakub Kuderski authored
Reviewed By: ThomasRaoux Differential Revision: https://reviews.llvm.org/D135234
-