- Aug 24, 2023
-
-
Johannes Doerfert authored
We used to have two separate implementations to derive the number of threads used in a target region. This lead us to sometimes miss out on user provided thread bounds (num_threads, or thread_limit) when we looked for "constant default values". If we might miss out on the presence of those bounds, we cannot set the thread_limit statically since the runtime will try to honor user input rather than cap it at the "preferred default". This patch replaces the secondary implementation with the primary in a mode that will not emit code but just look for the presence, and potentially upper bounds, of thread limiting clauses. The runtime test would not pass without this rewrite as we missed some clauses, set the static limit on the device to the preferred value, but then violated that value at runtime. Fixes: https://github.com/llvm/llvm-project/issues/64845 Differential Revision: https://reviews.llvm.org/D158381
-
Fangrui Song authored
-
Peter Rong authored
After recent patch D30189, #64323's error message become a new one. When DAGCombiner was optimizing `(vextract (scalar_to_vector val, 0) -> val`, it didn't consider the possibility that the inserted value type has less bit than the dest type. This patch fixes that. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D158355
-
Rahman Lavaee authored
-
Johannes Doerfert authored
Just re-running the script to make future updates easier
-
Johannes Doerfert authored
Without this we cannot update various clang OpenMP tests as the UTC_ARGS version of -global-value-regex is simply ignored. The handling of the flag should be changed to be in line with others, I left TODOs for now.
-
Johannes Doerfert authored
-
Vadim Paretsky authored
A few places in the loop collapse support code make small dynamic allocations that introduce a noticeable performance overhead when made on the heap. This change moves allocations up to 32 bytes to the stack instead of the heap. Differential Revision: https://reviews.llvm.org/D158220
-
Will Hawkins authored
Note that there is a Differential revision for the stride ranges view implementation and the names of the people who are developing it. Reviewed By: #libc, Mordante Differential Revision: https://reviews.llvm.org/D158446
-
Mark de Wever authored
Implements - P2497R0 Testing for success or failure of <charconv> functions Depends on D153192 Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D153199
-
Craig Topper authored
emitSubtargetFeatureBitEnumeration doesn't modify the map so we should use const reference.
-
Peter Klausler authored
Some vector indexing code in the preprocessor fails with empty tokens or token sequences in predefined macros. Fixes https://github.com/llvm/llvm-project/issues/64837. Differential Revision: https://reviews.llvm.org/D158451
-
David Green authored
After D157690 we are seeing some crashes from Global ISel, which seem to be related to the shift_of_shifted_logic_chain combine that can remove too many instructions if the shift amount is zero. This limits the fold to non-zero shifts, under the assumption that it is better in that case to fold away the shift to a COPY. Differential Revision: https://reviews.llvm.org/D158596
-
Dmitry Polukhin authored
See discussion in https://reviews.llvm.org/D89001 I updated test to match actual behavior of /Applications/Xcode_14.3.1_14E300b.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin/clang after that modified upstream to match the test. Differential Revision: https://reviews.llvm.org/D157283
-
Joseph Huber authored
This `MAX_LANE_SIZE` was a hack from the days when we used a single instance of the server and had some GPU state handle it. Now that we have everything templated this really shouldn't be used. This patch removes its use and replaces it with template arguments. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D158633
-
Yinying Li authored
Example: compressed-no -> compressed_no Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D158567
-
Mogball authored
Often times, large weights for ML models will be stored as resources in MLIR. It is sometimes advantageous to control whether to print these resources for debugging purposes. For example, some models contain very big weights with millions of characters in printed size, which may slow down whatever text editor you are using. This diff adds a flag which allows users to disable printing resources in these scenarios. Reviewed By: Mogball Differential Revision: https://reviews.llvm.org/D157928
-
Slava Zakharin authored
Related to https://github.com/llvm/llvm-project/issues/64866. This patch effectively disables CSE for identical hlfir.elemental operations, because it causes hlfir.destroy to be applied twice to the same temporary. Moreover, I think MemAlloc is correct for hlfir.elemental, in general. Reviewed By: tblah Differential Revision: https://reviews.llvm.org/D158565
-
Slava Zakharin authored
This effectively reverts D154715. The issue appears as the dialect conversion error because we try to erase an op that has already been erased. See the added LIT test case with HLFIR that may appear as a result of CSE. The `adaptor.getSource()` is an operation producing a tuple, which does not have users, so `allOtherUsesAreSafeForAssociate` just looks at the empty list of users. So we get completely wrong answers from it. This causes problems with the following `eraseAllUsesInDestroys` that tries to remove the `DestroyOp` twice during both `hflir.associate` processing. But we also cannot use `associate.getSource()` *efficiently*, because the original users may still hang around: one example is the original body of hlfir.elemental (see D154715), another example is other already converted AssociateOp's that are pending removal in the rewriter (that is why we have a temporary created for each hlfir.associate in the newly added LIT case). This patch just fixes the correctness issue. I think we have to separate the buffer reuse analysis from the conversion itself. I also tried to address the issues with the cloned bodies of `hlfir.elemental`, but this should not matter since D155778: if `hlfir.associate` is inside `hlfir.elemental`, it will end up inside a do-loop body region, so the early exit added in D155778 will prevent the buffer reuse. Reviewed By: tblah Differential Revision: https://reviews.llvm.org/D158471
-
Luke Lau authored
At some point a merge operand was added to the binary vl ops, so this combine was using the mask for the VL. This causes a crash when trying to select the vmv_v_x_vl, which showed up locally when messing about with selectVSplat, but thankfully in ToT the vmv_v_x_vl gets pattern matched away into the .vx and .vi operands every time, so there's no noticeable change. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D158634
-
Hao Jin authored
The option fuse-ld is not visible in Flang. Flang reports "Unknown argument: '-fuse-ld'" during link stage. Reviewed By: awarzynski, kiranchandramohan Differential Revision: https://reviews.llvm.org/D158430
-
Valentin Clement authored
Introduce the acc.set operation that models the acc set directive. Based on acc.init and acc.shutdown Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158554
-
Mark de Wever authored
This is a nicer way to suppress the diagnostic instead of using the pre-processor work-around. Reviewed By: ChuanqiXu Differential Revision: https://reviews.llvm.org/D158523
-
Felipe de Azevedo Piovezan authored
With D149881, we converted EntryValue MachineFunction table entries into `DbgVariables` initialized by a "DbgValue" intrinsic, which can only handle a single, non-fragment DIExpression. However, it is desirable to handle variables with multiple fragments and DIExpressions. To do this, we expand the `DbgVariable` class to handle the EntryValue case. This class can already operate under three different "modes" (stack slot, unchanging location described by a dbg value, changing location described by a loc list). A fourth case is added as a separate class entirely, but a subsequent patch should redesign `DbgVariable` with four subclasses in order to make the code more readable. This patch also exposed a bug in the `beginEntryValueExpression` function, which was not initializing the `LocationFlags` properly. Note how the `finalizeEntryValue` function resets that flag. We fix this bug here, as testing this changing in isolation would be tricky. Differential Revision: https://reviews.llvm.org/D158458
-
Felipe de Azevedo Piovezan authored
When we convert an EntryValue dbg.declare into an entry of the MF side table, we currently copy its DIExpression as is, and rely on subsequent layers to "know" that this expression is implicitly indirect. This is bad because it adds an implicit assumption to the IR representation, and requires subsequent layers to know about this assumption. This also limits the reusability of this table: what if, in the future, we want to use this table for dbg.values? This patch changes existing behavior so that the entities converting dbg_declares explicitly add an OP_deref when converting EntryValue dbg.declares. Differential Revision: https://reviews.llvm.org/D158437
-
Kazu Hirata authored
-
Kazu Hirata authored
This patch fixes: flang/lib/Semantics/resolve-directives.cpp:899:29: error: moving a temporary object prevents copy elision [-Werror,-Wpessimizing-move]
-
- Aug 23, 2023
-
-
Fangrui Song authored
Similar to D81116 (AArch64): separate the GISel components for organization purposes and match other targets ({AArch64,M68k,PowerPC,RISCV,X86}/GISel). Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D158489 -
Valentin Clement authored
This patch propagates the acc routine information to the module file so they can be used by the caller. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158541
-
Adrian Kuegel authored
Prefer to use .empty() instead of checking size().
-
Matthias Springer authored
Elementwise arith.select are currently not supported. Emit an error message instead of crashing. This fixes #61707. Differential Revision: https://reviews.llvm.org/D158617
-
LLVM GN Syncbot authored
-
Manna, Soumi authored
This reverts commit 9e150ada.
-
Adrian Kuegel authored
Prefer to use .empty() instead of checking size().
-
Victor Kingi authored
This commit addresses the comment at https://reviews.llvm.org/D158174#inline-1534036 which suggested adding a check for un-inlined remarks printed. Reviewed By: awarzynski Differential Revision: https://reviews.llvm.org/D158599
-
Manna, Soumi authored
Static Analyzer Tool complains about a large function call parameter which is is passed by value in CGBuiltin.cpp file. 1. In CodeGenFunction::EmitSMELdrStr(clang::SVETypeFlags, llvm::SmallVectorImpl<llvm::Value *> &, unsigned int): We are passing parameter TypeFlags of type clang::SVETypeFlags by value. 2. In CodeGenFunction::EmitSMEZero(clang::SVETypeFlags, llvm::SmallVectorImpl<llvm::Value *> &, unsigned int): We are passing parameter TypeFlags of type clang::SVETypeFlags by value. 3. In CodeGenFunction::EmitSMEReadWrite(clang::SVETypeFlags, llvm::SmallVectorImpl<llvm::Value *> &, unsigned int): We are passing parameter TypeFlags of type clang::SVETypeFlags by value. 4. In CodeGenFunction::EmitSMELd1St1(clang::SVETypeFlags, llvm::SmallVectorImpl<llvm::Value *> &, unsigned int): We are passing parameter TypeFlags of type clang::SVETypeFlags by value. I see many places in CGBuiltin.cpp file, we are passing parameter TypeFlags of type clang::SVETypeFlags by reference. clang::SVETypeFlags inherits several other types. This patch passes parameter TypeFlags by reference instead of by value in the function. Reviewed By: tahonermann, sdesmalen Differential Revision: https://reviews.llvm.org/D158522
-
Daniil Dudkin authored
This patch introduces new operations: `irdl.region` and `irdl.regions`. The former lets us to specify characteristics of a region, such as the arguments for the entry block and the number of blocks. The latter accepts all results of the former operations to define the set of the regions for the operation. Example: ``` irdl.dialect @example { irdl.operation @op_with_regions { %r0 = irdl.region %r1 = irdl.region() %v0 = irdl.is i32 %v1 = irdl.is i64 %r2 = irdl.region(%v0, %v1) %r3 = irdl.region with size 3 irdl.regions(%r0, %r1, %r2, %r3) } } ``` The above snippet demonstrates an operation named `@op_with_regions`, which is constrained to have four regions. * Region `%r0` doesn't have any constraints on the arguments or the number of blocks. * Region `%r1` should have an empty set of arguments. * Region `%r2` should have two arguments of types `i32` and `i64`. * Regio... -
Tom Stellard authored
Reviewed By: rnk Differential Revision: https://reviews.llvm.org/D158569
-
Vassil Vassilev authored
This reverts commit 3edd338a due to missing graphviz https://lab.llvm.org/buildbot/#/builders/92/builds/49520
-
Vassil Vassilev authored
This reverts commit eb0e6c31 due to failures in clangd such as https://lab.llvm.org/buildbot/#/builders/57/builds/29377
-