- Mar 01, 2023
-
-
Sander de Smalen authored
This work follows on from D142109 and addresses a possible regression when we know the loop iteration counter cannot overflow. When we know the overflow-check always evaluates to false, it's better to use the other style of tail folding where it assumes a runtime check was added, because that avoids having to calculate a modified trip-count. Reviewed By: paulwalker-arm Differential Revision: https://reviews.llvm.org/D142894
-
LLVM GN Syncbot authored
-
Sanjay Patel authored
This avoids the danger shown in issue #60906. There were no regression tests for these patterns, so these potential failures have been around for a long time. We freeze the condition and preserve the optimization because getting rid of a div/rem is always a win. Here are a couple of examples that can be corrected by freezing the condition: https://alive2.llvm.org/ce/z/sXHTTC Differential Revision: https://reviews.llvm.org/D144671
-
David Green authored
A LDR will implicitly zero the rest of the vector, so vector_insert(zeros, load, 0) can use a single load. This adds tablegen patterns for both scaled and unscaled loads, detecting where we are inserting a load into the lower element of a zero vector. Differential Revision: https://reviews.llvm.org/D144086
-
Nico Weber authored
-
Wu, Yingcong authored
`LogWriter::Close(LW)` is outside the null check if-else block, which, when `LW == nullptr`, will causing a NULL dereference. I think the close() means to be in else block, which is when `LW != nullptr`. Reviewed By: xgupta Differential Revision: https://reviews.llvm.org/D145039
-
Alex Bradbury authored
Check for size and alignment as we do for other types.
-
David Truby authored
To implement these we call the LLVM intrinsic is.fpclass indicating that we are checking for either a quiet or signalling NaN. Differential Revision: https://reviews.llvm.org/D144649
-
Nicolas Vasilache authored
This revision properly plumbs the subsitution of a padded op through iter_args in the case of an scf::ForOp consumer. Differential Revision: https://reviews.llvm.org/D145036
-
Ben Shi authored
Reviewed By: bcain Differential Revision: https://reviews.llvm.org/D145050
-
Marius Brehler authored
In addition to the component build, this enables the standalone example to be build as part of a monolithic LLVM build by using the LLVM external projects mechanism (`LLVM_EXTERNAL_PROJECTS`). Reviewed By: stephenneuendorffer, stellaraccident Differential Revision: https://reviews.llvm.org/D143718
-
Pavel Labath authored
The error message changed in D144664.
-
Manuel Klimek authored
Pull out common base class for formatting unit tests, removing duplicate code that accumulated over the years. Pull out macro expansion test into its own test file.
-
Ivan Kosarev authored
Reviewed By: foad Differential Revision: https://reviews.llvm.org/D144954
-
Christian Ulmann authored
This commit ensures that the sh script creates temporary files with mktmp to ensure they do not collide with existing files. The previous behaviour caused sporadic permission issues on a multi-user system. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D145054
-
Kiran Chandramohan authored
-
Stephen Tozer authored
The Visual Studio debugger currently uses blocking calls to Go and StepInto, which interferes with Dexter's ability to do any processing (e.g. checking for time outs) in between breakpoints. This patch updates these functions to use non-blocking calls. Reviewed By: Orlando Differential Revision: https://reviews.llvm.org/D144986
-
Alex Bradbury authored
Since D105001, HasFloat16 was unconditionally set to true for RISC-V. This patch adds test coverage for this.
-
Christian Ulmann authored
This commit make the name parameter of the DISubprogramAttr optional. LLVM will for example omit these subprogram names in initialization functions for globals. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D145046
-
Mariya Podchishchaeva authored
The tests can fail if wokring directory where the tests were launched has a `error` substring in its path. Reviewed By: benlangmuir Differential Revision: https://reviews.llvm.org/D144495
-
Sjoerd Meijer authored
We are missing patterns to generate vector splats using LD1R. A shuffle vector with all 0s is a vector splat if the operands are a load and undef for which we can generate a LD1R. Differential Revision: https://reviews.llvm.org/D145004
-
Sjoerd Meijer authored
-
Ingo Müller authored
The previous implementation did not notify the attached listener. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D145049
-
Jean Perier authored
The hlfir fir.box with the local lower bounds and type parameters must be generated conditionally when the entity is optional. Differential Revision: https://reviews.llvm.org/D144962
-
David Green authored
The AArch64 backend, during lowering, will convert an 64bit vector insert to a 128bit vector: vector_insert %dreg, %v, %idx => %qreg = insert_subvector undef, %dreg, 0 %ins = vector_insert %qreg, %v, %idx EXTRACT_SUBREG %ins, dsub This creates a bit of mess in the DAG, and the EXTRACT_SUBREG being a machine nodes makes it difficult to simplify. This patch removes that, treating the 64bit vector insert as legal and handling them with extra tablegen patterns. The end result is a simpler DAG that is easier to write tablegen patterns for. Differential Revision: https://reviews.llvm.org/D144550
-
Zhongyunde authored
Address the dominating condition, the urem fold is benefit from the analytics improvements. Fix https://github.com/llvm/llvm-project/issues/60546 NOTE: delete the calls in simplifyBinaryIntrinsic and foldICmpWithDominatingICmp is used to reduce compile time. Reviewed By: nikic, arsenm, erikdesjardins Differential Revision: https://reviews.llvm.org/D144248
-
Sander de Smalen authored
When using tail-folding and using the predicate for both data and control-flow (the next vector iteration's predicate is generated with the llvm.active.lane.mask intrinsic and then tested for the backedge), the LoopVectorizer still inserts a runtime check to see if the 'i + VF' may at any point overflow for the given trip-count. When it does, it falls back to a scalar epilogue loop. We can get rid of that runtime check in the pre-header and therefore also remove the scalar epilogue loop. This reduces code-size and avoids a runtime check. Consider the following loop: void foo(char * __restrict__ dst, char *src, unsigned long N) { for (unsigned long i=0; i<N; ++i) dst[i] = src[i] + 42; } If 'N' is e.g. ULONG_MAX, and the VF > 1, then the loop iteration counter will overflow when calculating the predicate for the next vector iteration at some point, because LLVM does: vector.ph: %active.lane.mask.entry = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 0, i64 %N) vector.body: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ] %active.lane.mask = phi <vscale x 16 x i1> [ %active.lane.mask.entry, %vector.ph ], [ %active.lane.mask.next, %vector.body ] ... %index.next = add i64 %index, 16 ; The add above may overflow, which would affect the lane mask and control flow. Hence a runtime check is needed. %active.lane.mask.next = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 %index.next, i64 %N) %8 = extractelement <vscale x 16 x i1> %active.lane.mask.next, i64 0 br i1 %8, label %vector.body, label %for.cond.cleanup, !llvm.loop !7 The solution: What we can do instead is calculate the predicate before incrementing the loop iteration counter, such that the llvm.active.lane.mask is calculated from 'i' to 'tripcount > VF ? tripcount - VF : 0', i.e. vector.ph: %active.lane.mask.entry = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 0, i64 %N) %N_minus_VF = select %N > 16 ? %N - 16 : 0 vector.body: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ] %active.lane.mask = phi <vscale x 16 x i1> [ %active.lane.mask.entry, %vector.ph ], [ %active.lane.mask.next, %vector.body ] ... %active.lane.mask.next = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 %index, i64 %N_minus_VF) %index.next = add i64 %index, %4 ; The add above may still overflow, but this time the active.lane.mask is not affected %8 = extractelement <vscale x 16 x i1> %active.lane.mask.next, i64 0 br i1 %8, label %vector.body, label %for.cond.cleanup, !llvm.loop !7 For N = 20, we'd then get: vector.ph: %active.lane.mask.entry = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 0, i64 %N) ; %active.lane.mask.entry = <1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1> %N_minus_VF = select 20 > 16 ? 20 - 16 : 0 ; %N_minus_VF = 4 vector.body: (1st iteration) ... ; using <1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1> as predicate in the loop ... %active.lane.mask.next = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 0, i64 4) ; %active.lane.mask.next = <1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0> %index.next = add i64 0, 16 ; %index.next = 16 %8 = extractelement <vscale x 16 x i1> %active.lane.mask.next, i64 0 ; %8 = 1 br i1 %8, label %vector.body, label %for.cond.cleanup, !llvm.loop !7 ; branch to %vector.body vector.body: (2nd iteration) ... ; using <1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0> as predicate in the loop ... %active.lane.mask.next = tail call <vscale x 16 x i1> @llvm.get.active.lane.mask.nxv16i1.i64(i64 16, i64 4) ; %active.lane.mask.next = <0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0> %index.next = add i64 16, 16 ; %index.next = 32 %8 = extractelement <vscale x 16 x i1> %active.lane.mask.next, i64 0 ; %8 = 0 br i1 %8, label %vector.body, label %for.cond.cleanup, !llvm.loop !7 ; branch to %for.cond.cleanup Reviewed By: fhahn, david-arm Differential Revision: https://reviews.llvm.org/D142109 -
Sander de Smalen authored
-
Nikita Popov authored
Per discussion on D144970, these are no longer necessary.
-
Valentin Clement authored
-
Valentin Clement authored
Update move_alloc to carry over the dyanmic type of `from` to `to` and reset the dynamic type of `from` to its declared type when it is polymorphic. Reviewed By: PeteSteinfeld Differential Revision: https://reviews.llvm.org/D144997
-
Lorenzo Chelini authored
`isScalar` only returns true if the operand is non-shaped. But we need to handle also rank zero tensors. Reviewed By: hanchung Differential Revision: https://reviews.llvm.org/D144989
-
Caroline Concatto authored
To make legalization easier, the operands and outputs have the same size for these ISD Nodes. When legalizing the results in SplitVectorResult the operands are legalized to the same size as the outputs. The ISD Node has two output/results, therefore the legalizing functions update both results/outputs. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D144744
-
Tobias Gysi authored
The revision introduces two interfaces that provide access to the alias analysis and access group metadata attributes. The AliasAnalysis interface combines all alias analysis related attributes (alias, noalias, and tbaa) similar to LLVM's getAAMetadata method, while the AccessGroup interface is dedicated to the access group metadata. Previously, only the load and store operations supported alias analysis and access group metadata. This revision extends this support to the atomic operations. A follow up revision will also add support for the memcopy, memset, and memove intrinsics. The interfaces then provide convenient access to the metadata attributes and eliminate the need of TypeSwitch or string based attribute access. The revision still relies on string based attribute access for the translation to LLVM IR (except for tbaa metadata). Only once the the memory access intrinsics also implement the new interfaces, the translation to LLVM IR can be fully switched to use interface based attribute accesses. Depends on D144875 Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D144851
-
Balázs Kéri authored
During AST import multiple different InjectedClassNameType objects could be created for a single class template. This can cause problems and failed assertions when these types are compared and found to be not the same (because the instance is different and there is no canonical type). The import of this type does not use the factory method in ASTContext, probably because the preconditions are not fulfilled at that state. The fix tries to make the code in ASTImporter work more like the code in ASTContext::getInjectedClassNameType. If a type is stored at the Decl or previous Decl object, it is reused instead of creating a new one. This avoids crash at least a part of the cases. Reviewed By: gamesh411, donat.nagy, vabridgers Differential Revision: https://reviews.llvm.org/D140562
-
Benjamin Chetioui authored
Differential Revision: https://reviews.llvm.org/D144972
-
Sander de Smalen authored
The C and C++ Language Extensions for AArch64 SME2 [1] adds a new type called `svcount_t` which describes a predicate. This is not a predicate vector mask, but rather a description of a predicate vector mask that can be expanded into a mask using explicit instructions. The type is a scalable opaque type. To implement `svcount_t` type this patch uses the existing Target Extension Type mechanism, but adds further support so that this type can be a scalable type. AArch64 CodeGen support will follow in a separate patch. [1] https://github.com/ARM-software/acle/pull/217 Reviewed By: jcranmer-intel, nikic Differential Revision: https://reviews.llvm.org/D136861
-
Christian Ulmann authored
This commit ensures that the LLVMIR export prioritizes existing DILocalScope attribute information as location scopes over files constructed from filenames. All DILocalScope attributes contain file information, so no information is lost. The previous implementation caused the introduction of superfluous DILexicalBlockFile nodes in certain cases. The old implementation remains as a fallback when no DILocalScope is present. Reviewed By: gysit Differential Revision: https://reviews.llvm.org/D144968
-
Ben Shi authored
Different AVR devices have different data regions. Current clang driver emits a default '-Tdata' option to the linker. This way works fine if there is no user specified linker script, but it will cause conflicts if there is one. A better solution for setting the default data region to GNU ld is defining symbol __DATA_REGION_ORIGIN__, which is expected by GNU ld's default AVR linker script. Fixes https://github.com/llvm/llvm-project/issues/60362 Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D144533
-
Emilio Cobos Alvarez authored
Let the branch fall through the error path like other functions here do. Differential Revision: https://reviews.llvm.org/D140074
-