- Oct 25, 2022
-
-
Alexander Belyaev authored
Differential Revision: https://reviews.llvm.org/D136688
-
gbreynoo authored
We were seeing an intermittent local test failure of utils\lit\tests\test-output.py in which the elapsed time output was being given in scientific notation. Python automatically represents small floating-point values in scientific notation so I have altered these tests regex to capture output in that format. Differential Revision: https://reviews.llvm.org/D136469
-
zhongyunde authored
Relax the constraint of wide/vectors types. Address the comment https://reviews.llvm.org/D136015?id=469189#inline-1314520 Reviewed By: spatel, chfast Differential Revision: https://reviews.llvm.org/D136661
-
Emmmer authored
After https://reviews.llvm.org/D135670 Reviewed By: DavidSpickett Differential Revision: https://reviews.llvm.org/D136674
-
Simon Pilgrim authored
[X86] combineConcatVectorOps - fold v4i64/v8x32 concat(broadcast(),broadcast()) -> permilps(concat()) Extend the existing v4f64 fold to handle v4i64/v8f32/v8i32 as well Fixes #58585
-
Paul Robinson authored
The test was added in D114846 but missed one place to introduce the 'httplib' feature keyword, so it has been UNSUPPORTED everywhere. Differential Revision: https://reviews.llvm.org/D136613
-
LLVM GN Syncbot authored
-
LLVM GN Syncbot authored
-
Juan Manuel MARTINEZ CAAMAÑO authored
getMergedLocation returns a 'line 0' DILocaiton if the two locations being merged don't perfecly match, even if they are in the same line but a different column. This commit adds support to keep the line number if it matches (but only the column differs). The merged column number is the leftmost between the two. Reviewed By: dblaikie, orlando Differential Revision: https://reviews.llvm.org/D135166
-
Jay Foad authored
Also avoid using AMDGPU::NoRegister when it's not neeeded.
-
Simon Pilgrim authored
Turns out we fail to do this for concat_v4i64(broadcast_v2i64,broadcast_v2i64) as well
-
David Green authored
This alters the 8.3 complex intrinsics to be target-gated, as opposed to hidden behind preprocessor macros. This is the last of arm_neon.h, and follows the same formula as before. Differential Revision: https://reviews.llvm.org/D135647
-
Simon Pilgrim authored
-
Johannes de Fine Licht authored
Use the `getReturnTypes()` API (which returns an `ArrayRef<Type>`) rather than the `getReturnType()` API (which returns a `Type`) to avoid returning a dangling reference in `LLVMFuncOp::getCallableResults()`. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D136669
-
Simon Pilgrim authored
If the inner broadcast scalar type is smaller/same width as the outer broadcast scalar type then we can broadcast using the same inner type directly. Works for vbroadcast_load as well.
-
Tiezhu Yang authored
Add as little code as possible to allow compiling lldb on LoongArch. Actual functionality will be implemented later. Reviewed By: SixWeining, DavidSpickett Differential Revision: https://reviews.llvm.org/D136578
-
Nico Weber authored
This reverts commit 04877284. Looks like this is still breaking the test Profile-x86_64 :: instrprof-darwin-dead-strip.c (see comment on https://reviews.llvm.org/D135340).
-
Michael Buch authored
-
Emmmer authored
Add: - RV64C instructions sets. - corresponding unittests. - `c.break` code for lldb and lldb-server Fix: - wrong decoding of imm in `DecodeSType` Reviewed By: DavidSpickett Differential Revision: https://reviews.llvm.org/D136362
-
Guillaume Chatelet authored
The new framework makes it explicit which processor feature is being used and allows for easier per platform customization: - ARM cpu now uses trivial implementations to reduce code size. - Memcmp, Bcmp and Memmove have been optimized for x86 - Bcmp has been optimized for aarch64. This is a reland of https://reviews.llvm.org/D135134 (b3f1d58a) Differential Revision: https://reviews.llvm.org/D136595
-
Daniel Grumberg authored
Adds a `--extract-api-ignores=` command line option that allows users to provide a file containing a new line separated list of symbols to unconditionally ignore when extracting API information. Differential Revision: https://reviews.llvm.org/D136450
-
David Sherwood authored
This patch adds the assembly/disassembly for the following instructions: SDOT : Signed integer 2-way dot product indexed and non-indexed UDOT : Unsigned integer 2-way dot product, indexed and non-indexed The reference can be found here: https://developer.arm.com/documentation/ddi0602/2022-09 Differential Revision: https://reviews.llvm.org/D136464
-
David Spickett authored
Without these it was rendering as one line for all three commands.
-
Sjoerd Meijer authored
This adds AArch64 TargetParser support to define CPU aliases, and ports the definition of Grace over to that. This is following up on D136425. Differential Revision: https://reviews.llvm.org/D136611
-
chenglin.bi authored
-
Adrian Kuegel authored
-
Cullen Rhodes authored
If OP in PTEST(PG, OP(PG, ...)) has a flag-setting variant change the opcode so the PTEST becomes redundant. This patch extends this existing optimization in AArch64::optimizePTestInstr to cover all flag-setting opcodes. Reviewed By: peterwaller-arm Differential Revision: https://reviews.llvm.org/D136083
-
Cullen Rhodes authored
A follow on patch will extend existing PTEST(PG, OP(PG, ...)) -> OP_FLAG_SETTING(PG, ...) optimization in AArch64InstrInfo::optimizePTestInstr to cover more of the flag-setting instructions Reviewed By: peterwaller-arm Differential Revision: https://reviews.llvm.org/D136161
-
Thomas Symalla authored
Nested BFI instruction with multiple uses.
-
David Sherwood authored
This patch adds the assembly/disassembly for the following instructions: BFMLSLB : BFloat16 floating-point multiply-subtract long from single-precision (bottom) BFMLSLT : BFloat16 floating-point multiply-subtract long from single-precision (top) Both the vector and indexed forms are added for each. The reference can be found here: https://developer.arm.com/documentation/ddi0602/2022-09 Differential Revision: https://reviews.llvm.org/D136439 -
Caroline Concatto authored
This patch adds the assembly/disassembly for the following instruction: FP: FMLA (multiple and indexed vector): Multi-vector floating-point fused multiply-add by indexed element. FMLS(multiple and indexed vector): Multi-vector floating-point fused multiply-subtract by indexed element. BFDOT (multiple and indexed vector): Multi-vector BFloat16 floating-point dot-product by indexed element. FDOT (multiple and indexed vector): Multi-vector half-precision floating-point dot-product by indexed element. BFVDOT: Multi-vector BFloat16 floating-point vertical dot-product by indexed element. FVDOT: Multi-vector half-precision floating-point vertical dot-product by indexed element. INT: SDOT (2-way, multiple and indexed vector): Multi-vector signed integer dot-product by indexed element. (4-way, multiple and indexed vector): Multi-vector signed integer dot-product by indexed element. SUDOT (multiple and indexed vector): Multi-vector signed by unsigned integer dot-product by indexed elements. SUVDOT: Multi-vector signed by unsigned integer vertical dot-product by indexed element. UDOT (2-way, multiple and indexed vector): Multi-vector unsigned integer dot-product by indexed element. (4-way, multiple and indexed vector): Multi-vector unsigned integer dot-product by indexed element. USDOT (multiple and indexed vector): Multi-vector unsigned by signed integer dot-product by indexed element. USVDOT: Multi-vector unsigned by signed integer vertical dot-product by indexed element. For the multi-vec ternary indexed with 2 and 4 ZA single-vectors for 32 and 64 bits according to the instruction The reference can be found here: https://developer.arm.com/documentation/ddi0602/2022-09 Depends on:D135563 Differential Revision: https://reviews.llvm.org/D135676 -
Jay Foad authored
Verify the LiveVariables analysis after a pass that claims to preserve it, even if there are no further passes (apart from the verifier itself) that would use the analysis. Differential Revision: https://reviews.llvm.org/D129213
-
Jay Foad authored
Following on from D129634, this patch fixes more X86 CodeGen test failures with D129213 applied, which adds verification of LiveIntervals after the TwoAddressInstruction pass runs. These failures only showed up with LLVM_ENABLE_EXPENSIVE_CHECKS=ON which adds the equivalent of an implicit -verify-machineinstrs on all tests. Differential Revision: https://reviews.llvm.org/D136596
-
Sander de Smalen authored
The Chain wasn't set correctly in the DAG for functions marked with aarch64_pstate_sm_body, which meant that SelectionDAG would dead-code some of the CopyToReg's. This didn't show up in the existing tests because all uses were in the same block, but when adding some control-flow, suddenly things would break. Reviewed By: kmclaughlin Differential Revision: https://reviews.llvm.org/D136579
-
Caroline Concatto authored
This patch adds the assembly/disassembly for the following instruction: Int: SCLAMP:Multi-vector signed clamp to minimum/maximum vector. UCLAMP:Multi-vector unsigned clamp to minimum/maximum vector. FP: FCLAMP: Multi-vector floating-point clamp to minimum/maximum number. The reference can be found here: https://developer.arm.com/documentation/ddi0602/2022-09 Depends on: D135563 Differential Revision: https://reviews.llvm.org/D135601 -
Tobias Hieta authored
-
Caroline Concatto authored
This patch adds the assembly/disassembly for the following instruction: INT: SMAX (multiple and single vector): Multi-vector signed maximum by vector. (multiple vectors): Multi-vector signed maximum. SMIN (multiple and single vector): Multi-vector signed minimum by vector. (multiple vectors): Multi-vector signed minimum. UMAX (multiple and single vector): Multi-vector unsigned maximum by vector. (multiple vectors): Multi-vector unsigned maximum. UMIN (multiple and single vector): Multi-vector unsigned minimum by vector. (multiple vectors): Multi-vector unsigned minimum. SRSHL (multiple and single vector): Multi-vector signed rounding shift left by vector. (multiple vectors): Multi-vector signed rounding shift left. URSHL (multiple and single vector): Multi-vector unsigned rounding shift left by vector. (multiple vectors): Multi-vector unsigned rounding shift left. FP: FMAX (multiple and single vector): Multi-vector floating-point maximum by vector. (multiple vectors): Multi-vector floating-point maximum. FMAXNM (multiple and single vector): Multi-vector floating-point maximum number by vector. (multiple vectors): Multi-vector floating-point maximum number. FMIN (multiple and single vector): Multi-vector floating-point minimum by vector. (multiple vectors): Multi-vector floating-point minimum. FMINNM (multiple and single vector): Multi-vector floating-point minimum number by vector. (multiple vectors): Multi-vector floating-point minimum number. The reference can be found here: https://developer.arm.com/documentation/ddi0602/2022-09 It also updates ADD and SQDMULH Depends on: D135563 Differential Revision: https://reviews.llvm.org/D135599 -
David Green authored
As a continuation of D132034, this switches the QRDMX v8.1a neon intrinsics over from preprocessor defines to be target-gated. As there is no "rdma" or "qrdmx" target feature, they use the "v8.1a" architecture feature directly. This works well for AArch64, but something needs to be done for Arm at the same time, as they both use the same header and tablegen emitter. This patch opts for adding "v8.1a" and all dependant target features to the Arm TargetParser, similar to what was recently done for AArch64 but through initFeatureMap when the Architecture is parsed. I attempted to make the code similar to the AArch64 backend. Otherwise this is similar to the changes made in D132034. Differential Revision: https://reviews.llvm.org/D135615
-
Tobias Hieta authored
In our code-base we auto-generate pragma regions the regions look like method signatures like: `#pragma region MYREGION(Foo: bar)` The problem here was that the rest of the line after region was not marked as stringliteral as in the case of pragma mark so clang-format tried to change the formatting based on the method signature. Added test and mark it similar as pragma mark. Reviewed By: owenpan Differential Revision: https://reviews.llvm.org/D136336
-
Kadir Cetinkaya authored
Summary: This brings IncludeCleaner's reference discovery from AST to the parity with current implementation in clangd. Some highlights: - Handling of MemberExprs, only the member declaration is marked as referenced and not the container, unlike clangd. - Constructor calls, only the constructor and not the container, unlike clangd. - All the possible candidates for unresolved overloads, same as clangd. - All the shadow decls for using-decls, same as clangd. - Declarations for definitions of enums with an underlying type and functions, same as clangd. - Using typelocs, using templatenames and typedefs only reference the found decl, same as clangd. - Template specializations only reference the primary template, not the explicit specializations, to be fixed. - Expr types aren't marked as used, unlike clangd. Going forward, we can consider having signals to indicate type of a reference (e.g. `implicit` signal for type of an expr) so that the applications can perform a filtering based on their needs. At the moment the biggest discrepancy is around type of exprs, i.e. not marking containers for member/constructor accesses. I believe this is the right model since the declaration of the member and the container should be available in a single file (modulo macros). Reviewers: sammccall Subscribers: Differential Revision: https://reviews.llvm.org/D132110
-