- Sep 22, 2023
-
-
Jeff Niu authored
Fixes https://github.com/llvm/llvm-project/issues/66402
-
Matthew Devereau authored
This patch separates PNR registers into their own register class instead of sharing a register class with PPR registers. This primarily allows us to return more accurate register classes when applying assembly constraints, but also more protection from supplying an incorrect predicate type to an invalid register operand.
-
michaelrj-google authored
Two major off-by-one errors are fixed in this patch. The first is in float_to_string.h with length_for_num, which wasn't accounting for the implicit leading bit when calculating the length of a number, causing a missing digit on 80 bit float max. The other off-by-one is the ryu_long_double_constants.h (a.k.a the Mega Table) not having any entries for the last POW10_OFFSET in POW10_SPLIT. This was also found on 80 bit float max. Finally, the integer calculation mode was using a slightly too short integer, again on 80 bit float max, not accounting for the mantissa width. All of these are fixed in this patch.
-
Alexey Bataev authored
artificial for better cost estimation. Need to use original source vector type, not the one artificially constructed, based on the number of vectorized scalars. It affect the cost significantly.
-
Andres Villegas authored
Introduce a new virtual class StackTracePrinter and an implementation FormattedStackTracePrinter in preparation of enabling symbolizer markup for linux. This change allows us to implement other behaviour under the same api for StackTracePrinter, for example, MarkupStackTracePrinter. Reason for revert: A missing header file for the sanitizer_symbolizer_markup.cpp files. This was not caught in local builds or pre-merge checks given that to trigger the error, the code has to be compiled for Fuchsia. For this reland I've build for the fuchsia targets as well as linux.
-
Arthur Eubanks authored
The spec doesn't allow splitting these strings and we're seeing compile issues with splitting it. String splitting was enabled for Verilog in https://reviews.llvm.org/D154093.
-
Kristof Beyls authored
MCPlusBuilder::getOrCreateAnnotationIndex(Name) can be called from different threads, for example when making use of ParallelUtilities::runOnEachFunctionWithUniqueAllocId. The race occurs when an Index for a particular annotation Name needs to be created for the first time. For example, this can easily happen when multiple "copies" of an analysis pass run on different BinaryFunctions, and the analysis pass creates a new Annotation Index to be able to store analysis results as annotations. This was found by using the ThreadSanitizer. No regression test was added; I don't think there is good way to write regression tests that verify the absence of data races? --------- Co-authored-by:Amir Ayupov <fads93@gmail.com>
-
Momchil Velikov authored
Generate a few of the relevant tests with `update_llc_test_checks.py` and pre-commit. Makes it easier to spot the differences in D152828. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D157116
-
Momchil Velikov authored
These are marked to be "as cheap as a move". According to publicly available Software Optimization Guides, they have one cycle latency and maximum throughput only on some microarchitectures, only for `LSL` and only for some shift amounts. This patch uses the subtarget feature `FeatureALULSLFast` to determine how cheap the instructions are. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D152827 Change-Id: I8f0d7e79bcf277ebf959719991c29a1bc7829486
-
Louis Dionne authored
-
Louis Dionne authored
-
Momchil Velikov authored
- remove `FeatureCustomCheapAsMoveHandling`: when you have target features affecting `isAsCheapAsAMove` that can be given on command line or passed via attributes, then every sub-target effectively has custom handling - remove special handling of `FMOVD0`/etc: `FVMOV` with an immediate zero operand is never[1] more expensive tha an `FMOV` with a register operand. - remove special handling of `COPY` - copy is trivially as cheap as itself - make the function default to the `MachineInstr` attribute `isAsCheapAsAMove` - remove special handling of `ANDWrr`/etc and of `ANDWri`/etc: the fallback `MachineInstr` attribute is already non-zero. - remove special handling of `ADDWri`/`SUBWri`/`ADDXri`/`SUBXri` - there are always[1] one cycle latency with maximum (for the micro-architecture) throughput - check if `MOVi32Imm`/`MOVi64Imm` can be expanded into a "cheap" sequence of instructions There is a little twist with determining whether a MOVi32Imm`/`MOVi64Imm` is "as-cheap-as-a-move". Even if one of these pseudo-instructions needs to be expanded to more than one MOVZ, MOVN, or MOVK instructions, materialisation may be preferrable to allocating a register to hold the constant. For the moment a cutoff at two instructions seems like a reasonable compromise. [1] according to 19 software optimisation manuals Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D154722 -
Kazu Hirata authored
This reverts commit d6f994ac. Several people have reported breakage resulting from this patch: - https://github.com/llvm/llvm-project/issues/65152 - https://github.com/llvm/llvm-project/issues/65205
-
Alexey Bataev authored
Need to change the order of the nodes vectorization to avoid too early insertion of the first node.
-
Ramkumar Ramachandra authored
70de0e ([VP][RISCV] Add vp.fshl/fshr and RISC-V support.) introduced VP_FSHL and VP_FSHR, by using a generic expansion for all targets: the core of this change is in TargetLowering. However, the commit erroneously introduced dead code in RISCVISelLowering. Remove this dead code.
-
Justin Lebar authored
This cmake rule is used by external clients, who may or may not have the LLVM_LIBRARY_OUTPUT_INTDIR variable set. If it is not set, then we pass `-Wl,-rpath-link,` to the compiler. It turns out that gcc and clang interpret this differently. * gcc passes `-rpath-link ""` to the linker, which is what we want. * clang passes `-rpath-link` to the linker. This is not what we want, because then the linker gobbles the next command-line argument, whatever it happens to be, and uses it as the -rpath-link target. Fix this by passing -rpath-link only if we actually have a path we want. -
Joseph Huber authored
Summary: I missed removing this now-unused function in the previous patch. Remove it to clean up the interface.
-
jeanPerier authored
There are currently several places that automatically deallocate allocatble if they are allocated: - INTENT(OUT) allocatable are deallocated on entry in the callee - INTENT(OUT) allocatable are also deallocated on the caller side of BIND(C) function in case the implementation is in C. - Results of function returning allocatable are deallocated after usage. - OPENMP privatized allocatable are deallocated at the end of OPENMP region. Introduce genDeallocateIfAllocated that centralize all this code, except for the function return that use genFreememIfAllocated since finalization is done separately currently. `fir::factory::genFinalization` and `fir::factory::genInlinedDeallocation` are removed and replaced by genFreemem since their name were misleading: finalization was not called. There is a fallout in the tests because previous generated code did not check the allocated status when doing inline deallocation. This was OK since free(null) is guaranteed to be a no-op, but this makes compiler code more complex, is a bit surprising in the generated IR IMHO, and it relied on knowing when genDeallocateBox inserts runtime calls or uses inlined code.
-
Amir Bishara authored
Replace the different reduce operations which is getting a constant tensor as an input argument with a constant tensor. As the arguement of the reduce operation is constant tensor and has only a single user we could calculate the resulted constant tensor in compilation time and replace it with reduced memory tensor This optimization has been implemented for: tosa.reduce_sum tosa.reduce_prod tosa.reduce_any tosa.reduce_all tosa.reduce_max tosa.reduce_min Reviewed By: rsuderman Differential Revision: https://reviews.llvm.org/D154832
-
Dávid Ferenc Szabó authored
A straight forward improvement which can already achieve 2x speed up in some cases like the one here: https://github.com/llvm/llvm-project/issues/65946.
-
Matthias Springer authored
Update outdated documentation and add an example.
-
Ingo Müller authored
This PR adds a new transform op that replaces `memref.alloca`s with `memref.get_global`s to newly inserted `memref.global`s. This is useful, for example, for allocations that should reside in the shared memory of a GPU, which have to be declared as globals.
-
Joseph Huber authored
Summary: This patch removes the `rpc_reset` function. This was previously used to initialize the RPC client on the device by setting up the pointers to communicate with the server. The purpose of this was to make it easier to initialize the device for testing. However, this prevented us from enforcing an invariant that the buffers are all read-only from the client side. The expected way to initialize the server is now to copy it from the host runtime. This will allow us to maintain that the RPC client is in the constant address space on the GPU, potentially through inference, and improving caching behaviour.
-
Ramkumar Ramachandra authored
There are several typos in fround.ll, persumably caused by copy-pasting, where there is a strange nvx5* type. From the surrounding code, it is clear that this was intended to be nvx4*. Fix these typos.
-
Matthias Springer authored
* "init" operands are specified with `MutableOperandRange` (which gives access to the underlying `OpOperand *`). No more magic numbers. * Remove most interface methods and make them helper functions. Only `getInitsMutable` should be implemented. * Provide separate helper functions for accessing mutable/immutable operands (`OpOperand`/`Value`, in line with #66515): `getInitsMutable` and `getInits` (same naming convention as auto-generated op accessors). `getInputOperands` was not renamed because this function cannot return a `MutableOperandRange` (because the operands are not necessarily consecutive). `OpOperandVector` is no longer needed. * The new `getDpsInits`/`getDpsInitsMutable` is more efficient than the old `getDpsInitOperands` because no `SmallVector` is created. The new functions return a range of operands. * Fix a bug in `getDpsInputOperands`: out-of-bounds operands were potentially returned.
-
- Sep 21, 2023
-
-
Louis Dionne authored
-
Sirish Pande authored
While simplifying some vector operators in DAG combine, we may need to create new instructions for simplified vectors. At that time, we need to make sure that all the flags of the new instruction are copied/modified from the old instruction. If "contract" is dropped from an instruction like FMUL, it may not generate FMA instruction which would impact performance. Here's an example where "contract" flag is dropped when FMUL is created. Replacing.2 t42: v2f32 = fmul contract t41, t38 With: t48: v2f32 = fmul t38, t38 Co-authored-by:Sirish Pande <sirish.pande@amd.com>
-
Scott Linder authored
Fold constructVariableDIEImpl into constructVariableDIE, simplify it and group related functions. Pull out the previously inline lambdas for visiting the active variant of the DbgVariable to add location and related attributes as an overload set for a private method applyConcreteDbgVariableAttributes. Rename applyVariableAttribute to reflect what kinds of attributes it applies, and to contrast it with the new applyConcreteDbgVariableAttributes. Move constructLabelDIE down in the implementation file, so all of the constructVariableDIE-related function impls are adjacent.
-
JingZe Cui authored
Reviewed By: guraypp Differential Revision: https://reviews.llvm.org/D159535
-
Guillaume Chatelet authored
This is the implementation of step 3 of https://discourse.llvm.org/t/rfc-customizable-namespace-to-allow-testing-the-libc-when-the-system-libc-is-also-llvms-libc/73079.
-
Alexey Bataev authored
Make add() function smart enough to understand that the shuffle of a single entry is requested, if it sees that the second node is the same as the first.
-
Mikhail R. Gadelha authored
This patch updates the siginfo_t struct definition to match the definition from the kernel here: https://github.com/torvalds/linux/blob/master/include/uapi/asm-generic/siginfo.h In particular, there are two main changes: 1. swap position of si_code and si_errno: si_code show come after si_errno in all systems except MIPS. Since we don't MIPS, the order is fixed for now, but can be easily \#ifdef'd if MIPS support is implemented in the future. 2. We add a union of structs that are filled depending on the signal raised. This change was required for the fork and spawn integration tests in rv32, since they fork/clone the running process, call wait/waitid/waitpid, and read the status, which was wrong in rv32 because wait/waitid/waitpid are implemented in rv32 using SYS_waitid. SYS_waitid takes a pointer to a siginfo_t and fills the proper fields in the struct. The previous siginfo_t definition was being incorrectly filled due to not taking into account the signal raised.
-
vic authored
This `Block` member function introduced in 87d77d3c may be misleading to users as the last operation in the block might have not been registered, so there would be no way to ensure that is a terminator. --------- Signed-off-by:
Victor Perez <victor.perez@codeplay.com>
-
Nikita Popov authored
While there also remove some UB from the test.
-
Jake Egan authored
Temporary workaround for the following CMake error: ``` CMake Error in /llvm/libcxx/benchmarks/CMakeLists.txt: The compiler feature "cxx_std_23" is not known to CXX compiler "IBMClang" ```
-
Alexey Bataev authored
scalar. No need to scan the whole graph when trying to find matching node for the scalar, vectorized in several nodes, better to store corresponding nodes along and scan just this small list.
-
Nikita Popov authored
Check the type of the phi node instead (as the comment already indicates).
-
Nikita Popov authored
We can directly use the type of Offset here.
-
Guray Ozen authored
-
Nikita Popov authored
A lot of SCEV expressions only work on integers -- in which case the effective type will always be the same as the type. There is a lot more cleanup to do here.
-