- Apr 28, 2021
-
-
Nico Weber authored
Commit 2a133224 extracted this code to a new function checkSectionName() and added a call to it, but didn't remove the original code. The original code is dead since the checkSectionName() early return would fire when it would trigger. (If it weren't dead, it'd make clang crash since err_attribute_section_invalid_for_target now takes two args instead of just the one that's passed.) No behavior change. Differential Revision: https://reviews.llvm.org/D101457
-
Krzysztof Parzyszek authored
Add a call to skipFunction().
-
David Sherwood authored
This patch fixes a crash encountered when vectorising the following loop: void foo(float *dst, float *src, long long n) { for (long long i = 0; i < n; i++) dst[i] = -src[i]; } using scalable vectors. I've added a test to Transforms/LoopVectorize/AArch64/sve-basic-vec.ll as well as cleaned up the other tests in the same file. Differential Revision: https://reviews.llvm.org/D98054 -
Arthur O'Dwyer authored
In particular, `span<int>::iterator` may be a raw pointer type and thus have no nested typedef `iterator::value_type`. However, we already know that the value_type we expect for `span<int>` is just `int`. Fix up all other iterator_concept_conformance tests in the same way. Differential Revision: https://reviews.llvm.org/D101420
-
David Goldman authored
Class properties are always implicit short-hands for the getter/setter class methods. We need to explicitly visit the interface decl `UIColor` in `UIColor.blueColor`, otherwise we instead show the method decl even while hovering over `UIColor` in the expression. Differential Revision: https://reviews.llvm.org/D99975
-
Nico Weber authored
-
Paul C. Anagnostopoulos authored
!find searches a source string for a target string and returns the position. Differential Revision: https://reviews.llvm.org/D101318
-
Tres Popp authored
-
Alexey Bataev authored
If the first tree element is vectorize and the second is gather, it still might be profitable to vectorize it if the gather node contains less scalars to vectorize than the original tree node. It might be profitable to use shuffles. Differential Revision: https://reviews.llvm.org/D101397
-
Roman Lebedev authored
-
Utkarsh Saxena authored
This is useful for running in batch mode. Getting the SymbolID from via getSymbolInfo may give SymbolID of a symbol different from that located by LocateSymbolAt (they have different semantics of choosing the symbol.) Differential Revision: https://reviews.llvm.org/D101388
-
Anton Zabaznov authored
Language options are not available when a target is being created, thus, a new method is introduced. Also, some refactoring is done, such as removing OpenCL feature macros setting from TargetInfo. Reviewed By: Anastasia Differential Revision: https://reviews.llvm.org/D101087
-
-
Matt Arsenault authored
This was picking a concrete size for a physical register, and enforcing exact match on the virtual register's type size. Some targets add multiple types to a register class, and some are smaller than the full bit width. For example x86 adds f32 to 128-bit xmm registers, and AMDGPU adds i16/f16 to 32-bit registers. It might be better to represent these cases as a copy of the full register and an extraction of the subpart, but a lot of code assumes you can directly copy. This will help fix the current usage of the DAG calling convention infrastructure which is incompatible with how GlobalISel is now using it. The API is somewhat cumbersome here, but I just mirrored the existing functions, except now with LLTs (and allow returning null on failure, unlike the MVT version). I think the concept of selecting register classes based on type is flawed to begin with, but I'm trying to keep this compatible with the existing handling.
-
David Sherwood authored
This patch simplifies the calculation of certain costs in getInstructionCost when isScalarAfterVectorization() returns a true value. There are a few places where we multiply a cost by a number N, i.e. unsigned N = isScalarAfterVectorization(I, VF) ? VF.getKnownMinValue() : 1; return N * TTI.getArithmeticInstrCost(... After some investigation it seems that there are only these cases that occur in practice: 1. VF is a scalar, in which case N = 1. 2. VF is a vector. We can only get here if: a) the instruction is a GEP/bitcast/PHI with scalar uses, or b) this is an update to an induction variable that remains scalar. I have changed the code so that N is assumed to always be 1. For GEPs the cost is always 0, since this is calculated later on as part of the load/store cost. PHI nodes are costed separately and were never previously multiplied by VF. For all other cases I have added an assert that none of the users needs scalarising, which didn't fire in any unit tests. Only one test required fixing and I believe the original cost for the scalar add instruction to have been wrong, since only one copy remains after vectorisation. I have also added a new test for the case when a pointer PHI feeds directly into a store that will be scalarised as we were previously never testing it. Differential Revision: https://reviews.llvm.org/D99718
-
Alexey Bataev authored
Need to respect mapping/privatization of declare target variables in the target regions if explicitly specified by the user. Differential Revision: https://reviews.llvm.org/D99530
-
Alexander Belyaev authored
Tensor inputs, if not used in the body of TiledLoopOp, can be removed. memref::CastOp can be folded into TiledLoopOp as well. Differential Revision: https://reviews.llvm.org/D101445
-
Adrian Kuegel authored
So far, only a conversion for complex::AbsOp is done, but more will be added. Differential Revision: https://reviews.llvm.org/D101442
-
Sander de Smalen authored
This patch also refactors the way the feasible max VF is calculated, although this is NFC for fixed-width vectors. After this change scalable VF hints are no longer truncated/clamped to a shorter scalable VF, nor does it drop the 'scalable flag' from the suggested VF to vectorize with a similar VF that is fixed. Instead, the hint is ignored which means the vectorizer is free to find a more suitable VF, using the CostModel to determine the best possible VF. Reviewed By: c-rhodes, fhahn Differential Revision: https://reviews.llvm.org/D98509
-
Alex Richardson authored
Previously printing R_386_RELATIVE relocations would trigger `error: can't read an entry at 0x40: it goes past the end of the section (0x40)` I found this while writing a test case for LLD (D100490). This also includes some minor cleanup in the elf-dynamic-relcos.test llvm-objdump test based on the newly added test. Reviewed By: jhenderson, MaskRay Differential Revision: https://reviews.llvm.org/D100489
-
Alex Richardson authored
The original page no longer works, so use a web.archive.org link instead. Reviewed By: atanasyan Differential Revision: https://reviews.llvm.org/D100949
-
Alex Richardson authored
While implementing support for the float128 routines on x86_64, I noticed that __builtin_isinf() was returning true for 128-bit floating point values that are not infinite when compiling with GCC and using the compiler-rt implementation of the soft-float comparison functions. After stepping through the assembly, I discovered that this was caused by GCC assuming a sign-extended 64-bit -1 result, but our implementation returns an enum (which then has zeroes in the upper bits) and therefore causes the comparison with -1 to fail. Fix this by using a CMP_RESULT typedef and add a static_assert that it matches the GCC soft-float comparison return type when compiling with GCC (GCC has a __libgcc_cmp_return__ mode that can be used for this purpose). Also move the 3 copies of the same code to a shared .inc file. Reviewed By: compnerd Differential Revision: https://reviews.llvm.org/D98205
-
Alex Richardson authored
This has been rather useful in our downstream CHERI target where we want to run tests both with addrspace(0) and addrspace(200) pointers. With this patch we can prefix the opt command with `sed -e 's/addrspace(200)/addrspace(0)/g' -e 's/-A200-P200-G200//g'` to test both cases using the same IR input. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D95137
-
David Spickett authored
'.' is used for unprintable chars (see NON_PRINTABLE_CHAR).
-
Tres Popp authored
This reverts commit 75d6b8bb. The reasoning is mentioned in https://reviews.llvm.org/D97667
-
Roman Lebedev authored
There are post-commit notest for e4c61d5f that suggest the test is failing on certain bots. It looks like the code there isn't being moved, which suggests cost-model involvement, which suggests that we need to hardcode the target triple. Hopefully this helps?
-
Roman Lebedev authored
-
Lorenzo Chelini authored
-
Kerry McLaughlin authored
When using the -enable-strict-reductions flag where UF>1 we generate multiple Phi nodes, though only one of these is used as an input to the vector.reduce.fadd intrinsics. The unused Phi nodes are removed later by instcombine. This patch changes widenPHIInstruction/fixReduction to only generate one Phi, and adds an additional test for unrolling to strict-fadd.ll Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D100570
-
Nathan James authored
Checks if introspection support is available set output kind parser. If it isn't present the auto complete will not suggest `srcloc` and an error query will be reported if a user tries to access it. Reviewed By: steveire Differential Revision: https://reviews.llvm.org/D101365
-
Jingu Kang authored
Prevent cases in which the start value of IV is bigger than bound for increasing. Prevent cases in which the start value of IV is smaller than bound for decreasing. Differential Revision: https://reviews.llvm.org/D101174
-
Benjamin Kramer authored
This lets clang diagnose unused statistics, so remove them.
-
Frederik Gossen authored
As a canonicalization, infer the resulting shape rank if possible. Differential Revision: https://reviews.llvm.org/D101377
-
Hans Wennborg authored
The /QIntel-jcc-erratum flag only works when targeting x86, so pass --target to the driver to do that also on non-x86 hosts.
-
Frederik Gossen authored
Both, `shape.broadcast` and `shape.cstr_broadcastable` accept dynamic and static extent tensors. If their operands are casted, we can use the original value instead. Differential Revision: https://reviews.llvm.org/D101376
-
Qiu Chaofan authored
This patch fixes the infinite loop in legalization of PPC32 SELECT_CC with 64-bit operand.
-
Stephen Tozer authored
This patch fixes a crash in LiveDebugVariables for inputs where a DBG_VALUE_LIST had 64 or more debug operands. This was triggering an assert, which was added under the assumption that only bad CodeGen would result in such a limit being hit, but relatively simple source files that result in these incredibly long debug values have been found, so this assert has been changed to a condition that drops the debug value if it is not met. Differential Revision: https://reviews.llvm.org/D101373
-
Hans Wennborg authored
-
Frederik Gossen authored
Also create all extent tensor constants with const_shape op. Differential Revision: https://reviews.llvm.org/D99197
-