- Feb 07, 2023
-
-
Simon Pilgrim authored
As mentioned on https://discourse.llvm.org/t/issues-in-llvm-tblgen-high-parallelized-build/68037, ItaniumManglingCanonicalizer is often slow to build, resulting in a bottleneck for distributed builds while waiting for LLVMSupport to complete. SymbolRemappingReader is the only current user of ItaniumManglingCanonicalizer, and this is only used by ProfileData and llvm-cxxmap - so I propose we move both files into the ProfileData library. Differential Revision: https://reviews.llvm.org/D143318
-
Fangrui Song authored
Driver::getToolChain called by Driver::BuildCompilation gets the `Triple` argument from a temporary. With delayed detection due to LazyDetector, we would reference a dangling `Triple`.
-
Noah Goldstein authored
Several cases where missing. 1. `(icmp eq/ne X*Z, Y*Z) [if Z % 2 != 0] -> (icmp eq/ne X, Y)` EQ: https://alive2.llvm.org/ce/z/6_HPZ5 NE: https://alive2.llvm.org/ce/z/c34qSU There was previously an implementation of this that work of `Y` was non-constant, but it was missing if `Y*Z` evaluated to a constant and/or `nsw`/`nuw` where both false. As well it only worked if `Z` was a constant but we can check 1s bit of `KnownBits` to cover more cases. 2. `(icmp eq/ne X*Z, Y*Z) [if Z != 0 and nsw(X*Y) and nsw(Y*Z)] -> (icmp eq/ne X, Y)` EQ: https://alive2.llvm.org/ce/z/6SdAG6 NE: https://alive2.llvm.org/ce/z/fjsq_b This was previously implemented only to work if `Z` was constant, but we can use `isKnownNonZero` to cover more cases. 3. `(icmp uPred X*Y, Y*Z) [if Z != 0 and nuw(X*Y) and nuw(X*Y)] -> (icmp uPred X, Y)` EQ: https://alive2.llvm.org/ce/z/FqWQLX NE: https://alive2.llvm.org/ce/z/2gHrd2 ULT: https://alive2.llvm.org/ce/z/MUAWgZ ULE: https://alive2.llvm.org/ce/z/szQQ2L UGT: https://alive2.llvm.org/ce/z/McVUdu UGE: https://alive2.llvm.org/ce/z/95uyC8 This was previously implemented only for `eq/ne` cases. As well only if `Z` was constant, but again we can use `isKnownNonZero` to cover more cases. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D142786 -
Noah Goldstein authored
We previously only did this if the `mul` was `nuw`, but it works for any odd value. Alive2 Links: EQ: https://alive2.llvm.org/ce/z/6_HPZ5 NE: https://alive2.llvm.org/ce/z/c34qSU Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D143026
-
Noah Goldstein authored
Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D142785
-
Noah Goldstein authored
Improve and enable folding of conditional branches with tail calls. 1. Make it so that conditional tail calls can be emitted even when there are multiple predecessors. 2. Don't guard the transformation behind -Os. The rationale for guarding it was static-prediction can be affected by whether the branch is forward of backward. This is no longer true for almost any X86 cpus (anything newer than `SnB`) so is no longer a meaningful concern. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D140931
-
Noah Goldstein authored
If the add/sub is not single use, it will need to be materialized later, in which case using the BMI instruction is a de-optimization in terms of code-size and throughput. i.e: ``` // Good leal -1(%rdi), %eax andl %eax, %eax xorl %eax, %esi ... ``` ``` // Unecessary BMI (lower throughput, larger code size) leal -1(%rdi), %eax blsr %edi, %eax xorl %eax, %esi ... ``` Note, this may cause more `mov` instructions to be emitted sometimes because BMI instructions only have 1 src and write-only to dst. A better approach may be to only avoid BMI for (and/xor X, (add/sub 0/-1, X)) if this is the last use of X but NOT the last use of (add/sub 0/-1, X). Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D141180
-
Noah Goldstein authored
(a & (-b)) & b is often lowered as: %sub = sub i32 0, %b %and0 = and i32 %sub, %a %and1 = and i32 %and0, %b Which won't get detected by the BLSI pattern as b & -b are never in the same SDNode. This patch will do a small search through associative operators and try and place BMI patterns in the same node so they will hit the pattern. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D141179 -
Noah Goldstein authored
Was previously de-optimizating if -march supported lzcnt as there is no reason to add the extra instruction. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D141464
-
Valentin Clement authored
The runtime function expects a 2 x newRank array and the code was passing a newRank x 2 array. This patch updates the creation of the array to fit the runtime expectation. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D143405
-
Ariel Burton authored
Currently when clang deals with a call to a builtin function that is supplied with an argument that has an explicit address space it rewrites the signature of the callee to make the types of the formal parameters match those of the actual arguments. This functionality was added to support OpenCL, and was introduced with commit b919c7d9. However, this does not work properly for "size" related address spaces such as those used for __ptr32. This affects platforms like Microsoft and z/OS. This change preserves the OpenCL functionality, but will use the formal parameter types when an address space is size-related. Reviewed By: akhuang Differential Revision: https://reviews.llvm.org/D142048
-
Haowei Wu authored
This patch adds llvm-mt and llvm-rc to the Clang bootstrap dependency when building the Clang under Windows. Differential Revision: https://reviews.llvm.org/D143025
-
Ron Lieberman authored
breaks amdgpu buildbot This reverts commit 402981ee.
-
Haowei Wu authored
This patch simplified the BOOTSTRAP_ flags, allowing them to be pass through from regular flags. Differential Revision: https://reviews.llvm.org/D143288
-
Joseph Huber authored
Summary: The wrapper bitcode currently only gets a temp file for the compiled object. This makes it more difficult to see what was actually generated.
-
Bjorn Pettersson authored
This reverts commit 525ed98b. Some buildbots are failing when linking bugpoint. Reverting to investigate that further.
-
Bjorn Pettersson authored
There are some helpers in the Lint analysis pass that will setup a pass manager and then run the Lint pass on a given Function/Module. Those have been using the LegacyPassManager, but as a small step towards removing the deprecated legacy pass manager this patch is changing those helpers into using the new pass manager instead. No idea if anyone is really is using those helpers. Maybe an alternative had been to just remove them. There is at least no unit tests or similar that verifies that they work, so I validated this patch by using a hacked opt binary that called those functions before running the normal pipeline. Differential Revision: https://reviews.llvm.org/D143388
-
Bjorn Pettersson authored
This patch is updating TailDuplicator::duplicateInstruction to fix some old bugs that has been found with an out-of-tree target. There are three different things being addressed: 1) In one situation two subregister indices are combined using the composeSubRegIndices helper. But the order in which those indices are combined has been incorrect. For this problem I managed to create some kind of reproducer using AArch64 (see the test case touched in this patch). 2) Another fault was found in the else branch for the above situation. Here we do not compose the two subregisters, instead we insert a COPY to replace the PHI, and then the subreg index in the using MO remains. Thus, the virtual register created for the COPY should always match with the size of the original register. Therefore the optimization that "constrain" (or rather relax) the register class by looking at the instruction desc must be limited to the situation when there is no subregister access. Otherwise we create a vreg with the wrong class. 3) Last problem addressed in this patch is that when a new register class is picked by looking at the instruction desc, then it isn't guaranteed that the isAllocatable property is set for that class. So one need to use the getAllocatableClass helper to find a subclass that is allocatable before using createVirualRegister, or alternatively (as in this patch) just use the OrigRC instead of relaxing the register class for the COPY destination. Haven't been able to find any in-tree reproducers for problem 2 and 3. The tricky part is to find a target that has register hierarchies that match with the problem to trigger those code paths (and with subreg accesses involved). Differential Revision: https://reviews.llvm.org/D140496
-
Bjorn Pettersson authored
Differential Revision: https://reviews.llvm.org/D140495
-
Gulfem Savrun Yeniceri authored
Originally, the following commit removed mapping coverage regions for system headers: https://github.com/llvm/llvm-project/commit/93205af066341a53733046894bd75c72c99566db It might be viable and useful to collect coverage from system headers in some systems. This patch adds --system-headers-coverage option (disabled by default) to enable collecting coverage from system headers. Differential Revision: https://reviews.llvm.org/D143304
-
Florian Hahn authored
This reverts commit 695ce48c. The compile-time regression causing the revert has been fixed. Recommit the original patch. Original commit message: The pass should help to close a functional gap when it comes to reasoning about related conditions in a relatively general way. It addresses multiple existing issues (linked below) and the need for a more powerful reasoning system was also discussed recently in https://discourse.llvm.org/t/rfc-alternative-approach-of-dealing-with-implications-from-comparisons-through-pos-analysis/65601/7 On AArch64, the new pass performs ~2000 simplifications on MultiSource,SPEC2006,SPEC2017 with -O3. Compile-time impact: NewPM-O3: +0.20% NewPM-ReleaseThinLTO: +0.32% NewPM-ReleaseLTO-g: +0.28% https://llvm-compile-time-tracker.com/compare.php?from=f01a3a893c147c1594b9a3fbd817456b209dabbf&to=577688758ef64fb044215ec3e497ea901bb2db28&stat=instructions:u Fixes #49344. Fixes #47888. Fixes #48253. Fixes #49229. Fixes #58074. Reviewed By: asbirlea Differential Revision: https://reviews.llvm.org/D135915
-
Benjamin Maxwell authored
-
Vassil Vassilev authored
In https://reviews.llvm.org/D119036 we fixed some of the infrastructure by removing the textual keyword. The underlying issue of PR50592 was that clang can re-export only submodules but under some conditions we needed to re-export the standalone module std_config via std. This patch provides a better fix to the symptom D119036 fixed. Differential revision: https://reviews.llvm.org/D142805
-
David Green authored
This removes a condition in the detection of AVG nodes, where we needn't be checking the LHS of an add node as any const will be canonicalized to the RHS.
-
David Green authored
This slightly extends the creation of hadd nodes to allow them to be generated with the original type size if wrapping flags allow. https://alive2.llvm.org/ce/z/bPjakD https://alive2.llvm.org/ce/z/fa_gzb Differential Revision: https://reviews.llvm.org/D143371
-
Benjamin Maxwell authored
Previously this would incorrectly return the raw offset into the .debug_addr section for the DW_FORM_addrx1/2/3/4 forms rather than the actual address. Note that this was handled correctly in the dump() function so this issue only occurs for users of this API and not in tools such as llvm-dwarfdump. The dump() method has now been updated to use this method to increase coverage. This also now adds a few unit tests for indexed addresses to DWARFDebugInfoTest. Differential Revision: https://reviews.llvm.org/D143073
-
Joseph Huber authored
Currently, the plan is to support testing on a single GPU architecture. We query the supported architectures from the user's system. However, there are times when the user would want to override this. This patch adds the `LIBC_GPU_TEST_ARCHITECTURE` option, which allows users to specify which GPU architecture to build for. Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D143400
-
Florian Hahn authored
This patch breaks up the solving step into 2 phases: 1. Collect all rows where the variable to eliminate is != 0 and remove it from the original system. 2. Process all collect rows to build new set of constraints, add them to the original system. This is much more efficient for excessive cases, as this avoids a large number of moves to the new system. This reduces the time spent in ConstraintElimination for the test case shared in D135915 from ~3s to 0.6s.
-
Shivam Gupta authored
Incorrect use of shared_ptr. found by PVS-Studio https://pvs-studio.com/en/blog/posts/cpp/1003/, N8 & N9. Differential Revision: https://reviews.llvm.org/D142309
-
Ivan Kosarev authored
v_swap_b32 is a VOP1-only instruction, meaning it neither encodes src1 nor has 64-bit encodings. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D143289
-
Chuanqi Xu authored
Close https://github.com/llvm/llvm-project/issues/60545. Previously, we would only pass the size parameter to the deallocation function if the type is completely the same. But it is good enough to make them unqualified the smae.
-
Dmitry Chernenkov authored
-
Florian Hahn authored
Move some accesses that are use multiple times to variables. This also will make updating them easier in the future.
-
- Feb 06, 2023
-
-
Jay Foad authored
-
Jay Foad authored
-
serge-sans-paille authored
Lazyly initialize uncommon toolchain detector Cuda and rocm toolchain detectors are currently run unconditionally, while their result may not be used at all. Make their initialization lazy so that the discovery code is not run in common cases. Reapplied since 77910ac3 landed and fixes the test ordering issue. Differential Revision: https://reviews.llvm.org/D142606
-
zhijian authored
Summary: since the class 'SymbolicFile ' do not have a is64Bit() API , when we need to check whether a SymbolicFile object is 64bit or not. we need to write a function to do it, it maybe cause duplication code. Reviewers: James Henderson, Fangrui Song Differential Revision: https://reviews.llvm.org/D143097
-
Simon Pilgrim authored
-
Simon Pilgrim authored
AVX1 doesn't benefit as nearly all integer ops will stay as 128-bit ops. This only exposes a couple of minor changes but will be a lot more useful in an upcoming shuffle combining patch.
-
John Brawn authored
Since commit 846b6767 SmallVectorBase<uint32_t> has been explicitly instantiated, which means that clang.exe must export it for a plugin to be able to link against it, but the constructor is not exported as currently no template constructors or destructors are exported. We can't just export all constructors and destructors, as that puts us over the symbol limit on Windows, so instead rewrite how we decide which templates need to be exported to be more precise. Currently we assume that templates instantiated many times have no explicit instantiations, but this isn't necessarily true and results also in exporting implicit template instantiations that we don't need to. Instead check for references to template members, as this indicates that the template must be explicitly instantiated (as if it weren't the template would just be implicitly instantiated on use). Doing this reduces the number of symbols exported from clang from 66011 to 53993 (in the build configuration that I've been testing). It also lets us get rid of the special-case handling of Type::getAs, as its explicit instantiations are now being detected as such. Differential Revision: https://reviews.llvm.org/D142989
-