- Aug 27, 2021
-
-
LLVM GN Syncbot authored
-
Jon Chesterfield authored
Lets the amdgpu plugin write to omptarget_device_environment to enable debugging. Intend to use in the near future to record the wavesize that a given deviceRTL was compiled with for running on hardware that supports 32 or 64. Patch sets all the attributes that are useful. Notably .data means the variable is set by writing to host memory before copying to the GPU instead of launching a kernel to update the image. Can simplify the plugin slightly to drop the code for patching after load if this is used consistently. NFC on nvptx, cuda plugin seems to work fine without any annotations. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D108698
-
Alexandre Rames authored
The `HashBuilder` interface allows conveniently building hashes of various data types, without relying on the underlying hasher type to know about hashed data types. Reviewed By: dexonsmith Differential Revision: https://reviews.llvm.org/D106910
-
Alexey Bataev authored
This reverts commit a28234e3 to investigate a compiler crash caused by the commit.
-
Balazs Benics authored
Previously by following the documentation it was not immediately clear what the capabilities of this checker are. In this patch, I add some clarification on when does the checker issue a report and what it's limitations are. I'm also advertising suppressing such reports by adding an assertion, as demonstrated by the test3(). I'm highlighting that this checker might produce an extensive amount of findings, but it might be still useful for code audits. Reviewed By: martong Differential Revision: https://reviews.llvm.org/D107756
-
Kazu Hirata authored
The last use was removed on Jun 17, 2021 in commit bd524955.
-
- Aug 26, 2021
-
-
Simon Wallis authored
A post-commit review comment on https://reviews.llvm.org/D107452 pointed out that https://llvm.org/docs/LangRef.html says: "In a function that uses the constrained intrinsics the strictfp attribute is required on all function calls." Although there are several files across several test directories which don't follow this guidance, it is straightforward to provide this attribute. Reviewed By: kpn Differential Revision: https://reviews.llvm.org/D107567
-
Yonghong Song authored
Generate btf_tag annotations for DISubprograms. The annotations are represented as an DINodeArray in DebugInfo. Differential Revision: https://reviews.llvm.org/D106618
-
Roman Lebedev authored
Even though https://bugs.llvm.org/show_bug.cgi?id=51615 appears to be introduced by D105390, the fix lies here. We can not replace a wide volatile load with a broadcast-from-memory, because that would narrow the load, which isn't legal for volatiles. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D108757
-
Florian Hahn authored
This patch reduces the bitwidth of types certain tests operate and gets rid of a number of @use(i1) calls and xor's the conditions together instead, which eliminates all timeouts when verifying the tests. See https://github.com/AliveToolkit/alive2/issues/744 for more details.
-
Anna Thomas authored
Since LICM has now unconditionally moved to MemorySSA based form, all passes that run in same LPM as LICM need to preserve MemorySSA (i.e. our downstream pipeline). Added loop-mssa to all tests and perform -verify-memoryssa within LoopPredication itself. Differential Revision: https://reviews.llvm.org/D108724
-
Arthur O'Dwyer authored
As requested in D107584. Differential Revision: https://reviews.llvm.org/D108743
-
Yonghong Song authored
Generate btf_tag annotations for DISubprogram types. A field "annotations" is introduced to DISubprogram, and annotations are represented as an DINodeArray, similar to DIComposite elements. The following example illustrates how annotations are encoded in IR: distinct !DISubprogram(..., annotations: !10) !10 = !{!11, !12} !11 = !{!"btf_tag", !"a"} !12 = !{!"btf_tag", !"b"} Differential Revision: https://reviews.llvm.org/D106618 -
Andrew Wei authored
In the combination of addressing modes, when replacing the matched phi nodes, sometimes the phi node to be replaced has been modified. For example, there’s matcher set [A, B] and [C, A], which will have cyclic dependency: A is replaced by B and C will be replaced by A. Because we tried to match new phi node to another new phi node, we should ignore new phi nodes when mapping new phi node to old one. Reviewed By: skatkov Differential Revision: https://reviews.llvm.org/D108635
-
Jacob Bramley authored
Following on from D102353, extend the fpto*i.sat intrinsics to use NEON fcvt* instructions. Differential Revision: https://reviews.llvm.org/D108460
-
Joe Loser authored
Fix typo in `#error` filepath. Differential Revision: https://reviews.llvm.org/D108764
-
Kent Ross authored
Mark the now-done [cmp.result] in spaceship projects as complete; normalize some status markers for papers and projects; fix alignment and line breaks in spaceship projects, add links to standard Differential Revision: https://reviews.llvm.org/D108502
-
Alexey Bataev authored
Reworked reordering algorithm. Originally, the compiler just tried to detect the most common order in the reordarable nodes (loads, stores, extractelements,extractvalues) and then fully rebuilding the graph in the best order. This was not effecient, since it required an extra memory and time for building/rebuilding tree, double the use of the scheduling budget, which could lead to missing vectorization due to exausted scheduling resources. Patch provide 2-way approach for graph reodering problem. At first, all reordering is done in-place, it doe not required tree deleting/rebuilding, it just rotates the scalars/orders/reuses masks in the graph node. The first step (top-to bottom) rotates the whole graph, similarly to the previous implementation. Compiler counts the number of the most used orders of the graph nodes with the same vectorization factor and then rotates the subgraph with the given vectorization factor to the most used order, if it is not empty. Then repeats the same procedure for the subgraphs with the smaller vectorization factor. We can do this because we still need to reshuffle smaller subgraph when buildiong operands for the graph nodes with lasrger vectorization factor, we can rotate just subgraph, not the whole graph. The second step (bottom-to-top) scans through the leaves and tries to detect the users of the leaves which can be reordered. If the leaves can be reorder in the best fashion, they are reordered and their user too. It allows to remove double shuffles to the same ordering of the operands in many cases and just reorder the user operations instead. Plus, it moves the final shuffles closer to the top of the graph and in many cases allows to remove extra shuffle because the same procedure is repeated again and we can again merge some reordering masks and reorder user nodes instead of the operands. Also, patch improves cost model for gathering of loads, which improves x264 benchmark in some cases. Gives about +2% on AVX512 + LTO (more expected for AVX/AVX2) for {625,525}x264, +3% for 508.namd, improves most of other benchmarks. The compile and link time are almost the same, though in some cases it should be better (we're not doing an extra instruction scheduling anymore) + we may vectorize more code for the large basic blocks again because of saving scheduling budget. Differential Revision: https://reviews.llvm.org/D105020 -
Kent Ross authored
These are the only files in libc++ that have CRLF line endings instead of LF. Differential Revision: https://reviews.llvm.org/D108748
-
Simon Pilgrim authored
dyn_cast can return nullptr, use cast<> to assert we have the correct type.
-
Simon Pilgrim authored
-
Balazs Benics authored
This reverts commit 6097a419.
-
Balazs Benics authored
Previously by following the documentation it was not immediately clear what the capabilities of this checker are. In this patch, I add some clarification on when does the checker issue a report and what it's limitations are. I'm also advertising suppressing such reports by adding an assertion, as demonstrated by the test3(). I'm highlighting that this checker might produce an extensive amount of findings, but it might be still useful for code audits. Reviewed By: martong Differential Revision: https://reviews.llvm.org/D107756
-
Andrew Wei authored
SCEVExpander::expandCodeFor may expand add recurrences for loop with a preheader, so we should make LoopDataPrefetch dependent on LoopSimplify. This patch will try to fix : https://bugs.llvm.org/show_bug.cgi?id=43784 Reviewed By: Meinersbur Differential Revision: https://reviews.llvm.org/D108448
-
Jessica Clarke authored
SelectADDRParam was discovered as being dead 5 years ago and removed in 7b4ef068 but the unused ComplexPattern definition was left behind. SelectADDRDWord has never existed as far as I can tell, even back when AMDGPU was R600-only and called that. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D108758
-
Jessica Clarke authored
This was used by CONVERT_RNDSAT, which was removed in def496c0, so the profile is now unused. Reviewed By: xgupta Differential Revision: https://reviews.llvm.org/D108508
-
Andrea Di Biagio authored
It turns out that SchedWrite WriteIMulH was always assigned to the low half of the result of a MULX (rather than to the high half). To avoid confusion, this patch swaps the two MULX writes in the tablegen definition of MULX32/64. That way, write names better describe what they actually refer to; this also avoids further complications if in future we decide to reuse the same MulH writes to also model other scalar integer multiply instructions. I also had to swap the latency values for the two MULX writes to make sure that the change is effectively an NFC. In fact, none of the existing x86 tests were affected by this small refactoring. This patch also fixes a bug in MCA: a wrong latency value was propagated for instructions that perform multiple writes to a same register. This last issue was found by Roman while testing MULX on targets that define a different latency for the Low/High part of the result. Differential Revision: https://reviews.llvm.org/D108727
-
Alex Richardson authored
We have to avoid calling renameat2 and clone on FreeBSD. Additionally, the mcontext structure has different members. Reviewed By: jrtc27, luismarques Differential Revision: https://reviews.llvm.org/D103886
-
Sindhu Chittireddy authored
Klocwork static code analysis exposed this concern: Pointer 'SubExpr' returned from call to getSubExpr() function which may return NULL from 'cast_or_null<Expr>(Operand)', which will be dereferenced in the statement following it Add an assert on SubExpr to make it clear this pointer cannot be null.
-
Matthew Devereau authored
Tell the cost model to use the scalable calculation for non-neon fixed vector. This results in a cheaper cost for fixed-length SVE masked gathers/scatters allowing the vectorizor to emit them more frequently.
-
Benjamin Kramer authored
-
Roman Lebedev authored
In LLVM IR, `AlignmentBitfieldElementT` is 5-bit wide But that means that the maximal alignment exponent is `(1<<5)-2`, which is `30`, not `29`. And indeed, alignment of `1073741824` roundtrips IR serialization-deserialization. While this doesn't seem all that important, this doubles the maximal supported alignment from 512MiB to 1GiB, and there's actually one noticeable use-case for that; On X86, the huge pages can have sizes of 2MiB and 1GiB (!). So while this doesn't add support for truly huge alignments, which i think we can easily-ish do if wanted, i think this adds zero-cost support for a not-trivially-dismissable case. I don't believe we need any upgrade infrastructure, and since we don't explicitly record the IR version, we don't need to bump one either. As @craig.topper speculates in D108661#2963519, this might be an artificial limit imposed by the original implementation of the `getAlignment()` functions. Differential Revision: https://reviews.llvm.org/D108661
-
Benjamin Kramer authored
These may not exist when CET isn't available.
-
Alex Richardson authored
This avoids references to the variables be generated when using e.g. max(). Reviewed By: dexonsmith Differential Revision: https://reviews.llvm.org/D95050
-
Alex Richardson authored
The existing code attempting to bitcast from a value in the default globals AS to i8 addrspace(0)* was triggering an assertion failure in our downstream fork. I found this while compiling poppler for CHERI-RISC-V (we use AS200 for all globals). The test case uses AMDGPU since that is one of the in-tree targets with a non-zero default globals address space. The new test previously triggered a "Invalid constantexpr bitcast!" assertion and now correctly generates code with addrspace(1) pointers. Reviewed By: rjmccall Differential Revision: https://reviews.llvm.org/D105972
-
Alex Richardson authored
We should be using #if instead of #ifdef here since LLVM_ENABLE_THREADS is set using #cmakedefine01 so is always defined. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D108110
-
Florian Hahn authored
This patch adds initial support to use facts from @llvm.assume calls. It intentionally does not handle all possible cases to keep things simple initially. For now, the condition from an assume is made available on entry to the containing block, if the assume is guaranteed to execute. Otherwise it is only made available in the successor blocks.
-
Florian Hahn authored
-
David Green authored
Like other similar instructions the xtn2 family do not have side effects, and explicitly marking them as such can help improve scheduling freedom.
-
David Green authored
-