- Aug 17, 2023
-
-
Aart Bik authored
This keeps the definitions closer together. Also removed some verbose comments for readability. Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D158109
-
Rahul Kayaith authored
Older python versions (e.g. 3.8) don't accept `tuple[...]` etc. in type hints.
-
Justin Bogner authored
When emitting assembly we don't particularly want the binary DXIL embedded in the output. This was mostly there for testing purposes, so we update those tests to run the test directly using `opt` and restrict the -dxil-embed and -dxil-globals passes to running normally only in the case where we're trying to emit a DXContainer. Differential Revision: https://reviews.llvm.org/D158051
-
Valentin Clement authored
Not all declare clause have an exit operation attach to them and therefore no dealloc function generated. Attach the pre/post deallocation attribute only for the clauses that have an exit operation. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158106
-
Valentin Clement authored
Lowering was missing to generate the pre/post alloc/dealloc functions for the acc declare variables. This patch adds the generation. These functions have the descriptor as their unique argument. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158103
-
Chris Bieneman authored
The pipeline state data captured in the PSV0 section of the DXContainer file encodes signature elements which are read by the runtime to map inputs and outputs from the GPU program. This change adds support for generating and parsing signature elements with testing driven through the ObjectYAML tooling. Reviewed By: bogner Differential Revision: https://reviews.llvm.org/D157671 Initially landed as 8c567e64, and reverted in 4d800633. ../llvm/include/llvm/BinaryFormat/DXContainerConstants.def ../llvm/test/ObjectYAML/DXContainer/PSVv1-amplification.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv1-compute.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv1-domain.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv1-geometry.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv1-vertex.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv2-amplification.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv2-compute.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv2-domain.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv2-geometry.yaml ../llvm/test/ObjectYAML/DXContainer/PSVv2-vertex.yaml
-
Jim Ingham authored
The TestEvents.py test I added for ShadowListeners fails on Windows. Since there's no reason to believe the ShadowListeners feature has different behavior from the other event-based tests here, I copied the skips & expected_flakey's from the other tests in that file to this one.
-
Lei Zhang authored
This commit starts enabling vector distruction over multiple dimensions. It requires delinearize the lane ID to match the expected rank. shape_cast and transfer_read now can properly handle multiple dimensions. Reviewed By: hanchung Differential Revision: https://reviews.llvm.org/D157931
-
Craig Topper authored
-
Chris Bieneman authored
This reverts commit 8c567e64.
-
Chris Bieneman authored
The pipeline state data captured in the PSV0 section of the DXContainer file encodes signature elements which are read by the runtime to map inputs and outputs from the GPU program. This change adds support for generating and parsing signature elements with testing driven through the ObjectYAML tooling. Reviewed By: bogner Differential Revision: https://reviews.llvm.org/D157671
-
Blue Gaston authored
Before refactoring this code, all arm64 were set to use the 32bit allocator. This patch reverts back that behavior for DriverKit. Because we target DriverKit as the target OS, rather than a specific platform, reverting back to the previous behavior is preferred to fix a failure we are seeing on embedded platforms. Though it may be more correct in the future to match the allocator to the platform being used. rdar://113649286 Differential Revision: https://reviews.llvm.org/D158028
-
Valentin Clement authored
Lower clauses to the routine info op. Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158007
-
Daniel Hoekwater authored
Machine function splitting will become available for AArch64; since MFS is no longer X86-only, the tests for generic behavior should live somewhere other than tests/CodeGen/X86. MFS implementation doesn't vary much across platforms, and most tests should be identical between X86 and AArch64 besides instruction selection, so the tests can live together in tests/CodeGen/Generic. Differential Revision: https://reviews.llvm.org/D157563
-
Valentin Clement authored
The wrong suffix was applied Reviewed By: razvanlupusoru Differential Revision: https://reviews.llvm.org/D158098
-
Hanhan Wang authored
We are able to fuse the pack op only if inner tiles are not tiled or they are fully used. Otherwise, it could generate a sequence of non-trivial ops. Differential Revision: https://reviews.llvm.org/D157932
-
Owen Pan authored
Fixes #63795. Differential Revision: https://reviews.llvm.org/D157568
-
Matt Arsenault authored
Drop unnecessary flags and metadata, add contract flags that should be necessary.
-
Jim Ingham authored
Before the addition of the process "Shadow Listener" you could only have one Listener observing the Process Broadcaster. That was necessary because fetching the Process event is what switches the public process state, and for the execution control logic to be manageable you needed to keep other listeners from causing this to happen before the main process control engine was ready. Ismail added the notion of a "ShadowListener" - which allowed you ONE extra process listener. This patch inverts that setup by designating the first listener as primary - and giving it priority in fetching events. Differential Revision: https://reviews.llvm.org/D157556
-
LLVM GN Syncbot authored
-
Nico Weber authored
-
Dhruv Chawla authored
[NFC][ValueTracking] Remove calls to computeKnownBits for non-intrinsic CallInsts in isKnownNonZeroFromOperator For non-intrinsic CallInsts, computeKnownBits only handles range metadata and checking getReturnedArgOperand(). Both of these are now handled in isKnownNonZero, so there is no need to fall through to a call to computeKnownBits anymore. Differential Revision: https://reviews.llvm.org/D158095
-
Kazushi (Jam) Marukawa authored
Avoid vectorizing store and load instructions in scalar mode. Reviewed By: efocht Differential Revision: https://reviews.llvm.org/D158049
-
V Donaldson authored
Generate a runtime error message for a reference to an invalid assigned format such as: if (.true.) print n end
-
Craig Topper authored
-
Kazu Hirata authored
This patch fixes: llvm/lib/Analysis/LoopAccessAnalysis.cpp:2001:12: error: unused variable 'MinDepDistBytesOld' [-Werror,-Wunused-variable]
-
V Donaldson authored
Accept "module procedure" (as well as module function/subroutine) in a separate module procedure definition, such as "bb1" in: module mm interface module subroutine mm1 end subroutine end interface end module submodule(mm) bb interface module subroutine bb1 end subroutine end interface contains module procedure mm1 call bb1 end procedure module procedure bb1 print*, 'bb1' end procedure end submodule use mm call mm1 end -
Michael Maitland authored
`MaxSafeDepDistBytes` was not correct based on its name an semantics in instances when there was a non-unit stride loop. For example, ``` for (int k = 0; k < len; k+=3) { a[k] = a[k+4]; a[k+2] = a[k+6]; } ``` Here, the smallest dependence distance is 24 bytes, but only vectorizing 8 bytes is safe. `MaxSafeVectorWidthInBits` reported the correct number of bits that could be vectorized as 64 bits. The semantics of of `MaxSafeDepDistBytes` should be: The smallest dependence distance in bytes in the loop. This may not be the same as the maximum number of bytes that are safe to operate on simultaneously. The name of this variable should reflect those semantics and its docstring should be updated accordingly, `MinDepDistBytes`. A debug message that used `MaxSafeDepDistBytes` to signify to the user how many bytes could be accessed in parallel is updated to use `MaxSafeVectorWidthInBits` instead. That way, the same message if communicated to the user, just in different units. This patch makes sure that when `MinDepDistBytes` is modified in a way that should impact `MaxSafeVectorWidthInBits`, that we update the latter accordingly. This patch also clarifies why `MaxSafeVectorWidthInBits` does not to be updated when `MinDepDistBytes` is (i.e. in the case of a forward dependency). Differential Revision: https://reviews.llvm.org/D156158 -
Nicholas Guy authored
Aligning functions yields small performance gains on embedded cores, moreso with numerous small function calls. Similar to aligning loops, if the function can fit within a single cache line then the performance overhead of fetching more instructions can be limited. Differential Revision: https://reviews.llvm.org/D157514
-
Ingo Müller authored
Fix forward bug in dac19b45, which uses the vertical bar operator for type hints, which is only supported by Python 3.10 and later, and thus breaks the builds on Python 3.8.
-
Siu Chi Chan authored
Change-Id: If4a830fdacf1b0e7b7634f48f648427d5ec7ea21 Reviewed By: kazu, arsenm Differential Revision: https://reviews.llvm.org/D158013
-
Jonas Devlieghere authored
Print an error message with instructions on how to install sphinx_automodapi. Differential revision: https://reviews.llvm.org/D158022
-
David Green authored
This allows us to select G_SHUFFLE_VECTOR with identity masks (possibly including undef elements), but avoid the actual EXT instruction if the shift amount is 0.
-
- Aug 16, 2023
-
-
Dhruv Chawla authored
This check is redundant as it is covered by the call to isPotentiallyReachable. Depends on D155726. Differential Revision: https://reviews.llvm.org/D155718
-
Dhruv Chawla authored
Differential Revision: https://reviews.llvm.org/D155726
-
Akash Banerjee authored
Migrate createForStaticInitFunction, createDispatchInitFunction, createDispatchNextFunction and createDispatchFiniFunction from Clang CodeGen to OMPIRBuilder. Differential Revision: https://reviews.llvm.org/D157994
-
Joseph Huber authored
This test hangs on AMDGPU sporadically, disable it for the time being. Fixes: https://github.com/llvm/llvm-project/issues/64733 Reviewed By: ronlieb Differential Revision: https://reviews.llvm.org/D158082
-
Benjamin Maxwell authored
This patch prevents `mlir-linalg-ods-yaml-gen` from adding extra whitespace around the summary and description fields. This broke the _italics_ of the summary as _ this _ is not recognised by markdown. It also meant the first line of the description was in a code block as it was indented two spaces. The separator between summary and description has also been updated to two newlines. This was already followed and prevents line wrapping the summary putting part of it in the description. These issues can be currently seen at: https://mlir.llvm.org/docs/Dialects/Linalg/ Reviewed By: awarzynski Differential Revision: https://reviews.llvm.org/D157853
-
Ingo Müller authored
I had forgotten to commit that test as part of https://reviews.llvm.org/D157638. Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D158074
-
Ingo Müller authored
Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D157735
-