- Jan 12, 2023
-
-
Jakub Kuderski authored
Fixes: https://github.com/llvm/llvm-project/issues/59938 Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D141524
-
Florian Hahn authored
Those CPUs can benefit from additional interleaving. Reviewed By: jroelofs Differential Revision: https://reviews.llvm.org/D141499
-
Ed Maste authored
This matches the configuration used by the prospective FreeBSD CI runner. Remove the (unused) ABI list created in my local environment. This reverts commit eca9196d. Reviewed by: Mordante Differential Revision: https://reviews.llvm.org/D141496
-
Aaron Ballman authored
It seems that changing Format.h is insufficient to get the docs to rebuild? Changing the .rst file directly to at least see if that gets the bot back to green finally.
-
Slava Zakharin authored
These are experimental changes in Flang AA to provide at least some means to disambiguate memory accesses in some simple cases. This AA is still not used by any transformation, so the LIT tests are the only way to trigger it currently. I will further look into applying this AA within Flang to address some of the known performance issues in the benchmarks. Credits to @Renaud-K for the initial implementation. Differential Revision: https://reviews.llvm.org/D141410
-
Thurston Dang authored
This fills in a gap in the tsan_shadow_test coverage: it is possible that the meta regions are aliased (e.g., the heap meta region overlaps the high app meta region). Indeed, the Aarch64_39 mapping has been silently broken in this way for quite some time. This CL checks whether the individual meta regions (for low/mid/high/heap) overlap. Note that (!kBrokenAliasedMetas && !kBrokenLinearity) implies that MemToMeta is invertible; we cannot directly test MetaToMem because that function does not exist. Differential Revision: https://reviews.llvm.org/D141445
-
Aaron Ballman authored
This is another attempt at getting the Sphinx build back to green. It seems Sphinx on the build server does not like any of these, likely due to the digit separators.
-
Craig Topper authored
Use isPhysical/isVirtual methods.
-
Jay Foad authored
Enabling this feature exposed some incorrect codegen, where a workgroup- scope barrier fails to properly synchronise two waves from the same workgroup running on different SIMDs of the same CU. Disabling FeatureBackOffBarrier causes an s_waitcnt to be emitted before the barrier which works around the problem. Differential Revision: https://reviews.llvm.org/D141379
-
Florian Hahn authored
-
Terry Wilmarth authored
When a team nested inside a teams construct is allocated, it is allocated to a size specified by the teams thread_limit. In the case where any mechanism that might not grant the full thread_limit is in use, we may get a smaller team. This possibility was not reflected in the code when using the th_teams_size.nth value stored on the master thread for the team. This value was never updated even when t_nproc on the team itself was different. I added a line to update it shortly before the team is forked. Added a simple teams test that uses KMP_DYNAMIC_MODE=random to mimic allocating teams with sizes <= thread_limit. Eventually, this will segfault without the fix in this commit. Differential Revision: https://reviews.llvm.org/D139960
-
Nico Weber authored
-
Markus Böck authored
See https://discourse.llvm.org/t/psa-new-improved-fold-method-signature-has-landed-please-update-your-downstream-projects/67618 for context Similar to the patch for the arith dialect, the math dialects fold implementations make heavy use of generic fold functions, hence the change being comparatively mechanical and mostly changing the function signature. Differential Revision: https://reviews.llvm.org/D141500
-
Joseph Huber authored
Summary: This capture isn't used, get rid of it and change the name since it's more generic now.
-
Sam McCall authored
(Previous commit assumed it was always off, which is the default)
-
Joseph Huber authored
Summary: This test uses command line tools but I forgot to add the requires clauses. Fix it.
-
bixia1 authored
Previously, we generate AOS subviews for indices buffers when constructing an immutable sparse tensor descriptor. We now only generate such subviews when getIdxMemRefOrView is requested. Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D141325
-
Siu Chi Chan authored
- Add support for finding device libraries in new ROCm directory structure - Simplify and remove the handling of legacy ROCm directory structure Change-Id: I04da3bc9da85ced4b56b0225efb6b94448b8c5a1 Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D140315
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
David Green authored
With glue nodes between the CMP and the CSINC, generating multiple uses of the CMP can fail when scheduling the DAG. This limits the fold from 90f24bef to a single use on the CSINC, to prevent the CMP being needed by two nodes.
-
Mark de Wever authored
D140653 has the same fix, without the extra tests. Fixes PR59763 Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D140819
-
OCHyams authored
This copies existing behaviour from other debug intrinsics to `dbg.assign`s. Reviewed By: scott.linder Differential Revision: https://reviews.llvm.org/D141140
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
Joseph Huber authored
Currently, the behaviour of `-save-temps` changes the generated output when offloading to AMDGPU. This is because we only have a single phase and it contains the `-disable-llvm-passes` flags which results in unoptimized bitcode. We need to make sure we generate another phase that produces both the optimized and unoptimized bitcode. There used to be a check that turned these phases into a no-op. But I believe it is more correct to not generate them this way in the first place. Doing this requires a bit of a hack, replacing an already generated phase action, but it should be fine. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D141440
-
Joseph Huber authored
This patch adds support for '--offload-arch=native' to OpenMP offloading. This will automatically generate the toolchains required to fulfil whatever GPUs the user has installed. Getting this to work requires a bit of a hack. The problem is that we need the ToolChain to launch its searching program. But we do not yet have that ToolChain built. I had to temporarily make the ToolChain and also add some logic to ignore regular warnings & errors. Depends on D141078 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D141105
-
Joseph Huber authored
This patch applies the same handling for the `--offload-arch=native' string to the new driver. The support for OpenMP will require some extra logic to infer the triples from the derived architecture strings. Depends on D141051 Reviewed By: tra Differential Revision: https://reviews.llvm.org/D141078
-
Joseph Huber authored
This patch adds basic support for `--offload-arch=native` to CUDA. This is done using the `nvptx-arch` tool that was introduced previously. Some of the logic for handling executing these tools was factored into a common helper as well. This patch does not add support for OpenMP or the "new" driver. That will be done later. Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D141051
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
bixia1 authored
Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D141335
-
Anshil Gandhi authored
As long as the memcpy occurs on a phi input (rather than the phi output), we can look through phi nodes in isOnlyCopiedFromConstantMemory(). This is split out of D136201, to only handle the case where the address spaces are the same, and no pointer rewrite is necessary.
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
Nikita Popov authored
This fold currently performs an unbounded recursive use walk. Make sure that we don't visit too many instructions (the limit is chosen arbitrarily). This is with an eye on also handling phi nodes, which will further extend the considered use graph.
-
Nikita Popov authored
I don't think this matters right now (because InstCombine cleans up unreachable code early), but this will help to make sure that we don't infinite loop once we handle phi nodes. The added test is an example where this would happen.
-
Matt Arsenault authored
This reverts commit 4f575620. I realized the test wasn't very good and when fixed, shows the reduction doesn't work correctly. Revert the change and keep the fixed version of the test.
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
- Jan 11, 2023
-
-
David Green authored
Invalid tail predicated loops could be formed by treating function arguments as FalseLanesZero due to getGlobalReachingDefs not returning any values. Make sure we check that the list of Defs is empty and if so treat it like a unknown value. Differential Revision: https://reviews.llvm.org/D141399
-
OCHyams authored
And link to the AssignmentTracking.md document which goes into more detail. Reviewed By: jryans Differential Revision: https://reviews.llvm.org/D141131
-
Anshil Gandhi authored
Tests for D136201.
-
Markus Böck authored
See https://discourse.llvm.org/t/psa-new-improved-fold-method-signature-has-landed-please-update-your-downstream-projects/67618 for context This simply ports all dialects in flang to use the new fold API. These were relatively little and basically just a function signature change, since in-tree folds did not make use of any of the constant operands values. Differential Revision: https://reviews.llvm.org/D141488
-
Louis Dionne authored
-