- Jan 12, 2023
-
-
Siu Chi Chan authored
- Add support for finding device libraries in new ROCm directory structure - Simplify and remove the handling of legacy ROCm directory structure Change-Id: I04da3bc9da85ced4b56b0225efb6b94448b8c5a1 Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D140315
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
David Green authored
With glue nodes between the CMP and the CSINC, generating multiple uses of the CMP can fail when scheduling the DAG. This limits the fold from 90f24bef to a single use on the CSINC, to prevent the CMP being needed by two nodes.
-
Mark de Wever authored
D140653 has the same fix, without the extra tests. Fixes PR59763 Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D140819
-
OCHyams authored
This copies existing behaviour from other debug intrinsics to `dbg.assign`s. Reviewed By: scott.linder Differential Revision: https://reviews.llvm.org/D141140
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
Joseph Huber authored
Currently, the behaviour of `-save-temps` changes the generated output when offloading to AMDGPU. This is because we only have a single phase and it contains the `-disable-llvm-passes` flags which results in unoptimized bitcode. We need to make sure we generate another phase that produces both the optimized and unoptimized bitcode. There used to be a check that turned these phases into a no-op. But I believe it is more correct to not generate them this way in the first place. Doing this requires a bit of a hack, replacing an already generated phase action, but it should be fine. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D141440
-
Joseph Huber authored
This patch adds support for '--offload-arch=native' to OpenMP offloading. This will automatically generate the toolchains required to fulfil whatever GPUs the user has installed. Getting this to work requires a bit of a hack. The problem is that we need the ToolChain to launch its searching program. But we do not yet have that ToolChain built. I had to temporarily make the ToolChain and also add some logic to ignore regular warnings & errors. Depends on D141078 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D141105
-
Joseph Huber authored
This patch applies the same handling for the `--offload-arch=native' string to the new driver. The support for OpenMP will require some extra logic to infer the triples from the derived architecture strings. Depends on D141051 Reviewed By: tra Differential Revision: https://reviews.llvm.org/D141078
-
Joseph Huber authored
This patch adds basic support for `--offload-arch=native` to CUDA. This is done using the `nvptx-arch` tool that was introduced previously. Some of the logic for handling executing these tools was factored into a common helper as well. This patch does not add support for OpenMP or the "new" driver. That will be done later. Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D141051
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
bixia1 authored
Reviewed By: Peiming Differential Revision: https://reviews.llvm.org/D141335
-
Anshil Gandhi authored
As long as the memcpy occurs on a phi input (rather than the phi output), we can look through phi nodes in isOnlyCopiedFromConstantMemory(). This is split out of D136201, to only handle the case where the address spaces are the same, and no pointer rewrite is necessary.
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
Nikita Popov authored
This fold currently performs an unbounded recursive use walk. Make sure that we don't visit too many instructions (the limit is chosen arbitrarily). This is with an eye on also handling phi nodes, which will further extend the considered use graph.
-
Nikita Popov authored
I don't think this matters right now (because InstCombine cleans up unreachable code early), but this will help to make sure that we don't infinite loop once we handle phi nodes. The added test is an example where this would happen.
-
Matt Arsenault authored
This reverts commit 4f575620. I realized the test wasn't very good and when fixed, shows the reduction doesn't work correctly. Revert the change and keep the fixed version of the test.
-
Guillaume Chatelet authored
This change is one of a series to implement the discussion from https://reviews.llvm.org/D141134.
-
- Jan 11, 2023
-
-
David Green authored
Invalid tail predicated loops could be formed by treating function arguments as FalseLanesZero due to getGlobalReachingDefs not returning any values. Make sure we check that the list of Defs is empty and if so treat it like a unknown value. Differential Revision: https://reviews.llvm.org/D141399
-
OCHyams authored
And link to the AssignmentTracking.md document which goes into more detail. Reviewed By: jryans Differential Revision: https://reviews.llvm.org/D141131
-
Anshil Gandhi authored
Tests for D136201.
-
Markus Böck authored
See https://discourse.llvm.org/t/psa-new-improved-fold-method-signature-has-landed-please-update-your-downstream-projects/67618 for context This simply ports all dialects in flang to use the new fold API. These were relatively little and basically just a function signature change, since in-tree folds did not make use of any of the constant operands values. Differential Revision: https://reviews.llvm.org/D141488
-
Louis Dionne authored
-
Philip Reames authored
This is a follow up to D141317 which extends the common code to include a target independent pseudo instruction. This is an alternative to (subset of) D92842 which tries to be as close to NFC as possible. A couple things to call out. * The test change in X86 is because we loose the scheduling information on the instruction. However, I think this was actually a bug in x86 since no instruction was emitted for a MEMBARRIER. Concluding that a meta instruction has latency just seems wrong? * I intentionally left some parts of D92842 out. Specifically, several of the changes in the X86 code (data independence and outlining) appear functional, and likely worthy of their own review. Additionally, I'm not handling ARM/AArch64 at all. Those targets need the ordering whereas none of the others do. I want to get this in and tested before retrofitting in ordering to support those targets. Differential Revision: https://reviews.llvm.org/D141408
-
Markus Böck authored
This is the dialect in-tree with the most `fold` method implementations by far. This patch simply changes all implementations to make use of the new signature. Admittedly, the code readability does not get a lot better in this case, simply due to most methods making use of `constFoldBinaryOp`. I did not modify that function or its interface as part of this patch, but might be something to consider in the future. Differential Revision: https://reviews.llvm.org/D141490
-
Markus Böck authored
These are the trivial cases which do not require any other code changes. Changing the default might not have any semantic changes but at least guarantees that no new fold methods may be added to these dialects while migrating. This commit is also revertible at the end of the migration Differential Revision: https://reviews.llvm.org/D141489
-
Louis Dionne authored
First, use __builtin_unreachable unconditionally. It is implemented by all the compilers that we support. Clang started supporting it around Clang 4, and GCC around GCC 4.10. Also add _LIBCPP_ASSERT so that we will actually get a guaranteed crash if we reached `std::unreachable()` and assertions have been enabled, since that's UB that's extremely easy to catch. Differential Revision: https://reviews.llvm.org/D131620
-
Sam McCall authored
This aligns with the default behavior of clang-tidy (which we offer no way to override). Fixes https://github.com/clangd/clangd/issues/1448 Differential Revision: https://reviews.llvm.org/D141495
-
Florian Hahn authored
Add extra tests for interleaving heuristics for different AArch64 CPUs.
-
NAKAMURA Takumi authored
-
NAKAMURA Takumi authored
-
Peixin Qiao authored
This supports the lowering of intrinsic IS_CONTIGUOUS for array argument. The argument of assumed rank is not supported since it is not implemented yet as the procedure argument. Add TODO for it. Reviewed By: PeteSteinfeld, jeanPerier Differential Revision: https://reviews.llvm.org/D141212
-
Simon Pilgrim authored
For the "all_of(setcc(x,y,eq)) -> PMOVMSKB(PCMPEQB())" fold, we failed to ensure that we could safely bitcast to <X x i8>, which in particular failed with boolean types Thanks to @lerno for catching this and providing the test case
-
Paul Walker authored
Instcombine prefers this canonical form (see getPreferredVectorIndex), as does IRBuilder when passing the index as an integer so we may as well use the prefered form from creation. NOTE: All test changes are mechanical with nothing else expected beyond a change of index type from i32 to i64. Differential Revision: https://reviews.llvm.org/D140983
-
NAKAMURA Takumi authored
It has been introduced since llvmorg-16-init-16838-gac1ffd3c
-
Dinar Temirbulatov authored
If both sides of AND operations are i1 splat_vectors or PTRUE node then we can produce just i1 splat_vector as the result. Differential Revision: https://reviews.llvm.org/D141043
-
Nikita Popov authored
Adjust the GEPs to be non-trivial, to preserve test intent.
-
Nikita Popov authored
Check lines were regenerated for these. The alignment changes in byval-2. look suspicious at first glance, but actually only propagate pre-existing UB.
-
Matt Arsenault authored
The current reduction logic tries to reproduce what a serial reduction would produce, and just takes the first one that is still interesting. We still have to wait for all others to complete though, which at that point is just a waste. This helps speed things up with long running reducers, which I frequently have. e.g. for the added sleep test on my system, it took about 8 seconds before this change and about 4 after. https://reviews.llvm.org/D138953
-
Sam McCall authored
HeaderSearch uses FileEntry::getName() to determine the best spelling of a header. FileEntry::getName() is now the name of the *last* retrieved ref. This means that when FileManager::getFile() hits an existing inode through a new path, it changes the spelling of that header. In the absence of explicit logic to track the preferred name(s) of header files, we should avoid gratuitously calling getFile() with paths different than how the header was originally included, such as the result of realpath(). The originally-specified path should be fine here: - if the same filemanager is being used for record/analysis, we'll hit the filename cache - if a different filemanager is being used e.g. preamble scenario, we should get the same result unless either the working directory has changed (which it shouldn't, else many other things will fail) or the file has gone/changed inode (in which case the old method doesn't work either) Needless to say this is fragile, but talking to @kadircet offline, it's good enough for our purposes for now. Differential Revision: https://reviews.llvm.org/D141478
-