- Feb 01, 2022
-
-
Marek Kurdej authored
Fixes https://github.com/llvm/llvm-project/issues/52772. This patch fixes the formatting of the code: ``` auto aaaaaaaaaaaaaaaaaaaaa = {}; auto b = g([] { return; }); ``` which should be left as is, but before this patch was formatted to: ``` auto aaaaaaaaaaaaaaaaaaaaa = {}; auto b = g([] { return; }); ``` Reviewed By: MyDeveloperDay, HazardyKnusperkeks Differential Revision: https://reviews.llvm.org/D115972
-
Fangrui Song authored
-
Siva Chandra Reddy authored
-
Fangrui Song authored
My x86-64 lld executable is 8KiB smaller.
-
Johannes Doerfert authored
-
Fangrui Song authored
My x86-64 lld executable is 1.1KiB smaller.
-
Marek Kurdej authored
Fixes https://github.com/llvm/llvm-project/issues/34626. Before, the include sorter would break the code: ``` #include <stdio.h> #include <stdint.h> /* long comment */ ``` and change it into: ``` #include <stdint.h> /* long #include <stdio.h> comment */ ``` This commit handles only the most basic case of a single block comment on an include line, but does not try to handle all the possible edge cases with multiple comments. Reviewed By: HazardyKnusperkeks Differential Revision: https://reviews.llvm.org/D118627
-
Johannes Doerfert authored
To make usage easier (compared to the many reachability related AAs), this patch introduces a helper API, `AA::isPotentiallyReachable`, which performs all the necessary steps. It also does the "backwards" reachability (see D106720) as that simplifies the AA a lot (backwards queries were somewhat different from the other query resolvers), and ensures we use cached values in every stage. To test inter-procedural reachability in a reasonable way this patch includes an extension to `AAPointerInfo::forallInterferingWrites`. Basically, we can exclude writes if they cannot reach a load "during the lifetime" of the allocation. That is, we need to go up the call graph to determine reachability until we can determine the allocation would be dead in the caller. This leads to new constant propagations (through memory) in `value-simplify-pointer-info-gpu.ll`. Note: The new code contains plenty debug output to determine how reachability queries are resolved. Parts extracted from D110078. Differential Revision: https://reviews.llvm.org/D118673
-
Johannes Doerfert authored
D106720 introduced features that did not work properly as we could add new queries after a fixpoint was reached and which could not be answered by the information gathered up to the fixpoint alone. As an alternative to D110078, which forced eager computation where we want to continue to be lazy, this patch fixes the problem. QueryAAs are AAs that allow lazy queries during their lifetime. They are never fixed if they have no outstanding dependences and always run as part of the updates in an iteration. To determine if we are done, all query AAs are asked if they received new queries, if not, we only need to consider updated AAs, as before. If new queries are present we go for another iteration. Differential Revision: https://reviews.llvm.org/D118669
-
Johannes Doerfert authored
This test shows how we can use alloca position and kernel+AS information to improve reachability queries and consequently store-load forwarding. The thirst argument passed to the @use function can be determined statically (a constant). The others cannot and are there for verification.
-
Kuter Dinel authored
This patch implement instruction reachability for AAFunctionReachability attribute. It is used to tell if a certain instruction can reach a function transitively. NOTE: I created a new commit based of D106720 and set the author back to Kuter. Other metadata, etc. is wrong. I also addressed the remaining review comments and fixed the unit test. Differential Revision: https://reviews.llvm.org/D106720 -
Johannes Doerfert authored
We missed out on AANoRecurse in the module pass because we had no call graph. With AAFunctionReachability we can simply ask if the function may reach itself. Differential Revision: https://reviews.llvm.org/D110099
-
Johannes Doerfert authored
genericValueTraversal can look through arguments and allow value simplification across function boundaries. In fact, the latter already happened unchecked. With this change we allow the user of genericValueTraversal to opt-out of interprocedural traversal if required. We explicitly look through arguments now which helps to do various things, incl. the propagation of constants into OpenMP parallel regions (on the host).
-
Christian Sigg authored
Following the discussion in D118318, mark `arith.addf/mulf` commutative. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D118600
-
Mogball authored
Part 2 of 3 of unifying the assembly formats of attributes/types and operations.The last patch that introduced attribute/type formats (D111594) factored out the format lexer entirely. This patch factors out most of the format parsers such that the attribute/type and op parsers only need to implement handling for specific elements. Certain things could be factored better (element verification, 'seen' variables) but the primary goal of factoring is so that features can be used across both assembly formats. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D117971
-
Johannes Doerfert authored
This fixes a conceptual problem with our AAIsDead usage which conflated call site liveness with call site return value liveness. Without the fix tests would obviously miscompile as we make genericValueTraversal more powerful (in a follow up). The effects on the tests are mixed but mostly marginal. The most prominent one is the lack of `noreturn` for functions. The reason is that we make entire blocks live at the same time (for time reasons). Now that we actually look at the block liveness, which we need to do, the return instructions are live and will survive. As an example, `noreturn_async.ll` has been modified to retain the `noreturn` even with block granularity. We could address this easily but there is little need in practice.
-
Johannes Doerfert authored
We moved to the edge API a while back, not all uses were adjusted. Edge liveness is more precise.
-
Johannes Doerfert authored
No tests as these were found browsing the code and I'm not sure how to test them properly.
-
Johannes Doerfert authored
-
Johannes Doerfert authored
We have two attributes that can answer readnone queries. While there is a dependence between them, it seems best to not force the users to know what AA to ask. The helpers also allow to check for readonly nicely. Test changes show where we now deduce readnone but haven't before, mostly because we only asked AAMemoryBehavior and not AAMemoryLocation. AANoAlias has not been ported to the new API yet.
-
Johannes Doerfert authored
Since D104432 we can look through memory by analyzing all writes that might interfere with a load. This patch provides some logic to exclude writes that cannot interfere with a location, due to CFG reasoning. We make sure to avoid multi-thread write-read situations properly while we ignore writes that cannot reach a load or writes that will be overwritten before the load is reached. Differential Revision: https://reviews.llvm.org/D106397
-
Johannes Doerfert authored
-
Johannes Doerfert authored
Due to num_threads (probably also other reasons) we cannot assume explicit barriers are always executed by all threads in an aligned fashion. We can optimize them if that property can be proven but that is different.
-
Johannes Doerfert authored
No-sync is a property that we need in more places as complex transformations emerge. To simplify the query we provide an `AA::isNoSyncInst` helper now and expose two existing helpers through the `AANoSync` class.
-
Johannes Doerfert authored
The namespaces made it more complicate to implement static helpers, among other things. We should not need them at all.
-
Johannes Doerfert authored
Patch originally by Giorgis Georgakoudis (@ggeorgakoudis), typos and bugs introduced later by me. This patch allows us to remove redundant barriers if they are part of a "consecutive" pair of barriers in a basic block with no impacted memory effect (read or write) in-between them. Memory accesses to local (=thread private) or constant memory are allowed to appear. Technically we could also allow any other memory that is not used to share information between threads, e.g., the result of a malloc that is also not captured. However, it will be easier to do more reasoning once the code is put into an AA. That will also allow us to look through phis/selects reasonably. At that point we should also deal with calls, barriers in different blocks, and other complexities. Differential Revision: https://reviews.llvm.org/D118002
-
Johannes Doerfert authored
We used to remove noinline from known OpenMP runtime functions (which are declared in OMPKinds.td). Now we remove noinline from all functions with the proper prefixes: __kmpc, _ZN4_OMP (= namespace omp), omp_
-
Amir Ayupov authored
Only enable --emit-relocs linker option for merge-fdata target if tests are enabled. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D118580
-
Siva Chandra Reddy authored
Reviewed By: michaelrj Differential Revision: https://reviews.llvm.org/D118641
-
Jez Ng authored
Reviewed By: keith Differential Revision: https://reviews.llvm.org/D118646
-
Serguei Katkov authored
PointerToBase is a mapping between potentially derived pointer to its base. As soon as we are in SSA form if there is a base of derived pointer and it is available at def of derived pointer, the same base will be available at any point where derived pointer is alive. So the mapping of derived pointer to base pointer is not a property of a call site but the same on function level. Reviewers: reames, yrouban Reviewed By: reames Subscribers: llvm-commits Differential Revision: https://reviews.llvm.org/D118604
-
Joseph Huber authored
Some of the new driver tests are flaky on AMDGPU, remove for now.
-
Joseph Huber authored
This patch adds a new target to the tests to run using the new driver as the method for generating offloading code. Depends on D116541 Differential Revision: https://reviews.llvm.org/D118637
-
Joseph Huber authored
The LTO support for OpenMP offloading allows us to run the OpenMPOpt pass during the LTO pipeline. This patch introduces an early run of the Module pass and a late run of the CGSCC pass. These are quick no-ops if there is no OpenMP in the module. Depends on D118198 Differential Revision: https://reviews.llvm.org/D118611
-
Joseph Huber authored
Summary: This patch removes the system call to the `clang-offload-wrapper` tool by replicating its functionality in a new file. This improves performance and makes the future wrapping functionality easier to change. Differential Revision: https://reviews.llvm.org/D118198
-
Joseph Huber authored
Summary: This patch replaces the system call to the `llc` binary with a library call to the target machine interface. This should be faster than relying on an external system call to compile the final wrapper binary. Differential Revision: https://reviews.llvm.org/D118197
-
Joseph Huber authored
Summary: Various changes and cleanup for the Linker Wrapper tool.
-
Joseph Huber authored
Summary: This parses the executable name out of the linker arguments so we can use it to give more informative temporary file names and so we don't accidentally use it for device linking.
-
Joseph Huber authored
Summary: This patch implements the `-save-temps` flag for the linker wrapper. This allows the user to inspect the intermeditary outpout that the linker wrapper creates.
-
Joseph Huber authored
Summary: Various changes to the linker wrapper, and the bitcode embedding is not done after the optimizations have run rather than after linking is done. This saves time when doing JIT.
-