- Oct 28, 2020
-
-
Kazu Hirata authored
I just removed thread-prob-{1,2}.ll in b2f05fae, so I am removing thread-prob-3.ll to thread-prob-1.ll. -
Kazu Hirata authored
This patch removes extraneous calls to setEdgeProbability introduced in c9148776. The follow-up patch, a7b662d0, has since fixed BranchProbabilityInfo::eraseBlock, so we don't need to worry about getting stale values from getEdgeProbability. Also, since getEdgeProbability(BB, BB->getSingleSuccessor()) returns edge probability 1/1 by default for BB with exactly one successor edge, we don't need to explicitly call setEdgeProbability. This patch introduces almost no functional change, but we do end up reducing debug messages from setEdgeProbability. Differential Revision: https://reviews.llvm.org/D90284
-
Mitch Phillips authored
This reverts commit 5b3bf8b4. This caused a regression in the ASan buildbot. See comments at https://reviews.llvm.org/D89817 for more information.
-
Mitch Phillips authored
This reverts commit a6336eab. This commit broke check-llvm under ASan: See http://lab.llvm.org:8011/#/builders/5/builds/446 for more details.
-
Alok Kumar Sharma authored
For any newly added parse function, clang-tidy complains. New parse functions are implicitly defined by a macro "Parse##CLASS(N, IsDistinct)". Now this macro and exising function definitions are corrected (lower case first character). Some other variable/function names are also corrected to comply LLVM coding style. Reviewed By: djtodoro Differential Revision: https://reviews.llvm.org/D90243
-
Carl Ritson authored
SIPreAllocateWWMRegs was being inserted after RegisterCoalescer but this pass does not exist during FastAlloc so pre-allocation pass was never being run. Insert pre-allocation after TwoAddressInstructionPass instead. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D90236
-
Johannes Doerfert authored
This matches the new PM model.
-
Johannes Doerfert authored
This will simplify debugging and tracking down problems.
-
Johannes Doerfert authored
-
Johannes Doerfert authored
Before we used to only mark unreachable static functions as dead if all uses were known dead. Now we optimistically assume uses to be dead until proven otherwise.
-
Johannes Doerfert authored
If we are looking at a call site argument it might be a load or call which is in a different context than the call site argument. We cannot simply use the call site argument range for the call or load. Bug reported and reduced by Whitney Tsang <whitneyt@ca.ibm.com>.
-
Johannes Doerfert authored
-
Johannes Doerfert authored
The call is not free, unsure if this is needed but it does not make it worse either.
-
Johannes Doerfert authored
In the AANoAlias logic we determine if a pointer may have been captured before a call. We need to look at other uses in the call not uses of the call. The new code is not perfect as it does not allow trivial cases where the call has multiple arguments but it is at least not unsound and a TODO was added.
-
Johannes Doerfert authored
-
John Demme authored
Enhance tblgen's declarative assembly format to allow `attr-dict` in custom directives. Reviewed By: rriddle Differential Revision: https://reviews.llvm.org/D89772
-
Valentin Clement authored
In the OpenACC specification, there are two different self clause. One for the update directive with a var-list argument. This clause is a synonym of the host clause. The second self clause is present for most of the compute construct and takes an optional condition. To solve this ambiguity, the self clause for the update directive is directly translated to a host clause during the parsing. The self clause in AccClause refers always to the compute construct clause. Reviewed By: kiranktp Differential Revision: https://reviews.llvm.org/D90185
-
Derek Schuff authored
This reverts commit bcb8a119.
-
Richard Smith authored
-
Eugene Zhulenev authored
+fix rocm runner Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D90274
-
Wei Wang authored
Summary: Propagate driver commandline remarks options to linker when LTO is enabled. This gives novice user a convenient way to collect and filter remarks throughout a typical toolchain invocation with sample profile and LTO using single switch from the clang driver. A typical use of this option from clang command-line: * Using -Rpass* options to print remarks to screen: clang -fuse-ld=lld -flto=thin -fprofile-sample-use=foo_sample.txt -Rpass=inline -Rpass-missed=inline -Rpass-analysis=inline -fdiagnostics-show-hotness -fdiagnostics-hotness-threshold=100 -o foo foo.cpp Remarks will be dumped to screen from both pre-lto and lto compilation. * Using serialized remarks options clang -fuse-ld=lld -flto=thin -fprofile-sample-use=foo_sample.txt -fsave-optimization-record -fdiagnostics-show-hotness -fdiagnostics-hotness-threshold=100 -o foo foo.cpp This will produce multiple yaml files containing optimization remarks: 1. foo.opt.yaml : remarks from pre-lto 2. foo.opt.ld.yaml.thin.1.yaml: remark during lto Differential Revision: https://reviews.llvm.org/D85810
-
Derek Schuff authored
Since Wasm comdat sections work similarly to ELF, we can use that mechanism to eliminate duplicate dwarf type information in the same way. Differential Revision: https://reviews.llvm.org/D88603
-
Johannes Doerfert authored
If `null_pointer_is_valid` is present, `dereferenceable` does not imply `nonnull`, make it clear. Came up in D17993. Reviewed By: aqjune Differential Revision: https://reviews.llvm.org/D89417
-
Johannes Doerfert authored
Reported by Colleen Bertoni <bertoni@anl.gov> after running the OvO test suite: https://github.com/TApplencourt/OvO/ The template overload is still hidden behind an ifdef for OpenMP. In the future we probably want to remove the ifdef but that requires further testing. Reviewed By: JonChesterfield, tra Differential Revision: https://reviews.llvm.org/D89971
-
Nico Weber authored
-
Siva Chandra Reddy authored
Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D90262
-
Fangrui Song authored
[BranchProbabilityInfo] Make MaxSuccIdx[Src] efficient and add a comment about the subtle eraseBlock. NFC Follow-up to D90272.
-
River Riddle authored
Resolves missed comments in D89103
-
River Riddle authored
At this point, these methods are just carbon copies of OpBuilder::create and aren't necessary given that PatternRewriter inherits from OpBuilder. Differential Revision: https://reviews.llvm.org/D90087
-
River Riddle authored
An InterfaceMap is generated for every single operation type, and is responsible for a large amount of the code size from MLIR given that its internals highly utilize templates. This revision refactors the internal implementation to use bare malloc/free for interface instances as opposed to static variables and moves as much code out of templates as possible. This led to a decrease of over >1mb (~12% of total MLIR related code size) for a downstream MLIR library with a large amount of operations. Differential Revision: https://reviews.llvm.org/D90086
-
River Riddle authored
When compiling for code size, the use of a vtable causes a destructor(and constructor in certain cases) to be generated for the class. Interface models don't need a complex constructor or a destructor, so this can lead to many megabytes of code size increase(even in opt). This revision switches to a simpler struct of function pointers approach that accomplishes the same API requirements as before. This change requires no updates to user code, or any other code aside from the generator, as the user facing API is still exactly the same. Differential Revision: https://reviews.llvm.org/D90085
-
River Riddle authored
[mlir][SIdeEffectInterface][NFC] Move several InterfaceMethods to the extraClassDeclarations instead All InterfaceMethods will have a corresponding entry in the interface model, and by extension have an implementation generated for every operation type. This can result in large binary size increases when a large amount of operations use an interface, such as the side effect interface. Differential Revision: https://reviews.llvm.org/D90084
-
MaheshRavishankar authored
This patch adds support for fusing linalg.indexed_generic op with linalg.tensor_reshape op by expansion, i.e. - linalg.indexed_generic op -> linalg.tensor_reshape op when the latter is expanding. - linalg.tensor_reshape op -> linalg.indexed_generic op when the former is folding. Differential Revision: https://reviews.llvm.org/D90082
-
Kazu Hirata authored
This patch ensures that BranchProbabilityInfo::eraseBlock(BB) deletes all entries in Probs associated with with BB. Without this patch, stale entries for BB may remain in Probs after eraseBlock(BB), leading to a situation where a newly created basic block has an edge probability associated with it even before the pass responsible for creating the basic block adds any edge probability to it. Consider the current implementation of eraseBlock(BB): for (const_succ_iterator I = succ_begin(BB), E = succ_end(BB); I != E; ++I) { auto MapI = Probs.find(std::make_pair(BB, I.getSuccessorIndex())); if (MapI != Probs.end()) Probs.erase(MapI); } Notice that it uses succ_begin(BB) and succ_end(BB), which are based on BB->getTerminator(). This means that if the terminator changes between calls to setEdgeProbability and eraseBlock, then we may not examine all pairs associated with BB. This is exactly what happens in MaybeMergeBasicBlockIntoOnlyPred, which merges basic blocks A into B if A is the sole predecessor of B, and B is the sole successor of A. It replaces the terminator of A with UnreachableInst before (indirectly) calling eraseBlock(A). The patch fixes the problem by keeping track of all edge probablities entered with setEdgeProbability in a map from BasicBlock* to a successor index. Differential Revision: https://reviews.llvm.org/D90272 -
Kazu Hirata authored
This patch teaches the jump threading pass to set edge probabilities whenever the pass creates new basic blocks. Without this patch, the compiler sometimes produces non-deterministic results. The non-determinism comes from the jump threading pass using stale edge probabilities in BranchProbabilityInfo. Specifically, when the jump threading pass creates a new basic block, we don't initialize its outgoing edge probability. Edge probabilities are maintained in: DenseMap<Edge, BranchProbability> Probs; in class BranchProbabilityInfo, where Edge is an ordered pair of BasicBlock * and a successor index declared as: using Edge = std::pair<const BasicBlock *, unsigned>; Probs maps edges to their corresponding probabilities. Now, we rarely remove entries from this map, so if we happen to allocate a new basic block at the same address as a previously deleted basic block with an edge probability assigned, the newly created basic block appears to have an edge probability, albeit a stale one. This patch fixes the problem by explicitly setting edge probabilities whenever the jump threading pass creates new basic blocks. Differential Revision: https://reviews.llvm.org/D90106
-
Sam Clegg authored
This field to represents the amount of static data needed by an dynamic library or executable it should not include things like heap or stack areas, which in the case of `-pie` are not determined until runtime (e.g. __stack_pointer is imported). Differential Revision: https://reviews.llvm.org/D90261
-
Eugene Zhulenev authored
Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D90264
-
Sanjay Patel authored
This was originally part of: f2c25c70 but that was reverted because there was an underlying bug in processing the vector type of these intrinsics. That was fixed with: 74ffc823 This is similar in spirit to 01ea93d8 (memcpy) except that here the underlying caller assumptions were created for vectorizer use (throughput) rather than other passes. That meant targets could have an enormous throughput cost with no corresponding size, latency, or blended cost increase. Paraphrasing from the previous commits: This may not make sense for some callers, but at least now the costs will be consistently wrong instead of mysteriously wrong. Targets should provide better overrides if the current modeling is not accurate.
-
Sanjay Patel authored
-
Martin Storsjö authored
On windows, wchar_t is 16 bit, while we might be widening chars to char32_t. This cast had been present since the initial commit, and removing it doesn't seem to make any tests fail. Differential Revision: https://reviews.llvm.org/D90228
-