- Feb 11, 2021
-
-
Rob Suderman authored
Added support for broadcasting size-1 dimensions for TOSA elemtnwise operations. Differential Revision: https://reviews.llvm.org/D96190
-
Arthur Eubanks authored
Reviewed By: ychen Differential Revision: https://reviews.llvm.org/D96452
-
Sean Silva authored
Differential Revision: https://reviews.llvm.org/D96391
-
Sean Silva authored
After discussion, it seems like we want to go with "inherent/discardable". These seem to best capture the relationship with the op semantics and don't conflict with other terms. Please let me know your preferences. Some of the other contenders are: ``` "intrinsic" side | "annotation" side -----------------+------------------ characteristic | annotation closed | open definitional | advisory essential | discardable expected | unexpected innate | acquired internal | external intrinsic | extrinsic known | unknown local | global native | foreign inherent | acquired ``` Rationale: - discardable: good. discourages use for stable data. - inherent: good - annotation: redundant and doesn't convey difference - intrinsic: confusable with "compiler intrinsics". - definitional: too much of a mounthful - extrinsic: too exotic of a word and hard to say - acquired: doesn't convey the relationship to the semantics - internal/external: not immediately obvious: what is internal to what? - innate: similar to intrinsic but worse - acquired: we don't typically think of an op as "acquiring" things - known/unknown: by who? - local/global: to what? - native/foreign: to where? - advisory: confusing distinction: is the attribute itself advisory or is the information it provides advisory? - essential: an intrinsic attribute need not be present. - expected: same issue as essential - unexpected: by who/what? - closed/open: whether the set is open or closed doesn't seem essential to the attribute being intrinsic. Also, in theory an op can have an unbounded set of intrinsic attributes (e.g. `arg<N>` for func). - characteristic: unless you have a math background this probably doesn't make as much sense Differential Revision: https://reviews.llvm.org/D96093
-
Dave Lee authored
Follow up to https://reviews.llvm.org/rG483ec136da7193de781a5284f1c37929cc27c05c
-
Nicolas Vasilache authored
-
Nicolas Vasilache authored
-
Jessica Paquette authored
GlobalISel was only doing this with minsize. SDAG does this with optsize. (See: `SelectionDAG::shouldOptForSize()`) This is a 0.3% code size improvement for CTMark at -Os. (Best: 1.1% improvements on lencod + pairlocalalign) Differential Revision: https://reviews.llvm.org/D96451
-
Hongtao Yu authored
-
Benjamin Kramer authored
There's no polymorphic deletion happening here.
-
Fangrui Song authored
-
Arthur Eubanks authored
Reviewed By: reames Differential Revision: https://reviews.llvm.org/D96449
-
Vitaly Buka authored
Depends on D96320. Reviewed By: eugenis Differential Revision: https://reviews.llvm.org/D96328
-
Vitaly Buka authored
Redundant check-prefixes is needed for folloup patches.
-
Jacques Pienaar authored
This reverts commit 5e77ea04. Causes a breakage on Windows buildbot.
-
Dave Lee authored
While learning about ThreadPlan, I did a bit of cleanup: * Remove unused code * Move functions to protected where applicable * Remove virtual for functions that are not overridden Differential Revision: https://reviews.llvm.org/D96277
-
Rong Xu authored
Break SampleProfileLoader into to a base and a derived class. Base class (SampleProfileLoaderBaseImpl) includes the common code for IR and MachineIR (CodeGen) sample loader. It will be templatelized in the later patch. Inline and Probe related code will remain in the derived class of SampleProfileLoader and stays in SampleProfile.cpp. We need to refactor some functions: (1) getInstWeight() to enable the code sharing -- put the core into getInstWeightImpl(). (2) emitAnnotation() and propagateWeights() to carve out the code specific to SampleProfileLoader. (3) make getInstWeight() and findFunctionSamples() virtual and override in SampleProfileLoader as they need to access the fields in the derived class. Differential Revision: https://reviews.llvm.org/D95832
-
Jessica Paquette authored
When we have a G_ADD which is fed by a G_ICMP on one side, we can fold it into the cset for the G_ICMP. e.g. Given ``` %cmp = G_ICMP ... %x, %y %add = G_ADD %cmp, %z ``` We would normally emit a cmp, cset, and add. However, `%add` is either `%z` or `%z + 1`. So, we can just use `%z` as the source of the cset rather than wzr, saving an instruction. This would probably be cleaner in AArch64PostLegalizerLowering, but we'd need to change the way we represent G_ICMP to do that, I think. For now, it's easiest to implement in selection. This is a 0.1% code size improvement on CTMark/pairlocalalign at -Os. Example: https://godbolt.org/z/7KdrP8 Differential Revision: https://reviews.llvm.org/D96388
-
Sam McCall authored
We now (since a while) turn this off centrally in ParsedAST and CodeComplete.
-
Sam McCall authored
This is obsoleted by the standard semanticTokens request family. As well as the protocol details, this allows us to remove a bunch of plumbing around pushing highlights to clients. This should not land until the new protocol has feature parity, see D77702. Differential Revision: https://reviews.llvm.org/D95576
-
Jacques Pienaar authored
If context is enabled/disabled and queried concurrently then this results in a data race/TSAN failure with RunSafely (where boolean variable was not locked). There doesn't seem to be a reasonable way to enable threads that enable and disable recovery in parallel (without also keeping gCrashRecoveryEnabled's lock held during Fn execution which seems undesirable). This makes enable checking if enabled thread local and consistent with other thread local usage of crash context here. Differential Revision: https://reviews.llvm.org/D93907
-
Hongtao Yu authored
The IR/MIR pseudo probe intrinsics don't get materialized into real machine instructions and therefore they don't incur runtime cost directly. However, they come with indirect cost by blocking certain optimizations. Some of the blocking are intentional (such as blocking code merge) for better counts quality while the others are accidental. This change unblocks perf-critical optimizations that do not affect counts quality. They include: 1. IR InstCombine, sinking load operation to shorten lifetimes. 2. MIR LiveRangeShrink, similar to #1 3. MIR TwoAddressInstructionPass, i.e, opeq transform 4. MIR function argument copy elision 5. IR stack protection. (though not perf-critical but nice to have). Reviewed By: wmi Differential Revision: https://reviews.llvm.org/D95982
-
Ilya Tokar authored
Not using builtins doesn't always imply worse code, but for e. g. isinf, this is 30%+ faster. Before: name time/op BM_isinf 2.14ns ± 2% After: name time/op BM_isinf 1.33ns ± 2% Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D88854
-
Adrian Prantl authored
salvageDebugInfoImpl() may fail and return a nullptr.
-
Philip Reames authored
The AssumptionCache mechanism is used to feed assumes into known bits computations. Most places in SCEV passed it in, but one place appears to have been missed. Spotted via inspection, don't have a test case which actually exercises this, but it seemed like an obvious fixit.
-
Sanjay Patel authored
This is a special-case multiply that replicates bits of the source operand. We need this fold to avoid regression if we make canonicalization to `mul` more aggressive for shl+or patterns. I did not see a way to make Alive generalize the bit width condition for even-number-of-bits only, but an example of the proof is: Name: i32 Pre: isPowerOf2(C1 - 1) && log2(C1) == C2 && (C2 * 2 == width(C2)) %m = mul nuw i32 %x, C1 %t = lshr i32 %m, C2 => %t = and i32 %x, C1 - 2 Name: i14 %m = mul nuw i14 %x, 129 %t = lshr i14 %m, 7 => %t = and i14 %x, 127 https://rise4fun.com/Alive/e52
-
Sanjay Patel authored
-
Mehdi Amini authored
Fix StridedMemRefType operator[] SFINAE to allow correctly selecting the `int64_t` overload for non-container operands
-
Pavel Labath authored
Although it is located under tools/lldb-server, this test is very different that other lldb-server tests. The most important distinction is that it does not test lldb-server directly, but rather interacts with it through the lldb client. It also tests the relevant client functionality (the platform connect command, which is even admitted in the test name). The fact that this test is structured as a lldb-server test means it cannot access most of the goodies available to the "normal" lldb tests (the runCmd function, which it reimplements; the run_break_set_by_symbol utility function; etc.). This patch makes it a full-fledged lldb this, and rewrites the relevant bits to make use of the standard features. I also move the test into the "commands" subtree to better reflect its new status.
-
Nawrin Sultana authored
This patch adds lower-bound and upper-bound to num_teams clause according to OpenMP 5.1 specification. The initial number of teams created is implementation defined, but it will be greater than or equal to lower-bound and less than or equal to upper-bound. If num_teams clause is not specified, the number of teams created is implementation defined, but it will be greater or equal to 1. Differential Revision: https://reviews.llvm.org/D95820
-
Jing Pu authored
Make the type contraint consistent with other shape dialect operations. Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D96377
-
Aart Bik authored
This revision connects the generated sparse code with an actual sparse storage scheme, which can be initialized from a test file. Lacking a first-class citizen SparseTensor type (with buffer), the storage is hidden behind an opaque pointer with some "glue" to bring the pointer back to tensor land. Rather than generating sparse setup code for each different annotated tensor (viz. the "pack" methods in TACO), a single "one-size-fits-all" implementation has been added to the runtime support library. Many details and abstractions need to be refined in the future, but this revision allows full end-to-end integration testing and performance benchmarking (with on one end, an annotated Lingalg op and, on the other end, a JIT/AOT executable). Reviewed By: nicolasvasilache, bixia Differential Revision: https://reviews.llvm.org/D95847
-
Christopher Di Bella authored
Implements parts of: - P0898R3 Standard Library Concepts - P1754 Rename concepts to standard_case for C++20, while we still can Differential Revision: https://reviews.llvm.org/D96235 -
Christopher Di Bella authored
Implements parts of: - P0898R3 Standard Library Concepts - P1754 Rename concepts to standard_case for C++20, while we still can Reviewed By: ldionne, #libc Differential Revision: https://reviews.llvm.org/D74292 -
Michael Kruse authored
Test the NewPM as well as the legacy PM.
-
Michael Kruse authored
-
Jameson Nash authored
This attempts to move all tools over to using `add_llvm_library` for better consistency. After doing this, I noticed it ended up as nearly a reimplementation of https://reviews.llvm.org/rL342148, which later got reverted in r342336 (b09a8c9b). With ccache and ninja on a large core machine (40), I haven't run into build errors, so I'm hopeful it's better now, though it doesn't seem to be any different / new. Reviewed By: stephenneuendorffer Differential Revision: https://reviews.llvm.org/D90970
-
Arthur Eubanks authored
It seems nicer to list passes given a flag rather than displaying all passes in opt --help. This is awkwardly structured because a PassBuilder is required, but reusing the PassBuilder in runPassPipeline() doesn't work because we read the input IR before getting to runPassPipeline(). So printing the list of passes needs to happen before reading the input IR. If we remove the legacy PM code in main() and move everything from NewPMDriver.cpp into opt.cpp, we can create the PassBuilder before reading IR and check if we should print the list of passes and exit. But until then this hack seems fine. Compared to the legacy PM, the new PM passes are lacking descriptions. We'll need to figure out a way to add descriptions if we think this is important. Also, this only works for passes specified in PassRegistry.def. If we want to print other custom registered passes, we'll need a different mechanism. Reviewed By: asbirlea Differential Revision: https://reviews.llvm.org/D96101
-
Craig Topper authored
-
Christopher Di Bella authored
Implements parts of: * P0898R3 Standard Library Concepts * P1754 Rename concepts to standard_case for C++20, while we still can Differential Revision: https://reviews.llvm.org/D88131
-