- Nov 12, 2021
-
-
Arthur Eubanks authored
Forgot to amend D113537 with these changes before committing.
-
Arthur Eubanks authored
Having a separate counting method runs the risk of a mismatch between the actual reduction method and the counting method. Instead, create an Oracle that always returns true for shouldKeep(), run the reduction, and count how many times shouldKeep() was called. The module should not be modified if shouldKeep() always returns true. Reviewed By: Meinersbur Differential Revision: https://reviews.llvm.org/D113537
-
Arthur Eubanks authored
Metadata operands tend to require special conditions, especially on dbg intrinsics. We also don't have a zero value for metadata. Replacing callee operands is a little weird, since calling undef/null doesn't make sense. It also causes tons of invalid reductions when reducing calls to intrinsics since only arguments to intrinsics can be of the metadata type. Reviewed By: Meinersbur Differential Revision: https://reviews.llvm.org/D113532
-
Arthur Eubanks authored
If a sysroot was specified, it would take precedence over the Android NDK sysroot since it would appear after in the command line. Also only build runtimes for enabled target arches. Many places have copied this around so create and use supported_android_toolchains. Reviewed By: pcc Differential Revision: https://reviews.llvm.org/D113606
-
Mehdi Amini authored
This reverts commit c7be8b75. Build is broken (multiple buildbots)
-
LLVM GN Syncbot authored
-
Michael Kruse authored
Add a new "operands-skip" pass whose goal is to remove instructions in the middle of dependency chains. For instance: ``` %baseptr = alloca i32 %arrayidx = getelementptr i32, i32* %baseptr, i32 %idxprom store i32 42, i32* %arrayidx ``` might be reducible to ``` %baseptr = alloca i32 %arrayidx = getelementptr ... ; now dead, together with the computation of %idxprom store i32 42, i32* %baseptr ``` Other passes would either replace `%baseptr` with undef (operands, instructions) or move it to become a function argument (operands-to-args), both of which might fail the interestingness check. In principle the implementation allows operand replacement with any value or instruction in the function that passes the filter constraints (same type, dominance, "more reduced"), but is limited in this patch to values that are directly or indirectly used to compute the current operand value, motivated by the example above. Additionally, function arguments are added...
-
Stella Laurenzo authored
* Depends on D111504, which provides the boilerplate for building aggregate shared libraries from installed MLIR. * Adds a full-fledged Python example dialect and tests to the Standalone example (need to do a bit of tweaking in the top level CMake and lit tests to adapt better to if not building with Python enabled). * Rips out remnants of custom extension building in favor of `pybind11_add_module` which does the right thing. * Makes python and extension sources installable (outputs to src/python/${name} in the install tree): Both Python and C++ extension sources get installed as downstreams need all of this in order to build a derived version of the API. * Exports sources targets (with our properties that make everything work) by converting them to INTERFACE libraries (which have export support), as recommended for the forseeable future by CMake devs. Renames custom properties to start with lower-case letter, as also recommended/required (groan). * Adds a ROOT_DIR argument to `declare_mlir_python_extension` since now all C++ sources for an extension must be under the same directory (to line up at install time). * Need to validate against a downstream or two and adjust, prior to submitting. Downstreams will need to adapt by: * Remove absolute paths from any SOURCES for `declare_mlir_python_extension` (I believe all downstreams are just using `${CMAKE_CURRENT_SOURCE_DIR}` here, which can just be ommitted). May need to set `ROOT_DIR` if not relative to the current source directory. * To allow further downstreams to install/build, will need to make sure that all C++ extension headers are also listed under SOURCES for `declare_mlir_python_extension`. Reviewed By: stephenneuendorffer, mikeurbach Differential Revision: https://reviews.llvm.org/D111513 -
Mogball authored
-
Phoebe Wang authored
This fixes the crash due to lacking VZEXT_MOVL support with i16. Reviewed By: LuoYuanke, RKSimon Differential Revision: https://reviews.llvm.org/D113661
-
Michael Kruse authored
This reverts commit fa4210a9. It causes compile failures, presumably because conflicting with another patch landed after I checked locally.
-
Matthias Springer authored
The remaining dialects will be decoupled from ComprehensiveBufferize in separate commits. Differential Revision: https://reviews.llvm.org/D113459
-
Mogball authored
With `-Os` turned on, results in 2-5% binary size reduction (depends on the original binary). Without it, the binary size is essentially unchanged. Depends on D113128 Differential Revision: https://reviews.llvm.org/D113331
-
Michael Kruse authored
Add a new "operands-skip" pass whose goal is to remove instructions in the middle of dependency chains. For instance: ``` %baseptr = alloca i32 %arrayidx = getelementptr i32, i32* %baseptr, i32 %idxprom store i32 42, i32* %arrayidx ``` might be reducible to ``` %baseptr = alloca i32 %arrayidx = getelementptr ... ; now dead, together with the computation of %idxprom store i32 42, i32* %baseptr ``` Other passes would either replace `%baseptr` with undef (operands, instructions) or move it to become a function argument (operands-to-args), both of which might fail the interestingness check. In principle the implementation allows operand replacement with any value or instruction in the function that passes the filter constraints (same type, dominance, "more reduced"), but is limited in this patch to values that are directly or indirectly used to compute the current operand value, motivated by the example above. Additionally, function arguments are added to the candidate set which helps reducing the number of relevant arguments mitigating a concern of too many arguments mentioned in https://reviews.llvm.org/D110274#3025013. Possible future extensions: * Instead of requiring the same type, bitcast/trunc/zext could be automatically inserted for some more flexibility. * If undef is added to the candidate set, "operands-skip"is able to produce any reduction that "operands" can do. Additional candidates might be zero and one, where the "reductive power" classification can prefer one over the other. If undefined behaviour should not be introduced, undef can be removed from the candidate set. Reviewed By: aeubanks Differential Revision: https://reviews.llvm.org/D111818
-
Matthias Springer authored
This helper struct allows users of ComprehensiveBufferize to inject "post analysis" steps that are implemented after the analysis but before the bufferization. Differential Revision: https://reviews.llvm.org/D113458
-
Matheus Izvekov authored
Signed-off-by:
Matheus Izvekov <mizvekov@gmail.com> Differential Revision: https://reviews.llvm.org/D113722
-
Matheus Izvekov authored
This implements the following changes: * AutoType retains sugared deduced-as-type. * Template argument deduction machinery analyses the sugared type all the way down. It would previously lose the sugar on first recursion. * Undeduced AutoType will be properly canonicalized, including the constraint template arguments. * Remove the decltype node created from the decltype(auto) deduction. As a result, we start seeing sugared types in a lot more test cases, including some which showed very unfriendly `type-parameter-*-*` types. Signed-off-by:
Matheus Izvekov <mizvekov@gmail.com> Reviewed By: rsmith Differential Revision: https://reviews.llvm.org/D110216
-
Alex Langford authored
-
Tue Ly authored
Combine two loops in decimalStringToFloat and hexadecimalStringToFloat that extract the digits and re-arrange them a little bit. This slightly improves the performance of strtof and strtod: Running libc_str_to_float_comparison_test parse-number-fxx-test_data/data/* on my machine (Ryzen 1700) - with glibc: ~1.92 seconds - with current implementation: ~1.78 seconds - with this change: ~1.67 seconds Differential Revision: https://reviews.llvm.org/D113681
-
Benjamin Kramer authored
-
Yaxun (Sam) Liu authored
The driver uses class SanitizerArgs to store parsed sanitizer arguments. It keeps a cached SanitizerArgs object in ToolChain and uses it for different jobs. This does not work if the sanitizer options are different for different jobs, which could happen when an offloading toolchain translates the options for different jobs. To fix this, SanitizerArgs should be created by using the actual arguments passed to jobs instead of the original arguments passed to the driver, since the toolchain may change the original arguments. And the sanitizer arguments should be diagnose once. This patch also fixes HIP toolchain for handling -fgpu-sanitize: a warning is emitted for GPU's not supporting sanitizer and skipped. This is for backward compatibility with existing -fsanitize options. -fgpu-sanitize is also turned on by default. Reviewed by: Artem Belevich, Evgenii Stepanov Differential Revision: https://reviews.llvm.org/D111443
-
Butygin authored
* Some long names were added and script decided to change whitespaces in a lot of places * `ImageOperand` was renamed to `ImageOperands` in spec * Some *NV enums were renamed to *KHR (spec actually maintains both variants with same value, but script pulled only *KHR versions) Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D113667
-
Peter Klausler authored
The labels of WHERE constructs were being created within the scope of the construct, not the scope of its parent, leading to incorrect error messages for branches to that label. Differential Revision: https://reviews.llvm.org/D113696
-
Thomas Raoux authored
Support load with broadcast, elementwise divf op and remove the hardcoded restriction on the vector size. Picking the right size should be enfored by user and will fail conversion to llvm/spirv if it is not supported. Differential Revision: https://reviews.llvm.org/D113618
-
Quinn Pham authored
[NFC] As part of using inclusive language within the llvm project, this patch renames master plan to controlling plan in lldb. Reviewed By: jingham Differential Revision: https://reviews.llvm.org/D113019
-
Chris Bieneman authored
Final macro diagnostics should log from system headers. As planned, final macros are hard-mode. They always log diagnostics.
-
Daniel McIntosh authored
Right now we drop the char_traits template argument, which presumes that string<_CharT, _Traits> and string<_CharT> are interchangeable. Reviewed By: Mordante, #libc, Quuxplusone Differential Revision: https://reviews.llvm.org/D112017
-
Aaron Ballman authored
This won't parse as either C or C++ according to Sphinx, so switched to text to appease Sphinx.
-
Snehasish Kumar authored
Set the default memprof serialization format as binary. 9 tests are updated to use print_text=true. Also fixed an issue with concatenation of default and test specified options (missing separator). Differential Revision: https://reviews.llvm.org/D113617
-
Snehasish Kumar authored
This change implements the raw binary format discussed in https://lists.llvm.org/pipermail/llvm-dev/2021-September/153007.html Summary of changes * Add a new memprof option to choose binary or text (default) format. * Add a rawprofile library which serializes the MIB map to profile. * Add a unit test for rawprofile. * Mark sanitizer procmaps methods as virtual to be able to mock them. * Extend memprof_profile_dump regression test. Differential Revision: https://reviews.llvm.org/D113317
-
Snehasish Kumar authored
The existing implementation uses a cache + eviction based scheme to record heap profile information. This design was adopted to ensure a constant memory overhead (due to fixed number of cache entries) along with incremental write-to-disk for evictions. We find that since the number to entries to track is O(unique-allocation-contexts) the overhead of keeping all contexts in memory is not very high. On a clang workload, the max number of unique allocation contexts was ~35K, median ~11K. For each context, we (currently) store 64 bytes of data - this amounts to 5.5MB (max). Given the low overheads for a complex workload, we can simplify the implementation by using a hashmap without eviction. Other changes: * Memory map is dumped at the end rather than startup. The relative order in the profile dump is unchanged since we no longer have evicted entries at runtime. * Added a test to check meminfoblocks are merged. Differential Revision: https://reviews.llvm.org/D111676
-
Snehasish Kumar authored
Move the memprof MemInfoBlock struct to it's own header as requested during the review of D111676. Differential Revision: https://reviews.llvm.org/D113315
-
Snehasish Kumar authored
This change adds a ForEach method to the AddrHashMap class which can then be used to iterate over all the key value pairs in the hash map. I intend to use this in an upcoming change to the memprof runtime. Added a unit test to cover basic insertion and the ForEach callback. Differential Revision: https://reviews.llvm.org/D111368
-
Nikita Popov authored
This handles a special case of foldAndOrOfICmpsUsingRanges() with two equality predicates.
-
Sanjay Patel authored
More coverage for D113603
-
Louis Dionne authored
This is part of https://wg21.link/P0355R7. I am adding these methods to provide an alternative for the {from,to}_time_t methods that were removed in https://llvm.org/D113027. Differential Revision: https://reviews.llvm.org/D113430
-
Louis Dionne authored
We are trying to remove duplication of third-party code in https://reviews.llvm.org/D112012, which will move the Google Benchmark code outside of the `libcxx/` directory. That breaks running the benchmarks in the Standalone build. Since we have deprecated the Standalone build anyway, this patch just removes support for the benchmark in Standalone mode until we remove that mode entirely. Differential Revision: https://reviews.llvm.org/D113503
-
Louis Dionne authored
Instead of hard-coding the target for our CI nodes, use the default compiler triple. Also, allow building compiler-rt for the single specified triple in case we're running on Darwin (otherwise, the bootstrapping build complains). Differential Revision: https://reviews.llvm.org/D113683
-
Bran Hagger authored
Differential Revision: https://reviews.llvm.org/D110354
-
Erich Keane authored
As discussed here: https://lwn.net/Articles/691932/ GCC6.0 adds target_clones multiversioning. This functionality is an odd cross between the cpu_dispatch and 'target' MV, but is compatible with neither. This attribute allows you to list all options, then emits a separately optimized version of each function per-option (similar to the cpu_specific attribute). It automatically generates a resolver, just like the other two. The mangling however, is... ODD to say the least. The mangling format is: <normal_mangling>.<option string>.<option ordinal>. Differential Revision:https://reviews.llvm.org/D51650
-