- Jul 11, 2023
-
-
Zi Xuan Wu (Zeson) authored
[RISCV] Don't fold RISCVISD::VMV_V_X_VL series node and scalar load to vector load when scalar load is update load We try to fold RISCVISD::VMV_V_X_VL series node + scalar load -> vector load. But if scalar load is indexed load (load update form), it's not profitable to fold because load update node can't be removed after fold. Differential Revision: https://reviews.llvm.org/D152222
-
pvanhout authored
Adds a new backend to power the GISel Combiners using the InstructionSelector's match tables. This does not depend on any of the data structures created for the current combiner and is intended to replace it entirely. See the RFC for more details: https://discourse.llvm.org/t/rfc-matchtable-based-globalisel-combiners/71457/6 Note: this would replace D141135. Reviewed By: aemerson, arsenm Differential Revision: https://reviews.llvm.org/D153757
-
pvanhout authored
Move all of the reusable logic out of `GlobalISelEmitter.cpp` into a `GlobalISelMatchTableExecutorEmitter` class so the future combiner backend can use it as well. Depends on D153755 Reviewed By: aemerson Differential Revision: https://reviews.llvm.org/D153756
-
pvanhout authored
Makes `InstructionSelector.h`/`InstructionSelectorImpl.h` generic so the match tables can also be used for the combiner. Some notes: - Coverage was made an optional parameter of `executeMatchTable`, combines won't use it for now. - `GIPFP_` -> `GICXXPred_` so it's more generic. Those are just C++ predicates and aren't PatFrag-specific. - Pass the MatcherState directly to testMIPredicate_MI, the combiner will need it. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D153755
-
Petr Hosek authored
On Linux crt is typically use in combination with builtins. In the Clang driver the use of builtins and crt is controlled by the --rtlib option. Both builtins and crt also have similar build requirements where they need to be built before any other runtimes and must avoid dependencies. We also want builtins and crt these to be buildable separately from the rest of compiler-rt for bootstrapping purposes. Given how simple crt is, rather than maintaining a separate directory with its own separate build setup, it's more efficient to just move crt into builtins. We still use separate CMake option to control whether to built crt same as before. This is an alternative to D89492 and D136664. Differential Revision: https://reviews.llvm.org/D153989
-
Job Noorman authored
BOLT used `ToolOutputFile::keep` to make sure the intermediary object file was written to disk for debugging purposes when `--keep-tmp` was passed. However, since and intermediary `buffer_ostream` was used to stream to, and this class only writes to its output stream in its destructor, the object file was lost whenever its destructor wouldn't run. This could happen, for example, if there is a crash while linking. This patch makes sure the object file is written to disk immediately after we're done creating it. This is very useful while debugging JITLink crashes. This patch also gets rid of creating a temporary file when `--keep-tmp` is not passed by streaming the object file directly to a `SmallString`. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D154826
-
Tobias Gysi authored
A distinct attribute associates a referenced attribute with a unique identifier. Every call to its create function allocates a new distinct attribute instance. The address of the attribute instance temporarily serves as its unique identifier. Similar to the names of SSA values, the final unique identifiers are generated during pretty printing. Examples: #distinct = distinct[0]<42.0 : f32> #distinct1 = distinct[1]<42.0 : f32> #distinct2 = distinct[2]<array<i32: 10, 42>> This mechanism is meant to generate attributes with a unique identifier, which can be used to mark groups of operations that share a common properties such as if they are aliasing. The design of the distinct attribute ensures minimal memory footprint per distinct attribute since it only contains a reference to another attribute. All distinct attributes are stored outside of the storage uniquer in a thread local store that is part of the context. It uses one bump pointer allocator per thread to ensure distinct attributes can be created in-parallel. Reviewed By: rriddle, Dinistro, zero9178 Differential Revision: https://reviews.llvm.org/D153360
-
Haojian Wu authored
Similar to the https://reviews.llvm.org/D86048 (it only sets the bit for C++ code), we propagate the contains-errors bit for C-code path. Fixes https://github.com/llvm/llvm-project/issues/50236 Fixes https://github.com/llvm/llvm-project/issues/50243 Fixes https://github.com/llvm/llvm-project/issues/48636 Fixes https://github.com/llvm/llvm-project/issues/50320 Differential Revision: https://reviews.llvm.org/D154861
-
Balazs Benics authored
When we construct a `NonParamVarRegion`, we canonicalize the decl to always use the same entity for consistency. At the moment that is the canonical decl - which is the first decl in the redecl chain. However, this can cause problems with tentative declarations and extern declarations if we declare an array with unknown bounds. Consider this C example: https://godbolt.org/z/Kdvr11EqY ```lang=C typedef typeof(sizeof(int)) size_t; size_t clang_analyzer_getExtent(const void *p); void clang_analyzer_dump(size_t n); extern const unsigned char extern_redecl[]; const unsigned char extern_redecl[] = { 1,2,3,4 }; const unsigned char tentative_redecl[]; const unsigned char tentative_redecl[] = { 1,2,3,4 }; const unsigned char direct_decl[] = { 1,2,3,4 }; void test_redeclaration_extent(void) { clang_analyzer_dump(clang_analyzer_getExtent(direct_decl)); // 4 clang_analyzer_dump(clang_analyzer_getExtent(extern_redecl)); // should be 4 instead of Unknown clang_analyzer_dump(clang_analyzer_getExtent(tentative_redecl)); // should be 4 instead of Unknown } ``` The `getType()` of the canonical decls for the forward declared globals, will return `IncompleteArrayType`, unlike the `getDefinition()->getType()`, which would have returned `ConstantArrayType` of 4 elements. This makes the `MemRegionManager::getStaticSize()` return `Unknown` as the extent for the array variables, leading to FNs. To resolve this, I think we should prefer the definition decl (if present) over the canonical decl when constructing `NonParamVarRegion`s. FYI The canonicalization of the decl was introduced by D57619 in 2019. Differential Revision: https://reviews.llvm.org/D154827
-
Yeting Kuo authored
The patch uses a way similiar to vp.load/store and consider the mask popcount as the effetive vector length. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D151713
-
Martin Braenne authored
Reviewed By: sammccall Differential Revision: https://reviews.llvm.org/D154834
-
Manish Kausik H authored
If modified attribute list is invalid, reverting the change is a low-cost maintainence solution as compared to examples like [this](https://github.com/llvm/llvm-project/blob/main/llvm/tools/bugpoint/CrashDebugger.cpp#L368). This will ensure that the ListReducer maintains the sanctity of any new attribute dependencies added in the future/already present. Reviewed By: modocache Differential Revision: https://reviews.llvm.org/D154348
-
wangpc authored
The code to resolve class/multiclass arguments are similar, we extract them to `resolveArguments`s to simplify code. Reviewed By: tra, reames Differential Revision: https://reviews.llvm.org/D154065
-
wangpc authored
For `CSR_Interrupt`, we can generate the register list via a single `sequence`. For `CSR_XLEN_F32_Interrupt` and `CSR_XLEN_F64_Interrupt`, I don't see the reason why we need to keep the order the same as how we used to allocate registers (and we have changed the order in D146488), so I fold them into one `sequence`. There are some *.ll changes because of the order change. Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D154837
-
Piyou Chen authored
Reviewed By: asb Differential Revision: https://reviews.llvm.org/D154690
-
Caroline Tice authored
In two calls to ReadMemory in DWARFExpression.cpp, the buffer size passed to ReadMemory is not checked and can be bigger than the actual size of the buffer. This caused a buffer overflow bug, which we found through Address Sanitizer. This patch fixes the problem by checking the address size when it is first read out of the DWARF, and setting an error and returning immediatley if the size is invalid. This is the second attempt to fix this issue; I reverted the first one, as it was not quite correct. Differential Revision: https://reviews.llvm.org/D154907
-
Brad Smith authored
Have ToolChain::IsIntegratedAssemblerDefault default to true. Almost all of the ToolChains are using IAS nowadays. There are a few exceptions like XCore, some NaCl archs, and NVPTX/XCore in Generic_GCC::IsIntegratedAssemblerDefault. Reviewed By: MaskRay Differential Revision: https://reviews.llvm.org/D154902
-
Rahul Kayaith authored
Since op `Attribute`s are automatically downcasted on access, these mappings aren't necessary anymore. Instead we just always generate the getters/setters for attributes even if there isn't a `PythonAttr` mapping. depends on D154462 Reviewed By: ftynse Differential Revision: https://reviews.llvm.org/D154468
-
Rahul Kayaith authored
Update remaining `PyAttribute`-returning APIs to return `MlirAttribute` instead, so that they go through the downcasting mechanism. Reviewed By: makslevental Differential Revision: https://reviews.llvm.org/D154462
-
Matt Arsenault authored
We need to replace the other uses of the call chain with the new load chain. Fixes not preserving the return def with unused x86_fp80 results. Regression reported here: https://reviews.llvm.org/rGb15bf305ca3e9ce63aaef7247d32fb3a75174531#1224999
-
Ashay Rane authored
Older versions of clang (for example, v12) throw an error when compiling CStringChecker.cpp that the initializers for `SourceArgExpr`, `DestinationArgExpr`, and `SizeArgExpr` are missing braces around initialization of subobject. Newer clang versions don't throw this error. This patch adds the initialization braces to satisfy clang. Reviewed By: steakhal Differential Revision: https://reviews.llvm.org/D154871
-
Peiming Liu authored
Reviewed By: jpienaar Differential Revision: https://reviews.llvm.org/D154912
-
Wang Rui authored
This revision explicitly specifies the machine instruction properties instead of relying on guesswork. This is because guessing instruction properties has proven to be inaccurate, such as the machine LICM not working: ``` void func(char *a, char *b) { int i; for (i = 0; i != 72526; i++) a[i] = b[i]; } ``` Guessing instruction properties: ``` func: # @func move $a2, $zero .LBB0_1: # =>This Inner Loop Header: Depth=1 ldx.b $a3, $a1, $a2 stx.b $a3, $a0, $a2 addi.d $a2, $a2, 1 lu12i.w $a3, 17 ori $a3, $a3, 2894 bne $a2, $a3, .LBB0_1 ret .Lfunc_end0: ``` Explicitly specify instruction properties: ``` func: # @func lu12i.w $a2, 17 ori $a2, $a2, 2894 move $a3, $zero .LBB0_1: # =>This Inner Loop Header: Depth=1 ldx.b $a4, $a1, $a3 stx.b $a4, $a0, $a3 addi.d $a3, $a3, 1 bne $a3, $a2, .LBB0_1 ret .Lfunc_end0: ``` Reviewed By: SixWeining, xen0n Differential Revision: https://reviews.llvm.org/D154192 -
Kai Sasaki authored
As TOSA does not support the tensor with zero dimensions, we can check the zero value for the static shape input. Ideally, we should be able to check the tensor shape more broadly, such as using `CPred` in the TOSA type definition. But based on [[ https://discourse.llvm.org/t/where-can-we-put-the-shared-verification-among-multiple-dialect-ops/71806 | the discussion here ]] It makes input type verification complicated and hard to maintain and still only applies to the case the input is statically shaped. Therefore, in this change, we have put the zero dimension check in the verification of each op, which would be flexible and maintainable. See: https://github.com/llvm/llvm-project/issues/63212 Reviewed By: eric-k256 Differential Revision: https://reviews.llvm.org/D154569
-
Teresa Johnson authored
Previously the MemProf profile was expected to be in the same profile file as a normal PGO profile, passed via the usual -fprofile-use= option, and was matched in the same pass. To simplify profile preparation, since the raw MemProf profile requires the binary for symbolization and may be simpler to index separately from the raw PGO profile, and also to enable providing a MemProf profile for a SamplePGO build, separate out the MemProf feedback option and matching pass. This patch adds the -fmemory-profile-use=${file} option, and the provided file is passed down to LLVM and ultimately used in a new MemProfUsePass which performs the matching of just the memory profile contents of that file. Note that a single profile file containing both normal PGO and MemProf profile data is still supported, and the relevant profile data is matched by the appropriate matching pass(es) based on which option(s) the profile is provided with (the same profile file can be supplied to both feedback options). Differential Revision: https://reviews.llvm.org/D154856 -
Caroline Tice authored
This reverts commit ee476996. That commit was not the right way to fix the issue (it could result in reading too many bytes). A better fix is in the works. Original review: https://reviews.llvm.org/D153840
-
Han Shen authored
The original MFS work D85368 shows good performance improvement with Instrumented FDO. However, AutoFDO or Flow-Sensitive AutoFDO (FSAFDO) does not show performance gain. This is mainly caused by a less accurate profile compared to the iFDO profile. For the past few months, we have been working to improve FSAFDO quality, like in D145171. Taking advantage of this improvement, MFS now shows performance improvements over FSAFDO profiles. That being said, 2 minor changes need to be made, 1) An FS-AutoFDO profile generation pass needs to be added right before MFS pass and an FSAFDO profile load pass is needed when FS-AutoFDO is enabled and the MFS flag is present. 2) MFS only applies to hot functions, because we believe (and experiment also shows) FS-AutoFDO is more accurate about functions that have plenty of samples than those with no or very few samples. With this improvement, we see a 1.2% performance improvement in clang benchmark, 0.9% QPS improvement in our internal search benchmark, and 3%-5% improvement in internal storage benchmark. This is #1 of the two patches that enables the improvement. Reviewed By: wenlei, snehasish, xur Differential Revision: https://reviews.llvm.org/D152399
-
Artem Dergachev authored
-
Kazu Hirata authored
This patch fixes: bolt/lib/Core/DIEBuilder.cpp:468:18: error: unused variable 'Ref' [-Werror,-Wunused-variable]
-
K-Wu authored
Differential Revision: https://reviews.llvm.org/D154772
-
Justin Bogner authored
These switch statements are fully covered. Add an llvm_unreachable so that compilers that don't recognize that don't warn. Differential Revision: https://reviews.llvm.org/D154882
-
Fangrui Song authored
If X86 target is not enabled, llc fails with a different error.
-
Fangrui Song authored
-
Jonas Devlieghere authored
Fix a crash when trying to complete an ambiguous subcommand. Take `set s tar` for example: for the subcommand `s` there's ambiguity between set and show. Pressing TAB after this input currently crashes LLDB. The problem is that we're trying to complete `tar` but give up at `s` because of the ambiguity. LLDB doesn't expect the completed string to be shorter than the current string and crashes when trying to eliminate the common prefix. rdar://111848598 Differential revision: https://reviews.llvm.org/D154643
-
Eduard Zingerman authored
There are plans to add some BTF processing to tools like objdump and readelf. This commit moves BTF.{h,def} files from BPF target specific location to include/llvm/DebugInfo/* to avoid tools including headers from lib/Target/*. Reviewed By: yonghong-song, MaskRay Differential Revision: https://reviews.llvm.org/D149501 -
David Green authored
Ensure we constrain the register class of the NewDef to that of OldDef, in case they do not match. Fixes #63777
-
Alexander Yermolovich authored
To reduce memory footprint changed so that we process and write out TUs first, reset DIEBuilder and process CUs. CUs are processed in buckets. First bucket contains all the CUs with cross CU references. Rest processd one at a time. clang-17 build in debug mode, by clang-17. before 8:25.81 real, 834.37 user, 86.03 sys, 0 amem, 79525064 mmem 8:02.20 real, 820.46 user, 81.81 sys, 0 amem, 79501616 mmem 7:52.69 real, 802.01 user, 83.99 sys, 0 amem, 79534392 mmem after 7:49.35 real, 822.04 user, 66.19 sys, 0 amem, 34934260 mmem 7:42.16 real, 825.46 user, 63.52 sys, 0 amem, 34951660 mmem 7:46.71 real, 821.11 user, 63.14 sys, 0 amem, 34981164 mmem Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D151909
-
Alexander Yermolovich authored
Changed how we handle writing out .dwo and .dwp files. We now write out DWO sections sooner and destroy DIEBuilder. This should decrease memory footprint. Ran on clang-17 build in debug mode with split-dwarf. before 8:07.49 real, 664.62 user, 69.00 sys, 0 amem, 41601612 mmem 8:07.06 real, 669.60 user, 68.75 sys, 0 amem, 41822588 mmem 8:00.36 real, 664.14 user, 66.36 sys, 0 amem, 41561548 mmem after 8:21.85 real, 682.23 user, 69.64 sys, 0 amem, 39379880 mmem 8:04.58 real, 671.62 user, 66.50 sys, 0 amem, 39735800 mmem 8:10.02 real, 680.67 user, 67.24 sys, 0 amem, 39662888 mmem Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D151908
-
Alexander Yermolovich authored
* Some cleanup and minor fixes for the new debug information re-writer before moving on to productatization. * The new rewriter wasn't handling binary with DWARF5 and DWARF4 with -fdebug-types-sections. * Removed dead cross cu reference code. * Added support for DW_AT_sibling. * With the new re-writer abbrev number can change which can lead to offset of Type Units changing. Before we would just copy raw data. Changed to write out Type Unit List. This is generated by gdb-add-index. * Fixed how bolt handles gdb-index generated by gdb-11 with types sections. Simplified logic that handles variations of gdb-index. * Clang can generate two type units with the same hash, but different content. LLD does not de-duplicate when ThinLTO is involved. Changed so that TU hash and offset are used to make TU's unique. * It is possible to have references within location expression to another DIE. Fixed it so that relative offset is updated correctly. * Removed all the code related to patching. * Removed dead code. Changed how we handling writting out TUs and TU Index. It now should fully work for DWARF4 and DWARF5. * Removed unused arguments from some APIs, changed return type to void, and other small cleanups. Reviewed By: maksfb Differential Revision: https://reviews.llvm.org/D151906
-
Rui Zhong authored
This revision implement new mechanism for DWARFRewriter. In the new mechanism, we adopt the same way with DWARFLinker did. By parsing Debug information into IR, we are allowed to handle debug information more flexible. Now the debug information updating process relies on IR and IR will be written out to binary once the updating finished. A new class was added: DIEBuilder. This class is responsible for parsing debug information and raising it to the IR level. This class is also used to write out the .debug_info and .debug_abbrev sections. Since we output brand new Abbrev section we won't need to always convert low_pc/high_pc into ranges. When conversion does happen we can also remove low_pc entry. Reviewed By: maksfb, ayermolo Differential Revision: https://reviews.llvm.org/D130315
-