- Apr 29, 2023
-
-
Arthur Eubanks authored
-
Michael Jones authored
The differential fuzzer for atof found a bug in glibc's handling of hexadecimal rounding. Since we can't easily update glibc and we want to avoid false positives when running the fuzzer, I've added an exception to skip all hexadecimal subnormal cases. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D149359
-
Vitaly Buka authored
Reviewed By: thurston Differential Revision: https://reviews.llvm.org/D149430
-
Teresa Johnson authored
Switch to the just updated versions of the API in tcmalloc that change the name of the hot cold paramter to a reserved identifier __hot_cold_t. This was based on feedback from Richard Smith, as I also need to add some follow-on handling to clang so they are annotated properly. Differential Revision: https://reviews.llvm.org/D149475
-
Jakub Kuderski authored
This enabled more optimization opportunities by moving zero/sign-extension closer to the use. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D149282
-
wlei authored
Part 2 of https://reviews.llvm.org/D147456 Use callee name on IR as an anchor to match the call target/inlinee name in the profile. The advantages of this in particular: - Different from the traditional way of encoding hash signatures to every block that would affect binary/profile size and build speed, it doesn't require any additional information for this, all the data is already in the IR and profiles. - Effective for current nested profile layout in which once a callsite is mismatched all the inlinee's profiles are dropped. **The input of the algorithm:** - IR locations: the anchor is the callee name of direct callsite. - Profile locations: the anchor is the call target name for `BodySample`s or inlinee's profile name for `CallsiteSamples`. The two lists are populated by parsing the IR and profile and both can be generalized as a sequence of locations with an optional anchor. For example: say location `1.2(foo)` refers to a callsite at `1.2` with callee name `foo` and `1.3` refers to a non-directcall location `1.3`. ``` // The current build source code: int main() { 1. ... 2. foo(); 3. ... 4 ... 5. ... 6. bar(); 7. ... } ``` IR locations are populated and simplified as: `[1, 2(foo), 3, 5, 6(bar), 7]`. ``` ; The "stale" profile: main:350:1 1: 1 2: 3 3: 100 foo:100 4: 2 7: 2 8: 200 bar:200 9: 30 ``` Profile locations are populated and simplified as `[1, 2, 3(foo), 4, 7, 8(bar), 9]` **Matching heuristic:** - Match all the anchors in lexical order first. - Match non-anchors evenly between two anchors: Split the non-anchor range, the first half is matched based on the start anchor, the second half is matched based on the end anchor. So the example above is matched like: ``` [1, 2(foo), 3, 5, 6(bar), 7] | | | | | | [1, 2, 3(foo), 4, 7, 8(bar), 9] ``` 3 -> 4 matching is based on anchor `foo`, 5 -> 7 matching is based on anchor `bar`. The output mapping of matching is [2->3, 3->4, 5->7, 6->8, 7->9]. For the implementation, the anchors are saved in a map for fast look-up. The result mapping is saved into `IRToProfileLocationMap`(see https://reviews.llvm.org/D147456) and distributed to all FunctionSamples(`distributeIRToProfileLocationMap`) **Clang-self build benchmark: ** Current build version: clang-10 The profiled version: clang-9 Results compared to a refresh profile(collected profile on clang-10) and to be fair, we invalidated new functions' profiles(both refresh and stale profile use the same profile list). 1) Regression to using refresh profile with this off : -3.93% 2) Regression to using refresh profile with this on : -1.1% So this algorithm can recover ~72% of the regression. **Internal(Meta) large-scale services.** we saw one real instance of a 3 week stale profile., it delivered a ~1.8% win. **Notes or future work:** - Classic AutoFDO support: the current version only supports pseudo-probe, but I believe it's not hard to extend to classic line-number based AutoFDO since pseudo-probe and line-number are shared the LineLocation structure. - The fuzzy matching is an open-ended area and there could be more heuristics to try out, but since the current version already recovers a reasonable percentage of regression(with some pseudo probe order change, it can recover close to 90%), I'm submitting the patch for review and we will try more heuristics in future. - Profile call target name are only available when the call is hit by samples, the missing anchor might mislead the matching, this can be mitigated in llvm-profgen to generate the call target for the zero samples. - This doesn't handle function name mismatch, we plan to solve it in future. Reviewed By: hoy, wenlei Differential Revision: https://reviews.llvm.org/D147545
-
wlei authored
AutoFDO/CSSPGO often has to deal with stale profiles collected on binaries built from several revisions behind release. It’s likely to get incorrect profile annotations using the stale profile, which results in unstable or low performing binaries. Currently for source location based profile, once a code change causes a profile mismatch, all the locations afterward are mismatched, the affected samples or inlining info are lost. If we can provide a matching framework to reuse parts of the mismatched profile - aka incremental PGO, it will make PGO more stable, also increase the optimization coverage and boost the performance of binary. This patch is the part 1 of stale profile matching, summary of the implementation: - Added a structure for the matching result:`LocToLocMap`, which is a location to location map meaning the location of current build is matched to the location of the previous build(to be used to query the “stale” profile). - In order to use the matching results for sample query, we need to pass them to all the location queries. For code cleanliness, we added a new pointer field(`IRToProfileLocationMap`) to `FunctionSamples`. - Added a wrapper(`mapIRLocToProfileLoc`) for the query to the location, the location from input IR will be remapped to the matched profile location. - Added a new switch `--salvage-stale-profile`. - Some refactoring for the staleness detection. Test case is in part 2 with the matching algorithm. Reviewed By: wenlei Differential Revision: https://reviews.llvm.org/D147456
-
Med Ismail Bennani authored
In 27f27d15, we added a new way to use textual (JSON) object files and symbol files with the interactive crashlog command, using the inlined symbols from the crash report. However, there was a missing piece after successfully adding the textual module to the target, we didn't mark it as available causing the module loading to exit early. This patch addresses that issue by marking the module as available when added successfully to the target. Differential Revision: https://reviews.llvm.org/D149477 Signed-off-by:
Med Ismail Bennani <ismail@bennani.ma>
-
Florian Hahn authored
-
Vasileios Porpodas authored
[NFC][SLP] Cleanup: Replace Value* operand with Instruction* in `vectorizeRootInstruction()` and `vectorizeHorReduction()` This makes it explicit that these functions work with instructions, and avoids calling them if the operand is not an instruction. Differential Revision: https://reviews.llvm.org/D149465
-
Louis Dionne authored
__use() can't be marked as constexpr in C++11 because it returns void, which is not a literal type. But it turns out that we just don't need to define it at all, since it's never called outside of decltype. Differential Revision: https://reviews.llvm.org/D149360
-
Manna, Soumi authored
Reported By Static Analyzer Tool, Coverity: Big parameter passed by value Copying large values is inefficient, consider passing by reference; Low, medium, and high size thresholds for detection can be adjusted. 1. Inside "CodeGenModule.cpp" file, in clang::CodeGen::CodeGenModule::EmitBackendOptionsMetadata(clang::CodeGenOptions): A very large function call parameter exceeding the high threshold is passed by value. pass_by_value: Passing parameter CodeGenOpts of type clang::CodeGenOptions const (size 2168 bytes) by value, which exceeds the high threshold of 512 bytes. 2. Inside "SemaType.cpp" file, in IsNoDerefableChunk(clang::DeclaratorChunk): A large function call parameter exceeding the low threshold is passed by value. pass_by_value: Passing parameter Chunk of type clang::DeclaratorChunk (size 176 bytes) by value, which exceeds the low threshold of 128 bytes. 3. Inside "CGNonTrivialStruct.cpp" file, in <unnamed>::getParamAddrs<1ull, <0ull...>>(std::integer_sequence<unsigned long long, T2...>, std::array<clang::CharUnits, T1>, clang::CodeGen::FunctionArgList, clang::CodeGen::CodeGenFunction *): A large function call parameter exceeding the low threshold is passed by value. .i. pass_by_value: Passing parameter Args of type clang::CodeGen::FunctionArgList (size 144 bytes) by value, which exceeds the low threshold of 128 bytes. 4. Inside "CGGPUBuiltin.cpp" file, in <unnamed>::containsNonScalarVarargs(clang::CodeGen::CodeGenFunction *, clang::CodeGen::CallArgList): A very large function call parameter exceeding the high threshold is passed by value. i. pass_by_value: Passing parameter Args of type clang::CodeGen::CallArgList (size 1176 bytes) by value, which exceeds the high threshold of 512 bytes. Reviewed By: tahonermann Differential Revision: https://reviews.llvm.org/D149163
-
Louis Dionne authored
This patch makes std::pair's constructors and assignment operators closer to conforming in C++23. The only missing bit I am aware of now is `reference_constructs_from_temporary_v` checks, which we don't have the tools for yet. This patch also refactors a long-standing non-standard extension where we'd provide constructors for tuple-like types in all standard modes. The criteria for being a tuple-like type are different from pair-like types as introduced recently in the standard, leading to a lot of complexity when trying to implement recent papers that touch the pair constructors. After this patch, the pre-C++23 extension is provided in a self-contained block so that we can easily deprecate and eventually remove the extension in future releases. Differential Revision: https://reviews.llvm.org/D143914
-
Valentin Clement authored
Remove copyinOperands, createOperands, createZeroOperands, attachOperands SmallVectors as they are not used anymore. The op itself cannot be cleanup yet because `mlir/lib/Target/LLVMIR/Dialect/OpenACC/OpenACCToLLVMIRTranslation.cpp` still depends on it. The final clean up on the op will be down once the translation uses the new data operand operations. Reviewed By: vzakhari Differential Revision: https://reviews.llvm.org/D149467
-
Kiran Chandramohan authored
A case was missed for the bfloat type in the getRealType function in FIRBuilder. This can cause a crash in some situations like in the provided test. The patch adds the missing case. Reviewed By: vzakhari Differential Revision: https://reviews.llvm.org/D149469
-
Shivam Gupta authored
There seems to be a lot of documentation is missing for different command line flags uses. This create confusion with same looking options, better to document clearly their uses. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D149405
-
Jakub Kuderski authored
This moves zero/sign-extension ops closer to their use and exposes more narrowing optimization opportunities. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D149233
-
Tom Praschan authored
This uses the logic added in D124690 Differential Revision: https://reviews.llvm.org/D147846
-
Mikhail Maltsev authored
According to the GNU ld manual https://sourceware.org/binutils/docs/ld/ARM.html#ARM the R_ARM_TARGET2 relocation (used in exception handling tables) is treated differently depending on the target. By default, LLD treats R_ARM_TARGET2 as R_ARM_GOT_PREL (--target2=got-rel), which is correct for Linux but not for embedded targets. This patch adds --target2=rel to linker options in the baremetal toolchain driver so that on baremetal targets, R_ARM_TARGET2 is treated as R_ARM_REL32. Such behavior is compatible with GNU ld and unwinding libraries (e.g., libuwind). Reviewed By: peter.smith, phosek Differential Revision: https://reviews.llvm.org/D149458
-
Erich Keane authored
We only need the list of constriant template arguments when we have a valid constraint. We give up on merging the auto-type constraints if the template arguments don't match, but neglected to clear the collection of template arguments. The result was we had an AutoType where we initialized the number of template arguments, but never initialized the template arguments themselves. This patch adds an assert to catch this in the future, plus ensures we clear out the vector so we don't try to create the AutoType incorrectly.
-
Vasileios Porpodas authored
This makes `matchAssociativeReduction()` a bit simpler. Differential Revision: https://reviews.llvm.org/D149452
-
Fangrui Song authored
From Florian Weimer's D144073 > On GNU/Linux (glibc), the crypt and crypt_r functions are not part of the main shared object (libc.so.6), but libcrypt (with multiple possible sonames). The sanitizer libraries do not depend on libcrypt, so it can happen that during sanitizer library initialization, no real implementation will be found because the crypt, crypt_r functions are not present in the process image (yet). If its interceptors are called nevertheless, this results in a call through a null pointer when the sanitizer library attempts to forward the call to the real implementation. > > Many distributions have already switched to libxcrypt, a library that is separate from glibc and that can be build with sanitizers directly (avoiding the need for interceptors). This patch disables building the interceptor for glibc targets. Let's remove crypt and crypt_r interceptors (D68431) to fix issues with newer glibc. For older glibc, msan will not know that an uninstrumented crypt_r call initializes `data`, so there is a risk for false positives. However, with some codebase survey, I think crypt_r uses are very few and the call sites typically have a `memset(&data, 0, sizeof(data));` anyway. Fix https://github.com/google/sanitizers/issues/1365 Related: https://bugzilla.redhat.com/show_bug.cgi?id=2169432 Reviewed By: #sanitizers, fweimer, thesamesam, vitalybuka Differential Revision: https://reviews.llvm.org/D149403
-
Saleem Abdulrasool authored
This generalises the GetXcodeSDKPath hook to a GetSDKRoot path which will be re-used for the Windows support to compute a language specific SDK path on the platform. Because there may be other options that we wish to use to compute the SDK path, sink the XcodeSDK parameter into a structure which can pass a disaggregated set of options. Furthermore, optionalise the parameter as Xcode is not available for all platforms. Differential Revision: https://reviews.llvm.org/D149397 Reviewed By: JDevlieghere
-
Slava Zakharin authored
The RHS cannot be casted to the LHS type, when LHS is polymorphic. With this change we will use the RHS type for emboxing with special hadling for i1 type. I created https://github.com/llvm/llvm-project/issues/62419 for the AllocaOp generated during HLFIRtoFir conversion. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D149392
-
Justin Lebar authored
If both the true and false operands of a `select` are poison, then the `select` is poison. Differential Revision: https://reviews.llvm.org/D149427
-
- Apr 28, 2023
-
-
Slava Zakharin authored
Assignment from a character dummy argument to a length-one character variable resulted in illegal fir.convert: %0 = fir.load %unboxed_dummy : !fir.ref<!fir.char<1,?>> %1 = fir.convert %0 : (!fir.char<1,?>) -> !fir.char<1> fir.store %1 to %local : !fir.ref<!fir.char<1>> This change fixes the length-one assignment code to use proper casts. For character dummy arguments with constant length we will now also type cast the unboxed reference to the character type with constant length during the lowering: fir.convert %x : (!fir.ref<!fir.char<1,?>>) -> !fir.ref<!fir.char<1,8>> I also adjusted the length-one assignment recognition so that in case of same-length assignment we recognize length-one from either LHS or RHS data types. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D149382
-
Slava Zakharin authored
The `none` type cannot be used for creating AssociateOp for the actual argument. I think it should be always okay to compute the storage data type based on the actual argument expression.
-
Jakub Kuderski authored
Fixes: https://github.com/llvm/llvm-project/issues/62368 Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D149376
-
Philip Reames authored
This allows us to model and thus test transforms which are legal only when a vector load with less than element alignment are supported. This was originally part of D126085, but was split out as we didn't have a good example of such a transform. As can be seen in the test diffs, we have the recently added concat_vector(loads) -> strided_load transform (from D147713) which now benefits from the unaligned support. While making this change, I realized that we actually *do* support unaligned vector loads and stores of all types via conversion to i8 element type. For contiguous loads and stores without masking, we actually already implement this in the backend - though we don't tell the optimizer that. For indexed, lowering to i8 requires complicated addressing. For indexed and segmented, we'd have to use indexed. All around, doesn't seem worthwhile pursuing, but makes for an interesting observation. Differential Revision: https://reviews.llvm.org/D149375
-
Alex Brachet authored
Differential Revision: https://reviews.llvm.org/D149283
-
Philip Reames authored
Note that the strided load from concat_vector combine was using the wrong legality test. It happened to work out as the alignment requirement is based on the scalar type either way, but unless I'm missing something allowsMisalignedAccess is expecting a contiguous memory access. Differential Revision: https://reviews.llvm.org/D149369
-
Florian Hahn authored
No SCEVs are formed for instructions with non-scevable types, so no other SCEV expressions can depend on them. Skip those instructions and their users when invalidating SCEV expressions. Depends on D144847. Reviewed By: mkazantsev Differential Revision: https://reviews.llvm.org/D144848
-
Florian Hahn authored
-
David Green authored
We already have tablegen patterns for a lot of these, but performing the combine earlier in DAG can help in a few extra cases. Differential Revision: https://reviews.llvm.org/D149269
-
Jay Foad authored
"convergent" is documented as meaning that the call cannot be made control-dependent on more values, but in practice we also require that it cannot be made control-dependent on fewer values, e.g. it cannot be hoisted out of the body of an "if" statement. In code like this, if we allow CSE to combine the two calls: x = convergent_call(); if (cond) { y = convergent_call(); use y; } then we get this: x = convergent_call(); if (cond) { use x; } This is conceptually equivalent to moving the second call out of the body of the "if", up to the location of the first call, so it should be disallowed. Differential Revision: https://reviews.llvm.org/D149348 -
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D149349
-
Joel E. Denny authored
Without this patch, if an incompatible libomptarget.so is present in a system directory, such as /usr/lib64, check-openmp fails many libomptarget tests with linking errors. The problem appears to have started at D129875, which landed as dc52712a. This patch extends the libomptarget test suite config with a -L for the current build directory of libomptarget.so. Reviewed By: jhuber6, JonChesterfield Differential Revision: https://reviews.llvm.org/D149391
-
Nikita Popov authored
-
Qiongsi Wu authored
On AIX, when the input files are LLVM bitcode files, `llvm-ar` should set the archive kind to `K_AIXBIG` as well, instead of leaving it to the default `K_GNU`. Reviewed By: daltenty Differential Revision: https://reviews.llvm.org/D149377
-
Daniel Kiss authored
Clang accepts preserve_all for AArch64 while it is missing form the backed. Fixes #58145 Reviewed By: efriedma Differential Revision: https://reviews.llvm.org/D135652
-