- May 23, 2023
-
-
Joseph Huber authored
The AMDGPU backend has a built-in pass to lower constructors. We do this manually in the `start.cpp` implementation so we can disable this to keep the binaries smaller. Differential Revision: https://reviews.llvm.org/D151213
-
Jonathan Peyton authored
While loop within task priority code did not have necessary update of variable which could lead to hangs if two threads collided when both attempted to execute the compare_and_exchange. Fixes: https://github.com/llvm/llvm-project/issues/62867 Differential Revision: https://reviews.llvm.org/D151138
-
Tue Ly authored
Make log10 correctly rounded for non-FMA targets and improve its performance. Implemented fast pass and accurate pass: **Fast Pass**: - Range reduction step 0: Extract exponent and mantissa ``` x = 2^(e_x) * m_x ``` - Range reduction step 1: Use lookup tables of size 2^7 = 128 to reduce the argument to: ``` -2^-8 <= v = r * m_x - 1 < 2^-7 where r = 2^-8 * ceil( 2^8 * (1 - 2^-8) / (1 + k * 2^-7) ) and k = trunc( (m_x - 1) * 2^7 ) ``` - Polynomial approximation: approximate `log(1 + v)` by a degree-7 polynomial generated by Sollya with: ``` > P = fpminimax((log(1 + x) - x)/x^2, 5, [|D...|], [-2^-8, 2^-7]); ``` - Combine the results: ``` log10(x) ~ ( e_x * log(2) - log(r) + v + v^2 * P(v) ) * log10(e) ``` - Perform additive Ziv's test with errors bounded by `P_ERR * v^2`. Return the result if Ziv's test passed. **Accurate Pass**: - Take `e_x`, `v`, and the lookup table index from the range reduction step of fast pass. - Perform 3 more range reduction steps: - Range reduction step 2: Use look-up tables of size 193 to reduce the argument to `[-0x1.3ffcp-15, 0x1.3e3dp-15]` ``` v2 = r2 * (1 + v) - 1 = (1 + s2) * (1 + v) - 1 = s2 + v + s2 * v where r2 = 2^-16 * round ( 2^16 / (1 + k * 2^-14) ) and k = trunc( v * 2^14 + 0.5 ). ``` - Range reduction step 3: Use look-up tables of size 161 to reduce the argument to `[-0x1.01928p-22 , 0x1p-22]` ``` v3 = r3 * (1 + v2) - 1 = (1 + s3) * (1 + v2) - 1 = s3 + v2 + s3 * v2 where r3 = 2^-21 * round ( 2^21 / (1 + k * 2^-21) ) and k = trunc( v * 2^21 + 0.5 ). ``` - Range reduction step 4: Use look-up tables of size 130 to reduce the argument to `[-0x1.0002143p-29 , 0x1p-29]` ``` v4 = r4 * (1 + v3) - 1 = (1 + s4) * (1 + v3) - 1 = s4 + v3 + s4 * v3 where r4 = 2^-28 * round ( 2^28 / (1 + k * 2^-28) ) and k = trunc( v * 2^28 + 0.5 ). ``` - Polynomial approximation: approximate `log10(1 + v4)` by a degree-4 minimax polynomial generated by Sollya with: ``` > P = fpminimax(log10(1 + x)/x, 3, [|128...|], [-0x1.0002143p-29 , 0x1p-29]); ``` - Combine the results: ``` log10(x) ~ e_x * log10(2) - log10(r) - log10(r2) - log10(r3) - log10(r4) + v * P(v) ``` - The combined results are computed using floating points of 128-bit precision. **Performance** - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.92%. - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log10 GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 20.402 + 0.589 clc/call; Median-Min = 0.277 clc/call; Max = 22.752 clc/call; -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 75.797 + 3.317 clc/call; Median-Min = 3.407 clc/call; Max = 79.371 clc/call; -- System LIBC reciprocal throughput -- [####################] 100 % Ntrial = 20 ; Min = 22.668 + 0.184 clc/call; Median-Min = 0.181 clc/call; Max = 23.205 clc/call; -- LIBC reciprocal throughput -- with FMA [####################] 100 % Ntrial = 20 ; Min = 25.977 + 0.183 clc/call; Median-Min = 0.138 clc/call; Max = 26.283 clc/call; -- LIBC reciprocal throughput -- without FMA [####################] 100 % Ntrial = 20 ; Min = 22.140 + 0.980 clc/call; Median-Min = 0.853 clc/call; Max = 23.790 clc/call; ``` - Latency from CORE-MATH's perf tool on Ryzen 5900X: ``` $ ./perf.sh log10 --latency GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 54.613 + 0.357 clc/call; Median-Min = 0.287 clc/call; Max = 55.701 clc/call; -- CORE-MATH latency -- without FMA (-march=x86-64-v2) [####################] 100 % Ntrial = 20 ; Min = 79.681 + 0.482 clc/call; Median-Min = 0.294 clc/call; Max = 81.604 clc/call; -- System LIBC latency -- [####################] 100 % Ntrial = 20 ; Min = 61.532 + 0.208 clc/call; Median-Min = 0.199 clc/call; Max = 62.256 clc/call; -- LIBC latency -- with FMA [####################] 100 % Ntrial = 20 ; Min = 41.510 + 0.205 clc/call; Median-Min = 0.244 clc/call; Max = 41.867 clc/call; -- LIBC latency -- without FMA [####################] 100 % Ntrial = 20 ; Min = 55.669 + 0.240 clc/call; Median-Min = 0.280 clc/call; Max = 56.056 clc/call; ``` - Accurate pass latency: ``` $ ./perf.sh log10 --latency --simple_stat GNU libc version: 2.35 GNU libc release: stable -- CORE-MATH latency -- with FMA 640.688 -- CORE-MATH latency -- without FMA (-march=x86-64-v2) 667.354 -- LIBC latency -- with FMA 495.593 -- LIBC latency -- without FMA 504.143 ``` Reviewed By: zimmermann6 Differential Revision: https://reviews.llvm.org/D150014 -
Manna, Soumi authored
Reported by Static Code Analyzer Tool, Coverity: Inside "SemaExprMember.cpp" file, in clang::Sema::BuildMemberReferenceExpr(clang::Expr *, clang::QualType, clang::SourceLocation, bool, clang::CXXScopeSpec &, clang::SourceLocation, clang::NamedDecl *, clang::DeclarationNameInfo const &, clang::TemplateArgumentListInfo const *, clang::Scope const *, clang::Sema::ActOnMemberAccessExtraArgs *): Return value of function which returns null is dereferenced without checking //Condition !Base, taking true branch. if (!Base) { TypoExpr *TE = nullptr; QualType RecordTy = BaseType; //Condition IsArrow, taking true branch. if (IsArrow) RecordTy = RecordTy->castAs<PointerType>()->getPointeeType(); //returned_null: getAs returns nullptr (checked 279 out of 294 times). //Condition TemplateArgs != NULL, taking true branch. //Dereference null return value (NULL_RETURNS) //dereference: Dereferencing a pointer that might be nullptr RecordTy->getAs() when calling LookupMemberExprInRecord. if (LookupMemberExprInRecord( *this, R, nullptr, RecordTy->getAs<RecordType>(), OpLoc, IsArrow, SS, TemplateArgs != nullptr, TemplateKWLoc, TE)) return ExprError(); if (TE) return TE; This patch uses castAs instead of getAs which will assert if the type doesn't match. Reviewed By: erichkeane Differential Revision: https://reviews.llvm.org/D151130 -
Nikita Popov authored
The test fails on the clang-ppc64le-rhel build bot, which has DEFAULT_LINKER set and an ld.lld binary in the LLVM build directory.
-
Joseph Huber authored
Currently AMDGPU offers extra ctor / dtor lowering by emitting a kernel that can be called. It's possible to handle ctors and dtors using the standard method as shown in D149340's commit message. In which case we on't need these extra kernels as they won't be called. This patch simply adds a way to conditionally turn off this handling if we do not want to get extra kernels in the output. Unrelated, but we could convert this handling to an ODR function that simply calls the code in D149340 constructed via LLVM-IR. That would handle priority correctly and would then be correct if not run in LTO mode. Reviewed By: yaxunl Differential Revision: https://reviews.llvm.org/D150565
-
Fangrui Song authored
[ubsan][test] Remove --check-prefix=UNIQUE for x86_64-apple from e215996a After switching to use a type hash instead of possibly-non-unique typeinfo objects, we no longer have unique/non-unique distinction.
-
Nikita Popov authored
Directly remove these dead extractelement instructions, rather than leaving them for the next InstCombine iteration to clean up. Should be mostly NFC, apart from worklist order differences.
-
Matthias Springer authored
This bug was recently introduced in D143927 and manifests as a dominance violation. Differential Revision: https://reviews.llvm.org/D151077
-
Aaron Ballman authored
These are showing up in MSVC builds.
-
Dinar Temirbulatov authored
Fixing last commit by adding actual change to AArch64TargetTransformInfo.cpp Differential Revision: https://reviews.llvm.org/D150336
-
Thomas Preud'homme authored
This will be required to allow arbitrary precision support to FileCheck's numeric variables and expressions. Note: as per getAsInteger(), this does not support negative value. If there is interest for that it can be added in a separate patch. Reviewed By: dblaikie Differential Revision: https://reviews.llvm.org/D150878
-
Dinar Temirbulatov authored
We noticed some runtime performance improvements by disabling maximising bandwidth for streaming compatible sve. Differential Revision: https://reviews.llvm.org/D150336
-
Thomas Preud'homme authored
Function valueFromStringRepr() throws an error on missing 0x prefix when parsing a number string into a value. However, getWildcardRegex() already ensures that only text with the 0x prefix will match and be parsed, making that error throwing code dead code. This commit turn the code into an assert and remove the unit tests exercising that test accordingly. Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D150797
-
Krasimir Georgiev authored
-
Pavel Iliin authored
Put features into function version name in increasing priority order. Differential Revision: https://reviews.llvm.org/D150800
-
Nikita Popov authored
For always poison shifts, any KnownBits return value is valid. Currently we return unknown, but returning zero is generally more profitable. We had some code in ValueTracking that tried to do this, but was actually dead code. Differential Revision: https://reviews.llvm.org/D150648
-
Kadir Cetinkaya authored
Underlying FS can store different file names inside the stat response (e.g. symlinks resolved, absolute paths, dots removed). But we store path names as requested inside the preamble, https://github.com/llvm/llvm-project/blob/main/clang/lib/Serialization/ASTWriter.cpp#L1635. This improves cache hit rates from ~30% to 90% in a build system that uses symlinks. Differential Revision: https://reviews.llvm.org/D151185
-
Vlad Serebrennikov authored
This reverts commit 85452b5f.
-
Nikita Popov authored
-
Leandro Lupori authored
Fixes gfortran test-suite regression. Differential Revision: https://reviews.llvm.org/D150686
-
Nikita Popov authored
Replace structured bindings with std::get, as they apparently break the modules build. ----- Store the end iterator on the VisitStack, instead of recomputing it every time, as doing so is not free.
-
Tim Northover authored
Since we're checking the triple directly, arm64_32 shows up differently and was still getting an attempt at asynchronous unwind that added lots more `__eh_frame` entries instead of the compact format.
-
Martin Braenne authored
This patch is part of the ongoing migration to strict handling of value categories (see https://discourse.llvm.org/t/70086 for details). Depends On D150775 Reviewed By: gribozavr2 Differential Revision: https://reviews.llvm.org/D150776
-
Nikita Popov authored
The test added in c5fe10f3 contains some typos in the check lines, due to which it never actually verified what was intended. Fix the test by adding the required input tree and adjusting the check lines appropriately. Differential Revision: https://reviews.llvm.org/D151195
-
LLVM GN Syncbot authored
-
Martin Storsjö authored
This allows all ExecutionEngine tests pass in MinGW build configurations. Differential Revision: https://reviews.llvm.org/D150555
-
Jun Zhang authored
This reverts commit 094ab478. Reland with changing `ParseAndExecute` to `Parse` in `Interpreter::create`. This avoid creating JIT instance everytime even if we don't really need them. This should fixes failures like https://lab.llvm.org/buildbot/#/builders/38/builds/11955 The original reverted patch also causes GN bot fails on M1. (https://lab.llvm.org/buildbot/#/builders/38/builds/11955) However, we can't reproduce it so let's reland it and see what happens. See discussions here: https://reviews.llvm.org/rGd71a4e02277a64a9dece591cdf2b34f15c3b19a0
-
Luo, Yuanke authored
-
Tom Weaver authored
This reverts commit eb5902ff. Caused buildbot failures on: https://lab.llvm.org/buildbot/#/builders/139/builds/41248 https://lab.llvm.org/buildbot/#/builders/216/builds/21637
-
Simon Pilgrim authored
-
Timm Bäder authored
Differential Revision: https://reviews.llvm.org/D151191
-
Jay Foad authored
RegScavenger::backward is preferred because it does not rely on accurate kill flags. Differential Revision: https://reviews.llvm.org/D150557
-
LLVM GN Syncbot authored
-
Simon Pilgrim authored
This patch analyzes AVX512 instructions for full vector width folded loads from the constant pool and attempts to determine if it can be replaced with a smaller broadcast folded variant. Typically the broadcast opportunities were missed by type-width mismatches or mulituse limitations which have been removed in later passes. As well as introducing broadcast fold tables (which can hopefully be extended/automated in the future), this also handles mismatches in the AND/ANDN/OR/XOR/TERNLOG type-widths, catching additional missed opportunities. This is patch is pulled from the ongoing work based on D150143, but without removing the existing DAG constant broadcast lowering code - this patch is currently a late stage cleanup only. The intention is to add additional broadcast/extension handling of constants in future patches, but it turned out that AVX512 broadcast handling was the easiest to start with. Differential Revision: https://reviews.llvm.org/D150526
-
Vlad Serebrennikov authored
CWG977 focus on point of /completeness/ of enums. Wording provided in CWG1482. CWG1482 and CWG2516 focus on locus (point) of /declaration/. Wording provided in CWG2516. Reviewed By: #clang-language-wg, shafik Differential Revision: https://reviews.llvm.org/D151042
-
Vlad Serebrennikov authored
[[https://wg21.link/p1787 | P1787]]: CWG2213 is resolved by allowing an elaborated-type-specifier to contain a simple-template-id without friend. Wording: see changes to [dcl.type.elab]]/1. The gist of the issue is that forward declaration of partial class template specialization was disallowed. Reviewed By: #clang-language-wg, shafik Differential Revision: https://reviews.llvm.org/D151032
-
Jay Foad authored
RegScavenger::backward is preferred because it does not rely on accurate kill flags. Differential Revision: https://reviews.llvm.org/D150558
-
Guillaume Chatelet authored
With more tests added to LLVM libc each week we want to keep track of unittest's runtime, especially for low end build bots. Top offender can be tracked with a bit of scripting (spoiler alert, mem function sweep tests are in the top ones) ``` ninja check-libc | grep "ms)" | awk '{print $(NF-1),$0}' | sort -nr | cut -f2- -d' ' ``` Unfortunately this doesn't work for hermetic tests since `clock` is unavailable. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D151097 -
Alex Bradbury authored
Our current approach is that if one extension requires another, we make LLVM treat it as implied. My initial zfbfmin patch failed to do this for the F extension (documented as a requirement of zfbfmin). This patch fixes that. Differential Revision: https://reviews.llvm.org/D151096
-