1. May 24, 2023
  2. May 23, 2023
    • Joseph Huber's avatar
      [libc] More efficiently send bytes via `send_n` and `recv_n` · e826762a
      Joseph Huber authored
      Currently we have the `send_n` and `recv_n` routines to stream data,
      such as a string to print, to the other side. The first operation is to
      send the size so the other side knows the number of bytes to recieve.
      However, this wasted 56 bytes that could've been sent. This meant that
      small values, like the arguments to a function to call on the host for
      example, needed to perform an extra send. This patch sends the first 56
      bytes in the first packet and continues if necessary.
      
      Depends on D150992
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D151041
      e826762a
    • Joseph Huber's avatar
      [libc] Fix the `send_n` and `recv_n` utilities under divergent lanes · 29d3da3b
      Joseph Huber authored
      We provide the `send_n` and `recv_n` utilities as a generic way to
      stream data between both sides of the process. This was previously
      tested and performed as expected when using a string of constant size.
      However, when the size was allowed to diverge between the threads in the
      warp or wavefront this could deadlock. This did not occur on NVPTX
      because of the use of the explicit warp sync. However, on AMD one of the
      work items in the wavefront could continue executing and hit the next
      `recv` call before the other threads, then we would deadlock as we
      violated the RPC invariants.
      
      This patch replaces the for loop with a thread ballot. This will cause
      every thread in the warp or wavefront to continue executing the loop
      until all of them can exit. This acts as a more explicit wavefront sync.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D150992
      29d3da3b
    • Nikolas Klauser's avatar
      [libc++] Remove tests from ranges.pass.cpp which violate semantic requirements · 75eb3bd1
      Nikolas Klauser authored
      This also removes some tests which we have grouped together into robust_from_*.pass.cpp tests.
      
      Specifically, checking that
      - `ranges::dangling` is returned is done in `libcxx/test/std/algorithms/ranges_robust_against_dangling.pass.cpp`
      - `std::invoke` is used is done in `libcxx/test/std/algorithms/ranges_robust_against_omitting_invoke.pass.cpp`.
      - implicit conversion to bool works is done in `libcxx/test/std/algorithms/ranges_robust_against_nonbool_predicates.pass.cpp`
      
      Checking the comparison order is invalid because the `operator==` isn't symmetric.
      Checking what the exact type of `operator==` is, is invalid because comparing the same object has to yield the same results if the objects are not modified.
      
      Reviewed By: ldionne, #libc
      
      Spies: EricWF, libcxx-commits
      
      Differential Revision: https://reviews.llvm.org/D150588
      75eb3bd1
    • Nikolas Klauser's avatar
      [libc++][NFC] Move basic_ios extern instantiations into <ios> · 9334a858
      Nikolas Klauser authored
      `basic_ios` is defined in `<ios>`, so it seems weird that we declare the explicit instantiation for it i `<streambuf>`, which is technically unrelated.
      
      Reviewed By: #libc, EricWF, ldionne
      
      Spies: ldionne, EricWF, libcxx-commits
      
      Differential Revision: https://reviews.llvm.org/D150912
      9334a858
    • Yaxun (Sam) Liu's avatar
      [HIP] Allow std::malloc in device function · f5033c37
      Yaxun (Sam) Liu authored
      D106463 caused a regression that prevents std::malloc to be
      called in the device function, which is allowed with nvcc.
      
      Basically the standard C++ header introducing malloc in
      std namespace by using ::malloc. The device ::malloc
      function needs to be declared before using ::malloc
      to be introduced into std namespace.
      
      Revert D106463 and add a test.
      
      Reviewed by: Artem Belevich
      
      Differential Revision: https://reviews.llvm.org/D150965
      f5033c37
    • Nikolas Klauser's avatar
      [libc++][NFC] Fix whitespace problems in the files added to ignore_format.txt in D151115 · 18c4695d
      Nikolas Klauser authored
      Reviewed By: ldionne, #libc, Mordante
      
      Spies: arichardson, Mordante, libcxx-commits
      
      Differential Revision: https://reviews.llvm.org/D151119
      18c4695d
    • Felipe de Azevedo Piovezan's avatar
      [lldb][NFCI] Use llvm's libDebugInfo for DebugRanges · 70aad4ec
      Felipe de Azevedo Piovezan authored
      In an effort to unify the different dwarf parsers available in the codebase,
      this commit removes LLDB's custom parsing for the `.debug_ranges` DWARF section,
      instead calling into LLVM's parser.
      
      Subsequent work should look into unifying `llvm::DWARDebugRangeList` (whose
      entries are pairs of (start, end) addresses) with `lldb::DWARFRangeList` (whose
      entries are pairs of (start, length)). The lists themselves are also different
      data structures, but functionally equivalent.
      
      Depends on D150363
      
      Differential Revision: https://reviews.llvm.org/D150366
      70aad4ec
    • Tue Ly's avatar
      [libc][math] Implement double precision log1p correctly rounded to all rounding modes. · b91e78da
      Tue Ly authored
      Implement double precision log1p function correctly rounded to all
      rounding modes.
      
      **Performance**
      
        - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.93%.
        - Benchmarks with `./perf.sh` tool from the CORE-MATH project, unit is (CPU clocks / call).
        - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log1p
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 39.792 + 1.011 clc/call; Median-Min = 0.940 clc/call; Max = 41.373 clc/call;
      
      -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 87.285 + 1.135 clc/call; Median-Min = 1.299 clc/call; Max = 89.715 clc/call;
      
      -- System LIBC reciprocal throughput --
      [####################] 100 %
      Ntrial = 20 ; Min = 20.666 + 0.123 clc/call; Median-Min = 0.125 clc/call; Max = 20.828 clc/call;
      
      -- LIBC reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 20.928 + 0.771 clc/call; Median-Min = 0.725 clc/call; Max = 22.767 clc/call;
      
      -- LIBC reciprocal throughput -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 31.461 + 0.528 clc/call; Median-Min = 0.602 clc/call; Max = 36.809 clc/call;
      
      ```
        - Latency from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log1p --latency
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 77.875 + 0.062 clc/call; Median-Min = 0.051 clc/call; Max = 78.003 clc/call;
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 101.958 + 1.202 clc/call; Median-Min = 1.325 clc/call; Max = 104.452 clc/call;
      
      -- System LIBC latency --
      [####################] 100 %
      Ntrial = 20 ; Min = 60.581 + 1.443 clc/call; Median-Min = 1.611 clc/call; Max = 62.285 clc/call;
      
      -- LIBC latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 48.817 + 1.108 clc/call; Median-Min = 1.300 clc/call; Max = 50.282 clc/call;
      
      -- LIBC latency -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 61.121 + 0.599 clc/call; Median-Min = 0.761 clc/call; Max = 62.020 clc/call;
      ```
        - Accurate pass latency:
      ```
      $ ./perf.sh log1p --latency --simple_stat
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      760.444
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      827.880
      
      -- LIBC latency -- with FMA
      711.837
      
      -- LIBC latency -- without FMA
      764.317
      ```
      
      Reviewed By: zimmermann6
      
      Differential Revision: https://reviews.llvm.org/D151049
      b91e78da
    • Nikita Popov's avatar
      [InstCombine] Add droppable users back to worklist (NFCI) · 18a5bd7a
      Nikita Popov authored
      When sinking and users are dropped, add the using instructions
      to the worklist, as they can likely be removed as well.
      
      This should be NFC apart from worklist order effects.
      18a5bd7a
    • Jean Perier's avatar
      [flang][NFC] Move Array constructor inlined temp management into a utility · 9ac452b2
      Jean Perier authored
      This patch moves the counter and storage management part of the array
      constructor inlined temporary strategy into its own utility so that it
      can be reused for the simple cases of temporary creations inside WHERE
      and FORALL.
      
      It actually fixes a bug where the counter first value  used for addressing
      was "2" leading to read/write after the allocated storage... It seems
      I ran the tests end-to-end without the HLFIR flag when previously testing
      this. So this may clear some segfaults.
      
      Differential Revision: https://reviews.llvm.org/D151106
      9ac452b2
    • Tom Eccles's avatar
      [flang] use greedy mlir driver for stack arrays pass · 74c2ec50
      Tom Eccles authored
      In upstream mlir, the dialect conversion infrastructure is used for
      lowering from one dialect to another: the passes are of the form
      XToYPass. Whereas, transformations within the same dialect tend to use
      applyPatternsAndFoldGreedily.
      
      In this case, the full complexity of applyPatternsAndFoldGreedily isn't
      needed so we can get away with the simpler applyOpPatternsAndFold.
      
      This change was suggested by @jeanPerier
      
      Differential Revision: https://reviews.llvm.org/D150853
      74c2ec50
    • Tue Ly's avatar
      [libc][math] Implement double precision log2 function correctly rounded to all rounding modes. · 111d2748
      Tue Ly authored
      Implement double precision log2 function correctly rounded to all
      rounding modes.
      
      See https://reviews.llvm.org/D150014 for a more detail description of the algorithm.
      
      **Performance**
      
        - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.91%.
      
        - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log2
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 15.458 + 0.204 clc/call; Median-Min = 0.224 clc/call; Max = 15.867 clc/call;
      
      -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 23.711 + 0.524 clc/call; Median-Min = 0.443 clc/call; Max = 25.307 clc/call;
      
      -- System LIBC reciprocal throughput --
      [####################] 100 %
      Ntrial = 20 ; Min = 14.807 + 0.199 clc/call; Median-Min = 0.211 clc/call; Max = 15.137 clc/call;
      
      -- LIBC reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 17.666 + 0.274 clc/call; Median-Min = 0.298 clc/call; Max = 18.531 clc/call;
      
      -- LIBC reciprocal throughput -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 26.534 + 0.418 clc/call; Median-Min = 0.462 clc/call; Max = 27.327 clc/call;
      
      ```
        - Latency from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log2 --latency
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 46.048 + 1.643 clc/call; Median-Min = 1.694 clc/call; Max = 48.018 clc/call;
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 62.333 + 0.138 clc/call; Median-Min = 0.119 clc/call; Max = 62.583 clc/call;
      
      -- System LIBC latency --
      [####################] 100 %
      Ntrial = 20 ; Min = 45.206 + 1.503 clc/call; Median-Min = 1.467 clc/call; Max = 47.229 clc/call;
      
      -- LIBC latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 43.042 + 0.454 clc/call; Median-Min = 0.484 clc/call; Max = 43.912 clc/call;
      
      -- LIBC latency -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 57.016 + 1.636 clc/call; Median-Min = 1.655 clc/call; Max = 58.816 clc/call;
      ```
        - Accurate pass latency:
      ```
      $ ./perf.sh log2 --latency --simple_stat
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      177.632
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      231.332
      
      -- LIBC latency -- with FMA
      459.751
      
      -- LIBC latency -- without FMA
      463.850
      ```
      
      Reviewed By: zimmermann6
      
      Differential Revision: https://reviews.llvm.org/D150374
      111d2748
    • Krzysztof Parzyszek's avatar
      [Hexagon] Fix safety check in moving instructions in HVC::AlignVectors · 34c7f2ac
      Krzysztof Parzyszek authored
      A prior commit accidentally affected a safety check allowing aliased memory
      instructions to be moved across one another.
      34c7f2ac
    • Manna, Soumi's avatar
      [NFC][CLANG] Fix static code analyzer concerns · 64e9ba70
      Manna, Soumi authored
      Reported by Static Code Analyzer Tool, Coverity:
      
      Dereference null return value
      
      Inside "ExprConstant.cpp" file, in <unnamed>::RecordExprEvaluator::VisitCXXStdInitializerListExpr(clang::CXXStdInitializerListExpr const *): Return value of function which returns null is dereferenced without checking.
      
        bool RecordExprEvaluator::VisitCXXStdInitializerListExpr(
         const CXXStdInitializerListExpr *E) {
             // returned_null: getAsConstantArrayType returns nullptr (checked 81 out of 93 times).
             //var_assigned: Assigning: ArrayType = nullptr return value from getAsConstantArrayType.
          const ConstantArrayType *ArrayType =
             Info.Ctx.getAsConstantArrayType(E->getSubExpr()->getType());
          LValue Array;
          //Condition !EvaluateLValue(E->getSubExpr(), Array, this->Info, false), taking false branch.
          if (!EvaluateLValue(E->getSubExpr(), Array, Info))
           return false;
      
          // Get a pointer to the first element of the array.
      
          //Dereference null return value (NULL_RETURNS)
          //dereference: Dereferencing a pointer that might be nullptr ArrayType when calling addArray.
          Array.addArray(Info, E, ArrayType);
      
      This patch adds an assert for unexpected type for array initializer.
      
      Reviewed By: erichkeane
      
      Differential Revision: https://reviews.llvm.org/D151040
      64e9ba70
    • Kadir Cetinkaya's avatar
    • Nikita Popov's avatar
      [InstCombine] Fix worklist management in select value equiv fold (NFCI) · 2938f9b4
      Nikita Popov authored
      Requeue the modified instruction.
      
      This should be NFC apart from worklist order effects.
      2938f9b4
    • Tue Ly's avatar
      [libc][math] Implement double precision log function correctly rounded to all rounding modes. · a68bbf42
      Tue Ly authored
      Implement double precision log function correctly rounded to all
      rounding modes.
      
      See https://reviews.llvm.org/D150014 for a more detail description of the algorithm.
      
      **Performance**
      
        - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.93%.
      
        - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 17.465 + 0.596 clc/call; Median-Min = 0.602 clc/call; Max = 18.389 clc/call;
      
      -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 54.961 + 2.606 clc/call; Median-Min = 2.180 clc/call; Max = 59.583 clc/call;
      
      -- System LIBC reciprocal throughput --
      [####################] 100 %
      Ntrial = 20 ; Min = 12.608 + 0.276 clc/call; Median-Min = 0.359 clc/call; Max = 13.147 clc/call;
      
      -- LIBC reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 20.952 + 0.468 clc/call; Median-Min = 0.602 clc/call; Max = 21.881 clc/call;
      
      -- LIBC reciprocal throughput -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 18.569 + 0.552 clc/call; Median-Min = 0.601 clc/call; Max = 19.259 clc/call;
      
      ```
        - Latency from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log --latency
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 48.431 + 0.699 clc/call; Median-Min = 0.073 clc/call; Max = 51.269 clc/call;
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 64.865 + 3.235 clc/call; Median-Min = 3.475 clc/call; Max = 71.788 clc/call;
      
      -- System LIBC latency --
      [####################] 100 %
      Ntrial = 20 ; Min = 42.151 + 2.090 clc/call; Median-Min = 2.270 clc/call; Max = 44.773 clc/call;
      
      -- LIBC latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 35.266 + 0.479 clc/call; Median-Min = 0.373 clc/call; Max = 36.798 clc/call;
      
      -- LIBC latency -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 48.518 + 0.484 clc/call; Median-Min = 0.500 clc/call; Max = 49.896 clc/call;
      ```
        - Accurate pass latency:
      ```
      $ ./perf.sh log --latency --simple_stat
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      598.306
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      632.925
      
      -- LIBC latency -- with FMA
      455.632
      
      -- LIBC latency -- without FMA
      488.564
      ```
      
      Reviewed By: zimmermann6
      
      Differential Revision: https://reviews.llvm.org/D150131
      a68bbf42
    • Manna, Soumi's avatar
      [NFC][Clang] Fix Coverity bug with dereference null return value in... · cc6a6c48
      Manna, Soumi authored
      [NFC][Clang] Fix Coverity bug with dereference null return value in clang::CodeGen::CodeGenFunction::EmitOMPArraySectionExpr()
      
      Reported by Coverity:
      
      Inside  "CGExpr.cpp" file, in clang::CodeGen::CodeGenFunction::EmitOMPArraySectionExpr(clang::OMPArraySectionExpr const *, bool): Return value of function which returns null is dereferenced without checking.
      
          } else {
        	//returned_null: getAsConstantArrayType returns nullptr (checked 83 out of 95 times).
        	// var_assigned: Assigning: CAT = nullptr return value from getAsConstantArrayType.
            auto *CAT = C.getAsConstantArrayType(ArrayTy);
        	//identity_transfer: Member function call CAT->getSize() returns an offset off CAT (this).
      
           // Dereference null return value (NULL_RETURNS)
           //dereference: Dereferencing a pointer that might be nullptr CAT->getSize() when calling APInt.
           ConstLength = CAT->getSize();
          }
      
      This patch adds an assert to resolve the bug.
      
      Reviewed By: erichkeane
      
      Differential Revision: https://reviews.llvm.org/D151137
      cc6a6c48
    • Nikita Popov's avatar
      [InstCombine] Regenerate test checks (NFC) · 42d97427
      Nikita Popov authored
      42d97427
    • Nikita Popov's avatar
      [InstCombine] Fix worklist management in replaceGEPIdxWithZero() fold (NFCI) · 08915751
      Nikita Popov authored
      Make sure the old load/store operand is queued for DCE.
      
      This should be NFC apart from worklist order effects.
      08915751
    • Joseph Huber's avatar
      [libc][AMDGPU] Disable the AMDGPU backend's ctor/dtor lowering for libc · ad00a3db
      Joseph Huber authored
      The AMDGPU backend has a built-in pass to lower constructors. We do this
      manually in the `start.cpp` implementation so we can disable this to
      keep the binaries smaller.
      
      Differential Revision: https://reviews.llvm.org/D151213
      ad00a3db
    • Jonathan Peyton's avatar
      [OpenMP] Insert missing variable update inside loop · d67c91b5
      Jonathan Peyton authored
      While loop within task priority code did not have necessary update of
      variable which could lead to hangs if two threads collided when both
      attempted to execute the compare_and_exchange.
      
      Fixes: https://github.com/llvm/llvm-project/issues/62867
      Differential Revision: https://reviews.llvm.org/D151138
      d67c91b5
    • Tue Ly's avatar
      [libc][math] Make log10 correctly rounded for non-FMA targets and improve itsperformance. · a0c92a38
      Tue Ly authored
      Make log10 correctly rounded for non-FMA targets and improve its
      performance.
      
      Implemented fast pass and accurate pass:
      
      **Fast Pass**:
      
        - Range reduction step 0: Extract exponent and mantissa
      ```
        x = 2^(e_x) * m_x
      ```
        - Range reduction step 1: Use lookup tables of size 2^7 = 128 to reduce the argument to:
      ```
         -2^-8 <= v = r * m_x - 1 < 2^-7
        where r = 2^-8 * ceil( 2^8 * (1 - 2^-8) / (1 + k * 2^-7) )
        and k = trunc( (m_x - 1) * 2^7 )
      ```
        - Polynomial approximation: approximate `log(1 + v)` by a degree-7 polynomial generated by Sollya with:
      ```
       > P = fpminimax((log(1 + x) - x)/x^2, 5, [|D...|], [-2^-8, 2^-7]);
      ```
        - Combine the results:
      ```
        log10(x) ~ ( e_x * log(2) - log(r) + v + v^2 * P(v) ) * log10(e)
      ```
        - Perform additive Ziv's test with errors bounded by `P_ERR * v^2`.  Return the result if Ziv's test passed.
      
      **Accurate Pass**:
      
        - Take `e_x`, `v`, and the lookup table index from the range reduction step of fast pass.
        - Perform 3 more range reduction steps:
          - Range reduction step 2: Use look-up tables of size 193 to reduce the argument to `[-0x1.3ffcp-15, 0x1.3e3dp-15]`
      ```
         v2 = r2 * (1 + v) - 1 = (1 + s2) * (1 + v) - 1 = s2 + v + s2 * v
        where r2 = 2^-16 * round ( 2^16 / (1 + k * 2^-14) )
        and k = trunc( v * 2^14 + 0.5 ).
      ```
          - Range reduction step 3: Use look-up tables of size 161 to reduce the argument to `[-0x1.01928p-22 , 0x1p-22]`
      ```
         v3 = r3 * (1 + v2) - 1 = (1 + s3) * (1 + v2) - 1 = s3 + v2 + s3 * v2
        where r3 = 2^-21 * round ( 2^21 / (1 + k * 2^-21) )
        and k = trunc( v * 2^21 + 0.5 ).
      ```
          - Range reduction step 4: Use look-up tables of size 130 to reduce the argument to `[-0x1.0002143p-29 , 0x1p-29]`
      ```
         v4 = r4 * (1 + v3) - 1 = (1 + s4) * (1 + v3) - 1 = s4 + v3 + s4 * v3
        where r4 = 2^-28 * round ( 2^28 / (1 + k * 2^-28) )
        and k = trunc( v * 2^28 + 0.5 ).
      ```
        - Polynomial approximation: approximate `log10(1 + v4)` by a degree-4 minimax polynomial generated by Sollya with:
      ```
        > P = fpminimax(log10(1 + x)/x, 3, [|128...|], [-0x1.0002143p-29 , 0x1p-29]);
      ```
        - Combine the results:
      ```
        log10(x) ~ e_x * log10(2) - log10(r) - log10(r2) - log10(r3) - log10(r4) + v * P(v)
      ```
        - The combined results are computed using floating points of 128-bit precision.
      
      **Performance**
      
        - For `0.5 <= x <= 2`, the fast pass hitting rate is about 99.92%.
      
        - Reciprocal throughput from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log10
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 20.402 + 0.589 clc/call; Median-Min = 0.277 clc/call; Max = 22.752 clc/call;
      
      -- CORE-MATH reciprocal throughput -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 75.797 + 3.317 clc/call; Median-Min = 3.407 clc/call; Max = 79.371 clc/call;
      
      -- System LIBC reciprocal throughput --
      [####################] 100 %
      Ntrial = 20 ; Min = 22.668 + 0.184 clc/call; Median-Min = 0.181 clc/call; Max = 23.205 clc/call;
      
      -- LIBC reciprocal throughput -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 25.977 + 0.183 clc/call; Median-Min = 0.138 clc/call; Max = 26.283 clc/call;
      
      -- LIBC reciprocal throughput -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 22.140 + 0.980 clc/call; Median-Min = 0.853 clc/call; Max = 23.790 clc/call;
      
      ```
        - Latency from CORE-MATH's perf tool on Ryzen 5900X:
      ```
      $ ./perf.sh log10 --latency
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 54.613 + 0.357 clc/call; Median-Min = 0.287 clc/call; Max = 55.701 clc/call;
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      [####################] 100 %
      Ntrial = 20 ; Min = 79.681 + 0.482 clc/call; Median-Min = 0.294 clc/call; Max = 81.604 clc/call;
      
      -- System LIBC latency --
      [####################] 100 %
      Ntrial = 20 ; Min = 61.532 + 0.208 clc/call; Median-Min = 0.199 clc/call; Max = 62.256 clc/call;
      
      -- LIBC latency -- with FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 41.510 + 0.205 clc/call; Median-Min = 0.244 clc/call; Max = 41.867 clc/call;
      
      -- LIBC latency -- without FMA
      [####################] 100 %
      Ntrial = 20 ; Min = 55.669 + 0.240 clc/call; Median-Min = 0.280 clc/call; Max = 56.056 clc/call;
      ```
        - Accurate pass latency:
      ```
      $ ./perf.sh log10 --latency --simple_stat
      GNU libc version: 2.35
      GNU libc release: stable
      
      -- CORE-MATH latency -- with FMA
      640.688
      
      -- CORE-MATH latency -- without FMA (-march=x86-64-v2)
      667.354
      
      -- LIBC latency -- with FMA
      495.593
      
      -- LIBC latency -- without FMA
      504.143
      ```
      
      Reviewed By: zimmermann6
      
      Differential Revision: https://reviews.llvm.org/D150014
      a0c92a38
    • Manna, Soumi's avatar
      [NFC][CLANG] Fix static code analyzer concerns with dereference null return value · 7586aeab
      Manna, Soumi authored
      Reported by Static Code Analyzer Tool, Coverity:
      
      Inside "SemaExprMember.cpp" file, in clang::Sema::BuildMemberReferenceExpr(clang::Expr *, clang::QualType, clang::SourceLocation, bool, clang::CXXScopeSpec &, clang::SourceLocation, clang::NamedDecl *, clang::DeclarationNameInfo const &, clang::TemplateArgumentListInfo const *, clang::Scope const *, clang::Sema::ActOnMemberAccessExtraArgs *): Return value of function which returns null is dereferenced without checking
      
        //Condition !Base, taking true branch.
        if (!Base) {
          TypoExpr *TE = nullptr;
          QualType RecordTy = BaseType;
      
           //Condition IsArrow, taking true branch.
           if (IsArrow) RecordTy = RecordTy->castAs<PointerType>()->getPointeeType();
          	//returned_null: getAs returns nullptr (checked 279 out of 294 times).
          	//Condition TemplateArgs != NULL, taking true branch.
      
           //Dereference null return value (NULL_RETURNS)
           //dereference: Dereferencing a pointer that might be nullptr RecordTy->getAs() when calling LookupMemberExprInRecord.
           if (LookupMemberExprInRecord(
                 *this, R, nullptr, RecordTy->getAs<RecordType>(), OpLoc, IsArrow,
                 SS, TemplateArgs != nullptr, TemplateKWLoc, TE))
              return ExprError();
           if (TE)
             return TE;
      
      This patch uses castAs instead of getAs which will assert if the type doesn't match.
      
      Reviewed By: erichkeane
      
      Differential Revision: https://reviews.llvm.org/D151130
      7586aeab
    • Nikita Popov's avatar
      [Driver] Try to fix linux-ld.c test with DEFAULT_LINKER set (NFC) · 28776d50
      Nikita Popov authored
      The test fails on the clang-ppc64le-rhel build bot, which has
      DEFAULT_LINKER set and an ld.lld binary in the LLVM build directory.
      28776d50
    • Joseph Huber's avatar
      [AMDGPU] Add an option to disable manual ctor / dtor lowering · 4a1236e0
      Joseph Huber authored
      Currently AMDGPU offers extra ctor / dtor lowering by emitting a kernel
      that can be called. It's possible to handle ctors and dtors using the
      standard method as shown in D149340's commit message. In which case we
      on't need these extra kernels as they won't be called. This patch simply
      adds a way to conditionally turn off this handling if we do not want to
      get extra kernels in the output.
      
      Unrelated, but we could convert this handling to an ODR function that simply
      calls the code in D149340 constructed via LLVM-IR. That would handle priority
      correctly and would then be correct if not run in LTO mode.
      
      Reviewed By: yaxunl
      
      Differential Revision: https://reviews.llvm.org/D150565
      4a1236e0
    • Fangrui Song's avatar
      [ubsan][test] Remove --check-prefix=UNIQUE for x86_64-apple from... · 39ccd573
      Fangrui Song authored
      [ubsan][test] Remove --check-prefix=UNIQUE for x86_64-apple from e215996a
      
      After switching to use a type hash instead of possibly-non-unique typeinfo
      objects, we no longer have unique/non-unique distinction.
      39ccd573
    • Nikita Popov's avatar
      [InstCombine] Remove dead extractelements (NFCI) · 4b832086
      Nikita Popov authored
      Directly remove these dead extractelement instructions, rather than
      leaving them for the next InstCombine iteration to clean up.
      
      Should be mostly NFC, apart from worklist order differences.
      4b832086
    • Matthias Springer's avatar
      [mlir][bufferization] Fix bug in findValueInReverseUseDefChain · aa909483
      Matthias Springer authored
      This bug was recently introduced in D143927 and manifests as a dominance violation.
      
      Differential Revision: https://reviews.llvm.org/D151077
      aa909483
    • Aaron Ballman's avatar
      Silence switch statement contains 'default' but no 'case' labels warning; NFC · 846bde48
      Aaron Ballman authored
      These are showing up in MSVC builds.
      846bde48
    • Dinar Temirbulatov's avatar
      [AArch64][LV] Disable maximising bandwidth for streaming compatible sve · 7489301c
      Dinar Temirbulatov authored
      Fixing last commit by adding actual change to AArch64TargetTransformInfo.cpp
      
      Differential Revision: https://reviews.llvm.org/D150336
      7489301c
    • Thomas Preud'homme's avatar
      Add StringRef::consumeInteger(APInt) · dd00421c
      Thomas Preud'homme authored
      This will be required to allow arbitrary precision support to
      FileCheck's numeric variables and expressions. Note: as per
      getAsInteger(), this does not support negative value. If there is
      interest for that it can be added in a separate patch.
      
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D150878
      dd00421c
    • Dinar Temirbulatov's avatar
      [AArch64][LV] Disable maximising bandwidth for streaming compatible sve · 1ff828c6
      Dinar Temirbulatov authored
      We noticed some runtime performance improvements by disabling maximising
      bandwidth for streaming compatible sve.
      
      Differential Revision: https://reviews.llvm.org/D150336
      1ff828c6
    • Thomas Preud'homme's avatar
      Turn unreachable error into assert · 13eb298d
      Thomas Preud'homme authored
      Function valueFromStringRepr() throws an error on missing 0x prefix when
      parsing a number string into a value. However, getWildcardRegex() already
      ensures that only text with the 0x prefix will match and be parsed,
      making that error throwing code dead code. This commit turn the code
      into an assert and remove the unit tests exercising that test
      accordingly.
      
      Reviewed By: jhenderson
      
      Differential Revision: https://reviews.llvm.org/D150797
      13eb298d
    • Krasimir Georgiev's avatar
      c37ced7d