1. Jul 10, 2020
    • Florian Hahn's avatar
      [DomTreeUpdater] Use const auto * when iterating over pointers (NFC). · ec00aa99
      Florian Hahn authored
      This silences the warning below:
      
      llvm-project/llvm/lib/Analysis/DomTreeUpdater.cpp:510:20: warning: loop variable 'BB' is always a copy because the range of type 'const SmallPtrSet<llvm::BasicBlock *, 8>' does not return a reference [-Wrange-loop-analysis]
        for (const auto &BB : DeletedBBs) {
                         ^
      llvm-project/llvm/lib/Analysis/DomTreeUpdater.cpp:510:8: note: use non-reference type 'llvm::BasicBlock *'
        for (const auto &BB : DeletedBBs) {
             ^~~~~~~~~~~~~~~~
      1 warning generated.
      ec00aa99
    • Florian Hahn's avatar
      [ARM] Add test with tcreturn and debug value. · eb5c7f6b
      Florian Hahn authored
      In the attached test case, a non-terminator instruction (DBG_VALUE) is
      inserted after a terminator, producing an invalid MBB.
      eb5c7f6b
    • Sanjay Patel's avatar
      d84b4e16
    • Sanjay Patel's avatar
      02fec9d2
    • Nicolas Vasilache's avatar
      [mlir][Vector] Add ExtractOp folding when fed by a TransposeOp · a490d387
      Nicolas Vasilache authored
      TransposeOp are often followed by ExtractOp.
      In certain cases however, it is unnecessary (and even detrimental) to lower a TransposeOp to either a flat transpose (llvm.matrix intrinsics) or to unrolled scalar insert / extract chains.
      
      Providing foldings of ExtractOp mitigates some of the unnecessary complexity.
      
      Differential revision: https://reviews.llvm.org/D83487
      a490d387
    • Joel E. Denny's avatar
      [FileCheck] Implement -dump-input-filter · 9fd4b5fa
      Joel E. Denny authored
      This makes the input dump filtering implemented by D82203 more
      configurable.  D82203 enables filtering out everything but the initial
      input lines of error diagnostics (plus some context).  This patch
      enables including any line with any kind of annotation.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D83097
      9fd4b5fa
    • Joel E. Denny's avatar
      [FileCheck] In input dump, elide only if ellipsis is shorter · 77b6ddf1
      Joel E. Denny authored
      For example, given `-dump-input-context=3 -vv`, the following now
      shows more leading context for the error than requested because a
      leading ellipsis would occupy the same number of lines as it would
      elide:
      
      ```
      <<<<<<
               1: foo6
               2: foo5
               3: foo4
               4: foo3
               5: foo2
               6: foo1
               7: hello world
      check:1     ^~~~~
      check:2           X~~~~ error: no match found
               8: foo1
      check:2     ~~~~
               9: foo2
      check:2     ~~~~
              10: foo3
      check:2     ~~~~
               .
               .
               .
      >>>>>>
      ```
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D83526
      77b6ddf1
    • Joel E. Denny's avatar
      [FileCheck] Implement -dump-input-context · bce8fced
      Joel E. Denny authored
      This patch is motivated by discussions at each of:
      
      * <https://reviews.llvm.org/D81422>
      * <http://lists.llvm.org/pipermail/llvm-dev/2020-June/142369.html>
      
      When input is dumped as specified by `-dump-input=fail`, this patch
      filters the dump to show only input lines that are the starting lines
      of error diagnostics plus the number of contextual lines specified
      `-dump-input-context` (defaults to 5).
      
      When `-dump-input=always`, there might be not be any errors, so all
      input lines are printed, as without this patch.
      
      Here's some sample output with `-dump-input-context=3 -vv`:
      
      ```
      <<<<<<
                 .
                 .
                 .
                13: foo
                14: foo
                15: hello world
      check:1       ^~~~~~~~~~~
                16: foo
      check:2'0     X~~ error: no match found
                17: foo
      check:2'0     ~~~
                18: foo
      check:2'0     ~~~
                19: foo
      check:2'0     ~~~
                 .
                 .
                 .
                27: foo
      check:2'0     ~~~
                28: foo
      check:2'0     ~~~
                29: foo
      check:2'0     ~~~
                30: goodbye word
      check:2'0     ~~~~~~~~~~~~
      check:2'1     ?            possible intended match
                31: foo
      check:2'0     ~~~
                32: foo
      check:2'0     ~~~
                33: foo
      check:2'0     ~~~
                 .
                 .
                 .
      >>>>>>
      ```
      
      Reviewed By: mehdi_amini, arsenm, jhenderson, rsmith, SjoerdMeijer, Meinersbur, lattner
      
      Differential Revision: https://reviews.llvm.org/D82203
      bce8fced
    • Alexandre Ganea's avatar
      [PDB] Fix out-of-bounds acces when sorting GSI buckets · 23cd70d7
      Alexandre Ganea authored
      When building in Debug on Windows-MSVC after b7402edc, a lot of tests were failing because we were dereferencing an element past the end of HashRecords. This happened towards the end of the table, in unused slots.
      23cd70d7
    • Sam McCall's avatar
      [clangd] Update semanticTokens support to reflect latest LSP draft · 5fea54bc
      Sam McCall authored
      Summary: Mostly a few methods and message names have been renamed.
      
      Reviewers: hokein
      
      Subscribers: ilya-biryukov, MaskRay, jkorous, arphaman, kadircet, usaxena95, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D83556
      5fea54bc
    • Roman Lebedev's avatar
      Reland "[InstCombine] Lower infinite combine loop detection thresholds"" · 7103c875
      Roman Lebedev authored
      This relands commit cd7f8051 that was
      reverted since lower threshold have successfully found an issue.
      Now that the issue is fixed, let's wait until the next one is reported.
      
      This reverts commit caa423ee.
      7103c875
    • Roman Lebedev's avatar
      [InstCombine] After merging store into successor, queue prev. store to be visited (PR46661) · 2655a70a
      Roman Lebedev authored
      We can happen to have a situation with many stores eligible for transform,
      but due to our visitation order (top to bottom), when we have processed
      the first eligible instruction, we would not try to reprocess the previous
      instructions that are now also eligible.
      
      So after we've successfully merged a store that was second-to-last instruction
      into successor, if the now-second-to-last instruction is also a such store
      that is eligible, add it to worklist to be revisited.
      
      Fixes https://bugs.llvm.org/show_bug.cgi?id=46661
      2655a70a
    • Roman Lebedev's avatar
      [NFCI][InstCombine] PR46661: multiple stores eligible for merging into successor - worklist issue · ef0ecb7b
      Roman Lebedev authored
      The testcase should pass with a single instcombine iteration.
      ef0ecb7b
    • Kevin P. Neal's avatar
      [FPEnv][Clang][Driver] Disable constrained floating point on targets lacking support." · 523a8513
      Kevin P. Neal authored
      Use the new -fexperimental-strict-floating-point flag in more cases to
      fix the arm and aarch64 bots.
      
      Differential Revision: https://reviews.llvm.org/D80952
      523a8513
    • Nicolas Vasilache's avatar
      [mlir][Linalg] Generalize Vectorization of Linalg contractions · 56c638b5
      Nicolas Vasilache authored
      This revision adds support for vectorizing named and generic contraction ops to vector.contract. Cases in which the memref is 0-D are special cased to emit std.load/std.store instead of vector.transfer. Relevant tests are added.
      
      Differential revision: https://reviews.llvm.org/D83307
      56c638b5
    • Haojian Wu's avatar
    • Nicolas Vasilache's avatar
      [mlir][Vector] Fold chains of ExtractOp · 22c8a08f
      Nicolas Vasilache authored
      This revision adds folding to ExtractOp by simply concatenating the position attributes.
      22c8a08f
    • Daniel Grumberg's avatar
      0555db0a
    • Kevin P. Neal's avatar
      Reland "[FPEnv][Clang][Driver] Disable constrained floating point on targets lacking support." · d4ce862f
      Kevin P. Neal authored
      We currently have strict floating point/constrained floating point enabled
      for all targets. Constrained SDAG nodes get converted to the regular ones
      before reaching the target layer. In theory this should be fine.
      
      However, the changes are exposed to users through multiple clang options
      already in use in the field, and the changes are _completely_ _untested_
      on almost all of our targets. Bugs have already been found, like
      "https://bugs.llvm.org/show_bug.cgi?id=45274".
      
      This patch disables constrained floating point options in clang everywhere
      except X86 and SystemZ. A warning will be printed when this happens.
      
      Use the new -fexperimental-strict-floating-point flag to force allowing
      strict floating point on hosts that aren't already marked as supporting
      it (X86 and SystemZ).
      
      Differential Revision: https://reviews.llvm.org/D80952
      d4ce862f
    • David Green's avatar
      Revert "[BasicAA] Enable -basic-aa-recphi by default" · e1135b48
      David Green authored
      This reverts commit af839a96.
      
      Some issues appear to be being caused by this. Reverting whilst we
      investigate.
      e1135b48
    • Sam McCall's avatar
      [clangd] Config: If.PathExclude · 86f13134
      Sam McCall authored
      Reviewers: hokein
      
      Subscribers: ilya-biryukov, MaskRay, jkorous, arphaman, kadircet, usaxena95, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D83511
      86f13134
    • Victor Huang's avatar
      [PowerPC] Implement R_PPC64_REL24_NOTOC calls, callee also has no TOC · 118366dc
      Victor Huang authored
      The PC Relative code allows for calls that are marked with the relocation
      R_PPC64_REL24_NOTOC. This indicates that the caller does not have a valid TOC
      pointer in R2 and does not require R2 to be restored after the call.
      
      This patch is added to support local calls to callees tha also do not have a TOC.
      
      Reviewed By: sfertile, MaskRay, stefanp
      
      Differential Revision: https://reviews.llvm.org/D82816
      118366dc
    • Ulrich Weigand's avatar
      [ABI] Handle C++20 [[no_unique_address]] attribute · 4c5a93bd
      Ulrich Weigand authored
      Many platform ABIs have special support for passing aggregates that
      either just contain a single member of floatint-point type, or else
      a homogeneous set of members of the same floating-point type.
      
      When making this determination, any extra "empty" members of the
      aggregate type will typically be ignored.  However, in C++ (at least
      in all prior versions), no data member would actually count as empty,
      even if it's type is an empty record -- it would still be considered
      to take up at least one byte of space, and therefore make those ABI
      special cases not apply.
      
      This is now changing in C++20, which introduced the [[no_unique_address]]
      attribute.  Members of empty record type, if they also carry this
      attribute, now do *not* take up any space in the type, and therefore
      the ABI special cases for single-element or homogeneous aggregates
      should apply.
      
      The C++ Itanium ABI has been updated accordingly, and GCC 10 has
      added support for this new case.  This patch now adds support to
      LLVM.  This is cross-platform; it affects all platforms that use
      the single-element or homogeneous aggregate ABI special case and
      implement this using any of the following common subroutines
      in lib/CodeGen/TargetInfo.cpp:
        isEmptyField
        isEmptyRecord
        isSingleElementStruct
        isHomogeneousAggregate
      4c5a93bd
    • Simon Pilgrim's avatar
      DomTreeUpdater::dump() - use const auto& iterator in for-range-loop. · b69e0f67
      Simon Pilgrim authored
      Avoids unnecessary copies and silences clang tidy warning.
      b69e0f67
    • Nathan James's avatar
    • Simon Pilgrim's avatar
      StackSafetyAnalysis.cpp - pass ConstantRange arg as const reference. · 9ce98312
      Simon Pilgrim authored
      Avoids unnecessary copies and silences clang tidy warning - we do this in most places, there are just a few that were missed.
      9ce98312
    • Simon Pilgrim's avatar
    • dstuttar's avatar
      [NFC] Change isFPPredicate comparison to ignore lower bound · 69a89b54
      dstuttar authored
      Summary:
      Since changing the Predicate to be an unsigned enum, the lower bound check for
      isFPPredicate no longer needs to check the lower bound, since
      it will always evaluate to true.
      
      Also fixed a similar issue in SIISelLowering.cpp by removing the need for
      comparing to FIRST and LAST predicates
      
      Added an assert to the isFPPredicate comparison to flag if the
      FIRST_FCMP_PREDICATE is ever changed to anything other than 0, in which case the
      logic will break.
      
      Without this change warnings are generated in VS.
      
      Change-Id: I358f0daf28c0628c7bda8ad4cab4e1757b761bab
      
      Subscribers: arsenm, jvesely, nhaehnle, hiraditya, kerbowa, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D83540
      69a89b54
    • Paul Walker's avatar
      [SVE] Code generation for fixed length vector truncates. · f78e6a30
      Paul Walker authored
       Lower fixed length vector truncates to a sequence of SVE UZP1 instructions.
      
      Differential Revision: https://reviews.llvm.org/D83395
      f78e6a30
    • Pavel Labath's avatar
      [lldb/pecoff] Use a different llvm createBinary overload for parsing · d372a8e8
      Pavel Labath authored
      Change the code the use the version which accepts a memory buffer,
      instead of the one taking a file name.
      
      This ensures we are not loading the file into memory twice
      (ObjectFilePECOFF also loads a copy), reducing our memory footprint, as
      well as enabling additional goodies in the future, like being able to
      open files which don't exist on disk (D83512).
      d372a8e8
    • Haojian Wu's avatar
      [clang-tidy] More strict on matching the standard memset function in memset-usage check. · 5f41ca48
      Haojian Wu authored
      The check assumed the matched function call has 3 arguments, but the
      matcher didn't guaranteed that.
      
      Differential Revision: https://reviews.llvm.org/D83301
      5f41ca48
    • Florian Hahn's avatar
      [LV] Pick vector loop body as insert point for SCEV expansion. · 264ab1e2
      Florian Hahn authored
      Currently the DomTree is not kept up to date for additional blocks
      generated in the vector loop, for example when vectorizing with
      predication. SCEVExpander relies on dominance checks when looking for
      existing instructions to re-use and in some cases that can lead to the
      expander picking instructions that do not actually dominate their insert
      point (e.g. as in PR46525).
      
      Unfortunately keeping the DT up-to-date is a bit tricky, because the CFG
      is only patched up after generating code for a block. For now, we can
      just use the vector loop header, as this ensures the inserted
      instructions dominate all uses in the vector loop. There should be no
      noticeable impact on the generated code, as other passes should sink
      those instructions, if profitable.
      
      Fixes PR46525.
      
      Reviewers: Ayal, gilr, mkazantsev, dmgreen
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D83288
      264ab1e2
    • Mirko Brkusanin's avatar
      [AMDGPU][GlobalISel] Fix G_AMDGPU_TBUFFER_STORE_FORMAT mapping · cf40db21
      Mirko Brkusanin authored
      Add missing mappings and tablegen definitions for TBUFFER_STORE_FORMAT.
      
      Differential Revision: https://reviews.llvm.org/D83240
      cf40db21
    • Simon Pilgrim's avatar
      extractConstantWithoutWrapping - use const APInt& returned by SCEVConstant::getAPInt() · 9a3e8b11
      Simon Pilgrim authored
      Avoids unnecessary APInt copies and silences clang tidy warning.
      9a3e8b11
    • Vitaly Buka's avatar
      c06417b2
    • Simon Pilgrim's avatar
      [X86][AVX] Attempt to fold PACK(SHUFFLE(X,Y),SHUFFLE(X,Y)) -> SHUFFLE(PACK(X,Y)). · 77133cc1
      Simon Pilgrim authored
      Truncations lowered as shuffles of multiple (concatenated) vectors often leave us with lane-crossing shuffles that feed a PACKSS/PACKUS, if both shuffles are fed from the same 2 vector sources, then we can PACK the sources directly and shuffle the result instead.
      
      This is currently limited to whole i128 lanes in a 256-bit vector, but we can extend this if the need arises (but I'm not seeing many examples in real world code).
      77133cc1
    • Valeriy Savchenko's avatar
      [analyzer][tests] Fix zip unpacking · 00997d1c
      Valeriy Savchenko authored
      Differential Revision: https://reviews.llvm.org/D83374
      00997d1c
    • Valeriy Savchenko's avatar
      9c7ff0a4
    • Valeriy Savchenko's avatar
    • Danila Kutenin's avatar
      [builtins] Optimize udivmodti4 for many platforms. · 68c011aa
      Danila Kutenin authored
      Summary:
      While benchmarking uint128 division we found out that it has huge latency for small divisors
      
      https://reviews.llvm.org/D83027
      
      ```
      Benchmark                                                   Time(ns)        CPU(ns)     Iterations
      --------------------------------------------------------------------------------------------------
      BM_DivideIntrinsic128UniformDivisor<unsigned __int128>            13.0           13.0     55000000
      BM_DivideIntrinsic128UniformDivisor<__int128>                     14.3           14.3     50000000
      BM_RemainderIntrinsic128UniformDivisor<unsigned __int128>         13.5           13.5     52000000
      BM_RemainderIntrinsic128UniformDivisor<__int128>                  14.1           14.1     50000000
      BM_DivideIntrinsic128SmallDivisor<unsigned __int128>             153            153        5000000
      BM_DivideIntrinsic128SmallDivisor<__int128>                      170            170        3000000
      BM_RemainderIntrinsic128SmallDivisor<unsigned __int128>          153            153        5000000
      BM_RemainderIntrinsic128SmallDivisor<__int128>                   155            155        5000000
      ```
      
      This patch suggests a more optimized version of the division:
      
      If the divisor is 64 bit, we can proceed with the divq instruction on x86 or constant multiplication mechanisms for other platforms. Once both divisor and dividend are not less than 2**64, we use branch free subtract algorithm, it has at most 64 cycles. After that our benchmarks improved significantly
      
      ```
      Benchmark                                                   Time(ns)        CPU(ns)     Iterations
      --------------------------------------------------------------------------------------------------
      BM_DivideIntrinsic128UniformDivisor<unsigned __int128>            11.0           11.0     64000000
      BM_DivideIntrinsic128UniformDivisor<__int128>                     13.8           13.8     51000000
      BM_RemainderIntrinsic128UniformDivisor<unsigned __int128>         11.6           11.6     61000000
      BM_RemainderIntrinsic128UniformDivisor<__int128>                  13.7           13.7     52000000
      BM_DivideIntrinsic128SmallDivisor<unsigned __int128>              27.1           27.1     26000000
      BM_DivideIntrinsic128SmallDivisor<__int128>                       29.4           29.4     24000000
      BM_RemainderIntrinsic128SmallDivisor<unsigned __int128>           27.9           27.8     26000000
      BM_RemainderIntrinsic128SmallDivisor<__int128>                    29.1           29.1     25000000
      ```
      
      If not using divq instrinsics, it is still much better
      
      ```
      Benchmark                                                   Time(ns)        CPU(ns)     Iterations
      --------------------------------------------------------------------------------------------------
      BM_DivideIntrinsic128UniformDivisor<unsigned __int128>            12.2           12.2     58000000
      BM_DivideIntrinsic128UniformDivisor<__int128>                     13.5           13.5     52000000
      BM_RemainderIntrinsic128UniformDivisor<unsigned __int128>         12.7           12.7     56000000
      BM_RemainderIntrinsic128UniformDivisor<__int128>                  13.7           13.7     51000000
      BM_DivideIntrinsic128SmallDivisor<unsigned __int128>              30.2           30.2     24000000
      BM_DivideIntrinsic128SmallDivisor<__int128>                       33.2           33.2     22000000
      BM_RemainderIntrinsic128SmallDivisor<unsigned __int128>           31.4           31.4     23000000
      BM_RemainderIntrinsic128SmallDivisor<__int128>                    33.8           33.8     21000000
      ```
      
      PowerPC benchmarks:
      
      Was
      ```
      BM_DivideIntrinsic128UniformDivisor<unsigned __int128>            22.3           22.3     32000000
      BM_DivideIntrinsic128UniformDivisor<__int128>                     23.8           23.8     30000000
      BM_RemainderIntrinsic128UniformDivisor<unsigned __int128>         22.5           22.5     32000000
      BM_RemainderIntrinsic128UniformDivisor<__int128>                  24.9           24.9     29000000
      BM_DivideIntrinsic128SmallDivisor<unsigned __int128>             394            394        2000000
      BM_DivideIntrinsic128SmallDivisor<__int128>                      397            397        2000000
      BM_RemainderIntrinsic128SmallDivisor<unsigned __int128>          399            399        2000000
      BM_RemainderIntrinsic128SmallDivisor<__int128>                   397            397        2000000
      ```
      
      With this patch
      ```
      BM_DivideIntrinsic128UniformDivisor<unsigned __int128>            21.7           21.7     33000000
      BM_DivideIntrinsic128UniformDivisor<__int128>                     23.0           23.0     31000000
      BM_RemainderIntrinsic128UniformDivisor<unsigned __int128>         21.9           21.9     33000000
      BM_RemainderIntrinsic128UniformDivisor<__int128>                  23.9           23.9     30000000
      BM_DivideIntrinsic128SmallDivisor<unsigned __int128>              32.7           32.6     23000000
      BM_DivideIntrinsic128SmallDivisor<__int128>                       33.4           33.4     21000000
      BM_RemainderIntrinsic128SmallDivisor<unsigned __int128>           31.1           31.1     22000000
      BM_RemainderIntrinsic128SmallDivisor<__int128>                    33.2           33.2     22000000
      ```
      
      My email: danilak@google.com, I don't have commit rights
      
      Reviewers: howard.hinnant, courbet, MaskRay
      
      Reviewed By: courbet
      
      Subscribers: steven.zhang, #sanitizers
      
      Tags: #sanitizers
      
      Differential Revision: https://reviews.llvm.org/D81809
      68c011aa