1. Mar 01, 2022
  2. Feb 28, 2022
    • Sanjay Patel's avatar
      2dc90eee
    • Erich Keane's avatar
      Clarify documentation of cpu_dispatch/cpu_specific · cd1489bb
      Erich Keane authored
      There has been some internal confusion lately as to how cpu_dispatch and
      cpu_specific dispatch to processors, so this patch clarifies the
      documentation to make it more clear that:
      
      1- Unlike ICC, we do not consider the vendor string (that is, an AMD
      processor might result in something other than generic)
      
      2- there are some processors that aren't really distinguishable thanks
      to the library limitation, so the variant being selected is unspecified.
      
      In reality, I believe the 'stable_sort' makes it so the 1st one that
      meets the requirements is the one that is selected (and that matches my
      experimented result), but I don't want to limit our implementation.
      cd1489bb
    • Alexey Bataev's avatar
      [SLP]Improve bottom-to-top reordering. · e4b96408
      Alexey Bataev authored
      Currently bottom-to-top reordering analysis counts orders of the
      operands and then adds natural order counts for the operand users. It is
      very conservative, this the user nodes themselves may require
      reordering. Patch improves bottom-to-top analysis by checking for the
      user nodes if they require/allows the reordring. If the user node must
      be reordered, has reused scalars, is an alternate op vectorization node,
      is a non-ordered gather node or may allow reordering because of the
      reordered operands, such node is considered as the node that allows
      reodring and is not counted as a node with the natural order.
      
      Differential Revision: https://reviews.llvm.org/D120492
      e4b96408
    • Dawid Jurczak's avatar
      [NFC][Lexer] Make Lexer::LangOpts const reference · a64d3c60
      Dawid Jurczak authored
      This change can be seen as code cleanup but motivation is more performance related.
      While browsing perf reports captured during Linux build we can notice unusual portion of instructions executed in std::vector<std::string> copy constructor like:
      
      0.59%     0.58%  clang-14    clang-14      [.] std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >,
                                                                      std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >::vector
      
      or even:
      
      1.42%     0.26%  clang    clang-14             [.] clang::LangOptions::LangOptions
             |
              --1.16%--clang::LangOptions::LangOptions
                        |
                         --0.74%--std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >,
                                  std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >::vector
      
      After more digging we can see that relevant LangOptions std::vector members (*Files, ModuleFeatures and NoBuiltinFuncs)
      are constructed when Lexer::LangOpts field is initialized on list:
      
      Lexer::Lexer(..., const LangOptions &langOpts, ...)
                  : ..., LangOpts(langOpts),
      
      Since LangOptions copy constructor is called by Lexer(..., const LangOptions &LangOpts,...) and local Lexer objects are created thousands times
      (in Lexer::getRawToken, Preprocessor::EnterSourceFile and more) during single module processing in frontend it makes std::vector copy constructors surprisingly hot.
      
      Unfortunately even though in current Lexer implementation mentioned std::vector members are unused and most of time empty,
      no compiler is smart enough to optimize their std::vector copy constructors out (take a look at test assembly): https://godbolt.org/z/hdoxPfMYY even with LTO enabled.
      However there is simple way to fix this. Since Lexer doesn't access *Files, ModuleFeatures, NoBuiltinFuncs and any other LangOptions fields (but only LangOptionsBase)
      we can simply get rid of redundant copy constructor assembly by changing LangOpts type to more appropriate const LangOptions reference: https://godbolt.org/z/fP7de9176
      
      Additionally we need to store LineComment outside LangOpts because it's written in SkipLineComment function.
      Also FormatTokenLexer need to be adjusted a bit to avoid lifetime issues related to passing local LangOpts reference to Lexer.
      
      After this change I can see more than 1% speedup in some of my microbenchmarks when using Clang release binary built with LTO.
      For Linux build gains are not so significant but still nice at the level of -0.4%/-0.5% instructions drop.
      
      Differential Revision: https://reviews.llvm.org/D120334
      a64d3c60
    • Nikita Popov's avatar
      [InlineCost] Use SmallPtrSet for DeadBlocks (NFC) · 3c53d3a7
      Nikita Popov authored
      This set is only used with contains operations, so there is no
      need to use a SetVector.
      3c53d3a7
    • Archibald Elliott's avatar
    • Sander de Smalen's avatar
      [AArch64][SVE] Fold away SETCC if original input was predicate vector. · eac2638e
      Sander de Smalen authored
      This adds the following two folds:
      
      Fold 1:
         setcc_merge_zero(
             all_active, extend(nxvNi1 ...), != splat(0))
        -> nxvNi1 ...
      
      Fold 2:
         setcc_merge_zero(
             pred, extend(nxvNi1 ...), != splat(0))
        -> nxvNi1 and(pred, ...)
      
      Reviewed By: david-arm
      
      Differential Revision: https://reviews.llvm.org/D119334
      eac2638e
    • Florian Hahn's avatar
      Recommit "[VPlan] Introduce recipe to build scalar steps." · b3e8ace1
      Florian Hahn authored
      This reverts the revert commit ff93260b.
      
      The underlying issue causing the PPC bot failures has been fixed in
      cbaac147 and a corresponding test case has been added in
      ad2cad1c.
      
      Original message:
      
          This patch adds a new VPScalarIVStepsRecipe to handle building scalar
          steps.
      
          In the first patch, it only handles the case where there is no vector
          induction variable needed.
      
          Reviewed By: Ayal
      
          Differential Revision: https://reviews.llvm.org/D115953
      b3e8ace1
    • Timm Bäder's avatar
      [clang][tests] Fix ve-toolchain tests with CLANG_DEFAULT_UNWINDLIB · 12d36792
      Timm Bäder authored
      Otherwise, the driver will insert e.g. -lgcc_s when
      CLANG_DEFAULT_UNWINDLIB=libgcc is set during the clang build.
      
      Differential Revision: https://reviews.llvm.org/D120644
      12d36792
    • Valentin Clement's avatar
      [flang] Lower power operations · 43c071fa
      Valentin Clement authored
      Lower the power operation for real, integer
      and complex.
      
      The power operation is lowered to library calls.
      
      This patch is part of the upstreaming effort from fir-dev branch.
      
      Depends on D120403
      
      Reviewed By: schweitz
      
      Differential Revision: https://reviews.llvm.org/D120556
      43c071fa
    • Momchil Velikov's avatar
      [AArch64] Async unwind - function prologues · 32e8b550
      Momchil Velikov authored
      This patch rearranges emission of CFI instructions, so the resulting
      DWARF and `.eh_frame` information is precise at every instruction.
      
      The current state is that the unwind info is emitted only after the
      function prologue. This is fine for synchronous (e.g. C++) exceptions,
      but the information is generally incorrect when the program counter is
      at an instruction in the prologue or the epilogue, for example:
      
      ```
      stp	x29, x30, [sp, #-16]!           // 16-byte Folded Spill
      mov	x29, sp
      .cfi_def_cfa w29, 16
      ...
      ```
      
      after the `stp` is executed the (initial) rule for the CFA still says
      the CFA is in the `sp`, even though it's already offset by 16 bytes
      
      A correct unwind info could look like:
      ```
      stp	x29, x30, [sp, #-16]!           // 16-byte Folded Spill
      .cfi_def_cfa_offset 16
      mov	x29, sp
      .cfi_def_cfa w29, 16
      ...
      ```
      
      Having this information precise up to an instruction is useful for
      sampling profilers that would like to get a stack backtrace. The end
      goal (towards this patch is just a step) is to have fully working
      `-fasynchronous-unwind-tables`.
      
      Reviewed By: danielkiss, MaskRay
      
      Differential Revision: https://reviews.llvm.org/D111411
      32e8b550
    • Adrian Kuegel's avatar
      a91ade0b
    • Sander de Smalen's avatar
      [AArch64][SVE] Handle more cases in findMoreOptimalIndexType. · 201e3686
      Sander de Smalen authored
      This patch addresses @paulwalker-arm's comment on D117900 to
      only update/write the by-ref operands iff the function returns
      true. It also handles a few more cases where a series of added
      offsets can be folded into the base pointer, rather than just looking
      at a single offset.
      
      Reviewed By: paulwalker-arm
      
      Differential Revision: https://reviews.llvm.org/D119728
      201e3686
    • David Spickett's avatar
      [compiler-rt] Disable coverage trace pc guard tests on Thumb · ee95fe5c
      David Spickett authored
      These are failing on our silent bot:
      https://lab.llvm.org/staging/#/builders/162/builds/358
      
      $ <run cmd>
      main
      foo
      bar
      baz
      SanitizerCoverage: ./sanitizer_coverage_trace_pc_guard-dso.cpp.tmp.2122517.sancov: 2 PCs written
      SanitizerCoverage: ./sanitizer_coverage_trace_pc_guard-dso.cpp.tmp_2.so.2122517.sancov: 1 PCs written
      SanitizerCoverage: ./sanitizer_coverage_trace_pc_guard-dso.cpp.tmp_1.so.2122517.sancov: 1 PCs written
      $ <sancov cmd>
      ERROR: Coverage points in binary and .sancov file do not match.
      
      Also reproduces if you build for Thumb on v8 hardware.
      
      Doesn't fail when built with Arm only code so I guess the Thumb mode bit
      in the PCs might be the issue.
      ee95fe5c
    • gysit's avatar
      [mlir][linalg] Check the iterator types are valid. · 11d144c5
      gysit authored
      Improve the LinalgOp verification to ensure the iterator types is known. Previously, unknown iterator types have been ignored without warning, which can lead to confusing bugs.
      
      Reviewed By: nicolasvasilache
      
      Differential Revision: https://reviews.llvm.org/D120649
      11d144c5
    • Florian Hahn's avatar
      [LV] Remove induction recipes only used outside vector loop. · cbaac147
      Florian Hahn authored
      Exit values of vector inductions are generated completely independent of
      the induction recipes. Consider them for removal, if they are not used
      in loop.
      
      This fixes a crash exposed by 49b23f45.
      cbaac147
    • David Green's avatar
      Partially revert "[SchedModels][CortexA55] Add ASIMD integer instructions" · 61b61675
      David Green authored
      The Cortex-A55 scheduling model is used for -mcpu=generic, meaning it
      can have a wider effect than just the A55. The changes to the A55
      scheduling model seems to have caused performance regressions on
      Cortex-A510 device which have latencies closer to the original and
      different forwarding paths.
      
      This partially reverts the changes from D117003, at least until we can
      do something to improve Cortex-A510. According to my results, this
      improves the A510 results without altering the A55 very much.
      61b61675
    • Luis Penagos's avatar
      [clang-format] Treat && followed by noexcept operator as a binary operator... · 24d4f601
      Luis Penagos authored
      [clang-format] Treat && followed by noexcept operator as a binary operator inside template arguments
      
      Fixes https://github.com/llvm/llvm-project/issues/44544.
      
      Reviewed By: curdeius, MyDeveloperDay
      
      Differential Revision: https://reviews.llvm.org/D120445
      24d4f601
    • Adrian Kuegel's avatar
      44adca60
    • Florian Hahn's avatar
      [LV] Add test with dead induction in vector loop used outside. · 8bbc5e17
      Florian Hahn authored
      Add test with a induction phi that is not used in the vector loop, but
      by an lcssa phi in the loop exit.
      8bbc5e17
    • Endre Fülöp's avatar
      [analyzer] Add more sources to Taint analysis · 34a73879
      Endre Fülöp authored
      Add more functions as taint sources to GenericTaintChecker.
      
      Reviewed By: steakhal
      
      Differential Revision: https://reviews.llvm.org/D120236
      34a73879
    • LLVM GN Syncbot's avatar
      [gn build] Port 61835d19 · a44c984d
      LLVM GN Syncbot authored
      a44c984d
    • Nikita Popov's avatar
      [InstCombine] Remove not of SPF min/max fold (NFCI) · 5423b0a5
      Nikita Popov authored
      This should no longer be necessary now that we canonicalize to
      intrinsics. Might not be strictly NFC due to worklist order.
      5423b0a5
    • esmeyi's avatar
      [llvm-objcopy] Initial XCOFF32 support. · 61835d19
      esmeyi authored
      Summary: This is an initial implementation of lvm-objcopy for XCOFF32.
      Currently only supports simple copying, op-passthrough to follow.
      
      Reviewed By: jhenderson, shchenz
      
      Differential Revision: https://reviews.llvm.org/D97656
      61835d19
    • Nikita Popov's avatar
      [InstCombine] Remove sub of SPF min/max fold (NFCI) · d5ea3b2f
      Nikita Popov authored
      This isn't necessary anymore, now that we canonicalize SPF min/max
      to intrinsics. Might not be strictly NFC due to worklist order
      changes.
      d5ea3b2f
    • Florian Hahn's avatar
      [LV] Add test with IV that needs scalar steps and user outside of loop. · ad2cad1c
      Florian Hahn authored
      Also add a run line to check interleaving only. This test covers the PPC
      buildbot failures caused by 49b23f45.
      ad2cad1c
    • Nikita Popov's avatar
      [InstCombine] Don't call matchSAddSubSat() for SPF (NFC) · 9353ed6a
      Nikita Popov authored
      Only call it for intrinsic min/max. The moved implementation is
      unchanged apart from the one-use check: It is now hardcoded to
      one-use, without the two-use special case for SPF.
      9353ed6a
    • Nikita Popov's avatar
      [InstCombine] Remove SPF moveAddAfterMinMax() (NFC) · 53602e4c
      Nikita Popov authored
      As SPF min/max is canonicalized to intrinsics before this point,
      this change should be entirely NFC.
      53602e4c
    • Nikita Popov's avatar
      [InstCombine] Remove SPF moveNotAfterMinMax() (NFC) · ee62dcdb
      Nikita Popov authored
      This happens after SPF -> intrinsic canonicalization, and as such
      should be entirely NFC.
      ee62dcdb
    • Nikita Popov's avatar
      [InstCombine] Remove SPF factorizeMinMaxTree() (NFC) · 0bc3e233
      Nikita Popov authored
      SPF integer min/max is canonicalized to min/max intrinsics before
      this code is reached, so this should be entirely NFC.
      0bc3e233