1. Jun 14, 2023
    • Simon Pilgrim's avatar
      [CostModel][X86] Tweak SSE2 v2i64 multiply costs based off D46276 script · 595a7439
      Simon Pilgrim authored
      It looks like we were trying to account for SLM costs, which are actually handled separately
      
      Fixes #62969
      595a7439
    • Simon Pilgrim's avatar
      [TTI][X86] Recognise PMULUDQ costs for vXi64 multiplies · e64f9140
      Simon Pilgrim authored
      Addresses part of Issue #62969 - if the upper 32-bits of the vXi64 elements are known to be zero, then a multiply simplifies to a single (fast) PMULUDQ instruction
      
      We still have the problem that minRequiredElementSize can't determine that the upper bits are zero for the test case from Issue #62969 - I'll take a look at that next.
      e64f9140
    • Théo Degioanni's avatar
      [mlir][llvm] Add memset support for mem2reg/sroa · 8404b23a
      Théo Degioanni authored
      This revision introduces support for memset intrinsics in SROA and
      mem2reg for the LLVM dialect. This is achieved for SROA by breaking
      memsets of aggregates into multiple memsets of scalars, and for mem2reg
      by promoting memsets of single integer slots into the value the memset
      operation would yield.
      
      The SROA logic supports breaking memsets of static size operating at the
      start of a memory slot. The intended most common case is for memsets
      covering the entirety of a struct, most often as a way to initialize it
      to 0.
      
      The mem2reg logic supports dynamic values and static sizes as input to
      promotable memsets. This is achieved by lowering memsets into
      `ceil(log_2(n))` LeftShift operations, `ceil(log_2(n))` Or operations
      and up to one ZExt operation (for n the byte width of the integer),
      computing in registers the integer value the memset would create. Only
      byte-aligned integers are supported, more types could easily be added
      afterwards.
      
      Reviewed By: gysit
      
      Differential Revision: https://reviews.llvm.org/D152367
      8404b23a
    • Cullen Rhodes's avatar
      Revert "[mlir][ArmSME] Add initial dialect with basic lowering of vector.transfer write to zero" · 1e41a29d
      Cullen Rhodes authored
      Apologies I shouldn't have comitted this, need to wait until the planned
      MLIR ODM:
      
        https://discourse.llvm.org/t/rfc-creating-a-armsme-dialect/67208/76
      
      This reverts commit a48fe898.
      1e41a29d
    • Cullen Rhodes's avatar
      [mlir][ArmSME] Add initial dialect with basic lowering of vector.transfer write to zero · a48fe898
      Cullen Rhodes authored
      This patch adds support for lowering a `vector.transfer_write` of zeroes
      and type `vector<[16x16]xi8>` to the SME `zero {za}` instruction [1],
      which zeroes the entire accumulator.
      
      This contributes to supporting a path from `linalg.fill` to SME.
      
      [1] https://developer.arm.com/documentation/ddi0602/2022-06/SME-Instructions/ZERO--Zero-a-list-of-64-bit-element-ZA-tiles-
      
      Reviewed By: awarzynski, dcaballe
      
      Differential Revision: https://reviews.llvm.org/D152508
      a48fe898
    • Guillaume Chatelet's avatar
      [libc] Dispatch memmove to memcpy when buffers are disjoint · 2cfae7cd
      Guillaume Chatelet authored
      Most of the time `memmove` is called on buffers that are disjoint, in that case we can use `memcpy` which is faster.
      The additional test is branchless on x86, aarch64 and RISCV with the zbb extension (bitmanip).
      On x86 this patch adds a latency of 2 to 3 cycles.
      
      Before
      ```
      --------------------------------------------------------------------------------
      Benchmark                      Time             CPU   Iterations UserCounters...
      --------------------------------------------------------------------------------
      BM_Memmove/0/0_median       5.00 ns         5.00 ns           10 bytes_per_cycle=1.25477/s bytes_per_second=2.62933G/s items_per_second=199.87M/s __llvm_libc::memmove,memmove Google A
      BM_Memmove/1/0_median       6.21 ns         6.21 ns           10 bytes_per_cycle=3.22173/s bytes_per_second=6.75106G/s items_per_second=160.955M/s __llvm_libc::memmove,memmove Google B
      BM_Memmove/2/0_median       8.09 ns         8.09 ns           10 bytes_per_cycle=5.31462/s bytes_per_second=11.1366G/s items_per_second=123.603M/s __llvm_libc::memmove,memmove Google D
      BM_Memmove/3/0_median       5.95 ns         5.95 ns           10 bytes_per_cycle=2.71865/s bytes_per_second=5.69687G/s items_per_second=167.967M/s __llvm_libc::memmove,memmove Google L
      BM_Memmove/4/0_median       5.63 ns         5.63 ns           10 bytes_per_cycle=2.28294/s bytes_per_second=4.78383G/s items_per_second=177.615M/s __llvm_libc::memmove,memmove Google M
      BM_Memmove/5/0_median       5.68 ns         5.68 ns           10 bytes_per_cycle=2.16798/s bytes_per_second=4.54295G/s items_per_second=176.015M/s __llvm_libc::memmove,memmove Google Q
      BM_Memmove/6/0_median       7.46 ns         7.46 ns           10 bytes_per_cycle=3.97619/s bytes_per_second=8.332G/s items_per_second=134.044M/s __llvm_libc::memmove,memmove Google S
      BM_Memmove/7/0_median       5.40 ns         5.40 ns           10 bytes_per_cycle=1.79695/s bytes_per_second=3.76546G/s items_per_second=185.211M/s __llvm_libc::memmove,memmove Google U
      BM_Memmove/8/0_median       5.62 ns         5.62 ns           10 bytes_per_cycle=3.18747/s bytes_per_second=6.67927G/s items_per_second=177.983M/s __llvm_libc::memmove,memmove Google W
      BM_Memmove/9/0_median        101 ns          101 ns           10 bytes_per_cycle=9.77359/s bytes_per_second=20.4803G/s items_per_second=9.9333M/s __llvm_libc::memmove,uniform 384 to 4096
      ```
      After
      ```
      BM_Memmove/0/0_median       3.57 ns         3.57 ns           10 bytes_per_cycle=1.71375/s bytes_per_second=3.59112G/s items_per_second=280.411M/s __llvm_libc::memmove,memmove Google A
      BM_Memmove/1/0_median       4.52 ns         4.52 ns           10 bytes_per_cycle=4.47557/s bytes_per_second=9.37843G/s items_per_second=221.427M/s __llvm_libc::memmove,memmove Google B
      BM_Memmove/2/0_median       5.70 ns         5.70 ns           10 bytes_per_cycle=7.37396/s bytes_per_second=15.4519G/s items_per_second=175.399M/s __llvm_libc::memmove,memmove Google D
      BM_Memmove/3/0_median       4.47 ns         4.47 ns           10 bytes_per_cycle=3.4148/s bytes_per_second=7.15563G/s items_per_second=223.743M/s __llvm_libc::memmove,memmove Google L
      BM_Memmove/4/0_median       4.53 ns         4.53 ns           10 bytes_per_cycle=2.86071/s bytes_per_second=5.99454G/s items_per_second=220.69M/s __llvm_libc::memmove,memmove Google M
      BM_Memmove/5/0_median       4.19 ns         4.19 ns           10 bytes_per_cycle=2.5484/s bytes_per_second=5.3401G/s items_per_second=238.924M/s __llvm_libc::memmove,memmove Google Q
      BM_Memmove/6/0_median       5.02 ns         5.02 ns           10 bytes_per_cycle=5.94164/s bytes_per_second=12.4505G/s items_per_second=199.14M/s __llvm_libc::memmove,memmove Google S
      BM_Memmove/7/0_median       4.03 ns         4.03 ns           10 bytes_per_cycle=2.47028/s bytes_per_second=5.17641G/s items_per_second=247.906M/s __llvm_libc::memmove,memmove Google U
      BM_Memmove/8/0_median       4.70 ns         4.70 ns           10 bytes_per_cycle=3.84975/s bytes_per_second=8.06706G/s items_per_second=212.72M/s __llvm_libc::memmove,memmove Google W
      BM_Memmove/9/0_median       90.7 ns         90.7 ns           10 bytes_per_cycle=10.8681/s bytes_per_second=22.7739G/s items_per_second=11.02M/s __llvm_libc::memmove,uniform 384 to 4096
      ```
      
      Reviewed By: courbet
      
      Differential Revision: https://reviews.llvm.org/D152811
      2cfae7cd
    • Vitaly Buka's avatar
      c93ca4bc
    • Carl Ritson's avatar
      [AMDGPU] Pre-commit test for D152892 (NFC) · 936c16a3
      Carl Ritson authored
      936c16a3
    • Nikita Popov's avatar
      [PhaseOrdering] Regenerate test checks (NFC) · bf977979
      Nikita Popov authored
      Just naming changes.
      bf977979
    • Nikita Popov's avatar
      [InstCombine] Handle use count decrement in more cases · 09d6fa79
      Nikita Popov authored
      These two helpers also decrement the use count of the replaced
      operand, so give them the same treatment as eraseInstruction().
      09d6fa79
    • Vitaly Buka's avatar
      [test][hwasan] Rename constants in test · 4a69e0a0
      Vitaly Buka authored
      4a69e0a0
    • Nikita Popov's avatar
      [Clang] Rename getElementBitCast() -> withElementType() (NFC) · 0211a75e
      Nikita Popov authored
      This no longer creates a bitcast, just changes the element type
      of the ConstantAddress.
      0211a75e
    • Joshua Cao's avatar
      [SimpleLoopUnswitch] Unswitch AND/OR conditions of selects · 2f171b27
      Joshua Cao authored
      If a select's condition is a AND/OR, we can unswitch invariant operands.
      This patch uses existing logic from unswitching AND/OR's for branch
      conditions.
      
      This patch fixes the Cost computation for unswitching selects to have
      the cost of the entire loop, since unswitching selects do not remove
      branches. This is required for this patch because otherwise, there are
      cases where unswitching selects of AND/OR is beating out unswitching of
      branches.
      
      This patch also prevents unswitching of logical AND/OR selects. This
      should instead be done by unswitching of AND/OR branch conditions.
      
      Reviewed By: nikic
      
      Differential Revision: https://reviews.llvm.org/D151677
      2f171b27
    • Joshua Cao's avatar
    • Matthias Springer's avatar
      [mlir][vector][bufferize] Better analysis for vector.transfer_write · 80853a16
      Matthias Springer authored
      The destination operand does not bufferize to a memory read if it is completely overwritten.
      
      Differential Revision: https://reviews.llvm.org/D152823
      80853a16
    • Michael Platings's avatar
      Fix test Driver/mips-mti-linux.c · 01b6f062
      Michael Platings authored
      01b6f062
    • Nikita Popov's avatar
      [InstCombine] Revisit user of newly one-use instructions · f7a977c7
      Nikita Popov authored
      Many folds in InstCombine are limited to one-use instructions. For
      that reason, if the use-count of an instruction drops to one, it
      makes sense to revisit that one user. This is one of the most
      common reasons why InstCombine fails to finish in a single iteration.
      
      Doing this revisit actually slightly improves compile-time, because
      we save an extra InstCombine iteration in enough cases to make a
      visible difference.
      
      This is conceptually NFC, but not NFC in practice, because differences
      in worklist order can result in slightly different folding behavior.
      
      The regressed tests in or-shifted-masks.ll now require a sequence of
      instcombine,early-cse,instcombine to fold fully. D152876 would make
      these fold in a single instcombine run again.
      
      Differential Revision: https://reviews.llvm.org/D151807
      f7a977c7
    • eopXD's avatar
      [11/11][Clang][RISCV] Expand all variants for vset on tuple types · 804306b3
      eopXD authored
      This is the 11th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152078.
      
      This patch also fixes the suffix for non-overloaded variants for
      vset on tuple types.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152079
      804306b3
    • eopXD's avatar
      [10/11][Clang][RISCV] Expand all variants for vget on tuple types · d1b6b893
      eopXD authored
      This is the 10th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152077.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152078
      d1b6b893
    • eopXD's avatar
      [9/11][Clang][RISCV] Expand all variants for indexed strided segment store · 9828c895
      eopXD authored
      This is the 9th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152076.
      
      This patch expands all variants of indexed strided segment store.
      This patch also fixes the trailing suffix in the intrinsics' function
      name that representing the return type, adding `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vsuxseg2ei32, vsoxseg2ei32 is added for now.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152077
      9828c895
    • eopXD's avatar
      [8/11][Clang][RISCV] Expand all variants for indexed strided segment load · 9588b189
      eopXD authored
      This is the 8th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152075.
      
      This patch expands all variants of indexed strided segment load,
      including the policy variants. This patch also fixes the trailing suffix
      in the intrinsics' function name that representing the return type,
      adding `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vluxseg2ei32, vloxseg2ei32 is added for now.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152076
      9588b189
    • eopXD's avatar
      [7/11][Clang][RISCV] Expand all variants for strided segment store · 91b90e0e
      eopXD authored
      This is the 7th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152074.
      
      This patch expands all variants for strided segment store. The store
      intrinsics does not have any policy variants. This patch also fixes the
      trailing suffix in the intrinsics' function name that representing the
      return type, adding `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vssseg2e32 is added for now.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152075
      91b90e0e
    • eopXD's avatar
      [6/11][Clang][RISCV] Expand all variants for strided segment load · 06c0b0ac
      eopXD authored
      This is the 6th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152073.
      
      This patch expands all variants of strided segment load, including the
      policy variants. This patch also fixes the trailing suffix in the
      intrinsics' function name that representing the return type, adding
      `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vlsseg2e32 is added.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152074
      06c0b0ac
    • Chuanqi Xu's avatar
      [NFC] skip the test modules-vtable.cppm on windows · baf0b12c
      Chuanqi Xu authored
      The new added test has problems on windows since the patch is about ABI
      and MSVC ABI is not covered. Skip the test on windows to make the CI
      green.
      baf0b12c
    • eopXD's avatar
      [5/11][Clang][RISCV] Expand all variants for unit stride fault-first segment load · ff6f7902
      eopXD authored
      This is the 5th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152072.
      
      This patch expands all variants of unit stride fault-first segment
      load, including the policy variants. This patch also fixes the
      trailing suffix in the intrinsics' function name that representing
      the return type, adding `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vlseg2e32ff is added.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152073
      ff6f7902
    • Zi Xuan Wu (Zeson)'s avatar
      [CSKY] Add support for half-precision floats · d90468d1
      Zi Xuan Wu (Zeson) authored
      Complete fp16 support by ensuring that load extension / truncate store operations are properly expanded.
      d90468d1
    • Piyou Chen's avatar
      [NFC][RISCV] rename findFirstNonVersionCharacter with findLastNonVersionCharacter · fc4350cb
      Piyou Chen authored
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152506
      fc4350cb
    • Matthias Springer's avatar
      [mlir][IR] Improve listener notifications for ops without results · 71d50c89
      Matthias Springer authored
      `RewriterBase::Listener::notifyOperationReplaced` notifies observers that an op is about to be replaced with a range of values. This notification is not very useful for ops without results, because it does not specify the replacement op (and it cannot be deduced from the replacement values). It provides no additional information over the `notifyOperationRemoved` notification.
      
      This revision adds an additional notification when a rewriter replaces an op with another op. By default, this notification triggers the original "op replaced with values" notification, so there is no functional change for existing code.
      
      This new API is useful for the transform dialect, which needs to track op replacements. (Updated in a subsequent revision.)
      
      Also includes minor documentation improvements.
      
      Differential Revision: https://reviews.llvm.org/D152814
      71d50c89
    • eopXD's avatar
      [4/11][Clang][RISCV] Expand all variants for unit stride segment store · b59ca492
      eopXD authored
      This is the 4th patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152071.
      
      This patch expands all variants for unit stride segment store. The
      store intrinsics does not have any policy variants. This patch also
      fixes the trailing suffix in the intrinsics' function name that
      representing the return type, adding `x{NF}`.
      
      For the same reason mentioned in [3/11], only full test case for
      vsseg2e32 is added.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152072
      b59ca492
    • eopXD's avatar
      [3/11][Clang][RISCV] Expand all variants for unit stride segment load · 0e9548bb
      eopXD authored
      This is the 3rd patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152070.
      
      This patch expands all variants of unit stride segment load, including
      the policy variants. This patch also fixes the trailing suffix in the
      intrinsics' function name that representing the return type, adding
      `x{NF}`.
      
      Currently the tuple type co-exists with the non-tuple type intrinsics.
      Since the co-existance is temporary, this patch only adds test cases of
      all variants for vlseg2e32 to show the capability done.
      
      Test cases of other data type and NF will be added in the patch-set
      when the replacement happens.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152071
      0e9548bb
    • eopXD's avatar
      [2/11][Clang][RISCV] Expand all variants of RVV intrinsic tuple types · 5847ec4d
      eopXD authored
      This is the 2nd patch of the patch-set. For the cover letter, please
      checkout D152069.
      
      Depends on D152069.
      
      This patch also removes redundant checks related to tuples and dedicate
      the check to happen in `RVVType::verifyType`.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D152070
      5847ec4d
    • Valentin Clement's avatar
      [flang][openacc] Lower gang dim to MLIR · 66546d94
      Valentin Clement authored
      Lower gang dim from the parse tree to the new MLIR
      representation.
      
      Depends on D151972
      
      Reviewed By: razvanlupusoru, jeanPerier
      
      Differential Revision: https://reviews.llvm.org/D151973
      66546d94
    • River Riddle's avatar
      [Orc][Coff] Skip registration of voltbl sections · 639e411c
      River Riddle authored
      We're getting asserts for duplicate section registration during
      linking which stems back to these sections. From previous
      discussions, it seems like these are metadata sections that can
      be dropped. See the discussion in D116474 and
      https://bugs.llvm.org/show_bug.cgi?id=45111.
      
      Differential Revision: https://reviews.llvm.org/D152574
      639e411c
    • Michael Platings's avatar
      [Docs] Multilib design · a5aeba73
      Michael Platings authored
      Reviewed By: peter.smith, MaskRay
      
      Differential Revision: https://reviews.llvm.org/D143587
      a5aeba73
    • Michael Platings's avatar
      [Driver] BareMetal ToolChain multilib layering · ab2c80bb
      Michael Platings authored
      This enables layering baremetal multilibs on top of each other.
      For example a multilib containing only a no-exceptions libc++ could be
      layered on top of a multilib containing C libs. This avoids the need
      to duplicate the C library for every libc++ variant.
      
      Differential Revision: https://reviews.llvm.org/D143075
      ab2c80bb
    • Michael Platings's avatar
      [Driver] Enable selecting multiple multilibs · edc1130c
      Michael Platings authored
      This will enable layering multilibs on top of each other.
      For example a multilib containing only a no-exceptions libc++ could be
      layered on top of a multilib containing C libs. This avoids the need
      to duplicate the C library for every libc++ variant.
      
      This change doesn't expose the functionality externally, it only opens
      the functionality up to be potentially used by ToolChain classes.
      
      Differential Revision: https://reviews.llvm.org/D143059
      edc1130c
    • Michael Platings's avatar
      [Driver] Enable multilib.yaml in the BareMetal ToolChain · b4eebc86
      Michael Platings authored
      The default location for multilib.yaml is lib/clang-runtimes, without
      any target-specific suffix. This will allow multilibs for different
      architectures to share a common include directory.
      
      To avoid breaking the arm-execute-only.c CHECK-NO-EXECUTE-ONLY-ASM
      test, add a ForMultilib argument to getARMTargetFeatures.
      
      Since the presence of multilib.yaml can change the exact location of a
      library, relax the baremetal.cpp test.
      
      Differential Revision: https://reviews.llvm.org/D142986
      b4eebc86
    • Michael Platings's avatar
      [Driver] Add -print-multi-flags-experimental option · a794ab92
      Michael Platings authored
      This option causes the flags used for selecting multilibs to be printed.
      This is an experimental feature that is documented in detail in D143587.
      
      Differential Revision: https://reviews.llvm.org/D142933
      a794ab92
    • Michael Platings's avatar
      [Driver] Multilib YAML parsing · 4794bdab
      Michael Platings authored
      The format includes a ClangMinimumVersion entry to avoid a potential
      source of subtle errors if an older version of Clang were to be used
      with a multilib.yaml that requires a newer Clang to work correctly.
      This feature is comparable to CMake's cmake_minimum_required.
      
      Reviewed By: peter.smith
      
      Differential Revision: https://reviews.llvm.org/D142932
      4794bdab
    • Kirill Stoimenov's avatar
      [HWASAN] Implement munmap interceptor for HWASAN · fba9fd1a
      Kirill Stoimenov authored
      Reviewed By: vitalybuka
      
      Differential Revision: https://reviews.llvm.org/D152763
      fba9fd1a