1. Mar 17, 2021
  2. Mar 16, 2021
    • Thomas Preud'homme's avatar
      [MemDepAnalysis] Remove redundant comment. · f12433f1
      Thomas Preud'homme authored
      Exact same comment is found 2 lines above.
      f12433f1
    • Aart Bik's avatar
      [mlir][amx] blocked tilezero integration test · b388bbd3
      Aart Bik authored
      This adds a new integration test. However, it also
      adapts to a recent memref.XXX change for existing tests
      
      Reviewed By: ftynse
      
      Differential Revision: https://reviews.llvm.org/D98680
      b388bbd3
    • Josh Berdine's avatar
      [OCaml] Add missing TypeKinds, Opcode, and AtomicRMWBinOps · ece6d8e7
      Josh Berdine authored
      There are several enum values that have been added to LLVM-C that are
      missing from the OCaml bindings. The types defined in
      bindings/ocaml/llvm/llvm.ml should be in sync with the corresponding
      enum definitions in include/llvm-c/Core.h. The enum values are passed
      from C to OCaml unmodified, and clients of the OCaml bindings
      interpret them as tags of the corresponding OCaml types. So the only
      changes needed are to add the missing constructors to the type
      definitions, and to change the name of the maximum opcode in an
      assertion.
      
      Differential Revision: https://reviews.llvm.org/D98578
      ece6d8e7
    • Joe Ellis's avatar
      [AArch64][SVE] Fold vector ZExt/SExt into gather loads where possible · ff2dd8a2
      Joe Ellis authored
      This commit folds sxtw'd or uxtw'd offsets into gather loads where
      possible with a DAGCombine optimization.
      
      As an example, the following code:
      
           1	#include <arm_sve.h>
           2
           3	svuint64_t func(svbool_t pred, const int32_t *base, svint64_t offsets) {
           4	  return svld1sw_gather_s64offset_u64(
           5	    pred, base, svextw_s64_x(pred, offsets)
           6	  );
           7	}
      
      would previously lower to the following assembly:
      
          sxtw	z0.d, p0/m, z0.d
          ld1sw	{ z0.d }, p0/z, [x0, z0.d]
          ret
      
      but now lowers to:
      
          ld1sw   { z0.d }, p0/z, [x0, z0.d, sxtw]
          ret
      
      Differential Revision: https://reviews.llvm.org/D97858
      ff2dd8a2
    • Max Kazantsev's avatar
      [SCEV][NFC] Move check up the stack · 5097143f
      Max Kazantsev authored
      One of (and primary) callers of isBasicBlockEntryGuardedByCond is
      isKnownPredicateAt, which makes isKnownPredicate check before it.
      It already makes non-recursive check inside. So, on this execution
      path this check is made twice. The only other caller is
      isLoopEntryGuardedByCond. Moving the check there should save some
      compile time.
      5097143f
    • Craig Topper's avatar
      [RISCV] Look through copies when trying to find an implicit def in addVSetVL. · 229eeb18
      Craig Topper authored
      The InstrEmitter can sometimes insert a copy after an IMPLICIT_DEF
      before connecting it to the vector instruction. This occurs when
      constrainRegClass reduces to a class with less than 4 registers.
      I believe LMUL8 on masked instructions triggers this since the
      result can only use the v8, v16, or v24 register group as the mask
      is using v0.
      
      Reviewed By: frasercrmck
      
      Differential Revision: https://reviews.llvm.org/D98567
      229eeb18
    • David Zarzycki's avatar
      [lit testing] Mark reorder.py as unavailable on Windows · 61ca7064
      David Zarzycki authored
      The test file has embedded slashes. This is fine for normal users that
      are just recording and reordering paths, but not great when the trace
      data is committed back to a repository that should work on both Unix and
      Windows.
      61ca7064
    • Joe Ellis's avatar
      [AArch64][SVEIntrinsicOpts] Factor out redundant SVE mul/fmul intrinsics · 14bd44ed
      Joe Ellis authored
      This commit implements an IR-level optimization to eliminate idempotent
      SVE mul/fmul intrinsic calls. Currently, the following patterns are
      captured:
      
          fmul  pg  (dup_x  1.0)  V  =>  V
          mul   pg  (dup_x  1)    V  =>  V
      
          fmul  pg  V  (dup_x  1.0)  =>  V
          mul   pg  V  (dup_x  1)    =>  V
      
          fmul  pg  V  (dup  v  pg  1.0)  =>  V
          mul   pg  V  (dup  v  pg  1)    =>  V
      
      The result of this commit is that code such as:
      
          1  #include <arm_sve.h>
          2
          3  svfloat64_t foo(svfloat64_t a) {
          4    svbool_t t = svptrue_b64();
          5    svfloat64_t b = svdup_f64(1.0);
          6    return svmul_m(t, a, b);
          7  }
      
      will lower to a nop.
      
      This commit does not capture all possibilities; only the simple cases
      described above. There is still room for further optimisation.
      
      Differential Revision: https://reviews.llvm.org/D98033
      14bd44ed
    • Craig Topper's avatar
      [RISCV] Improve i32 UADDSAT/USUBSAT on RV64. · a33ce06c
      Craig Topper authored
      The default promotion uses zero extends that become shifts. We
      cam use sign extend instead which is better for RISCV.
      
      I've used two different implementations based on whether we
      have minu/maxu instructions.
      
      Differential Revision: https://reviews.llvm.org/D98683
      a33ce06c