1. Jun 18, 2022
  2. Jun 17, 2022
    • Nico Weber's avatar
      [gn build] (manually) port 7cca33b4 · c2bb2e59
      Nico Weber authored
      Not really needed for anything as far as I can tell (?),
      more for completeness.
      c2bb2e59
    • Christopher Bate's avatar
      [mlir][nvgpu] fix missing build dependency for NVGPUTransforms · d089d68a
      Christopher Bate authored
      Fixes build failure caused by 51b925df
      d089d68a
    • Aart Bik's avatar
      [mlir][sparse] move from by-value to by-reference for data types · aef20f59
      Aart Bik authored
      This fixes all sorts of ABI issues due to passing by-value
      (using by-reference with memref's exclusively).
      
      Reviewed By: bkramer
      
      Differential Revision: https://reviews.llvm.org/D128018
      aef20f59
    • Christopher Bate's avatar
      [mlir][nvgpu] shared memory access optimization pass · 51b925df
      Christopher Bate authored
      This change adds a transformation and pass to the NvGPU dialect that
      attempts to optimize reads/writes from a  memref representing GPU shared
      memory in order to avoid bank conflicts. Given a value representing a
      shared memory memref, it traverses all reads/writes within the parent op
      and, subject to suitable conditions, rewrites all last dimension index
      values such that element locations in the final (col) dimension are
      given by
      `newColIdx = col % vecSize + perm[row](col/vecSize,row)`
      where `perm` is a permutation function indexed by `row` and `vecSize`
      is the vector access size in elements (currently assumes 128bit
      vectorized accesses, but this can be made a parameter). This specific
      transformation can help optimize typical distributed & vectorized accesses
      common to loading matrix multiplication operands to/from shared memory.
      
      Differential Revision: https://reviews.llvm.org/D127457
      51b925df
    • Guillaume Chatelet's avatar
    • Quinn Pham's avatar
      [PowerPC] Fix PPCVSXSwapRemoval pass to include MTVSCR and MFVSCR as not swappable. · deb76552
      Quinn Pham authored
      This patch adds the instructions `MTVSCR` and `MFVSCR` as not swappable to the
      PPCVSXSwapRemoval pass because they are not lane-insensitive. This will prevent
      the compiler from optimizing out required swaps when using `lxvd2x` and
      `stxvd2x`.
      
      Reviewed By: #powerpc, nemanjai
      
      Differential Revision: https://reviews.llvm.org/D128062
      deb76552
    • Philip Reames's avatar
      [RISCV] Avoid changing etype for splat of 0 or -1 · 755c84c6
      Philip Reames authored
      A splat of the values 0 and -1 as sign extended 12 bit immediates are always the same bit pattern regardless of the etype used to perform the operation. As a result, we can sometimes avoid introducing a vsetvli just for the purposes of a splat.
      
      Looking at the diffs, we don't get a huge amount of immediate value out of this. We mostly push the vsetvli one instruction down, usually in front of a vmerge. We also don't get the corresponding fixed length vector cases because VL typically is changed despite the actual bits written being the same. Both of these are areas I plan to explore in future patches.
      
      Interestingly, this makes a great example of why we need the forward and backward implementation to be consistent. Before we merged the demanded field handling, if we implement only the forward direction, we lost the ability to mutate a prior vsetvli and eliminate a later one entirely. This resulted in practical regressions instead of improvements. It's always nice when practice matches theory. :)
      
      Differential Revision: https://reviews.llvm.org/D128006
      755c84c6
    • Ben Langmuir's avatar
      [clang][deps] Sort submodules when calculating dependencies · 4a3a9a5f
      Ben Langmuir authored
      Dependency scanning does not care about the order of submodules for
      correctness, so sort the submodules so that we get the same
      command-lines to build the module across different TUs. The order of
      inferred submodules can vary depending on the order of #includes in the
      including TU.
      
      Differential Revision: https://reviews.llvm.org/D128008
      4a3a9a5f