1. Dec 10, 2021
  2. Dec 09, 2021
    • Krzysztof Drewniak's avatar
      [MLIR][GPU] Define gpu.printf op and its lowerings · e1da6291
      Krzysztof Drewniak authored
      - Define a gpu.printf op, which can be lowered to any GPU printf() support (which is present in CUDA, HIP, and OpenCL). This op only supports constant format strings and scalar arguments
      - Define the lowering of gpu.pirntf to a call to printf() (which is what is required for AMD GPUs when using OpenCL) as well as to the hostcall interface present in the AMD Open Compute device library, which is the interface present when kernels are running under HIP.
      - Add a "runtime" enum that allows specifying which of the possible runtimes a ROCDL kernel will be executed under or that the runtime is unknown. This enum controls how gpu.printf is lowered
      
      This change does not enable lowering for Nvidia GPUs, but such a lowering should be possible in principle.
      
      And:
      [MLIR][AMDGPU] Always set amdgpu-implicitarg-num-bytes=56 on kernels
      
      This is something that Clang always sets on both OpenCL and HIP kernels, and failing to include it causes mysterious crashes with printf() support.
      
      In addition, revert the max-flat-work-group-size to (1, 256) to avoid triggering bugs in the AMDGPU backend.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D110448
      e1da6291
    • David Sherwood's avatar
      [LoopVectorize][AArch64] Add vectoriser cost model tests for gathers/scatters · def8b952
      David Sherwood authored
      I've added some tests that were previously missing for the gather-scatter costs
      being calculated by the vectorizer for AArch64:
      
        Transforms/LoopVectorize/AArch64/sve-gather-scatter-cost.ll
      
      The costs are sometimes different to the ones in
      
        Analysis/CostModel/AArch64/sve-gather.ll
      
      because the vectorizer also adds on the address computation cost.
      def8b952
    • Brian Cain's avatar
      Revert "[xray] add support for hexagon" · ab28cb1c
      Brian Cain authored
      This reverts commit 543a9ad7.
      ab28cb1c
    • Eugene Zhulenev's avatar
      [mlir] AsyncParallelFor: align block size to be a multiple of inner loops iterations · 49ce40e9
      Eugene Zhulenev authored
      Depends On D115263
      
      By aligning block size to inner loop iterations parallel_compute_fn LLVM can later unroll and vectorize some of the inner loops with small number of trip counts. Up to 2x speedup in multiple benchmarks.
      
      Reviewed By: bkramer
      
      Differential Revision: https://reviews.llvm.org/D115436
      49ce40e9