1. Sep 21, 2021
  2. Sep 20, 2021
    • Tobias Gysi's avatar
      [mlir][linalg] Add IndexOp support to fusion on tensors. · 7be28d82
      Tobias Gysi authored
      This revision depends on https://reviews.llvm.org/D109761 and https://reviews.llvm.org/D109766.
      
      Reviewed By: nicolasvasilache
      
      Differential Revision: https://reviews.llvm.org/D109774
      7be28d82
    • Morten Borup Petersen's avatar
      [MLIR][SCF] Add for-to-while loop transformation pass · 644b55d5
      Morten Borup Petersen authored
      This pass transforms SCF.ForOp operations to SCF.WhileOp. The For loop condition is placed in the 'before' region of the while operation, and indctuion variable incrementation + the loop body in the 'after' region. The loop carried values of the while op are the induction variable (IV) of the for-loop + any iter_args specified for the for-loop.
      Any 'yield' ops in the for-loop are rewritten to additionally yield the (incremented) induction variable.
      
      This transformation is useful for passes where we want to consider structured control flow solely on the basis of a loop body and the computation of a loop condition. As an example, when doing high-level synthesis in CIRCT, the incrementation of an IV in a for-loop is "just another part" of a circuit datapath, and what we really care about is the distinction between our datapath and our control logic (the condition variable).
      
      Differential Revision: https://reviews.llvm.org/D108454
      644b55d5
    • Tobias Gysi's avatar
      [mlir][linalg] Fix typo (NFC). · 09100c75
      Tobias Gysi authored
      09100c75
    • Alexey Bataev's avatar
      [SLP]Improve graph reordering. · bc69dd62
      Alexey Bataev authored
      Reworked reordering algorithm. Originally, the compiler just tried to
      detect the most common order in the reordarable nodes (loads, stores,
      extractelements,extractvalues) and then fully rebuilding the graph in
      the best order. This was not effecient, since it required an extra
      memory and time for building/rebuilding tree, double the use of the
      scheduling budget, which could lead to missing vectorization due to
      exausted scheduling resources.
      
      Patch provide 2-way approach for graph reodering problem. At first, all
      reordering is done in-place, it doe not required tree
      deleting/rebuilding, it just rotates the scalars/orders/reuses masks in
      the graph node.
      
      The first step (top-to bottom) rotates the whole graph, similarly to the previous
      implementation. Compiler counts the number of the most used orders of
      the graph nodes with the same vectorization factor and then rotates the
      subgraph with the given vectorization factor to the most used order, if
      it is not empty. Then repeats the same procedure for the subgraphs with
      the smaller vectorization factor. We can do this because we still need
      to reshuffle smaller subgraph when buildiong operands for the graph
      nodes with lasrger vectorization factor, we can rotate just subgraph,
      not the whole graph.
      
      The second step (bottom-to-top) scans through the leaves and tries to
      detect the users of the leaves which can be reordered. If the leaves can
      be reorder in the best fashion, they are reordered and their user too.
      It allows to remove double shuffles to the same ordering of the operands in
      many cases and just reorder the user operations instead. Plus, it moves
      the final shuffles closer to the top of the graph and in many cases
      allows to remove extra shuffle because the same procedure is repeated
      again and we can again merge some reordering masks and reorder user nodes
      instead of the operands.
      
      Also, patch improves cost model for gathering of loads, which improves
      x264 benchmark in some cases.
      
      Gives about +2% on AVX512 + LTO (more expected for AVX/AVX2) for {625,525}x264,
      +3% for 508.namd, improves most of other benchmarks.
      The compile and link time are almost the same, though in some cases it
      should be better (we're not doing an extra instruction scheduling
      anymore) + we may vectorize more code for the large basic blocks again
      because of saving scheduling budget.
      
      Differential Revision: https://reviews.llvm.org/D105020
      bc69dd62
    • peter klausler's avatar
      [flang] Put intrinsic function table back into order · 5661317f
      peter klausler authored
      Some intrinsic functions weren't findable because the table
      wasn't strictly in order of names.
      
      And complete a missing generalization of the extension DCONJG
      to accept any kind of complex argument, like DREAL and DIMAG
      were.
      
      Differential Revision: https://reviews.llvm.org/D110002
      5661317f
    • Wang, Pengfei's avatar
      [X86] Always check the size of SourceTy before getting the next type · 22767339
      Wang, Pengfei authored
      D109607 results in a regression in llvm-test-suite.
      The reason is we didn't check the size of SourceTy, so that we will
      return wrong SSE type when SourceTy is overlapped.
      
      Reviewed By: Meinersbur
      
      Differential Revision: https://reviews.llvm.org/D110037
      22767339
    • Wang, Pengfei's avatar
      5b47256f
    • Justas Janickas's avatar
      [OpenCL] Supports atomics in C++ for OpenCL 2021 · 228dd20c
      Justas Janickas authored
      Atomics in C++ for OpenCL 2021 are now handled the same way as in
      OpenCL C 3.0. This is a header-only change.
      
      Differential Revision: https://reviews.llvm.org/D109424
      228dd20c