1. Aug 17, 2022
    • Dmitri Gribenko's avatar
      [clang][dataflow] Use llvm::is_contained() · 941959d6
      Dmitri Gribenko authored
      Reviewed By: samestep, xazax.hun
      
      Differential Revision: https://reviews.llvm.org/D131975
      941959d6
    • Nicolas Miller's avatar
      Fix subrange liveness checking at rematerialization · ccfabfbb
      Nicolas Miller authored
      This patch fixes an issue where an instruction reading a whole register would be moved during register allocation into a spot where one of the subregisters was dead.
      
      The code to check whether an instruction can be rematerialized at a given point or not was already checking for subranges to ensure that subregisters are live, but only when the instruction being moved was using a subregister, this patch changes that so the subranges are checked even when the moved instruction uses the full register.
      
      This patch also adds a case to the original test for the subrange checking that trigger the issue described above.
      
      The original subrange checking code was introduced in this revision: https://reviews.llvm.org/D115278
      
      And I've encountered this issue on AMDGPUs while working with DPC++: https://github.com/intel/llvm/issues/6209
      
      Essentially the greedy register allocator attempts to move the following instruction:
      
      ```
      %3961:vreg_64 = V_LSHLREV_B64_e64 3, %3078:vreg_64, implicit $exec
      ```
      
      From `@3440` into the body of a loop `@16312`, but `%3078` has the following live ranges:
      
      ```
      %3078 [2224r,2240r:0)[2240r,3488B:1)[16192B,38336B:1) 0@2224r 1@2240r  L0000000000000003 [2224r,3440r:0) 0@2224r  L000000000000000C [2240r,3488B:0)[16192B,38336B:0) 0@2240r
      ```
      
      So `@16312e` `%3078.sub1` is alive but `%3078.sub0` is dead, so this instruction being moved there leads to invalid memory accesses as `3078.sub0` ends up being trashed and the result of this instruction is used as part of an address calculation for a load.
      
      On the original ticket this issue showed up on gfx906 and gfx90a but not on gfx908, this turned out to be because on gfx908 instead of moving the shift instruction into the loop, its value is spilled into an ACC register, gfx906 doesn't have ACC registers and for gfx90a ACC registers are used like regular vector registers and so aren't used for spilling.
      
      With this patch the original application from the DPC++ ticket works properly on gfx906, and the result of the shift instruction is correctly spilled instead of moving the instruction in the loop.
      
      Original Author: npmiller
      
      Reviewed by: rampitec
      
      Submitted by: rampitec
      
      Differential Revision: https://reviews.llvm.org/D131884
      ccfabfbb
    • David Blaikie's avatar
      Revert "flang: Fix flang build with -Wctad-maybe-unsupported" · fe7450b3
      David Blaikie authored
      -Wctad-maybe-unsupported is now disabled for flang so these explicit
      deduction guides are not required.
      
      This reverts commit 248591aa.
      fe7450b3
    • David Blaikie's avatar
      Revert "Some more from-the-hip ctad-maybe-unsupported fixes for flang" · e2333b55
      David Blaikie authored
      -Wctad-maybe-unsupported is now disabled for flang so these explicit
      deduction guides are not required.
      
      This reverts commit ec3956b6.
      e2333b55
    • David Blaikie's avatar
    • Zequan Wu's avatar
      [LLDB][NativePDB] Add nullptr checking. · 7ebbef2b
      Zequan Wu authored
      7ebbef2b
    • Mark de Wever's avatar
      [libc++] Improve updating data files. · 130b1816
      Mark de Wever authored
      This changes makes it easier to update the Unicode data files used for
      the Extended Graphme Clustering as added in D126971.
      
      Reviewed By: ldionne, #libc
      
      Differential Revision: https://reviews.llvm.org/D129668
      130b1816
    • Mark de Wever's avatar
      [libc++][format] Improve format buffer. · f7c0df00
      Mark de Wever authored
      Allow bulk output operations on the buffer instead of adding one
      code unit at a time. This has a huge performance benefit at the cost of
      larger binary. This doesn't implement @vitaut's earlier suggestion to
      avoid buffering for std::string when writing a strings. That can be done
      in a follow-up patch.
      
      There are some minor complications for the non-buffered format_to_n.
      When writing one character at a time it's easy to detect when reaching
      the limit n. This is solved by adding a small overhead for format_to_n.
      When the next write would overflow it stores the data in the internal
      buffer and copies that up-to n code units. The overhead isn't measured,
      but it's expected to only be an issue for small values of n; for larger
      values the general improvements will outweight the new overhead.
      
      ```
         text	   data	    bss	    dec	    hex	filename
       349081	   6096	    440	 355617	  56d21	format.libcxx.out-baseline
       344442	   6088	    440	 350970	  55afa	formatted_size.libcxx.out...
      f7c0df00
    • Vitaly Buka's avatar
      69c09d11
    • Slava Zakharin's avatar
      [mlir][math] Added basic support for FPowI operation. · f9d988f1
      Slava Zakharin authored
      The operation computes pow(b, p), where 'b' is floating point
      and 'p' is a signed integer. The result's type matches 'b' type.
      The operands must have the same shape.
      
      Differential Revision: https://reviews.llvm.org/D129811
      f9d988f1
    • Steven Wu's avatar
      [CMake] Cleanup the descriptions for gRPC options · 07c2f592
      Steven Wu authored
      As a followup to https://reviews.llvm.org/D131593, clean up gRPC related
      option names and messages to make them more generic.
      07c2f592
    • David Blaikie's avatar
  2. Aug 16, 2022