1. Oct 28, 2022
  2. Oct 27, 2022
    • Kito Cheng's avatar
      [RISCV] Drop single letter b extension support · ae116f43
      Kito Cheng authored
      It splited into several zb* extensions, and `b` is dropped after
      0.93, so it time to retired that as other non-ratified zb* extensions.
      
      Currntly clang can accept that with warning:
      
      $ clang -target riscv64-elf ~/hello.c -S  -march=rv64gcb
      '+b' is not a recognized feature for this target (ignoring feature)
      '+b' is not a recognized feature for this target (ignoring feature)
      '+b' is not a recognized feature for this target (ignoring feature)
      
      Reviewed By: asb, luismarques
      
      Differential Revision: https://reviews.llvm.org/D136812
      ae116f43
    • Dave Lee's avatar
      995d556f
    • Alexandros Lamprineas's avatar
      [FuncSpec] Do not overestimate the specialization bonus for users inside loops. · dbeaf6ba
      Alexandros Lamprineas authored
      When calculating the specialization bonus for a given function argument,
      we recursively traverse the chain of (certain) users, accumulating the
      instruction costs. Then we exponentially increase the bonus to account
      for loop nests. This is problematic for two reasons: (a) the users might
      not themselves be inside the loop nest, (b) if they are we are accounting
      for it multiple times. Instead we should be adjusting the bonus before
      traversing the user chain.
      
      This reduces the instruction count for CTMark (newPM-O3) when Function
      Specialization is enabled without actually reducing the amount of
      specializations performed (geomean: -0.001% non-LTO, -0.406% LTO).
      
      Differential Revision: https://reviews.llvm.org/D136692
      dbeaf6ba
    • Sanjay Patel's avatar
      [InstCombine] improve demanded bits for Sub operand 0 · d2d23795
      Sanjay Patel authored
      This is copying the code that was added for 'add' with D130075.
      (That patch removed a fallthrough in the cases, but we can
      probably still share at least some code again as a follow-up
      cleanup, but I didn't want to risk it here.)
      
      The reasoning is similar to the carry propagation for 'add':
      if we don't demand low bits of the subtraction and the
      subtrahend (aka RHS or operand 1) is known zero in those low
      bits, then there can't be any borrowing required from the
      higher bits of operand 0, so the low bits don't matter.
      
      Also, the no-wrap flags can be propagated (and I think that
      should be true for add too).
      
      Here's an attempt to prove that in Alive2:
      https://alive2.llvm.org/ce/z/xqh7Pa
      (can add nsw or nuw to src and tgt, and it should still pass)
      
      Differential Revision: https://reviews.llvm.org/D136788
      d2d23795
    • Liming Liu's avatar
      [P0857R0 Part-B] Allows `require' clauses appearing in · ae48d1c7
      Liming Liu authored
      template-template parameters. Although it effects whether a template can be
      used as an argument for another template, the constraint seems not to
      be checked, nor other major implementations (GCC, MSVC, et al.) check it.
      
      Additionally, Part-A of the document seems to have been implemented.
      So mark P0857R0 as completed.
      
      Differential Revision: https://reviews.llvm.org/D134128
      ae48d1c7
    • gonglingqin's avatar
      bba97c3c
    • John Brawn's avatar
      [MachineCSE] Allow PRE of instructions that read physical registers · 628467e5
      John Brawn authored
      Currently MachineCSE forbids PRE when the instruction reads a physical
      register. Relax this so that it's allowed when the value being read is
      the same as what would be read in the place the instruction would be
      hoisted to.
      
      This is being done in preparation for adding FPCR handling to the
      AArch64 backend, in order to prevent it to from worsening the
      generated code, but for targets that already have a similar register
      it should improve things.
      
      This patch affects code generation in several tests. The new code
      looks better except for in Thumb2/LowOverheadLoops/memcall.ll where
      we perform PRE but the LowOverheadLoops transformation then undoes
      it. Also in AMDGPU/selectcc-opt.ll the CHECK makes things look worse,
      but actually the function as a whole is better (as a MOV is PRE'd).
      
      Differential Revision: https://reviews.llvm.org/D136675
      628467e5
    • LLVM GN Syncbot's avatar
      [gn build] Port b51b90d6 · 5ed5b2a2
      LLVM GN Syncbot authored
      5ed5b2a2
    • LLVM GN Syncbot's avatar
      [gn build] Port 17059753 · 93736d41
      LLVM GN Syncbot authored
      93736d41