1. Dec 01, 2023
  2. Nov 30, 2023
    • Nikita Popov's avatar
      [InstSimplify] Fix select bit test miscompile with disjoint · 07c18a05
      Nikita Popov authored
      The select condition ensures the disjointness here. The transform
      is not valid without dropping the flag, which InstSimplify can't
      do.
      07c18a05
    • Nikita Popov's avatar
      c89553ae
    • Lucas Prates's avatar
      8eb70532
    • Philip Reames's avatar
      [RISCV] Add combines to form binop from tail insert idioms (#72675) · ff5e536b
      Philip Reames authored
      This patch contains two related combines:
      1) If we have an scalar vector insert into the result of a
      concat_vector,
         sink the insert into the operand of the concat.
      2) If we have a insert of a scalar binop into a vector binop of the
         same opcode and the RHS of both are constant, perform the insert
         and then the binop.
      
      The common theme to both is pushing inserts closer to the sources of the
      computation graph. The goal is to enable forming vector bin ops from
      inserts of scalar binops at the end of another vector.
      
      For RISCV specifically, the concat_vector transform will push inserts to
      smaller vectors. This will have the effect of reducing lmul for the
      vslides, and usually doesn't require an additional vsetvli since
      the source vectors are already working in the narrower VL.   I tried
      that one as a target independent combine first, and it doesn't appear
      profitable on all targets.
      
      This is only one approach to the problem. Another idea would be to
      aggressively form build_vectors and subvector inserts from the
      individual scalar inserts, and then have a transform which sunk a
      subvector_insert down through the concat. The advantage of the alternate
      approach is that we expose parallelism in the insert sequence, even if
      the source vector isn't a concat_vector. If reviewers are okay with it,
      I'd like to start with this approach, and then explore that direction in
      a follow up patch.
      ff5e536b
    • Dinar Temirbulatov's avatar
      [AArch64][SME2] Add multi-vector SEL (x2, x4) ACLE builtins & intrinsics (#73188) · 0ef013c8
      Dinar Temirbulatov authored
      Add multi-vector SEL (x2, x4) ACLE builtins & intrinsics
      Patch by: David Sherwood <david.sherwood@arm.com>
      0ef013c8
    • Jeremy Morse's avatar
      [DebugInfo][RemoveDIs] Support maintaining DPValues in CodeGenPrepare (#73660) · 3ef98bcd
      Jeremy Morse authored
      CodeGenPrepare needs to support the maintenence of DPValues, the
      non-instruction replacement for dbg.value intrinsics. This means there are
      a few functions we need to duplicate or replicate the functionality of:
       * fixupDbgValue for setting users of sunk addr GEPs,
       * The remains of placeDbgValues needs a DPValue implementation for sinking
       * Rollback of RAUWs needs to update DPValues
       * Rollback of instruction removal needs supporting (see github #73350)
       * A few places where we have to use iterators rather than instructions.
      
      There are three places where we have to use the setHeadBit call on
      iterators to indicate which portion of debug-info records we're about to
      splice around. This is because CodeGenPrepare, unlike other optimisation
      passes, is very much concerned with which block an operation occurs in and
      where in the block instructions are because it's preparing things to be in
      a format that's good for SelectionDAG.
      
      Ther...
      3ef98bcd
    • Piotr Sobczak's avatar
      [AMDGPU] Add test for GCNRegPressure tracker bug (#73786) · 73d9f5fd
      Piotr Sobczak authored
      Add a test to document an existing problem in GCNRegPressure tracker.
      
      The upward tracker does not count the registers used (16 of them) in
      movrel instruction (for example V_INDIRECT_REG_WRITE_MOVREL_B32_V16).
      
      The downward tracker counts the registers but reports a mismatch:
      %0:L0000000000000C00 isn't found in LIS reported set
      73d9f5fd
    • Nikita Popov's avatar
      [InstCombine] Fix phi or icmp fold with disjoint flag · d8bc5465
      Nikita Popov authored
      We're changing the operand of the or here, such that the disjoint
      flag may no longer hold. Clear it.
      d8bc5465
    • Nikita Popov's avatar
      b7af286a
    • Simon Pilgrim's avatar
      [X86] X86InstrFoldTables.cpp - add Op4 Broadcast Fold/Unfold table entries · b8bbd5fe
      Simon Pilgrim authored
      Prep work for #73509 (missed in #73654)
      b8bbd5fe
    • Alexey Bataev's avatar
      [SLP][NFC] Unify code for cost estimation/codegen for buildvector, NFC. (#73182) · ba523106
      Alexey Bataev authored
      This just moves towards reusing same function for both cost
      estimation/codegen for buildvector.
      ba523106
    • Sam Tebbs's avatar
      [AArch64] Warn when calling a NEON builtin in a streaming function (#73672) · 5234fe31
      Sam Tebbs authored
      This patch introduces a warning that is emitted when a Neon builtin is
      called from a streaming function, as that situation is not supported.
      
      Uses work by Kerry McLaughlin.
      5234fe31
    • Alexey Bataev's avatar
      [SLP]Fix/improve minbitwidth mapping to use TreeEntry as a key. · 1f88e62d
      Alexey Bataev authored
      Currently, MinBWs map uses Value* as a key and stores mapping for each
      value to be demoted. It make is it hard to get the actual MinBWs value
      for the buildvector scalars(constants), since same constant might be
        used in different nodes with the different MinBWs values/decisions.
      Also, it consumes extra memory for the vectorized values/instructions
       from the same nodes.
      Better to map actual nodes. It fixes the bitwidth data fetching for
      buildvector scalars and improves memory consumption/analysis time for
      other instructions.
      1f88e62d
    • Dominik Wójt's avatar
      4c338f80
    • Shengchen Kan's avatar
      [X86][tablgen] Auto-gen broadcast tables (#73654) · a4e1aa25
      Shengchen Kan authored
      1. Add TB_BCAST_SH for FP16
      2. Auto-gen 4 broadcast tables BroadcastTable[1-4]
      
      issue: https://github.com/llvm/llvm-project/issues/66360
      a4e1aa25
    • David Spickett's avatar
      [llvm][AArch64] Preserve regmask when expanding the BLR_BTI pseudo instruction (#73927) · 99d48591
      David Spickett authored
      Fixes #73787
      
      Not doing so lead to us making use of a register after the call, which
      has been clobbered by the call.
      
      Added an MIR test that runs only the pseudo expansion pass.
      99d48591
    • frgossen's avatar
      Fix bazel build (#73942) · 402591a4
      frgossen authored
      402591a4
    • Hui's avatar
      [libc++] Workaround linker errors in floating-point atomic tests (#73398) · 6677f029
      Hui authored
      We now add -latomic whenever we detect that it's supported on the platform,
      and we mark the tests as UNSUPPORTED on platforms where non-lockfree
      atomics are not supported.
      6677f029
    • Lucas Duarte Prates's avatar
      [AArch64] Fix predicates for FEAT_CPA's SVE-specific instructions (#73923) · 78237b70
      Lucas Duarte Prates authored
      Following up on #73777, this fixes the predicate for the SVE-specific
      FEAT_CPA instructions to require SVE instead of SVE or SME. These
      instructions should not be availabe if only SME is enabled.
      78237b70
    • Adam Paszke's avatar
      [MLIR][CUDA] Update export macros in CudaRuntimeWrappers (#73932) · 1c2a0768
      Adam Paszke authored
      This fixes a few issues present in the current version:
      1) The macro doesn't enforce the default visibility on exported
         functions, causing compilation to fail when using
         `-fvisibility=hidden`
      2) Not all functions are exported
      3) Sometimes the macro ended up weirdly interleaved with `extern "C"`
         declarations
      1c2a0768
    • Nikita Popov's avatar
      [InstCombine] Add KnownBits consistency assertion behind option (NFC) · 10b44fb6
      Nikita Popov authored
      I'm occasionally using this to find cases where computeKnownBits()
      and SimplifyDemandedBits() went out of sync.
      
      This option is not enabled by default (even under EXPENSIVE_CHECKS)
      because it has a number of known failures in our tests. The reason
      for this failures is that computeKnownBits() performs recursive
      queries using the original context instruction, while
      SimplifyDemandedBits() uses the current instruction. This is
      something we can improve, but using the original context wouldn't
      always be safe in this context (when non-speculatable instructions
      are involved).
      10b44fb6
    • Jeremy Morse's avatar
      [DebugInfo] Set all dbg.value intrinsics to be tail-calls (#73661) · cd02e4b8
      Jeremy Morse authored
      This change has no meaningful effect on the compiler, although it has a
      functional effect of dbg.value intrinsics being printed differently. The
      tail-call flag is meaningless for debug-intrinsics and doesn't serve a
      purpose, it's just extra baggage that dbg.values are built on top of.
      Some facilities create debug-intrinsics with the flag, others don't.
      However, the RemoveDIs project to represent debug-info without
      intrinsics doesn't have a corresponding flag, which can cause spurious
      test differences.
      
      Specifically: we can convert a dbg.value to a DPValue, run an
      optimisation pass, then convert the DPValue back to dbg.value form.
      Right now, we always set the "tail" flag when converting it back. This
      causes the auto-update-tests script to fail sometimes because in one
      mode (dbg.value) intrinsics might not have a tail flag, but in the other
      they do have a tail flag. Consistently picking one or the other in the
      conversion routine doesn't help, because the rest of LLVM is
      inconsistent about it anyway.
      
      Thus: whenever we make a dbg.value intrinsic, create it as a tail call,
      so that we get consistent output behaviours no matter which debug-info
      mode we're in, DPValue or dbg.value. No tests fail as a result of this
      patch because the extra 'tail' generated in numerous tests is
      automatically ignored by FileCheck as being leading-rubbish before the
      CHECK match.
      cd02e4b8
    • leecheechen's avatar
      [LoongArch] Add some binary IR instructions testcases for LSX (#73929) · 29a0f3ec
      leecheechen authored
      The IR instructions include:
      - Binary Operations: add fadd sub fsub mul fmul udiv sdiv fdiv
      - Bitwise Binary Operations: shl lshr ashr
      29a0f3ec
    • Matt Arsenault's avatar
      MachineVerifier: Reject extra non-register operands on instructions (#73758) · c44dca15
      Matt Arsenault authored
      We were allowing extra immediate arguments, and only bothering to check
      if registers were implicit or not.
      
      Also consolidate extra operand checks in verifier, to make this
      testable. We had 3 different places checking if you were trying to build
      an instruction with more operands than allowed by the definition. We had
      an assertion in addOperand, a direct check in the MIRParser to avoid the
      assertion, and the machine verifier checks. Remove the assert and parser
      check so the verifier can provide a consistent verification experience,
      which will also handle instructions modified in place.
      c44dca15
    • Nikita Popov's avatar
      [InstCombine] Require non-demanded known bits to be accurate (NFC) · 2031e722
      Nikita Popov authored
      In practice this is already true, and having this as an explicit
      guarantee is useful for #72912. I don't think there is any good
      reason why we would want to produce incorrect KnownBits results
      for non-demanded bits.
      2031e722
    • Nikita Popov's avatar
      [InstCombine] Use analyzeKnownBitsFromAndXorOr() in multi-use demanded bits · 062058ef
      Nikita Popov authored
      We were using this helper in single-use demanded bits but not
      multi-use demanded bits.
      
      This fixes an assertion failure when asserting consistency between
      computeKnownBits() and SimplifyDemandedBits().
      062058ef
    • Antonio Frighetto's avatar
      [InstCombine] Switch to use FileCheck as UTC was favoured (NFC) · 3f6a8e9b
      Antonio Frighetto authored
      FileCheck was previously missing while moving to UTC, as part of
      regenerating other tests within InstCombine.
      3f6a8e9b
    • Nikita Popov's avatar
      [ValueTracking] Switch analyzeKnownBitsFromAndXorOr() to use SimplifyQuery (NFC) · 1566380e
      Nikita Popov authored
      It already used it internally, make the public API use it as well.
      1566380e
    • Jeremy Morse's avatar
      [DebugInfo][RemoveDIs] Have LICM insert at iterator positions (#73671) · 5ba5211a
      Jeremy Morse authored
      Because we're storing some extra debug-info information in the iterator
      class, we need to insert new LICM-created stores using such iterators.
      Switch LICM to storing iterators instead of pointers when it promotes
      variables in loops, add a test for the desired behaviour, and enable
      RemoveDIs instrumentation on a variety of other LICM tests for good
      measure.
      
      (This would appear to be the only pass in LLVM that needs to store
      iterators on the heap).
      5ba5211a
    • Nikita Popov's avatar
      [InstCombine] Use pointer alignment in SimplifyDemandedBits · b8a5a015
      Nikita Popov authored
      For parity with computeKnownBits(). This came up when adding a
      consistency assertion.
      b8a5a015
    • Guillaume Chatelet's avatar
      [libc] Add more functions in CPP/bit.h (#73814) · b703bd82
      Guillaume Chatelet authored
      Once this is submitted we can remove `include/__support/bit.h` that
      duplicates some of this functionality.
      b703bd82
    • Jeremy Morse's avatar
      [DebugInfo][RemoveDIs] Emulate inserting insts in dbg.value sequences (#73350) · 2ec0283c
      Jeremy Morse authored
      Here's a problem for the RemoveDIs project to make debug-info not be
      stored in instructions -- in the following sequence:
          dbg.value(foo
          %bar = add i32 ...
          dbg.value(baz
      It's possible for rare passes (only CodeGenPrepare) to remove the add
      instruction, and then re-insert it back in the same place. When
      debug-info is stored in instructions and there's a total order on "when"
      things happen this is easy, but by moving that information out of the
      instruction stream we start having to do manual maintenance.
      
      This patch adds some utilities for re-inserting an instruction into a
      sequence of DPValue objects. Someday we hope to design this away, but
      for now it's necessary to support all the things you can do with
      dbg.values. The two unit tests show how DPValues get shuffled around
      using the relevant function calls. A follow-up patch adds
      instrumentation to CodeGenPrepare.
      2ec0283c
    • Guillaume Chatelet's avatar