1. Feb 03, 2024
    • Simon Pilgrim's avatar
      [AMDGPU] Regenerate ctpop64.ll test checks · faeb3d1f
      Simon Pilgrim authored
      faeb3d1f
    • Simon Pilgrim's avatar
      [X86] Allow i8 CTPOP expansion to work with a 'shifted' active bits value of 8 bits or less · 3a758076
      Simon Pilgrim authored
      Shift down the value so the active bits are at the lsb
      3a758076
    • Krystian Stasiowski's avatar
      [Clang][Sema] Correctly look up primary template for variable template specializations (#80359) · 7ecfb66c
      Krystian Stasiowski authored
      Consider the following:
      ```
      namespace N0 {
        namespace N1 {
          template<typename T>
          int x1 = 0;
        }
        using namespace N1;
      }
      template<>
      int N0::x1<int>;
      ```
      
      According to [dcl.meaning.general] p3.3:
      > - If the _declarator_ declares an explicit instantiation or a partial
      or explicit specialization, the _declarator_ does not bind a name. If it
      declares a class member, the terminal name of the _declarator-id_ is not
      looked up; otherwise, **only those lookup results that are nominable in
      `S` are considered when identifying any function template specialization
      being declared**.
      
      In particular, the requirement for lookup results to be nominal in the
      lookup context of the terminal name of the _declarator-id_ only applies
      to function template specializations -- not variable template
      specializations. We currently reject the above declaration, but we do
      (correctly) accept it if the using-directive is replaced with a `using`
      declaration naming `N0::N1::x1`. This patch makes it so the above
      specialization is (correctly) accepted.
      7ecfb66c
    • Nathan Gauër's avatar
      [SPIR-V] add convergence region analysis (#78456) · 7b08b436
      Nathan Gauër authored
      This new analysis returns a hierarchical view of the convergence regions
      in the given function.
      This will allow our passes to query which basic block belongs to which
      convergence region, and structurize the code in consequence.
      
      Definition
      ----------
      
      A convergence region is a CFG with:
       - a single entry node.
       - one or multiple exit nodes (different from LLVM's regions).
       - one back-edge
       - zero or more subregions.
      
      Excluding sub-regions nodes, the nodes of a region can only reference a
      single convergence token. A subregion uses a different convergence
      token.
      
      Algorithm
      ---------
      
      This algorithm assumes all loops are in the Simplify form.
      
      Create an initial convergence region for the whole function.
        - the convergence token is the function entry token.
        - the entry is the function entrypoint.
      - Exits are all the basic blocks terminating with a return instruction.
      
      Take the function CFG, and process it in DAG order (ignoring
      back-edges). If a basic block is a loop header:
       - Create a new region.
      - The parent region is the parent's loop region if any, otherwise, the
      top level region.
         - The region blocks are all the blocks belonging to this loop.
      - For each loop exit: - visit the rest of the CFG in DAG order (ignore
      back-edges). - if the region's convergence token is found, add all the
      blocks dominated by the exit from which the token is reachable to the
      region.
         - continue the algorithm with the loop headers successors.
      7b08b436
    • Manish Kausik H's avatar
      [SelectionDAG] Use unaligned store to move AVX registers onto stack for `extractelement` (#78422) · a768bc6e
      Manish Kausik H authored
      Prior to this patch, SelectionDAG generated aligned move onto stacks for
      AVX registers when the function was marked as a no-realign-stack
      function. This lead to misalignment between the stack and the
      instruction generated. This patch fixes the issue.
      
      Fixes #77730
      a768bc6e
    • Philip Reames's avatar
      [TTI] Add costing for vp.strided.load and vp.strided.store (#80360) · b78b2645
      Philip Reames authored
      The primary motivation of this patch is to add testing infrastructure
      atop the recently landed 8ad14b6d, so
      that we can separate the costing aspects of strided memory operations
      from the SLP implementation details.
      
      I want to be clear that I am *not* proposing that we use the
      vp.strided.* forms as our canonical IR representation. I'm merely using
      them as a testing vehicle to exercise the costing machinery. The
      canonical IR form remains a masked.gather or masked.scatter. I do want
      to explore adding a non-vp strided load/store intrinsic, but that's a
      separate line of work.
      
      There is one costing change included in this. As I wrote my test, I
      discovered that the default implementation was scalarized (if invoked
      via generic routines such as getInstructionCost), and when adding the
      call into the strided specific costing discovered that we hadn't modeled
      the fallback to scalarization properly in the initial patch. After
      fixing that, there is a minor difference in scalarization cost reported
      for the unaligned case but I believe that to be uninteresting.
      
      For the record, I did confirm that vp.strided.store is lowered to a
      strided store on RISCV. :)
      b78b2645
    • Rushi Bhamani's avatar
    • Philip Reames's avatar
      [LSR][term-fold] Adjust expansion budget based on trip count (#80304) · 28865da3
      Philip Reames authored
      Follow up to https://github.com/llvm/llvm-project/pull/74747
      
      
      
      This change extends the previously added fixed expansion threshold by
      scaling down the cost allowed for an expansion for a loop with either a
      small known trip count or a profile which indicates the trip count is
      likely small. The goal here is to improve code generation for a loop
      nest where the outer loop has a high trip count, and the inner loop runs
      only a handful of iterations.
      
      ---------
      
      Co-authored-by: default avatarNikita Popov <github@npopov.com>
      28865da3
    • Guillaume Chatelet's avatar
      b629414a
    • LLVM GN Syncbot's avatar
      [gn build] Port 67eee4a0 · 30503116
      LLVM GN Syncbot authored
      30503116
    • Nikolas Klauser's avatar
      [libc++] Optimize vector growing of trivially relocatable types (#76657) · 67eee4a0
      Nikolas Klauser authored
      This patch introduces a new trait to represent whether a type is
      trivially
      relocatable, and uses that trait to optimize the growth of a std::vector
      of trivially relocatable objects.
      
      ```
      --------------------------------------------------
      Benchmark                           old        new
      --------------------------------------------------
      bm_grow<int>                    1354 ns    1301 ns
      bm_grow<std::string>            5584 ns    3370 ns
      bm_grow<std::unique_ptr<int>>   3506 ns    1994 ns
      bm_grow<std::deque<int>>       27114 ns   27209 ns
      ```
      
      This also changes to order of moving and destroying the objects when
      growing the vector. This should not affect our conformance.
      67eee4a0
    • Natalie Chouinard's avatar
      7524b037
  2. Feb 02, 2024