1. Jun 03, 2024
    • Nathan Gauër's avatar
      [SPIR-V] Add pass to merge convergence region exit targets (#92531) · a5641f10
      Nathan Gauër authored
      
      
      The structurizer required regions to be SESE: single entry, single exit.
      This new pass transforms multiple-exit regions into single-exit regions.
      
      ```
            +---+
            | A |
            +---+
            /   \
         +---+ +---+
         | B | | C |  A, B & C belongs to the same convergence region.
         +---+ +---+
           |     |
         +---+ +---+
         | D | | E |  C & D belongs to the parent convergence region.
         +---+ +---+  This means B & C are the exit blocks of the region.
            \   /     And D & E the targets of those exits.
             \ /
              |
            +---+
            | F |
            +---+
      ```
      
      This pass would assign one value per exit target:
      B = 0
      C = 1
      
      Then, create one variable per exit block (B, C), and assign it to the
      correct value: in B, the variable will have the value 0, and in C, the
      value 1.
      
      Then, we'd create a new block H, with a PHI node to gather those 2
      variables, and a switch, to route to the correct target.
      
      Finally, the branches in B and C are updated to exit to this new block.
      
      ```
            +---+
            | A |
            +---+
            /   \
         +---+ +---+
         | B | | C |
         +---+ +---+
            \   /
            +---+
            | H |
            +---+
            /   \
         +---+ +---+
         | D | | E |
         +---+ +---+
            \   /
             \ /
              |
            +---+
            | F |
            +---+
      ```
      
      Note: the variable is set depending on the condition used to branch. If
      B's terminator was conditional, the variable would be set using a
      SELECT.
      All internal edges of a region are left intact, only exiting edges are
      updated.
      
      ---------
      
      Signed-off-by: default avatarNathan Gauër <brioche@google.com>
      a5641f10
    • Kareem Ergawy's avatar
      [flang] Emit `argNo` debug info only for `func` block args (#93921) · 5bfc4445
      Kareem Ergawy authored
      Fixes a bug uncovered by
      [pr43337.f90](https://github.com/llvm/llvm-test-suite/blob/main/Fortran/gfortran/regression/gomp/pr43337.f90)
      in the test suite.
      
      In particular, this emits `argNo` debug info only if the parent op of a
      block is a `func.func` op. This avoids DI conflicts when a function
      contains a nested OpenMP region that itself has block arguments with DI
      attached to them; for example, `omp.parallel` with delayed privatization
      enabled.
      5bfc4445
    • AlexGhiti's avatar
      [RISCV] Remove experimental from Zabha (#93831) · 6b744496
      AlexGhiti authored
      
      
      The Zabha extension was ratified in April 2024.
      
      Co-authored-by: default avatarAlexandre Ghiti <alexghiti@rivosinc.com>
      6b744496
    • Pavel Labath's avatar
    • David Spickett's avatar
      [lldb][test] Fix D lang mangling test on Windows (#94196) · 6abf3619
      David Spickett authored
      On Windows the function does not have a symbol associated with it:
      Function: id = {0x000001c9}, name = "_Dfunction", range =
      [0x0000000140001000-0x0000000140001004)
            LineEntry: <...>
      
      Whereas it does on Linux:
      Function: id = {0x00000023}, name = "_Dfunction", range =
      [0x0000000000000734-0x0000000000000738)
            LineEntry: <...>
      Symbol: id = {0x00000058}, range =
      [0x0000000000000734-0x0000000000000738), name="_Dfunction"
      
      This means that frame.symbol is not valid on Windows.
      
      However, frame.function is valid and it also has a "mangled" attribute.
      
      So I've updated the test to check the symbol if we've got it, and the
      function always.
      
      In both cases we check that mangled is empty (meaning it has not been
      treated as mangled) and that the display name matches the original
      symbol name.
      6abf3619
    • David Spickett's avatar
      770b6c79
    • Tom Eccles's avatar
      [flang][OpenMP][NFC] Reduce FunctionFiltering pass boilerplate (#93951) · 6a217307
      Tom Eccles authored
      The pass constructor can be generated automatically.
      
      This pass doesn't need to be adapted to support other top level
      operations because it is specifically supposed to filter functions. We
      don't need to filter non-function top level operations because without
      use inside of functions they shouldn't lead to any codegen.
      6a217307
    • David Stone's avatar
    • David Spickett's avatar
      [lldb][test] Skip D lang mangling test on Windows · 09c06079
      David Spickett authored
      While the fix is reviewed.
      09c06079
    • Pavel Labath's avatar
      [lldb] Avoid (unlimited) GetNumChildren calls when printing values (#93946) · 763b96c8
      Pavel Labath authored
      For some data formatters, even getting the number of children can be an
      expensive operations (e.g., needing to walk a linked list to determine
      the number of elements). This is then wasted work when we know we will
      be printing only small number of them.
      
      This patch replaces the calls to GetNumChildren (at least those on the
      "frame var" path) with the calls to the capped version, passing the
      value of `max-children-count` setting (plus one)
      763b96c8
    • Vyacheslav Levytskyy's avatar
      [SPIR-V] Validate type of the last parameter of OpGroupWaitEvents (#93661) · ce73e17e
      Vyacheslav Levytskyy authored
      This PR fixes invalid OpGroupWaitEvents emission to ensure that SPIR-V
      Backend inserts a bitcast before OpGroupWaitEvents if the last argument
      is a pointer that doesn't point to OpTypeEvent.
      ce73e17e
    • David Green's avatar
      [ARM] Convert vector fdiv+fcvt fixed-point combine to fmul. · 264b1b24
      David Green authored
      Instcombine will convert fdiv by a power-2 to fmul, this converts the
      PerformVDIVCombine that converts fdiv+fcvt to fixed-point fcvt to fmul+fcvt.
      The fdiv tests will look worse, but won't appear in practice (and should be
      improved again by #93882).
      264b1b24
    • Sander de Smalen's avatar
      [AArch64] Avoid NEON ORR when NEON and SVE are unavailable (#93940) · b71434f8
      Sander de Smalen authored
      For streaming-compatible functions with only +sme, we can't use
      a NEON ORR (aliased as 'mov') for copies of Q-registers, so
      we need to use a spill/fill instead.
      
      This also fixes the fill, which should use the post-incrementing
      addressing mode.
      b71434f8
    • Chuanqi Xu's avatar
      [serialization] no transitive decl change (#92083) · ccb73e88
      Chuanqi Xu authored
      Following of https://github.com/llvm/llvm-project/pull/86912
      
      #### Motivation Example
      
      The motivation of the patch series is that, for a module interface unit
      `X`, when the dependent modules of `X` changes, if the changes is not
      relevant with `X`, we hope the BMI of `X` won't change. For the specific
      patch, we hope if the changes was about irrelevant declaration changes,
      we hope the BMI of `X` won't change. **However**, I found the patch
      itself is not very useful in practice, since the adding or removing
      declarations, will change the state of identifiers and types in most
      cases.
      
      That said, for the most simple example,
      
      ```
      // partA.cppm
      export module m:partA;
      
      // partA.v1.cppm
      export module m:partA;
      export void a() {}
      
      // partB.cppm
      export module m:partB;
      export void b() {}
      
      // m.cppm
      export module m;
      export import :partA;
      export import :partB;
      
      // onlyUseB;
      export module onlyUseB;
      import m;
      export inline void onluUseB() {
          b();
      }
      ```
      
      the BMI of `onlyUseB` will change after we change the implementation of
      `partA.cppm` to `partA.v1.cppm`. Since `partA.v1.cppm` introduces new
      identifiers and types (the function prototype).
      
      So in this patch, we have to write the tests as:
      
      ```
      // partA.cppm
      export module m:partA;
      export int getA() { ... }
      export int getA2(int) { ... }
      
      // partA.v1.cppm
      export module m:partA;
      export int getA() { ... }
      export int getA(int) { ... }
      export int getA2(int) { ... }
      
      // partB.cppm
      export module m:partB;
      export void b() {}
      
      // m.cppm
      export module m;
      export import :partA;
      export import :partB;
      
      // onlyUseB;
      export module onlyUseB;
      import m;
      export inline void onluUseB() {
          b();
      }
      ```
      
      so that the new introduced declaration `int getA(int)` doesn't introduce
      new identifiers and types, then the BMI of `onlyUseB` can keep
      unchanged.
      
      While it looks not so great, the patch should be the base of the patch
      to erase the transitive change for identifiers and types since I don't
      know how can we introduce new types and identifiers without introducing
      new declarations. Given how tightly the relationship between
      declarations, types and identifiers, I think we can only reach the ideal
      state after we made the series for all of the three entties.
      
      #### Design details
      
      The design of the patch is similar to
      https://github.com/llvm/llvm-project/pull/86912, which extends the
      32-bit DeclID to 64-bit and use the higher bits to store the module file
      index and the lower bits to store the Local Decl ID.
      
      A slight difference is that we only use 48 bits to store the new DeclID
      since we try to use the higher 16 bits to store the module ID in the
      prefix of Decl class. Previously, we use 32 bits to store the module ID
      and 32 bits to store the DeclID. I don't want to allocate additional
      space so I tried to make the additional space the same as 64 bits. An
      potential interesting thing here is about the relationship between the
      module ID and the module file index. I feel we can get the module file
      index by the module ID. But I didn't prove it or implement it. Since I
      want to make the patch itself as small as possible. We can make it in
      the future if we want.
      
      Another change in the patch is the new concept Decl Index, which means
      the index of the very big array `DeclsLoaded` in ASTReader. Previously,
      the index of a loaded declaration is simply the Decl ID minus
      PREDEFINED_DECL_NUMs. So there are some places they got used
      ambiguously. But this patch tried to split these two concepts.
      
      #### Overhead
      
      As https://github.com/llvm/llvm-project/pull/86912 did, the change will
      increase the on-disk PCM file sizes. As the declaration ID may be the
      most IDs in the PCM file, this can have the biggest impact on the size.
      In my experiments, this change will bring 6.6% increase of the on-disk
      PCM size. No compile-time performance regression observed. Given the
      benefits in the motivation example, I think the cost is worthwhile.
      ccb73e88
    • Stephan T. Lavavej's avatar
      [libc++][test] Mark `optional` test functions as `TEST_CONSTEXPR_CXX20` (#94172) · 84742cd8
      Stephan T. Lavavej authored
      [P2231R1](https://wg21.link/P2231R1) "Missing `constexpr` in
      `std::optional` and `std::variant`" was accepted as a C++20 Defect
      Report, not a C++17 Defect Report. Accordingly, `test_empty_emplace()`
      and `check_reset()` should be marked as `TEST_CONSTEXPR_CXX20`. Note
      that their `static_assert`s are properly guarded:
      
      
      https://github.com/llvm/llvm-project/blob/4ce65423be0ba1d90c11b6a79981d6314e1cf36d/libcxx/test/std/utilities/optional/optional.object/optional.object.assign/emplace.pass.cpp#L270-L272
      
      
      https://github.com/llvm/llvm-project/blob/4ce65423be0ba1d90c11b6a79981d6314e1cf36d/libcxx/test/std/utilities/optional/optional.object/optional.object.mod/reset.pass.cpp#L53-L55
      
      Found while running libc++'s tests with MSVC's STL, as we activate our
      `constexpr` here for C++20 and above.
      84742cd8
    • Owen Pan's avatar
      [clang-format][doc] Minor cleanup · f4a7f81a
      Owen Pan authored
      f4a7f81a
    • David Green's avatar
    • Chuanqi Xu's avatar
      [NFC] [C++20] [Modules] [Reduced BMI] Reorder Emitting reduced BMI and normal BMI for named modules · a41a20bd
      Chuanqi Xu authored
      When we generate the reduced BMI on the fly, the order of the emitting
      phase is different within `-emit-obj` and `-emit-module-interface`.
      Although this is meant to be fine, we observed it in
      https://github.com/llvm/llvm-project/issues/93859 (that the different phase order may cause problems).
      Also it turns out to be a different fundamental reason to the orders.
      
      But it might be fine to make the order of emitting reducing BMI at first
      to avoid such confusions in the future.
      a41a20bd
    • Haojian Wu's avatar
      [bazel]: port for the libc change 142afde0 · 12c85cd3
      Haojian Wu authored
      12c85cd3
    • David Green's avatar
      [ARM] Rewrite vdiv_combine.ll test. NFC · ef4c91c4
      David Green authored
      Instcombine will convert the fdiv by constant to fmul. This cleans up the
      vdiv_combine.ll test and adds fmul variants of the existing fdiv test.
      ef4c91c4
    • Chuanqi Xu's avatar
      [C++20] [Modules] [Reduced BMI] Handling Deduction Guide in reduced BMI · a68638bf
      Chuanqi Xu authored
      carefully
      
      Close https://github.com/llvm/llvm-project/issues/93859
      
      The direct pattern of the issue is that, in a reduced BMI, we're going
      to wrtie a class but we didn't write the deduction guide. Although we
      handled deduction guide, but we tried to record the found deduction
      guide from `noload_lookup` directly.
      
      It is slightly problematic if the found deduction guide is from AST.
      e.g.,
      
      ```
      module;
      export module m;
      import xxx; // Also contains the class and the deduction guide
      ...
      ```
      
      Then when we writes the class in the current file, we tried to record
      the deduction guide, but `noload_lookup` returns the deduction guide
      from the AST file then we didn't record the local deduction guide. Then
      mismatch happens.
      
      To mitiagte the problem, we tried to record the canonical declaration
      for the decution guide.
      a68638bf
    • martinboehme's avatar
      [clang][dataflow] Rewrite `getReferencedDecls()` with a `RecursiveASTVisitor`. (#93461) · 5161a3f6
      martinboehme authored
      We previously had a hand-rolled recursive traversal here that was
      exactly what
      `RecursiveASTVistor` does anyway. Using the visitor not only eliminates
      the
      explicit traversal logic but also allows us to introduce a common
      visitor base
      class for `getReferencedDecls()` and `ResultObjectVisitor`, ensuring
      that the
      two are consistent in terms of the nodes they visit. Inconsistency
      between these
      two has caused crashes in the past when `ResultObjectVisitor` tried to
      propagate
      result object locations to entities that weren't modeled becasue
      `getReferencedDecls()` didn't visit them.
      5161a3f6
    • David Stone's avatar
    • AtariDreams's avatar
      Reland "[InstCombine] Fold (sub nuw X, (Y << nuw Z)) >>u exact Z --> (X >>u... · 10e7671d
      AtariDreams authored
      Reland "[InstCombine] Fold (sub nuw X, (Y << nuw Z)) >>u exact Z --> (X >>u exact Z) sub nuw Y" (#93571)
      
      This is the same fold as ((X << nuw Z) sub nuw Y) >>u exact Z --> X sub
      nuw (Y >>u exact Z), but with the sub operands swapped.
      
      Alive2 Proof:
      https://alive2.llvm.org/ce/z/pT-RxG
      10e7671d
    • Kazu Hirata's avatar
      [memprof] Add accessors to Frame::SymbolName (#94085) · f367eaa4
      Kazu Hirata authored
      This patch adds accessors to Frame::SymbolName so that we can change
      the underlying type of SymbolName without affecting downstream users
      once they switch to the new accessors.
      
      Note that SymbolName is only used for debugging.  Changing the type of
      SymbolName from std::optional<std::string> to
      std::unique_ptr<std::string> cuts down sizeof(Frame) by half -- from
      64 bytes to 32 bytes.  (std::optional<T> sets aside the storage in
      case T is instantiated.)
      
      During deserialization, the memory usage is dominated by Frames.
      Shrinking the type cuts down the memory usage and deserialization time
      nearly by half.
      f367eaa4
    • hev's avatar
    • Dhruv Chawla's avatar
    • Enna1's avatar
      [BPI] Cache LoopExitBlocks to improve compile time (#93451) · f779ec7c
      Enna1 authored
      The `LoopBlock` stored in `LoopWorkList` consist of basic block and its
      loop data information. When iterate `LoopWorkList`, if estimated weight
      of a loop is not stored in `EstimatedLoopWeight`, `getLoopExitBlocks()`
      is called to get all exit blocks of the loop. The estimated weight of a
      loop is calculated by iterating over edges leading from basic block to
      all exit blocks of the loop. If at least one edge has unknown estimated
      weight, the estimated weight of loop is unknown and will not be stored
      in `EstimatedLoopWeight`. `LoopWorkList` can contain different blocks in
      a same loop, so there is wasted work that calls `getLoopExitBlocks()`
      for same loop multiple times.
      
      Since computing the exit blocks of loop is expensive and the loop
      structure is not mutated in Branch Probability Analysis, we can cache
      the result and improve compile time.
      
      With this change, the overall compile time for a file containing a very
      large loop is dropped by around 82%.
      f779ec7c
    • Owen Pan's avatar
    • Owen Pan's avatar
      [clang-format][doc] Clean up quotes, etc. · 2fbc9f21
      Owen Pan authored
      2fbc9f21
    • Joshua Cao's avatar
      [IR] Do not set `none` for function uwtable (#93387) · ab08df22
      Joshua Cao authored
      This avoids the pitfall where we set the uwtable to none:
      ```
      func.setUWTableKind(llvm::UWTableKind::None)
      ```
      `Attribute::getAsString()` would see an unknown attribute and fail an
      assertion. In this patch, we assert that we do not see a None uwtable
      kind.
      
      This also skips the check of `UWTableKind::Async`. It is dominated by
      the check of `UWTableKind::Default`, which has the same enum value
      (nfc).
      ab08df22
    • Kazu Hirata's avatar
      [memprof] Use const ref for IndexedRecord (#94114) · 4ce65423
      Kazu Hirata authored
      The type of *Iter here is "const IndexedMemProfRecord &" as defined in
      RecordLookupTrait.  Assigning *Iter to a variable of type
      "const IndexedMemProfRecord &" avoids a copy, reducing the cycle and
      instruction counts by 1.8% and 0.2%, respectively, with
      "llvm-profdata show" modified to deserialize all MemProfRecords.
      
      Note that RecordLookupTrait has an internal copy of
      IndexedMemProfRecord, so we don't have to worry about a dangling
      reference to a temporary.
      4ce65423
    • Owen Pan's avatar
    • Owen Pan's avatar
      [clang-format] Handle attributes before lambda return arrow (#94119) · 80303cb2
      Owen Pan authored
      Fixes #92657.
      80303cb2
    • Nikolas Klauser's avatar
      [libc++] Don't give functions C linkage (#94102) · 5367b2c8
      Nikolas Klauser authored
      There is no reason to give any of the functions C linkage. This makes
      all of the libc++ functions have C++ linkage, removing the need for
      `_LIBCPP_HIDE_FROM_ABI_C`.
      5367b2c8
    • Kazu Hirata's avatar
      [TableGen] Use llvm::unique (NFC) (#94163) · d9293519
      Kazu Hirata authored
      d9293519
    • Vlad Serebrennikov's avatar
      [clang][NFC] Update CWG issues list · c26a9938
      Vlad Serebrennikov authored
      c26a9938
    • Stephan T. Lavavej's avatar
      [libc++] [test] Cleanup compile-only tests (#94121) · df9167bf
      Stephan T. Lavavej authored
      I noticed that these tests had empty `main` functions. Dropping them and
      renaming the tests to `MEOW.compile.pass.cpp` will slightly improve test
      throughput.
      df9167bf
  2. Jun 02, 2024