1. Jan 12, 2024
    • Yingwei Zheng's avatar
      [InstCombine] Fold `icmp pred (inttoptr X), (inttoptr Y) -> icmp pred X, Y` (#77832) · 2aae304c
      Yingwei Zheng authored
      NOTE: Alive2 proofs are unavailable because `inttoptr` is unsupported.
      2aae304c
    • Alexander Yermolovich's avatar
      [LLVM][DWARF] Fix accelerator table switching between CU and TU (#77511) · d199ab46
      Alexander Yermolovich authored
      Bug 1 is triggered when a TU is already created, and we process the same
      DICompositeType at a top level. We would switch to TU accelerator table,
      but
      would not switch back on early exit. As the result we would add CU
      entries to the TU
      accelerator table. When we try to write out TUs and normalize entries,
      the
      offsets for DIEs that are part of a CU would not have been computed, and
      it
      would assert on getOffset().
      
      Bug 2 is triggered when processing nested TUs. When we exit from
      addDwarfTypeUnitType we switched back to CU accelerator table. If we
      were processing nested TUs, the rest of the entries from TUs would be
      added to CU accelerator table. When we write out TUs, all the DIE
      pointers will become invalid. Eventually it will assert during
      normalization step after CU is processed.
      d199ab46
    • Benjamin Maxwell's avatar
      [mlir][ArmSME] Add rudimentary support for tile spills to the stack (#76086) · 5417a5fe
      Benjamin Maxwell authored
      This adds very basic (and inelegant) support for something like spilling
      and reloading tiles, if you use more SME tiles than physically exist.
      
      This is purely implemented to prevent the compiler from aborting if a
      function uses too many tiles (i.e. due to bad unrolling), but is
      expected to perform very poorly.
      
      Currently, this works in two stages:
      
      During tile allocation, if we run out of tiles instead of giving up, we
      switch to allocating 'in-memory' tile IDs. These are tile IDs that start
      at 16 (which is higher than any real tile ID). A warning will also be
      emitted for each (root) tile op assigned an in-memory tile ID:
      
      ```
      warning: failed to allocate SME virtual tile to operation, all tile operations will go through memory, expect degraded performance
      ```
      
      Everything after this works like normal until `-convert-arm-sme-to-llvm`
      
      Here the in-memory tile op:
      
      ```mlir
      arm_sme.tile_op { tile_id = <IN MEMORY TILE> }
      ```
      
      Is lowered to:
      
      ```mlir
      // At function entry:
      %alloca = memref.alloca ... : memref<?x?xty>
      
      // Around the op:
      // Swap the contents of %alloca and tile 0.
      scf.for %slice_idx {
        %current_slice = "arm_sme.intr.read.horiz" ... <{tile_id = 0 : i32}>
        "arm_sme.intr.ld1h.horiz"(%alloca, %slice_idx)  <{tile_id = 0 : i32}>
        vector.store %current_slice, %alloca[%slice_idx, %c0]
      }
      // Execute op using tile 0.
      arm_sme.tile_op { tile_id = 0 }
      // Swap the contents of %alloca and tile 0.
      // This restores tile 0 to its original state.
      scf.for %slice_idx {
        %current_slice = "arm_sme.intr.read.horiz" ... <{tile_id = 0 : i32}>
        "arm_sme.intr.ld1h.horiz"(%alloca, %slice_idx)  <{tile_id = 0 : i32}>
        vector.store %current_slice, %alloca[%slice_idx, %c0]
      }
      ```
      
      This is inserted during the lowering to LLVM as spilling/reloading
      registers is a very low-level concept, that can't really be modeled
      correctly at a high level in MLIR.
      
      Note: This is always doing the worst case full-tile swap. This could be
      optimized to only spill/load data the tile op will use, which could be
      just a slice. It's also not making any use of liveness, which could
      allow reusing tiles. But these is not seen as important as correct code
      should only use the available number of tiles.
      5417a5fe
    • Louis Dionne's avatar
      [libc++] Deprecate the _LIBCPP_ENABLE_CXX20_REMOVED_ALLOCATOR_MEMBERS macro (#77692) · 8751bbe7
      Louis Dionne authored
      As described in #69994, using the escape hatch makes us non-conforming
      in C++20 due to incorrect constexpr-ness. It also leads to bad
      diagnostics as reported by #63900. We discussed the issue in the libc++
      monthly meeting and we agreed that we should deprecate the macro in LLVM
      18, and then remove it in LLVM 19 since it causes too many problems.
      
      This patch does the first part of this -- it deprecates the macro.
      
      Fixes #69994
      Fixes #63900
      Partially addresses #75975
      8751bbe7
    • Matthias Springer's avatar
      [mlir][Transforms] `GreedyPatternRewriteDriver`: log successful folding (#77796) · dec908a2
      Matthias Springer authored
      Similar to successful pattern applications, dump the rewritten IR after
      each successful folding when running with `-debug`.
      dec908a2
    • erichkeane's avatar
      [OpenACC] Implement the 'rest' of the simple 'var-list' clauses · 7700ea10
      erichkeane authored
      A large number of clauses are simple, required parens with a var-list.
      This patch adds them all, as adding them is quite trivial.
      7700ea10
    • Matthias Springer's avatar
      [mlir][vector] Fix dominance error in warp vector distribution (#77771) · ad100b36
      Matthias Springer authored
      This commit fixes a test in `vector-warp-distribute.mlir` when
      `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS` is enabled.
      
      ```
      within split at /usr/local/google/home/springerm/mlir_public/llvm-project/mlir/test/Dialect/Vector/vector-warp-distribute.mlir:1 offset :18:10: error: operand #0 does not dominate this use
          %1 = vector.extract %0[9] : f32 from vector<64xf32>
               ^
      within split at /usr/local/google/home/springerm/mlir_public/llvm-project/mlir/test/Dialect/Vector/vector-warp-distribute.mlir:1 offset :18:10: note: see current operation: %1 = "affine.apply"(%8) <{map = affine_map<()[s0] -> (s0 ceildiv 2)>}> : (index) -> index
      within split at /usr/local/google/home/springerm/mlir_public/llvm-project/mlir/test/Dialect/Vector/vector-warp-distribute.mlir:1 offset :18:10: note: operand defined here (op in a child region)
      "func.func"() <{function_type = (index) -> f32, sym_name = "vector_extract_1d"}> ({
      ^bb0(%arg0: index):
        %0:2 = "vector.warp_execute_on_lane_0"(%arg0) <{warp_size = 32 : i64}> ({
          %7 = "some_def"() : () -> vector<64xf32>
          %8 = "arith.constant"() <{value = 9 : index}> : () -> index
          %9 = "vector.extractelement"(%7, %8) : (vector<64xf32>, index) -> f32
          "vector.yield"(%9, %7) : (f32, vector<64xf32>) -> ()
        }) : (index) -> (f32, vector<2xf32>)
        %1 = "affine.apply"(%8) <{map = affine_map<()[s0] -> (s0 ceildiv 2)>}> : (index) -> index
        %2 = "affine.apply"(%8) <{map = affine_map<()[s0] -> (s0 mod 2)>}> : (index) -> index
        %3 = "vector.extractelement"(%0#1, %2) : (vector<2xf32>, index) -> f32
        %4 = "arith.index_cast"(%1) : (index) -> i32
        %5 = "arith.constant"() <{value = 32 : i32}> : () -> i32
        %6:2 = "gpu.shuffle"(%3, %4, %5) <{mode = #gpu<shuffle_mode idx>}> : (f32, i32, i32) -> (f32, i1)
        "func.return"(%6#0) : (f32) -> ()
      }) : () -> ()
      LLVM ERROR: IR failed to verify after pattern application
      ```
      
      The position at which `vector.extractelement` extracts must also be
      distributed. The fix in `WarpOpExtractElement` is similar to
      `WarpOpInsertElement`.
      ad100b36
    • Utkarsh Saxena's avatar
      [clang] Reapply Handle templated operators with reversed arguments (#72213) · 460ff58f
      Utkarsh Saxena authored
      Re-applies https://github.com/llvm/llvm-project/pull/69595 with extra
      [diff](https://github.com/llvm/llvm-project/pull/72213/commits/79181efd0d7aef1b8396d44cdf40c0dfa4054984)
      ### New changes
      
      Further relax ambiguities with a warning for member operators of a
      template class (primary templates of such ops do not match). Eg:
      ```cpp
      template <class T>
      struct S {
          template <typename OtherT>
          bool operator==(const OtherT &rhs); 
      };
      struct A : S<int> {};
      struct B : S<bool> {};
      bool x = A{} == B{}; // accepted with a warning.
      ```
      
      This is important for making llvm build using previous clang versions in
      C++20 mode (eg: this makes the commit
      e558be51 keep working with a warning
      instead of an error).
      
      ### Description from https://github.com/llvm/llvm-project/pull/69595
      
      https://github.com/llvm/llvm-project/pull/68999 correctly computed
      conversion sequence for reversed args to a template operator. This was a
      breaking change as code, previously accepted in C++17, starts to break
      in C++20.
      
      Example:
      ```cpp
      struct P {};
      template<class S> bool operator==(const P&, const S &);
      
      struct A : public P {};
      struct B : public P {};
      bool check(A a, B b) { return a == b; }  // This is now ambiguous in C++20.
      ```
      
      In order to minimise widespread breakages, as a clang extension, we had
      previously accepted such ambiguities with a warning
      (`-Wambiguous-reversed-operator`) for non-template operators. Due to the
      same reasons, we extend this relaxation for template operators.
      
      Fixes https://github.com/llvm/llvm-project/issues/53954
      460ff58f
    • antangelo's avatar
      [Sema] Use lexical DC for friend functions when getting constraint instantiation args (#77552) · 45568135
      antangelo authored
      Fixes a crash where the template argument depth computed in the semantic
      context for a friend FunctionDecl with a constrained parameter is
      compared against arguments in the lexical context for the purpose of
      checking if the constraint depends on enclosing template parameters.
      
      Since getTemplateInstantiationArgs in this case follows the semantic DC
      for friend FunctionDecls, the resulting depth is incorrect and trips an
      assertion.
      
      Fixes #75426
      45568135
    • Qiongsi Wu's avatar
      [SimplifyCFG] `switch`: Do Not Transform the Default Case if the Condition is Too Wide (#77831) · 39bb790b
      Qiongsi Wu authored
      https://github.com/llvm/llvm-project/pull/76669 taught SimplifyCFG to
      handle switches when `default` has only one case. When the `switch`'s
      condition is wider than 64 bit, the current implementation can calculate
      the wrong default value. This PR skips cases where the condition is too
      wide.
      39bb790b
    • Alex Zinenko's avatar
      b32001a2
    • Nikita Popov's avatar
      [IRBuilder] Add CreatePtrAdd() method (NFC) (#77582) · 6c2fbc3a
      Nikita Popov authored
      This abstracts over the common pattern of creating a gep with i8 element
      type.
      6c2fbc3a
    • Guray Ozen's avatar
      ae5d6392
    • Florian Hahn's avatar
      [VPlan] Support narrowing widened loads in truncateToMinimimalBitwidths. · 59d6f033
      Florian Hahn authored
      MinBWs may also contain widened load instructions, handle them by only
      narrowing their result.
      
      Fixes https://github.com/llvm/llvm-project/issues/77468
      59d6f033
    • Alexey Bataev's avatar
      [SLP]Fix a crash for reduced values with minbitwidth, which are reused. · 39b2104b
      Alexey Bataev authored
      If the reduced values are additionally affected by minbitwidth analysis,
      need to cast them to a proper type before doing any math, if they are
      reused.
      39b2104b
    • Matthias Springer's avatar
      [mlir][vector] Fix rewrite pattern API violation in `VectorToSCF` (#77909) · aa2dc792
      Matthias Springer authored
      A rewrite pattern is not allowed to change the IR if it returns
      "failure". This commit fixes
      `test/Conversion/VectorToSCF/vector-to-scf.mlir` when running with
      `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS`.
      
      ```
      Processing operation : 'vector.transfer_read'(0x55823a409a60) {
        %5 = "vector.transfer_read"(%arg0, %0, %0, %2, %4) <{in_bounds = [true, true], operandSegmentSizes = array<i32: 1, 2, 1, 1>, permutation_map = affine_map<(d0, d1) -> (d0, d1)>}> : (memref<?x4xf32>, index, index, f32, vector<[4]x4xi1>) -> vector<[4]x4xf32>
      
        * Pattern (anonymous namespace)::lowering_n_d_unrolled::UnrollTransferReadConversion : 'vector.transfer_read -> ()' {
      Trying to match "(anonymous namespace)::lowering_n_d_unrolled::UnrollTransferReadConversion"
          ** Insert  : 'vector.splat'(0x55823a445640)
      "(anonymous namespace)::lowering_n_d_unrolled::UnrollTransferReadConversion" result 0
        } -> failure : pattern failed to match
      
      LLVM ERROR: pattern returned failure but IR did change
      ```
      aa2dc792
    • Alexey Lapshin's avatar
      [DWARFLinker][NFC] Rename libraries to match with directories name. (#77592) · 35708b07
      Alexey Lapshin authored
      It was noted that new DWARFLinker libraries do not follow naming
      agreement -
      https://github.com/llvm/llvm-project/pull/75925#issuecomment-1883301659
      This patch rename libraries to match with the agreement.
      
      Rename LLVMDWARFLinkerBase library into the LLVMDWARFLinker. Rename
      LLVMDWARFLinker library into the LLVMDWARFLinkerClassic. Correct include
      path according to the new directory structure.
      35708b07
    • Oleksandr "Alex" Zinenko's avatar
      [mlir] introduce debug transform dialect extension (#77595) · 2798b72a
      Oleksandr "Alex" Zinenko authored
      Introduce a new extension for simple print-debugging of the transform
      dialect scripts. The initial version of this extension consists of two
      ops that are printing the payload objects associated with transform
      dialect values. Similar ops were already available in the test extenion
      and several downstream projects, and were extensively used for testing.
      2798b72a
    • jeanPerier's avatar
      [flang] Get ProvenanceRange from CharBlock starting with expanded macro (#77791) · c87e94b0
      jeanPerier authored
      When a CharBlock starts with an expanded macro but does not end in this
      macro expansion, GetProvenanceRange fails to return a ProvenanceRange
      which may cause error message to be emitted without location or lowering
      to emit code without source location (which is problematic if this code
      contains calls to procedures defined in the same file since LLVM will
      later crash with the error:
      "inlinable function call in a function with a DISubprogram location must
      have a debug location"
      
      Fix this situation by returning the ProvenanceRange starting at the
      replaced macro reference.
      c87e94b0
    • Dan McGregor's avatar
      [flang] include sys/wait.h for EXECUTE_COMMAND_LINE (#77675) · a762cc21
      Dan McGregor authored
      Linux defines WEXITSTATUS in stdlib.h, but at least FreeBSD and NetBSD
      only define it in sys/wait.h. Include this header unconditionally, since
      it is required on the BSDs and should be harmless on other platforms.
      
      Fixes FreeBSD build after #74077.
      a762cc21
    • Matthias Springer's avatar
      [mlir][vector] Support warp distribution of `transfer_read` with dependencies (#77779) · 35c19fdd
      Matthias Springer authored
      Support distribution of `vector.transfer_read` ops when operands are
      defined inside of the region of `warp_execute_on_lane_0` (except for the
      buffer from which the op is reading).
      
      Such IR was previously not supported. This commit changes the
      implementation such that indices and the padding value are also
      distributed.
      
      This commit simplifies the implementation considerably: the original
      implementation created a new `transfer_read` op and then checked if this
      new op is valid. If not, the rewrite pattern failed. This was a bit
      hacky. It was also a violation of the rewrite pattern API (detected by
      `MLIR_ENABLE_EXPENSIVE_PATTERN_API_CHECKS`) because the IR was modified,
      but the pattern returned "failure".
      35c19fdd
    • Corentin Jabot's avatar
    • Mirko Brkušanin's avatar
      [AMDGPU][NFC] Rename DotIUVOP3PMods to VOP3PModsNeg (#77785) · 2adbf254
      Mirko Brkušanin authored
      
      
      This is used to select the source modifier (neg) from the immediate
      operand. After a follow up commit this will no longer be DOTIU specific.
      
      Co-authored-by: default avatarChangpeng Fang <changpeng.fang@amd.com>
      2adbf254
    • Matthew Devereau's avatar
      [AArch64][SME2] Fix SME2 mla/mls tests (#76711) · 42fe3bc1
      Matthew Devereau authored
      The ACLE defines these builtins as svmla[_single]_za32[_f32]_vg1x2,
      which means the SVE_ACLE_FUNC macro should test the overloaded forms as
      
      SVE_ACLE_FUNC(svmla,_single,_za32,_f32,_vg1x2)
      
      
      https://github.com/ARM-software/acle/blob/b88cbf7e9c104100bb5016c848763171494dee44/main/acle.md?plain=1#L10170-L10205
      42fe3bc1
    • Matthew Devereau's avatar
      [AArch64][SME] Fix multi vector cvt builtins (#77656) · a8f83cc1
      Matthew Devereau authored
      This fixes cvt multi vector builtins that erroneously had inverted
      return vectors and vector parameters. This caused the incorrect
      instructions to be emitted.
      a8f83cc1
    • Jie Fu's avatar
      fbac3b0d
    • ostannard's avatar
      [AArch64] Disable FP loads/stores when fp-armv8 not enabled (#77817) · 9c9bffe2
      ostannard authored
      Most of the floating-point instructions are already gated on the
      fp-armv8 subtarget feature (or some other feature), but most of the load
      and store instructions, and one move instruction, were not.
      
      I found this list of instructions with a script which consumes the
      output of llvm-tblgen --dump-json, looking for instructions which have
      an FPR operand but no predicate. That script now finds zero
      instructions.
      
      This only affects assembly, not codegen, because the floating-point
      types and registers are already not marked as legal when the FPU is
      disabled, so it is impossible for any of these to be selected.
      9c9bffe2
    • Frederik Carlier's avatar
      [ObjC]: Make type encoding safe in symbol names (#77797) · 3168192d
      Frederik Carlier authored
      Type encodings are part of symbol names in the Objective C ABI. Replace
      characters which are reseved in symbol names:
      
      - ELF: avoid including '@' characters in type encodings
      - Windows: avoid including '=' characters in type encodings
      3168192d
    • Matthias Springer's avatar
      [mlir][Interfaces] `DestinationStyleOpInterface`: Rename `hasTensor/BufferSemantics` (#77574) · 0a8e3dd4
      Matthias Springer authored
      Rename interface functions as follows:
      * `hasTensorSemantics` -> `hasPureTensorSemantics`
      * `hasBufferSemantics` -> `hasPureBufferSemantics`
      
      These two functions return "true" if the op has tensor/buffer operands
      but not buffer/tensor operands.
      
      Also drop the "ranked" part from the interface, i.e., do not distinguish
      between ranked/unranked types.
      
      The new function names describe the functions more accurately. They also
      align their semantics with the notion of "tensor semantics" with the
      bufferization framework. (An op is supposed to be bufferized if it has
      tensor operands, and we don't care if it also has memref operands.)
      
      This change is in preparation of #75273, which adds
      `BufferizableOpInterface::hasTensorSemantics`. By renaming the functions
      in the `DestinationStyleOpInterface`, we can avoid name clashes between
      the two interfaces.
      0a8e3dd4
    • martinboehme's avatar
      1aacdfe4
    • jeanPerier's avatar
      [flang] finish BIND(C) VALUE derived type passing ABI on X86-64 (#77742) · 011ba725
      jeanPerier authored
      Derived type passed with VALUE in BIND(C) context must be passed like C
      struct and LLVM is not implementing the ABI for this (it is up to the
      frontends like clang).
      
      Previous patch #75802 implemented the simple cases where the derived
      type have one field, this patch implements the general case. Note that
      the generated LLVM IR is compliant from a X86-64 C ABI point of view and
      compatible with clang generated assembly, but that it is not guaranteed
      to match the LLVM IR signatures generated by clang for the C equivalent
      functions because several LLVM IR signatures may lead to the same X86-64
      signature.
      011ba725
    • MyDeveloperDay's avatar
      [clang-format] SpacesInSquareBrackets not working for Java (#77833) · c65b939f
      MyDeveloperDay authored
      
      
      spaces in [] needs to be handled the same in Java the same as C#.
      
      Co-authored-by: default avatarpaul_hoad <paul_hoad@amat.com>
      c65b939f
    • Jie Fu's avatar
      [clang][dataflow] Remove unused private field 'StmtToEnv' (NFC) · ef156f91
      Jie Fu authored
      llvm-project/clang/lib/Analysis/FlowSensitive/TypeErasedDataflowAnalysis.cpp:148:23:
       error: private field 'StmtToEnv' is not used [-Werror,-Wunused-private-field]
        148 |   const StmtToEnvMap &StmtToEnv;
            |                       ^
      1 error generated.
      ef156f91
    • martinboehme's avatar
      [clang][dataflow] Process terminator condition within `transferCFGBlock()`. (#77750) · 537bbb46
      martinboehme authored
      In particular, it's important that we create the "fallback" atomic at
      this point
      (which we produce if the transfer function didn't produce a value for
      the
      expression) so that it is placed in the correct environment.
      
      Previously, we processed the terminator condition in the
      `TerminatorVisitor`,
      which put the fallback atomic in a copy of the environment that is
      produced as
      input for the _successor_ block, rather than the environment for the
      block
      containing the expression for which we produce the fallback atomic.
      
      As a result, we produce different fallback atomics every time we process
      the
      successor block, and hence we don't have a consistent representation of
      the
      terminator condition in the flow condition.
      
      This patch includes a test (authored by ymand@) that fails without the
      fix.
      537bbb46
    • Jie Fu's avatar
      [clangd] Use starts_with instead of startswith in CompileCommands.cpp (NFC) · dabc9018
      Jie Fu authored
      llvm-project/clang-tools-extra/clangd/CompileCommands.cpp:324:52:
       error: 'startswith' is deprecated: Use starts_with instead [-Werror,-Wdeprecated-declarations]
        324 |         Cmd, [&](llvm::StringRef Arg) { return Arg.startswith(Flag); });
            |                                                    ^~~~~~~~~~
            |                                                    starts_with
      dabc9018
    • Kon's avatar
      [clangd] Fix sysroot flag handling in CommandMangler to prevent duplicates (#75694) · f489fb3d
      Kon authored
      CommandMangler should guess the sysroot path of the host system and add
      that through `-isysroot` flag only when there is no `--sysroot` or
      `-isysroot` flag in the original compile command to avoid duplicate
      sysroot.
      
      Previously, CommandMangler appropriately avoided adding a guessed
      sysroot flag if the original command had an argument in the form of
      `--sysroot=<sysroot>`, `--sysroot <sysroot>`, or `-isysroot <sysroot>`.
      However, when presented as `-isysroot<sysroot>` (without spaces after
      `-isysroot`), CommandMangler mistakenly appended the guessed sysroot
      flag, resulting in duplicated sysroot in the final command.
      
      This commit fixes it, ensuring the final command has no duplicate
      sysroot flags. Also adds unit tests for this fix.
      f489fb3d
    • Jie Fu's avatar
    • Guray Ozen's avatar
      [mlir][nvgpu] Improve verifier of `ldmatrix` (#77807) · 24918670
      Guray Ozen authored
      PR improves the verifier of `nvgpu.ldmatrix` Op, so `nvgpu-to-nvvm`
      lowering does not crash.
      24918670
    • Craig Topper's avatar
      [RISCV] Simplify the description for ssaia and smaia. (#77870) · 2e78c220
      Craig Topper authored
      It feels more important to expand out Advanced Interrupt Architecture
      for users than to have a description that explains how one extension is
      different from the other.
      2e78c220
    • Fangrui Song's avatar
      [test] Improve x86 inline asm tests · 7e604485
      Fangrui Song authored
      Reorganize *asm-modifier* and make other cleanups.
      7e604485