1. Jun 30, 2023
  2. Jun 29, 2023
    • sstwcw's avatar
      [clang-format] Indent Verilog struct literal on new line · 6bf66d83
      sstwcw authored
      Before:
      ```
      c = //
      '{default: 0};
      ```
      
      After:
      ```
      c = //
          '{default: 0};
      ```
      
      If the line has to be broken, the continuation part should be
      indented.  Before this fix, it was not the case if the continuation
      part was a struct literal.  The rule that caused the problem was added
      in 783bac6b.  It was intended for aligning the field labels in
      ProtoBuf.  The type `TT_DictLiteral` was only for colons back then, so
      the program didn't have to check whether the token was a colon when it
      was already type `TT_DictLiteral`.  Now the type applies to more
      things including the braces enclosing a dictionary literal.  In
      Verilog, struct literals start with a quote.  The quote is regarded as
      an identifier by the program.  So the rule for aligning the fields in
      ProtoBuf applied to this situation by mistake.
      
      Reviewed By: HazardyKnusperkeks
      
      Differential Revision: https://reviews.llvm.org/D152623
      6bf66d83
    • Arnold Schwaighofer's avatar
      Add a type_checked_load_relative to support relative function pointer tables · 98eb8abf
      Arnold Schwaighofer authored
      This adds a type_checked_load_relative intrinsic whose semantics it is to
      load a relative function pointer.
      
      A relative function pointer is a pointer to a 32bit value that when
      added to its address yields the address of the function.
      
      Differential Revision: https://reviews.llvm.org/D143204
      98eb8abf
    • Peter Klausler's avatar
      [flang] Honor #line and related preprocessing directives · e12ffe6a
      Peter Klausler authored
      Extend the SourceFile class to take account of #line directives
      when computing source file positions for error messages.
      Adjust the output of #line directives to -E output so that they
      reflect any #line directives that were in the input.
      
      Differential Revision: https://reviews.llvm.org/D153910
      e12ffe6a
    • Jeffrey Byrnes's avatar
      [HIP]: Add -fhip-emit-relocatable to override link job creation for -fno-gpu-rdc · be8a65b5
      Jeffrey Byrnes authored
      Differential Revision: https://reviews.llvm.org/D153667
      
      Change-Id: Idcc5c7c25dc350b8dc9a1865fd67982904d06ecd
      be8a65b5
    • Haojian Wu's avatar
      [clangd] Don't show header for namespace decl in Hover · 21b6da35
      Haojian Wu authored
      The header for namespace symbol is barely useful.
      
      Differential Revision: https://reviews.llvm.org/D154068
      21b6da35
    • Arnold Schwaighofer's avatar
      WholeProgramDevirt: Fix call target propagation for ptrauth architectures · 200a1cce
      Arnold Schwaighofer authored
      We can't have a call with a constant target with a ptrauth bundle. Remove the
      ptrauth bundle operand in such a case
      
      rdar://105696396
      
      Differential Revision: https://reviews.llvm.org/D144581
      200a1cce
    • Zahira Ammarguellat's avatar
      Fix test regression on 32-bit x86. · 9fd3321d
      Zahira Ammarguellat authored
      Differential Revision: https://reviews.llvm.org/D153770
      9fd3321d
    • Matthias Springer's avatar
      [mlir][Transforms][NFC] CSE: Add non-pass entry point · 189033e6
      Matthias Springer authored
      Add an additional entry point so that CSE can be used without a pass. This allows CSE to be used from the Transform dialect without invalidating all handles.
      
      * All IR modifications are done with a rewriter.
      * The C++ entry point takes a `RewriterBase &`, which may have a listener attached to it. This allows users to track all IR modifications.
      
      Differential Revision: https://reviews.llvm.org/D145226
      189033e6
    • Philip Reames's avatar
      [RISCV] Remove legacy TA/TU pseudo distinction for unary instructions · 92b5a340
      Philip Reames authored
      This change continues with the line of work discussed in https://discourse.llvm.org/t/riscv-transition-in-vector-pseudo-structure-policy-variants/71295. In D153155, we started removing the legacy distinction between unsuffixed (TA) and _TU pseudos. This patch continues that effort for the unary instruction families.
      
      The change consists of a few interacting pieces:
      * Adding a vector policy operand to VPseudoUnaryNoMaskTU.
      * Then using VPseudoUnaryNoMaskTU for all cases where VPseudoUnaryNoMask was previously used and deleting the unsuffixed form.
      * Then renaming VPseudoUnaryNoMaskTU to VPseudoUnaryNoMask, and adjusting the RISCVMaskedPseudo table to use the combined pseudo.
      * Fixing up two places in C++ code which manually construct VMV_V_* instructions.
      
      Normally, I'd try to factor this into a couple of changes, but in this case, the table structure is tied to naming and thus we can't really separate the otherwise NFC bits.
      
      As before, we see codegen changes (some improvements and some regressions) due to scheduling differences caused by the extra implicit_def instructions.
      
      Differential Revision: https://reviews.llvm.org/D153899
      92b5a340
    • Simon Pilgrim's avatar
      [X86] LowerTRUNCATE - attempt to use PACKSS/PACKUS on AVX512 targets if the... · 34961c60
      Simon Pilgrim authored
      [X86] LowerTRUNCATE - attempt to use PACKSS/PACKUS on AVX512 targets if the truncation source is concatenating from smaller subvectors
      
      Don't just use AVX512 truncation ops if PACKSS/PACKUS can do this more cheaply
      34961c60
    • Simon Pilgrim's avatar
      [X86] Add isFreeToSplitVector helper to detect nodes that we can freely... · 0501c160
      Simon Pilgrim authored
      [X86] Add isFreeToSplitVector helper to detect nodes that we can freely split/extract subvectors from.
      
      Helper wrapper around the existing collectConcatOps method.
      0501c160
    • LLVM GN Syncbot's avatar
      [gn build] Port cfa096d9 · d662865f
      LLVM GN Syncbot authored
      d662865f
    • Christian Trott's avatar
      [libc++][mdspan] Implement layout_right · cfa096d9
      Christian Trott authored
      This commit implements layout_right in support of C++23 mdspan
      (https://wg21.link/p0009
      
      ). layout_right is a layout mapping policy
      whose index mapping corresponds to the memory layout of multidimensional
      C-arrays, and is thus also referred to as the C-layout.
      
      Co-authored-by: default avatarDamien L-G <dalg24@gmail.com>
      
      Differential Revision: https://reviews.llvm.org/D151267
      cfa096d9
    • Takuya Shimizu's avatar
      [clang][Sema] Remove dead diagnostic for loss of __unaligned qualifier · 8038086a
      Takuya Shimizu authored
      D120936 has made the loss of `__unaligned` qualifier NOT a bad-conversion.
      Because of this, the bad-conversion note about the loss of this qualifier does not take effect.
      e.g.
      ```
      void foo(int *ptr);
      
      void func(const __unaligned int *var) { foo(var); }
      ```
      BEFORE this patch:
      ```
      source.cpp:3:41: error: no matching function for call to 'foo'
          3 | void func(const __unaligned int *var) { foo(var); }
            |                                         ^~~
      source.cpp:1:6: note: candidate function not viable: 1st argument ('const __unaligned int *') would lose __unaligned qualifier
          1 | void foo(int *ptr);
            |      ^
          2 |
          3 | void func(const __unaligned int *var) { foo(var); }
            |                                             ~~~
      ```
      AFTER this patch:
      ```
      source.cpp:3:41: error: no matching function for call to 'foo'
          3 | void func(const __unaligned int *var) { foo(var); }
            |                                         ^~~
      source.cpp:1:6: note: candidate function not viable: 1st argument ('const __unaligned int *') would lose const qualifier
          1 | void foo(int *ptr);
            |      ^
          2 |
          3 | void func(const __unaligned int *var) { foo(var); }
            |                                             ~~~
      ```
      Please note the different mentions of `__unaligned` and `const` in notes.
      
      Reviewed By: cjdb, rnk
      Differential Revision: https://reviews.llvm.org/D153690
      8038086a
    • Mike Crowe's avatar
      [clang-tidy] Fix modernize-use-std-print check when return value used · 09ed2102
      Mike Crowe authored
      The initial implementation of the modernize-use-std-print check was
      capable of converting calls to printf (etc.) which used the return value
      to calls to std::print which has no return value, thus breaking the
      code.
      
      Use code inspired by the implementation of bugprone-unused-return-value
      check to ignore cases where the return value is used. Add appropriate
      lit test cases and documentation.
      
      Reviewed By: PiotrZSL
      
      Differential Revision: https://reviews.llvm.org/D153860
      09ed2102
    • Alexey Lapshin's avatar
      [DWARFv5][DWARFLinker] Remove dsymutil-classic compatibility feature as it leads to an error. · 4546015f
      Alexey Lapshin authored
      DWARFLinker has a compatibility feature with dsymutil-classic.
      It may keep location expression attribute even if does not
      reference live address. Current llvm-dwarfdump --verify
      reports a error if variable references an address but is not
      added into the .debug_names table.
      
      error: Name Index @ 0x0: Entry for DIE @ 0xf35 (DW_TAG_variable) with name seed missing.
      
      DW_TAG_variable
        DW_AT_name      ("seed")
        DW_AT_type      (0x00000000000047b7 "uint64_t")
        DW_AT_location  (DW_OP_addr 0x9ff8)  <<<< dead address
      
      DWARFLinker does not add the variable into .debug_names table
      because it references dead address. To have a valid variable and
      consistent accelerator table it is necessary to remove location expression
      referencing dead address. This patch removes dsymutil-classic
      compatibilty feature.
      
      Differential Revision: https://reviews.llvm.org/D153988
      4546015f
    • Nikita Popov's avatar
      [X86] Add tests for PR63475 (NFC) · d95c2c27
      Nikita Popov authored
      d95c2c27
    • pvanhout's avatar
      [MCP] Optimize copies from undef · c59f9ead
      pvanhout authored
      Revert D152502 and instead optimize away copy from undefs, but clear the undef flag on the original copy.
      Apparently, not optimizing the COPY can cause performance issues in some cases.
      
      Fixes SWDEV-405813, SWDEV-405899
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D153838
      c59f9ead
    • Sean Perry's avatar
      [SystemZ][z/OS] Add support for z/OS link step (executable and shared libs) · 5e87ec1e
      Sean Perry authored
      Add support for performing a link step on z/OS.  This will support C & C++ building executables and shared libs.
      
      Reviewed By: zibi, abhina.sreeskantharajan
      
      Differential Revision: https://reviews.llvm.org/D153580
      5e87ec1e
    • Luke Lau's avatar
      [RISCV] Add tests for cost modelling constants in phis · b87a0930
      Luke Lau authored
      Reviewed By: ABataev
      
      Differential Revision: https://reviews.llvm.org/D149168
      b87a0930
    • pvanhout's avatar
      [AMDGPU] Handle Additional Cases in tryFoldPhiAGPR · 026fc9e9
      pvanhout authored
       Sometimes PHI have different incoming values, such as:
       ```
      %1:vgpr_256 = COPY %0:agpr_256
      %2:vgpr_32 = COPY %1:vgpr_256.sub0
      ```
      
      Those weren't handled, which could lead to massive performance issues if break-large-PHIs kicked in + AGPRs were used (MFMA)
      
      Fixes SWDEV-407986
      
      Reviewed By: #amdgpu, arsenm
      
      Differential Revision: https://reviews.llvm.org/D153879
      026fc9e9
    • Luke Lau's avatar
      [TTI] Use users of GEP to guess access type in getGEPCost · a68dcd09
      Luke Lau authored
      Currently getGEPCost uses the target type of the GEP as a heuristic for
      the type that will be accessed, to pass onto isLegalAddressingMode.
      Targets use this to work out if a GEP can then be folded into the
      load/store instruction that uses the GEP.
      For example, on RISC-V loads and stores can have an offset added to a
      base register folded into a single instruction, so the following GEP is
      free:
      
      %p = getelementptr i32, ptr %base, i32 42       ; getInstructionCost = 0
      %x = load i32, ptr %p                           ; getInstructionCost = 1
      ------------------------------------------------------------------------
      lw t0, a0(42)
      
      However vector loads and stores cannot have an offset folded into them,
      so the following GEP is costed:
      
      %p = getelementptr <2 x i32>, ptr %base, i32 42 ; getInstructionCost = 1
      %x = load <2 x i32>, ptr %p                     ; getInstructionCost = 1
      ------------------------------------------------------------------------
      addi  a0, 42
      vle32 v8, (a0)
      
      The issue arises whenever there is a mismatch between the target type of
      the GEP and the type that is actually accessed:
      
      %p = getelementptr i32, ptr %base, i32 42       ; getInstructionCost = 0
      %x = load <2 x i32>, ptr %p                     ; getInstructionCost = 1
      ------------------------------------------------------------------------
      addi  a0, 42
      vle32 v8, (a0)
      
      Even though this GEP will result in an add instruction, because TTI
      thinks it's loading an i32, it will think it can be folded and not
      charge for it.
      
      The target type can become mismatched with the memory access during
      transformations, noticeably during SLP where a scalar base pointer will
      be reused to perform a vector load or store.
      
      This patch adds an optional AccessType argument to getGEPCost which
      allows the type of memory accessed by users to be passed in as a hint,
      so that we can more accurately determine if the GEP can be folded into
      its users.
      
      If AccessType is not provided, getGEPCost falls back to the old
      behaviour of using the PointeeType to guess the memory access type. This
      can be revisited in a later patch.
      
      Also for now, only GEPs with exactly one user use the access type hint.
      Whilst we could look through all users and use all access types to
      determine if we can fold the GEP, this patch avoids doing so to prevent
      O(N) behaviour.
      
      Differential Revision: https://reviews.llvm.org/D149889
      a68dcd09
    • Luke Lau's avatar
      [RISCV][SLP] Add tests for GEP costs · cb941f92
      Luke Lau authored
      This patch updates the tests in gep.ll to have explicitly memory
      accesses using them, to illustrate the new behaviour in D149889.
      New tests have also been added for mismatched pointer types and memory
      access types, and gep-zero-indices.ll has also been added to make sure
      that we always cost GEPs with all zero indices as free.
      cb941f92
    • Ivan Butygin's avatar
      [mlir][memref] Add some missing interfaces to memref ops. · 1f91fe32
      Ivan Butygin authored
      Add `ViewLikeOpInterface` to `ExtractStridedMetadataOp` as it returns its buffer as one of the results.
      Add mem Read/Write attributes to atomic ops.
      
      Differential Revision: https://reviews.llvm.org/D153647
      1f91fe32