1. Jun 01, 2023
    • David Green's avatar
      [ARM][AArch64] Add tests for shuffles load patterns. NFC · 8d82f12a
      David Green authored
      See D151029
      8d82f12a
    • Krzysztof Drewniak's avatar
    • Teresa Johnson's avatar
      Revert "[ThinLTO] Disable partial sample profile scaling by default" · 96fb18a3
      Teresa Johnson authored
      This reverts commit aae8524b, which was
      found to cause a few unexpected benchmark performance differences that
      need investigation.
      96fb18a3
    • Peter Klausler's avatar
      [flang] CUDA Fortran - part 2/5: symbols & scopes · 27f71807
      Peter Klausler authored
      Add representations of CUDA Fortran data and subprogram attributes
      to the symbol table and scopes of semantics.  Set them in name
      resolution, and emit them to module files.
      
      Depends on https://reviews.llvm.org/D150159.
      
      Differential Revision: https://reviews.llvm.org/D150161
      27f71807
    • Luke Lau's avatar
      [RISCV][InsertVSETVLI] Relax tail policy more often for vmv.s.x · f3b39cea
      Luke Lau authored
      If a vm.s.x pseudo has an undef passthru operand, then we're free to use
      whatever tail policy we want for VL > 1. We previously relaxed the tail
      policy for this but only when we could also expand the SEW.
      This patch changes it to relax the tail policy even if the SEW can't be
      expanded and removes a few more toggles, as well as fully moving the
      vmv.s.x logic into getDemanded.
      f3b39cea
    • Luke Lau's avatar
      [RISCV][InsertVSETVLI] Avoid vmv.s.x SEW toggle if at start of block · badf11de
      Luke Lau authored
      vmv.s.x/vfmv.s.f instructions that only write to the first destination
      element can use any SEW greater than or equal to its original SEW,
      provided that it's writing to an implicit_def operand where we can
      clobber the other lanes.
      
      We were already handling this in needVSETVLI, which meant that when
      scanning the instructions from top to bottom we could detect this and
      avoid the toggle:
      
      	vsetivli	zero, 4, e64, mf2, ta, ma
      	li	a0, 11
      	vsetivli	zero, 1, e8, mf8, ta, ma
      	vmv.s.x	v0, a0
      
      ->
      	vsetivli	zero, 4, e64, mf2, ta, ma
      	li	a0, 11
      	vmv.s.x	v0, a0
      The issue that this patch aims to solve is arises when the vmv.s.x is
      the first vector instruction in the block and doesn't have any prior
      predecessor info:
      
      entry_bb:
      	li	a0, 11
      	; No previous state here: forced to set VL/VTYPE
      	vsetivli	zero, 1, e8, mf8, ta, ma
      	vmv.s.x	v0, a0
      	vsetivli	zero, 4, e16, mf2, ta, ma
      	vmerge.vvm	v8, v9, v8, v0
      doLocalPostpass can work backwards from bottom to top and work out if
      an earlier vsetvli can be mutated to avoid a toggle. It uses
      DemandedFields and getDemanded for this, which previously didn't take
      into account the possibility of going to a larger SEW.
      
      A previous patch consolidated the vmv.s.x logic from needVSETVLI logic
      into getDemanded, and this patch removes the gate around it so that
      doLocalPostpass can now delete vsetvlis like in the scenario below:
      
      entry_bb:
      	li	a0, 11
      	; Previous vsetivli mutated: second one deleted
      	vsetivli	zero, 4, e16, mf2, ta, ma
      	vmv.s.x	v0, a0
      	vmerge.vvm	v8, v9, v8, v0
      
      Differential Revision: https://reviews.llvm.org/D151561
      badf11de
    • Luke Lau's avatar
      [RISCV][InsertVSETVLI] Move vmv.s.x SEW check into getDemandedBits. NFC · 257cc049
      Luke Lau authored
      This patch restructures the logic that checks if vmv.s.x's SEW can be
      expanded into getDemandedBits, so that it can be shared by both the
      top-to-bottom and bottom-to-top passes.
      
      It adds a third option for SEW in DemandedFields, that's weaker than
      demanded but stronger than not demanded, that states that it the new SEW
      must be greater than or equal to the current SEW.
      
      Note that we now need to take care of the order of operands in
      areCompatibleVTYPEs as the relation is no longer commutative.
      
      A later patch will remove the gating on the bottom-to-top pass
      (dolocalPostpass) and another one will relax the demands on the tail
      policy further.
      257cc049
    • Manna, Soumi's avatar
      [NFC][CLANG] Fix nullptr dereference issue in SetValueDataBasedOnQualType() · 5bb06924
      Manna, Soumi authored
      This patch uses castAs instead of getAs which will assert if the type doesn't match in SetValueDataBasedOnQualType(clang::Value &, unsigned long long).
      
      Reviewed By: erichkeane
      
      Differential Revision: https://reviews.llvm.org/D151770
      5bb06924
    • Craig Topper's avatar
      Revert "[RISCV] Add Zvfhmin extension for clang." · a5e5eea2
      Craig Topper authored
      This reverts commit 35a00792.
      
      The backend support is not present yet. The intrinsics will crash
      the compiler if compiled to assembly or binary.
      a5e5eea2
    • Luke Lau's avatar
      319adf5d
    • Luke Lau's avatar
      [RISCV][InsertVSETVLI] Avoid vmv.s.x SEW toggle if at start of block · 0ba41dd3
      Luke Lau authored
      vmv.s.x and friends that only write to the first destination element can
      use any SEW greater than or equal to its original SEW, provided that
      it's writing to an implicit_def operand where we can clobber the other
      lanes.
      
      We were already handling this in needVSETVLI, which meant that when
      scanning the instructions from top to bottom we could detect this and
      avoid the toggle:
      
      ```
      	vsetivli	zero, 4, e64, mf2, ta, ma
      	li	a0, 11
      	vsetivli	zero, 1, e8, mf8, ta, ma
      	vmv.s.x	v0, a0
      
      ->
      	vsetivli	zero, 4, e64, mf2, ta, ma
      	li	a0, 11
      	vmv.s.x	v0, a0
      
      ```
      The issue that this patch aims to solve is whenever vmv.s.x arises when
      the first vector instruction in the block and doesn't have any prior
      predecessor info:
      
      ```
      entry_bb:
      	li	a0, 11
      	; No previous state here: forced to set VL/VTYPE
      	vsetivli	zero, 1, e8, mf8, ta, ma
      	vmv.s.x	v0, a0
      	vsetivli	zero, 4, e16, mf2, ta, ma
      	vmerge.vvm	v8, v9, v8, v0
      ```
      
      doLocalPostpass can work backwards from bottom to top and work out if
      an earlier vsetvli can be mutated to avoid a toggle. It uses
      DemandedFields and getDemanded for this, which previously didn't take
      into account the possibility of going to a larger SEW.
      
      This patch adds a third option for SEW in DemandedFields, that's weaker
      than demanded but stronger than not demanded, that states that it the
      new SEW must be greater than or equal to the current SEW.
      
      We can then use this option to move that vmv.s.x specific logic from
      needVSETVLI into getDemanded, making it available for both phase 2 and
      3, i.e. we can now mutate the earlier vsetivli going from bottom to top:
      
      ```
      entry_bb:
      	li	a0, 11
      	; Previous vsetivli mutated: second one deleted
      	vsetivli	zero, 4, e16, mf2, ta, ma
      	vmv.s.x	v0, a0
      	vmerge.vvm	v8, v9, v8, v0
      ```
      
      Reviewed By: reames
      
      Differential Revision: https://reviews.llvm.org/D151561
      0ba41dd3
    • Manna, Soumi's avatar
      [NFC][CLANG] Fix nullptr dereference issue in HandleRISCVRVVVectorBitsTypeAttr() · 071d4ab3
      Manna, Soumi authored
      This patch uses castAs instead of getAs which will assert if the type doesn't match in HandleRISCVRVVVectorBitsTypeAttr(clang::QualType &, clang::ParsedAttr &, clang::Sema &)
      
      Reviewed By: erichkeane
      
      Differential Revision: https://reviews.llvm.org/D151769
      071d4ab3
    • Nick Desaulniers's avatar
      [Demangle] fix deref of std::string_view::end() · 54a2994f
      Nick Desaulniers authored
      In D148546, I replaced much of the use of llvm::StringView w/
      std::string_view.  There's one important semantic difference between the
      two:
      
      In most STL containers, end() returns an iterator that refers to one
      past the end of the container. But llvm::StringView::end() refers to the
      last element.
      
      Expressions such as `&*my_std_string_view.end()` produce the failed
      assertion:
      
        include/c++/v1/__iterator/bounded_iter.h:93: assertion
        __in_bounds(__current_) failed: __bounded_iter::operator*: Attempt to
        dereference an out-of-range iterator
      
      This was caught when copying the recent downstream changes back upstream
      in D148566, and is reproducible via:
      
        $ libcxx/utils/ci/run-buildbot generic-debug-mode
      
      when compiled with clang and clang++. The correct way to get the same
      value as before without dereferencing invalid iterators is to prefer
      `&*my_std_string_view.rbegin() + 1`.
      
      Fix this downstream so that I might copy it back upstream in D148566.
      
      The other instance of `&*my_std_string_view.end()` that I introduced in
      D148546 has been fixed already in D149061.
      
      Reviewed By: ashay-github
      
      Differential Revision: https://reviews.llvm.org/D151760
      54a2994f
    • Peter Klausler's avatar
      [flang] CUDA Fortran - part 1/5: parsing · 4ad72793
      Peter Klausler authored
      Begin upstreaming of CUDA Fortran support in LLVM Flang.
      
      This first patch implements parsing for CUDA Fortran syntax,
      including:
       - a new LanguageFeature enum value for CUDA Fortran
       - driver change to enable that feature for *.cuf and *.CUF source files
       - parse tree representation of CUDA Fortran syntax
       - dumping and unparsing of the parse tree
       - the actual parsers for CUDA Fortran syntax
       - prescanning support for !@CUF and !$CUF
       - basic sanity testing via unparsing and parse tree dumps
      
      ... along with any minimized changes elsewhere to make these
      work, mostly no-op cases in common::visitors instances in
      semantics and lowering to allow them to compile in the face
      of new types in variant<> instances in the parse tree.
      
      Because CUDA Fortran allows the kernel launch chevron syntax
      ("call foo<<<blocks, threads>>>()") only on CALL statements and
      not on function references, the parse tree nodes for CallStmt,
      FunctionReference, and their shared Call were rearranged a bit;
      this caused a fair amount of one-line changes in many files.
      
      More patches will follow that implement CUDA Fortran in the symbol
      table and name resolution, and then semantic checking.
      
      Differential Revision: https://reviews.llvm.org/D150159
      4ad72793
    • Shubham Sandeep Rastogi's avatar
      Fix -u option in dsymutil, to not emit an extra DW_LNE_set_address if the... · 5a010894
      Shubham Sandeep Rastogi authored
      Fix -u option in dsymutil, to not emit an extra DW_LNE_set_address if the original line table was empty
      
      With dsymutil's -u option, only the accelerator tables should be
      updated, but with https://reviews.llvm.org/D150554 the -u option will
      still re-generate the line table. If the line table was empty, that is,
      it was a dummy line table, with no entries in it, dsymutil will always
      generate a line table with a DW_LNE_end_sequence, a funky side effect of
      this is that when the line table is re-generated, it will always emit a
      DW_LNE_set_address first, which will change the line table total size.
      This patch addresses this by making sure that if all the line table has
      in it is a DW_LNE_end_sequence, it is the same as a dummy entry.
      
      Differential Revision: https://reviews.llvm.org/D151579
      5a010894
    • LLVM GN Syncbot's avatar
      [gn build] Port 8e728adc · 142eccd9
      LLVM GN Syncbot authored
      142eccd9
    • Lorenzo Chelini's avatar
      [MLIR][Linalg] (NFC) Improve RUN command in `generalize-pad-tensor.mlir` · 9656c87b
      Lorenzo Chelini authored
      There is no need to specify any `check-prefix` here.
      9656c87b
    • Tom Stellard's avatar
      workflows/release-tasks: Upload lit releases to pypi · 3e984182
      Tom Stellard authored
      Reviewed By: thieta, kwk
      
      Differential Revision: https://reviews.llvm.org/D146491
      3e984182
    • Peter Klausler's avatar
      [flang] Fix interpretations of x87 80-bit Inf/NaN · 0e05ab67
      Peter Klausler authored
      Current implementations of x87 80-bit extended precision floating
      point interpret 7FFF8000000000000000 as +Inf, not a Nan.  The explicit
      MSB in the significand must be set for an infinity.
      
      Differential Revision: https://reviews.llvm.org/D151739
      0e05ab67
    • Tom Stellard's avatar
      clang/openmp: Fix alignment for ThreadID Address variables · e88fe818
      Tom Stellard authored
      There are places in the runtime, like __kmp_init_indirect_csptr, which
      assume these pointers are aligned to sizeof(void*), so make sure we emit
      them with the correct alignment.
      
      Fixes #62668
      
      Reviewed By: jlpeyton
      
      Differential Revision: https://reviews.llvm.org/D150723
      e88fe818
    • Arthur Eubanks's avatar
      [compiler-rt][CMake] Properly set COMPILER_RT_HAS_LLD · 395a614d
      Arthur Eubanks authored
      LLVM_TOOL_LLD_BUILD is a relic of the pre-monorepo times. This causes us to never set COMPILER_RT_HAS_LLD.
      
      Instead, set it from the runtimes build if lld is being built and lld is used as the compiler-rt linker.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D144660
      395a614d
    • Peter Klausler's avatar
      [flang] Detect output field width overflow for Inf/NaN · 763e036c
      Peter Klausler authored
      The output editing code paths for F and E/D output that handle
      IEEE-754 infinities and NaNs fail to check for overflow of the
      output field, which should cause the field to be filled with
      asterisks instead.  Catch these cases.
      
      Differential Revision: https://reviews.llvm.org/D151738
      763e036c
  2. May 31, 2023