1. Mar 08, 2024
  2. Mar 07, 2024
    • sylvain-audi's avatar
      [NFC][Asan] Prepare AddressSanitizer to detect inserted runtime calls (#84223) · d6b3be37
      sylvain-audi authored
      This is in preparation for an upcoming commit that will add "funclet"
      OpBundle to the inserted runtime calls where the function's EH
      personality requires it.
      
      See PR https://github.com/llvm/llvm-project/pull/82533
      d6b3be37
    • Joseph Huber's avatar
      [LinkerWrapper] Use the correct empty file on Windows (#84322) · a213df5d
      Joseph Huber authored
      Summary:
      The clang-offload-bundler uses an empty file to control the bundles made
      for embedding. Previously this still used `/dev/null` by mistake even on
      Windows.
      a213df5d
    • Alexey Bataev's avatar
      [SLP]Improve minbitwidth analysis. · 4ce52e2d
      Alexey Bataev authored
      This improves overall analysis for minbitwidth in SLP. It allows to
      analyze the trees with store/insertelement root nodes. Also, instead of
      using single minbitwidth, detected from the very first analysis stage,
      it tries to detect the best one for each trunc/ext subtree in the graph
      and use it for the subtree.
      Results in better code and less vector register pressure.
      
      Metric: size..text
      
      Program                                                                                                                                                size..text
                                                                                                                                                             results     results0    diff
                                                                            test-suite :: SingleSource/Benchmarks/Adobe-C++/simple_types_loop_invariant.test    92549.00    92609.00  0.1%
                                                                                        test-suite :: External/SPEC/CINT2017speed/625.x264_s/625.x264_s.test   663381.00   663493.00  0.0%
                                                                                         test-suite :: External/SPEC/CINT2017rate/525.x264_r/525.x264_r.test   663381.00   663493.00  0.0%
                                                                                                     test-suite :: MultiSource/Benchmarks/Bullet/bullet.test   307182.00   307214.00  0.0%
                                                                                   test-suite :: External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s.test  1394420.00  1394484.00  0.0%
                                                                                    test-suite :: External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r.test  1394420.00  1394484.00  0.0%
                                                                                      test-suite :: External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test  2040257.00  2040273.00  0.0%
      
                                                                                    test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 12396098.00 12395858.00 -0.0%
                                                                                               test-suite :: External/SPEC/CINT2006/445.gobmk/445.gobmk.test   909944.00   909768.00 -0.0%
      
      SingleSource/Benchmarks/Adobe-C++/simple_types_loop_invariant - 4 scalar
      instructions remain scalar (good).
      Spec2017/x264 - the whole function idct4x4dc is vectorized using <16
      x i16> instead of <16 x i32>, also zext/trunc are removed. In other
      places last vector zext/sext removed and replaced by
      extractelement + scalar zext/sext pair.
      MultiSource/Benchmarks/Bullet/bullet - reduce or <4 x i32> replaced by
      reduce or <4 x i8>
      Spec2017/imagick - Removed extra zext from 2 packs of the operations.
      Spec2017/parest - Removed extra zext, replaced by extractelement+scalar
      zext
      Spec2017/blender - the whole bunch of vector zext/sext replaced by
      extractelement+scalar zext/sext, some extra code vectorized in smaller
      types.
      Spec2006/gobmk - fixed cost estimation, some small code remains scalar.
      
      Reviewers: RKSimon
      
      Pull Request: https://github.com/llvm/llvm-project/pull/84334
      4ce52e2d
    • Michael Maitland's avatar
      Revert "[GISEL] Add IRTranslation for shufflevector on scalable vector types" (#84330) · 552da248
      Michael Maitland authored
      Reverts llvm/llvm-project#80378
      
      causing Buildbot failures that did not show up with check-llvm or CI.
      552da248
    • SahilPatidar's avatar
      [DAG] Fix Failure to reassociate SMAX/SMIN/UMAX/UMIN (#82175) · 9e0f5909
      SahilPatidar authored
      Resolve #58110
      9e0f5909
    • David Spickett's avatar
      [lldb][test][FreeBSD] xfail TestPlatformConnect on AArch64 · 03588a27
      David Spickett authored
      Details in the linked issue. Might fail on other architectures
      but I can't confirm, they can add to this if it does.
      03588a27
    • Jordan Rupprecht's avatar
      [lldb][test] Enforce `pexpect` system availability by default (#84270) · 3239b4dc
      Jordan Rupprecht authored
      This switches the default of `LLDB_TEST_USE_VENDOR_PACKAGES` from `ON`
      to `OFF` in preparation for eventually deleting it. All known LLDB
      buildbots have this package installed, so flipping the default will
      uncover any other users.
      
      If this breaks anything, the preferred fix is to install `pexpect` on
      the host system. The second fix is to build with cmake option
      `-DLLDB_TEST_USE_VENDOR_PACKAGES=ON` as a temporary measure until
      `pexpect` can be installed. If neither of those work, reverting this
      patch is OK.
      3239b4dc
    • Michael Maitland's avatar
      [GISEL] Add IRTranslation for shufflevector on scalable vector types (#80378) · 2b8aaef0
      Michael Maitland authored
      This patch is stacked on
      https://github.com/llvm/llvm-project/pull/80372,
      https://github.com/llvm/llvm-project/pull/80307, and
      https://github.com/llvm/llvm-project/pull/80306.
      
      ShuffleVector on scalable vector types gets IRTranslate'd to
      G_SPLAT_VECTOR since a ShuffleVector that has operates on scalable
      vectors is a splat vector where the value of the splat vector is the 0th
      element of the first operand, because the index mask operand is the
      zeroinitializer (undef and poison are treated as zeroinitializer here).
      This is analogous to what happens in SelectionDAG for ShuffleVector.
      
      `buildSplatVector` is renamed to`buildBuildVectorSplatVector`. I did not
      make this a separate patch because it would cause problems to revert
      that change without reverting this change too.
      2b8aaef0
    • Joseph Huber's avatar
      [libc] Fix missing standard definitions in the GPU config · 043a0206
      Joseph Huber authored
      Summary:
      Some dependencies on the standard C extensions are added transitively.
      This patch adds the new values.
      043a0206
    • Marius Brehler's avatar
      [mlir][EmitC] Allow further ops within expressions (#84284) · f355cd6f
      Marius Brehler authored
      This adds the `CExpression` trait to additional ops to allow to use
      these ops within the expression operation. Furthermore, the operator
      precedence is defined for those ops.
      f355cd6f
    • Lei Huang's avatar
      [libc++] Fixes time formatter test output for Linux on PowerPC (#75526) · e4d4cfa5
      Lei Huang authored
      Fix output to match actual.
      e4d4cfa5
    • Yaxun (Sam) Liu's avatar
      [ClangOffloadBundler] fix unbundling archive (#84195) · 61b13e0d
      Yaxun (Sam) Liu authored
      When unbundling an archive, need to save the content of each object file
      to a temporary file before passing it to llvm-objcopy, instead of
      passing the original input archive file to llvm-objcopy.
      
      Also allows extracting host bundles for archives.
      
      Fixes: https://github.com/llvm/llvm-project/issues/83509
      61b13e0d
    • Florian Hahn's avatar
      4cfd4a78
    • Joseph Huber's avatar
      [Clang][NFC] Remove '--' separator in the linker wrapper usage (#84253) · 597be90f
      Joseph Huber authored
      Summary:
      The very first version of the `clang-linker-wrapper` used `--` as a
      separator for the host and device arguments. I moved away from this
      towards a commandline parsing implementation years ago but never got
      around to officially removing this.
      597be90f
    • cor3ntin's avatar
      [Clang] Fix approved revision of P2266 · 48dd118f
      cor3ntin authored
      48dd118f
    • Simon Pilgrim's avatar
      [TTI] SK_ExtractSubvector - Ensure we use the src / subvector types in the correct order · c669c038
      Simon Pilgrim authored
      Fixes typo in #84156, fixes buildbot assertion (most targets don't seem to care so tricky to create a testcase).
      c669c038
    • aniplcc's avatar
      [Clang] Update value for __cpp_implicit_move (#84216) (#84228) · 2acccf67
      aniplcc authored
      Fixes #84216
      2acccf67
    • Stefan Gränitz's avatar
      [clang-repl] Names declared in if conditions and for-init statements are local... · 4b70d17b
      Stefan Gränitz authored
      
      [clang-repl] Names declared in if conditions and for-init statements are local to the inner context (#84150)
      
      Make TopLevelStmtDecl a DeclContext so that variables defined in statements
      are attached to the TopLevelDeclContext. This fixes redefinition errors
      from variables declared in if conditions and for-init statements. These
      must be local to the inner context (C++ 3.3.2p4), but they had generated
      definitions on global scope instead.
      
      This PR makes the TopLevelStmtDecl looking more like a FunctionDecl and
      that's fine because the FunctionDecl is very close in terms of semantics.
      
      Additionally, ActOnForStmt() requires a CompoundScope when processing a
      NullStmt body.
      
      ---------
      
      Co-authored-by: default avatarVassil Vassilev <v.g.vassilev@gmail.com>
      4b70d17b
    • Stephen Tozer's avatar
      [RemoveDIs][DebugInfo][IR] Add parsing for non-intrinsic debug values (#79818) · 464d9d96
      Stephen Tozer authored
      This patch adds support for parsing the proposed non-instruction debug
      info ("RemoveDIs") from textual IR, and adds a test for the parser as well
      as a set of verifier tests that are dependent on parsing to fire.
      
      An important detail of this patch is the fact that although we can now
      parse in the RemoveDIs (new) and Intrinsic (old) debug info formats, we
      will always convert back to the old format at the end of parsing - this
      is done for two reasons: firstly to ensure that every tool is able to
      process IR printed in the new format, regardless of whether that tool
      has had RemoveDIs support added, and secondly to maintain the effect of
      the existing flags: for the tools where support for the new format has
      been added, we will run LLVM passes in the new format iff
      `--try-experimental-debuginfo-iterators=true`, and we will print in the
      new format iff `--write-experimental-debuginfo-iterators=true`; the
      format of the textual IR input should have no effect on either of these
      features.
      464d9d96
    • Kareem Ergawy's avatar
      [flang][OpenMP] Add `%flang_fc1` `RUN` to delayed privatization tests (#84296) · 59e405b3
      Kareem Ergawy authored
      I did not know how `-mmlir` flag works and was deferring the addition of
      `--openm-enabled-delayed-privatization` until later because I thought
      some work needs to be done to do that. This commit just adds some extra
      `RUN` lines to delayed privatization tests to run them from `flang` as
      well.
      59e405b3
    • martinboehme's avatar
      [clang][nullability] Don't discard expression state before end of full-expression. (#82611) · d5aecf0c
      martinboehme authored
      In https://github.com/llvm/llvm-project/pull/72985, I made a change to
      discard
      expression state (`ExprToLoc` and `ExprToVal`) at the beginning of each
      basic
      block. I did so with the claim that "we never need to access entries
      from these
      maps outside of the current basic block", noting that there are
      exceptions to
      this claim when control flow happens inside a full-expression (the
      operands of
      `&&`, `||`, and the conditional operator live in different basic blocks
      than the
      operator itself) but that we already have a mechanism for retrieving the
      values
      of these operands from the environment for the block they are computed
      in.
      
      It turns out, however, that the operands of these operators aren't the
      only
      expressions whose values can be accessed from a different basic block;
      when
      control flow happens within a full-expression, that control flow can be
      "interposed" between an expression and its parent. Here is an example:
      
      ```cxx
      void f(int*, int);
      bool cond();
      
      void target() {
        int i = 0;
        f(&i, cond() ? 1 : 0);
      }
      ```
      
      ([godbolt](https://godbolt.org/z/hrbj1Mj3o))
      
      In the CFG[^1] , note how the expression for `&i` is computed in block
      B4,
      but the parent of this expression (the `CallExpr`) is located in block
      B1.
      The the argument expression `&i` and the `CallExpr` are essentially
      "torn apart"
      into different basic blocks by the conditional operator in the second
      argument.
      In other words, the edge between the `CallExpr` and its argument `&i`
      straddles
      the boundary between two blocks.
      
      I used to think that this scenario -- where an edge between an
      expression and
      one of its children straddles a block boundary -- could only happen
      between the
      expression that triggers the control flow (`&&`, `||`, or the
      conditional
      operator) and its children, but the example above shows that other
      expressions
      can be affected as well; the control flow is still triggered by `&&`,
      `||` or
      the conditional operator, but the expressions affected lie outside these
      operators.
      
      Discarding expression state too soon is harmful. For example, an
      analysis that
      checks the arguments of the `CallExpr` above would not be able to
      retrieve a
      value for the `&i` argument.
      
      This patch therefore ensures that we don't discard expression state
      before the
      end of a full-expression. In other cases -- when the evaluation of a
      full-expression is complete -- we still want to discard expression state
      for the
      reasons explained in https://github.com/llvm/llvm-project/pull/72985
      (avoid
      performing joins on boolean values that are no longer needed, which
      unnecessarily extends the flow condition; improve debuggability by
      removing
      clutter from the expression state).
      
      The impact on performance from this change is about a 1% slowdown in the
      Crubit nullability check benchmarks:
      
      ```
      name                              old cpu/op   new cpu/op   delta
      BM_PointerAnalysisCopyPointer     71.9µs ± 1%  71.9µs ± 2%    ~     (p=0.987 n=15+20)
      BM_PointerAnalysisIntLoop          190µs ± 1%   192µs ± 2%  +1.06%  (p=0.000 n=14+16)
      BM_PointerAnalysisPointerLoop      325µs ± 5%   324µs ± 4%    ~     (p=0.496 n=18+20)
      BM_PointerAnalysisBranch           193µs ± 0%   192µs ± 4%    ~     (p=0.488 n=14+18)
      BM_PointerAnalysisLoopAndBranch    521µs ± 1%   525µs ± 3%  +0.94%  (p=0.017 n=18+19)
      BM_PointerAnalysisTwoLoops         337µs ± 1%   341µs ± 3%  +1.19%  (p=0.004 n=17+19)
      BM_PointerAnalysisJoinFilePath    1.62ms ± 2%  1.64ms ± 3%  +0.92%  (p=0.021 n=20+20)
      BM_PointerAnalysisCallInLoop      1.14ms ± 1%  1.15ms ± 4%    ~     (p=0.135 n=16+18)
      ```
      
      [^1]:
      ```
       [B5 (ENTRY)]
         Succs (1): B4
      
       [B1]
         1: [B4.9] ? [B2.1] : [B3.1]
         2: [B4.4]([B4.6], [B1.1])
         Preds (2): B2 B3
         Succs (1): B0
      
       [B2]
         1: 1
         Preds (1): B4
         Succs (1): B1
      
       [B3]
         1: 0
         Preds (1): B4
         Succs (1): B1
      
       [B4]
         1: 0
         2: int i = 0;
         3: f
         4: [B4.3] (ImplicitCastExpr, FunctionToPointerDecay, void (*)(int *, int))
         5: i
         6: &[B4.5]
         7: cond
         8: [B4.7] (ImplicitCastExpr, FunctionToPointerDecay, _Bool (*)(void))
         9: [B4.8]()
         T: [B4.9] ? ... : ...
         Preds (1): B5
         Succs (2): B2 B3
      
       [B0 (EXIT)]
         Preds (1): B1
      ```
      d5aecf0c
    • martinboehme's avatar
    • Diana Picus's avatar
      [AMDGPU] Rename getNumVGPRBlocks. NFC (#84161) · 0086cc95
      Diana Picus authored
      Rename getNumVGPRBlocks to getEncodedNumVGPRBlocks, to clarify that it's
      using the encoding granule. This is used to program the hardware. In
      practice, the hardware will use the alloc granule instead, so this patch
      also adds a new helper, getAllocatedNumVGPRBlocks, which can be useful
      when driving heuristics.
      0086cc95
    • Nikolas Klauser's avatar
    • Jay Foad's avatar
      [AMDGPU] Simplify EXP Real instruction definitions. NFC. · 4119042d
      Jay Foad authored
      Pass the Pseudo (instead of its name) into EXP_Real_Row and
      EXP_Real_ComprVM since it is already available in all subclasses.
      4119042d
    • martinboehme's avatar
      Revert "[dataflow][nfc] Fix u8 string usage with c++20" (#84301) · 5830d1a2
      martinboehme authored
      Reverts llvm/llvm-project#84291
      
      The patch broke Windows builds.
      5830d1a2
    • Simon Pilgrim's avatar
      [CostModel] getInstructionCost - improve estimation of costs for length changing shuffles (#84156) · 55304d0d
      Simon Pilgrim authored
      Fix gap in the cost estimation for length changing shuffles, by adjusting the shuffle mask and either widening the shuffle inputs or extracting the lower elements of the result.
      
      A small step towards moving some of this implementation inside improveShuffleKindFromMask and/or target getShuffleCost handlers (and reduce the diffs in cost estimation depending on whether coming from a ShuffleVectorInst or the raw operands / mask components)
      55304d0d
    • Guillaume Chatelet's avatar
      [reland][libc] Remove UB specializations of type traits for `BigInt` (#84299) · 245d669f
      Guillaume Chatelet authored
      Note: This is a reland of #84035.
      
      The standard specifies that it it UB to specialize the following traits:
       - `std::is_integral`
       - `std::is_unsigned`
       - `std::make_unsigned`
       - `std::make_signed`
      
      This patch:
       - Removes specializations for `BigInt`
       - Transforms SFINAE for `bit.h` functions from template parameter to
         return type (This makes specialization easier).
       - Adds `BigInt` specialization for `bit.h` functions.
       - Fixes code depending on previous specializations.
      245d669f