1. Apr 11, 2024
    • Alexey Bataev's avatar
      [SLP]Buildvector for alternate instructions with non-profitable gather operands. · 2b00a73f
      Alexey Bataev authored
      If the operands of the potentially alternate node are going to produce
      buildvector sequences, which result in more instructions, than the
      original code, then suhinstructions should be vectorized as alternate
      node, better to end up with the buildvector node.
      
      Left column - experimental, Right - reference.
      
      Metric: size..text
      
      Program                                                                                                                                                size..text
                                                                                                                                                             results     results0    diff
                                                                                            test-suite :: SingleSource/Benchmarks/Adobe-C++/loop_unroll.test   413680.00   416272.00  0.6%
                                                                                    test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 12351788.00 12354844.00  0.0%
                                                                                        test-suite :: External/SPEC/CINT2017speed/625.x264_s/625.x264_s.test   664901.00   664949.00  0.0%
                                                                                         test-suite :: External/SPEC/CINT2017rate/525.x264_r/525.x264_r.test   664901.00   664949.00  0.0%
                                                                                      test-suite :: External/SPEC/CFP2017rate/511.povray_r/511.povray_r.test  1171371.00  1171355.00 -0.0%
                                                                                               test-suite :: MultiSource/Benchmarks/7zip/7zip-benchmark.test  1036396.00  1036284.00 -0.0%
                                                                               test-suite :: MultiSource/Benchmarks/MiBench/consumer-jpeg/consumer-jpeg.test   111280.00   111248.00 -0.0%
                                                                                    test-suite :: External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r.test  1392113.00  1391361.00 -0.1%
                                                                                   test-suite :: External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s.test  1392113.00  1391361.00 -0.1%
                                                                              test-suite :: MultiSource/Benchmarks/Prolangs-C/TimberWolfMC/timberwolfmc.test   281676.00   281452.00 -0.1%
                                                                                          test-suite :: MultiSource/Benchmarks/VersaBench/ecbdes/ecbdes.test     3025.00     3019.00 -0.2%
                                                                                      test-suite :: MultiSource/Benchmarks/Prolangs-C/plot2fig/plot2fig.test     6351.00     6335.00 -0.3%
      
      Metric: SLP.NumVectorInstructions
      
      Program                                                                                                                                                SLP.NumVectorInstructions
                                                                                                                                                             results                   results0 diff
                                                                                          test-suite :: MultiSource/Benchmarks/VersaBench/ecbdes/ecbdes.test    15.00                     16.00   6.7%
                                                                                         test-suite :: External/SPEC/CINT2017rate/525.x264_r/525.x264_r.test  1703.00                   1707.00   0.2%
                                                                                        test-suite :: External/SPEC/CINT2017speed/625.x264_s/625.x264_s.test  1703.00                   1707.00   0.2%
                                                                                    test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test 26241.00                  26239.00  -0.0%
                                                                                      test-suite :: External/SPEC/CFP2017rate/510.parest_r/510.parest_r.test 11761.00                  11754.00  -0.1%
                                                                              test-suite :: MultiSource/Benchmarks/Prolangs-C/TimberWolfMC/timberwolfmc.test   824.00                    822.00  -0.2%
                                                                                   test-suite :: External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s.test  5668.00                   5654.00  -0.2%
                                                                                    test-suite :: External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r.test  5668.00                   5654.00  -0.2%
                                                                                           test-suite :: External/SPEC/CINT2017rate/502.gcc_r/502.gcc_r.test   792.00                    790.00  -0.3%
                                                                                          test-suite :: External/SPEC/CINT2017speed/602.gcc_s/602.gcc_s.test   792.00                    790.00  -0.3%
                                                                                             test-suite :: MultiSource/Benchmarks/FreeBench/pifft/pifft.test  1389.00                   1384.00  -0.4%
                                                                                               test-suite :: MultiSource/Benchmarks/7zip/7zip-benchmark.test   596.00                    590.00  -1.0%
                                                                                      test-suite :: MultiSource/Benchmarks/Prolangs-C/plot2fig/plot2fig.test     6.00                      5.00 -16.7%
      
      Metric: exec_time
      
      Program                                                                                                                                                exec_time
                                                                                                                                                             results   results0  diff
                                                                                     test-suite :: External/SPEC/CFP2017rate/526.blender_r/526.blender_r.test     99.14    100.00    0.9%
      
      Other changes are not significant (less than 0.1% percent with exectime
      less 5 secs).
      
      SingleSource/Benchmarks/Adobe-C++/loop_unroll - same small patterns
      remain scalar, smaller code.
      External/SPEC/CFP2017rate/526.blender_r/526.blender_r - many small
      changes, some extra stores gets vectorized.
      External/SPEC/CINT2017speed/625.x264_s/625.x264_s
      External/SPEC/CINT2017rate/525.x264_r/525.x264_r
      x264 has one change in a loop body, in function ssim_end4, some code
      remain scalar, resulting in less code size.
      External/SPEC/CFP2017rate/511.povray_r/511.povray_r - some extra code
      gets vectorized, looks like some other patterns were matched.
      MultiSource/Benchmarks/7zip/7zip-benchmark - extra stores were
      vectorized (looks like the graphs become profitable)
      MultiSource/Benchmarks/MiBench/consumer-jpeg/consumer-jpeg - small
      changes in vectorized code (some small part remain scalar).
      External/SPEC/CFP2017rate/538.imagick_r/538.imagick_r
      External/SPEC/CFP2017speed/638.imagick_s/638.imagick_s
      Many changes cause by the fact that the code of one function becomes
      smaller (onvertLCHabToRGB) and this functions gets inlined after that.
      MultiSource/Benchmarks/Prolangs-C/TimberWolfMC/timberwolfmc - some small
      changes here and there, some extra code is vectorized, some remain
      scalar (2 x vectors)
      MultiSource/Benchmarks/VersaBench/ecbdes/ecbdes - emits 2 scalars
      + 2 insertelems instead of insert, broadcast, alt code (3 instructions,
        total 5 insts)
      MultiSource/Benchmarks/Prolangs-C/plot2fig/plot2fig - small graph
      becomes profitable and gets vectorized.
      External/SPEC/CINT2017rate/502.gcc_r/502.gcc_r
      External/SPEC/CINT2017speed/602.gcc_s/602.gcc_s
      Some small graph becomes profitable and gets vectorized.
      MultiSource/Benchmarks/FreeBench/pifft/pifft - no changes in final code.
      
      Reviewers: RKSimon, dtcxzyw
      
      Reviewed By: RKSimon
      
      Pull Request: https://github.com/llvm/llvm-project/pull/84978
      2b00a73f
    • Noah Goldstein's avatar
      [ValueTracking] Add support for `xor`/`disjoint or` in `isKnownNonZero` · 81cdd35c
      Noah Goldstein authored
      Handles cases like `X ^ Y == X` / `X disjoint| Y == X`.
      
      Both of these cases have identical logic to the existing `add` case,
      so just converting the `add` code to a more general helper.
      
      Proofs: https://alive2.llvm.org/ce/z/Htm7pe
      
      Closes #87706
      81cdd35c
    • Noah Goldstein's avatar
    • Noah Goldstein's avatar
      [ValueTracking] Add support for `xor`/`disjoint or` in `getInvertibleOperands` · 0c57a2e4
      Noah Goldstein authored
      This strengthens our `isKnownNonEqual` logic with some fairly
      trivial cases.
      
      Proofs: https://alive2.llvm.org/ce/z/4pxRTj
      
      Closes #87705
      0c57a2e4
    • Noah Goldstein's avatar
    • Noah Goldstein's avatar
      [ValueTracking] Add support for `insertelement` in `isKnownNonZero` · 9c545a14
      Noah Goldstein authored
      Inserts don't modify the data, so if all elements that end up in the
      destination are non-zero the result is non-zero.
      
      Closes #87703
      9c545a14
    • Noah Goldstein's avatar
    • Noah Goldstein's avatar
      [ValueTracking] Add support for `shufflevector` in `isKnownNonZero` · 87528bfe
      Noah Goldstein authored
      Shuffles don't modify the data, so if all elements that end up in the
      destination are non-zero the result is non-zero.
      
      Closes #87702
      87528bfe
    • Noah Goldstein's avatar
    • Kevin P. Neal's avatar
      [FPEnv][BitcodeReader] Correct strictfp test. · b9a3551c
      Kevin P. Neal authored
      Correct a strictfp test to follow the rules documented in the LangRef:
      https://llvm.org/docs/LangRef.html#constrained-floating-point-intrinsics
      
      This test needed the strictfp attribute added to a function definition.
      
      Test changes verified with D146845.
      b9a3551c
    • martinboehme's avatar
      [clang][dataflow] Propagate locations from result objects to initializers. (#87320) · 21009f46
      martinboehme authored
      Previously, we were propagating storage locations the other way around,
      i.e.
      from initializers to result objects, using `RecordValue::getLoc()`. This
      gave
      the wrong behavior in some cases -- see the newly added or fixed tests
      in this
      patch.
      
      In addition, this patch now unblocks removing the `RecordValue` class
      entirely,
      as we no longer need `RecordValue::getLoc()`.
      
      With this patch, the test `TransferTest.DifferentReferenceLocInJoin`
      started to
      fail because the framework now always uses the same storge location for
      a
      `MaterializeTemporaryExpr`, meaning that the code under test no longer
      set up
      the desired state where a variable of reference type is mapped to two
      different
      storage locations in environments being joined. Rather than trying to
      modify
      this test to set up the test condition again, I have chosen to replace
      the test
      with an equivalent test in DataflowEnvironmentTest.cpp that sets up the
      test
      condition directly; because this test is more direct, it will also be
      less
      brittle in the face of future changes.
      21009f46
    • Aaron Ballman's avatar
      int -> uintptr_t to silence diagnostics · 4d80dff8
      Aaron Ballman authored
      'int' may not be sufficiently large to store a pointer representation
      anyway, so this is also a correctness fix.
      4d80dff8
    • Jun Wang's avatar
      [AMDGPU] New clang option for emitting a waitcnt instruction after each memory instruction (#79236) · 86842e1f
      Jun Wang authored
      
      
      This patch introduces a new command-line option for clang, namely,
      amdgpu-precise-mem-op (or precise-memory in the backend). When this option is specified, a waitcnt
      instruction is generated after each memory load/store instruction. The
      counter values are always 0, but which counters are involved depends on
      the memory instruction.
      
      ---------
      
      Co-authored-by: default avatarJun Wang <jun.wang7@amd.com>
      86842e1f
    • Craig Topper's avatar
      [RISCV] Remove interrupt handler special case from... · f27f3697
      Craig Topper authored
      [RISCV] Remove interrupt handler special case from RISCVFrameLowering::determineCalleeSaves. (#88069)
      
      This code was trying to save temporary argument registers in interrupt
      handler functions that contain calls. With the exception that all FP
      registers are saved including the normally callee saved registers.
      
      If all of the callees use an FP ABI and the interrupt handler doesn't
      touch the normally callee saved FP registers, we don't need to save
      them.
      
      It doesn't appear that we need to special case functions with calls. The
      normal callee saved register handling will already check each of the calls
      and consider a register clobbered if the call doesn't explicitly say it is preserved.
      
      All of the test changes are from the removal of the FP callee saved
      registers. There are tests for interrupt handlers with F and D extension
      that use ilp32 or lp64 ABIs that are not affected by this change. They
      still save the FP callee saved registers as they should.
      
      gcc appears to have a bug where the D extension being enabled with the
      ilp32f or lp64f ABI does not save the FP callee saved regs. The callee
      would only save/restore the lower 32 bits and clobber the upper bits.
      LLVM saves the FP callee saved regs in this case and there is an
      unchanged test for it.
      
      The unnecessary save/restore was raised in this thread
      https://discourse.llvm.org/t/has-bugs-when-optimizing-save-restore-csrs-by-changing-csr-xlen-f32-interrupt/78200/1
      f27f3697
    • higher-performance's avatar
      Fix quadratic slowdown in AST matcher parent map generation (#87824) · c54afe5c
      higher-performance authored
      Avoids the need to linearly re-scan all seen parent nodes to check for
      duplicates, which previously caused a slowdown for ancestry checks in
      Clang AST matchers.
      
      Fixes: #86881
      c54afe5c
    • Kojo Acquah's avatar
      Update `LowerContractionToSMMLAPattern` to ingnore matvec (#88288) · 04bf1a40
      Kojo Acquah authored
      Patterns in `LowerContractionToSMMLAPattern` are designed to handle
      vector-to-matrix multiplication but not matrix-to-vector. This leads to
      the following error when processing `rhs` with rank < 2:
      
      ```
      iree-compile: /usr/local/google/home/kooljblack/code/iree-build/llvm-project/tools/mlir/include/mlir/IR/BuiltinTypeInterfaces.h.inc:268: int64_t mlir::detail::ShapedTypeTrait<mlir::VectorType>::getDimSize(unsigned int) const [ConcreteType = mlir::VectorType]: Assertion `idx < getRank() && "invalid index for shaped type"' failed.
      ```
      
      Updates to explicitly check the rhs rank and fail cases that cannot
      process.
      04bf1a40
    • David Green's avatar
      4dcf33b6
    • Vyacheslav Levytskyy's avatar
      [SPIRV] Tweak parsing of base type name in builtins (#88255) · 335d5d5f
      Vyacheslav Levytskyy authored
      This PR is a small improvement of parsing of base type name in builtins,
      allowing to understand `unsigned ...` types. The test case that fails
      without the fix is attached.
      335d5d5f
    • Simon Pilgrim's avatar
    • Aart Bik's avatar
      [mlir][sparse] update doc and examples of the [dis]assemble operations (#88213) · f388a3a4
      Aart Bik authored
      The doc and examples of the [dis]assemble operations did not reflect all
      the recent changes on order of the operands. Also clarified some of the
      text.
      f388a3a4
    • erichkeane's avatar
      [NFC] Remove unneeded 'maybe_unused' attributes · 3d468566
      erichkeane authored
      This was added while we only had a partial implementation of clauses, so
      we don't need these anymore.
      3d468566
    • erichkeane's avatar
      [NFC] Update SemaRef.Diag to just Diag in OpenACC implementation · 48c5c70f
      erichkeane authored
      I missed these two in my last patch as the two patches crossed in
      review, so correct this now.
      48c5c70f
    • Mehdi Amini's avatar
      Revert "Fix complex log1p accuracy with large abs values." (#88290) · 43b2b2eb
      Mehdi Amini authored
      Reverts llvm/llvm-project#88260
      
      The test fails on the GCC7 buildbot.
      43b2b2eb
    • Evgenii Stepanov's avatar
      [msan] Overflow intrinsics. (#88210) · e72c949c
      Evgenii Stepanov authored
      e72c949c
    • Craig Topper's avatar
      [RISCV] Optimize undef Even vector in getWideningInterleave. (#88221) · 323d3ab2
      Craig Topper authored
      We recently optimized the code when the Odd vector was undef to fix a
      poison bug.
      
      There are additional optimizations we can do if the even vector is
      undef. With Zvbb, we can use a single vwsll. Without Zvbb, we can use a
      vzext.vf2 and a vsll.
      323d3ab2
    • Jan Svoboda's avatar
      [clang][modules] Only compute affecting module maps with implicit search (#87849) · 51786eb5
      Jan Svoboda authored
      When writing out a PCM, we compute the set of module maps that did
      affect the compilation and we strip the rest to make the output
      independent of them. The most common way to read a module map that is
      not affecting is with implicit module map search. The other option is to
      pass a bunch of unnecessary `-fmodule-map-file=<path>` arguments on the
      command-line, in which case the client should probably not give those to
      Clang anyway.
      
      This makes serialization of explicit modules faster, mostly due to
      reduced file system traffic.
      51786eb5
    • Jan Svoboda's avatar
      [clang][modules] Stop eagerly reading files with diagnostic pragmas (#87442) · fc3dff9b
      Jan Svoboda authored
      This makes it so that the importer doesn't need to stat all input files
      of a module that contain diagnostic pragmas, reducing file system
      traffic.
      fc3dff9b
  2. Apr 10, 2024