1. Oct 02, 2023
    • Martijn Vels's avatar
      [libc++] Optimize vector push_back to avoid continuous load and store of end pointer · 6fe4e033
      Martijn Vels authored
      Credits: this change is based on analysis and a proof of concept by
      gerbens@google.com.
      
      Before, the compiler loses track of end as 'this' and other references
      possibly escape beyond the compiler's scope. This can be see in the
      generated assembly:
      
           16.28 │200c80:   mov     %r15d,(%rax)
           60.87 │200c83:   add     $0x4,%rax
                 │200c87:   mov     %rax,-0x38(%rbp)
            0.03 │200c8b: → jmpq    200d4e
            ...
            ...
            1.69 │200d4e:   cmp     %r15d,%r12d
                 │200d51: → je      200c40
           16.34 │200d57:   inc     %r15d
            0.05 │200d5a:   mov     -0x38(%rbp),%rax
            3.27 │200d5e:   mov     -0x30(%rbp),%r13
            1.47 │200d62:   cmp     %r13,%rax
                 │200d65: → jne     200c80
      
      We fix this by always explicitly storing the loaded local and pointer
      back at the end of push back. This generates some slight source 'noise',
      but creates nice and compact fast path code, i.e.:
      
           32.64 │200760:   mov    %r14d,(%r12)
            9.97 │200764:   add    $0x4,%r12
            6.97 │200768:   mov    %r12,-0x38(%rbp)
           32.17 │20076c:   add    $0x1,%r14d
            2.36 │200770:   cmp    %r14d,%ebx
                 │200773: → je     200730
            8.98 │200775:   mov    -0x30(%rbp),%r13
            6.75 │200779:   cmp    %r13,%r12
                 │20077c: → jne    200760
      
      Now there is a single store for the push_back value (as before), and a
      single store for the end without a reload (dependency).
      
      For fully local vectors, (i.e., not referenced elsewhere), the capacity
      load and store inside the loop could also be removed, but this requires
      more substantial refactoring inside vector.
      
      Differential Revision: https://reviews.llvm.org/D80588
      6fe4e033
    • Timm Bäder's avatar
    • Kiran Chandramohan's avatar
      [Flang][OpenMP] NFC: Port worksharing loop tests to HLFIR lowering · 5b66987c
      Kiran Chandramohan authored
      These are copies of tests in ../flang/test/Lower/OpenMP/FIR/wsloop-*
      5b66987c
    • Corentin Jabot's avatar
      [C++] Implement "Deducing this" (P0847R7) · af475173
      Corentin Jabot authored
      This patch implements P0847R7 (partially),
      CWG2561 and CWG2653.
      
      Reviewed By: aaron.ballman, #clang-language-wg
      
      Differential Revision: https://reviews.llvm.org/D140828
      af475173
    • Matt Arsenault's avatar
      CodeGen: Disable isCopyInstrImpl if there are implicit operands · bc7d88fa
      Matt Arsenault authored
      This is a conservative workaround for broken liveness tracking of
      SUBREG_TO_REG to speculatively fix all targets. The current reported
      failures are on X86 only, but this issue should appear for all targets
      that use SUBREG_TO_REG. The next minimally correct refinement would be
      to disallow only implicit defs.
      
      The coalescer now introduces implicit-defs of the super register to
      track the dependency on other subregisters. If we see such an implicit
      operand, we cannot simply treat the subregister def as the result
      operand in case downstream users depend on the implicitly defined
      parts. Really target implementations should be considering the
      implicit defs and trying to interpret them appropriately (maybe with
      some generic helpers). The full implicit def could possibly be
      reported as the move result, rather than the subregister def but that
      requires additional work.
      
      Hopefully fixes #64060 as well.
      
      This needs to be applied to the release branch.
      
      https://reviews.llvm.org/D156346
      bc7d88fa
    • Florian Hahn's avatar
      [ConstraintElim] Add extra tests with nested loops and iv decrements. · e9b33d08
      Florian Hahn authored
      Add extra test coverage for induction logic to cover nested loops and
      loops with induction decrements. This adds coverage for upcoming
      patches.
      e9b33d08
    • Krasimir Georgiev's avatar
      Revert "[asan] Ensure __asan_register_elf_globals is called in COMDAT asan.module_ctor (#67745)" · 6420d330
      Krasimir Georgiev authored
      This reverts commit 16eed8c9.
      
      Causes some failures internally, will share privately with the author.
      6420d330
    • Simon Pilgrim's avatar
    • Simon Pilgrim's avatar
      [X86] matchIndexRecursively - fold zext(addlike(shl_nuw(x,c1),c2) patterns into LEA · 2984e352
      Simon Pilgrim authored
      Pulled out of D155472 - handle zeroextended scaled address indices
      2984e352
    • Simon Pilgrim's avatar
      [X86] Add test coverage for zext(or(shl_nuw(x,c1),c2)) pointer math · 29081420
      Simon Pilgrim authored
      Additional test coverage for D155472
      29081420
    • JP Lehr's avatar
    • Mats Petersson's avatar
      [flang]Add vscale argument parsing (#67676) · 11e68c7e
      Mats Petersson authored
      Support for vector scale range arguments, for AArch64 scalable vector
      extension (SVE) support.
      
      Adds -msve-vector-bits to the flang frontend, and for flang fc1 the
      options are -mvscale-min and -mvscale-max (optional). These match the
      clang and clang cc1 options for the same purposes.
      
      A further patch will actually USE these arguments.
      11e68c7e
    • Matt Arsenault's avatar
      RegisterCoalescer: Add implicit-def of super register when coalescing SUBREG_TO_REG · 414ff812
      Matt Arsenault authored
      Currently coalescing with SUBREG_TO_REG introduces an invisible load
      bearing undef. There is liveness for the super register not
      represented in the MIR.
      
      This is part 1 of a fix for regressions that appeared after
      b7836d85. The allocator started
      recognizing undef-def subregister MOVs as copies. Since there was no
      representation for the dependency on the high bits, different undef
      segments of the super register ended up disconnected and downstream
      users ended up observing different undefs than they did previously.
      
      This does not yet fix the regression. The isCopyInstr handling needs
      to start handling implicit-defs on any instruction.
      
      I wanted to include an end to end IR test since the actual failure
      only appeared with an interaction between the coalescer and the
      allocator. It's a bit bigger than I'd like but I'm having a bit of
      trouble reducing it to something which definitely shows a diff that's
      meaningful.
      
      The same problem likely exists everywhere trying to do anything with
      SUBREG_TO_REG. I don't understand how this managed to be broken for so
      long.
      
      This needs to be applied to the release branch.
      
      https://reviews.llvm.org/D156345
      414ff812
    • Martin Storsjö's avatar
      [LLD] [COFF] Restore the current dir as the first entry in the search path (#67857) · f906fd53
      Martin Storsjö authored
      Before af744f0b, the first entry
      among the search paths was the empty string, indicating searching
      in (or starting from) the current directory. After
      af744f0b, the toolchain/clang
      specific lib directories were added at the head of the search path.
      
      This would cause lookups of literal file names or relative paths
      to match paths in the toolchain, if there are coincidental files
      with similar names there, even if they would be find in the current
      directory as well.
      
      Change addClangLibSearchPaths to append to the list like all other
      operations on searchPaths - but move the invocation of the
      function to the right place in the sequence.
      
      This fixes #67779.
      f906fd53
    • Martin Storsjö's avatar
      [LLD] [COFF] Clarify -print-search-path for the empty string element (#67856) · 7d7d9e46
      Martin Storsjö authored
      Also switch the test case to use -NEXT to strictly match all lines.
      7d7d9e46
    • Nikita Popov's avatar
      [IR] Mark zext/sext constant expressions as undesirable · 3b25407d
      Nikita Popov authored
      Introduce isDesirableCastOp() which determines whether IR builder
      and constant folding should produce constant expressions for a
      given cast type. This mirrors what we do for binary operators.
      
      Mark zext/sext as undesirable, which prevents most creations of such
      constant expressions. This is still somewhat incomplete and there
      are a few more places that can create zext/sext expressions.
      
      This is part of the work for
      https://discourse.llvm.org/t/rfc-remove-most-constant-expressions/63179.
      
      The reason for the odd result in the constantexpr-fneg.c test is
      that initially the "a[]" global is created with an [0 x i32] type,
      at which point the icmp expression cannot be folded. Later it is
      replaced with an [1 x i32] global and the icmp gets folded away.
      But at that point we no longer fold the zext.
      3b25407d
    • Jie Fu's avatar
      [CodeGen] Fix -Wunused-variable in RegisterCoalescer.cpp (NFC) · 2214026e
      Jie Fu authored
      /llvm-project/llvm/lib/CodeGen/RegisterCoalescer.cpp:1320:18: error: unused variable 'DefSubIdx' [-Werror,-Wunused-variable]
        const unsigned DefSubIdx = DefMI->getOperand(0).getSubReg();
                       ^
      1 error generated.
      2214026e
    • Matt Arsenault's avatar
      RegisterCoalescer: Avoid redundant implicit-def on rematerialize · e28708d4
      Matt Arsenault authored
      If this was coalescing a def of a subregister with a def of the super
      register, it was introducing a redundant super-register def and
      marking the subregister def as dead.
      
      Resulting in something like:
      
        dead $eax = MOVr0, implicit-def $rax, implicit-def $rax
      
      Avoid this by checking if the new instruction already has the super
      def, so we end up with this instead:
      
        dead $eax = MOVr0, implicit-def $rax
      
      The dead flag looks suspicious to me, seems like it's easy to buggily
      interpret dead def of subreg and a non-dead def of an aliasing
      register. It seems to be intentional though.
      
      https://reviews.llvm.org/D156343
      e28708d4
    • Chuanqi Xu's avatar
      [C++20] [Modules] Fix crash when emitting module inits for duplicated modules · 3f092736
      Chuanqi Xu authored
      Close https://github.com/llvm/llvm-project/issues/67893
      
      The root cause of the crash is an oversight that we missed the point
      that the same module can be imported multiple times. And we should use
      `SmallSetVector` instead of `SmallVector` to filter the case.
      3f092736
    • Martin Storsjö's avatar
      [compiler-rt] Reinstate removal of CRT choice flags from CMAKE_*_FLAGS* (#67935) · 7bc09a47
      Martin Storsjö authored
      This reverts one part of commit
      9f4dfcb7, with a modified comment added
      about the code.
      
      Ideally, this would only be reinstated temporarily - but given the
      situation in vcpkg, it looks likely that they would keep passing the
      duplicate options for quite some time. The conflicting CRT choice
      usually are benign but only would cause warnings about one option
      overriding the other, if passing e.g. "/MDd /MT".
      
      However when vcpkg currently sets these options in CMAKE_*_FLAGS_DEBUG,
      it passes the redundant option /D_DEBUG; thus the compiler finally ends
      up with e.g. "/D_DEBUG /MDd /MT", which has the effect of defining
      _DEBUG while using a release mode CRT, which allegedly breaks the build.
      
      There's a PR up for removing this redundant /D_DEBUG option in vcpkg in
      https://github.com/microsoft/vcpkg/pull/34123. With that in place, this
      change wouldn't be strictly needed.
      7bc09a47
    • Matt Arsenault's avatar
      RegisterCoalescer: Handle implicit-def of a super register when rematerializing · b1295dd5
      Matt Arsenault authored
      Permit an implicit-def of a virtual register when rematerializing if
      it defines a super register of a subregister def. The
      rematerialization pre-legality check should really have been checking
      the implicit operands, but that should be fixed separately.
      
      https://reviews.llvm.org/D156331
      b1295dd5
    • Uday Bondhugula's avatar
      [LIT] NFC. Add missing punctuation on a LIT driver message (#67941) · 4533d474
      Uday Bondhugula authored
      Add missing punctuation on a LIT driver message.
      4533d474
    • David Sherwood's avatar
      [Analysis][SVE] Improve cost model for some extending masked loads (#65957) · fad69a50
      David Sherwood authored
      When performing a masked load of an unpacked SVE vector type, i.e.
      nxv8i8, followed by a zero- or sign-extend to an illegal wide type
      such as nxv8i32 we typically end up with a combination of an
      extending masked load and pair(s) of uunpklo/hi or sunpklo/hi
      instructions. For example, see test @masked_sload_8i8_8i32 in file
      
        CodeGen/AArch64/sve-masked-ldst-sext.ll
      
      where
      
        %aval = call <vscale x 8 x i8> @llvm.masked.load.nxv8i8(...
        %aext = sext <vscale x 8 x i8> %aval to <vscale x 8 x i32>
      
      gets lowered to
      
        ld1sb { z1.h }, ...
        sunpklo z0.s, z1.h
        sunpkhi z1.s, z1.h
      
      Currently the cost for the 'sext' operation in the example above is
      1, whereas this patch changes it to 2 to reflect the pair of
      instructions required. Similarly, when doing a masked load of a
      nxv8i8 and extending to nxv8i64 the cost is changed to 6 to reflect
      the 6 unpacks required.
      fad69a50
    • Matt Arsenault's avatar
      RegisterCoalescer: Add new rematerializing with subregister tests · 274ba2c9
      Matt Arsenault authored
      None of the existing MIR tests seem to be directly targeting this
      situation.
      274ba2c9
    • Nikita Popov's avatar
      [InstCombine] Avoid use of ConstantExpr::getSExt() (NFC) · 9ace23c9
      Nikita Popov authored
      Use the constant folding API instead.
      9ace23c9
    • Nikita Popov's avatar
      [AMDGPUInstCombine] Avoid use of ConstantExpr::getSExt() (NFC) · bc7ca917
      Nikita Popov authored
      Let the IRBuilder handle the constant folding instead.
      bc7ca917
    • Matt Arsenault's avatar
      RegisterCoalescer: Forcibly leave SSA to avoid MIR test errors · 32a23aec
      Matt Arsenault authored
      Not sure how to produce a test that demonstrates the problem
      today. The coalescer would have to introduce a verifier caught SSA
      violation, like multiple defs of a virtual register. I'm not sure what
      would do that now, but an upcoming patch will.
      
      https://reviews.llvm.org/D156271
      32a23aec
    • David Spickett's avatar
      Revert "[Flang] [FlangRT] Introduce FlangRT project as solution to Flang's... · ffc67bb3
      David Spickett authored
      Revert "[Flang] [FlangRT] Introduce FlangRT project as solution to Flang's runtime LLVM integration"
      
      This reverts commit 6403287e.
      
      This is failing on all but 1 of Linaro's flang builders.
      CMake Error at /home/tcwg-buildbot/worker/clang-aarch64-full-2stage/llvm/flang-rt/unittests/CMakeLists.txt:37 (message):
        Target llvm_gtest not found.
      ffc67bb3
    • Matthias Springer's avatar
      [mlir][bufferization] Better analysis around allocs and block arguments (#67923) · 43198b0a
      Matthias Springer authored
      Values that are the result of buffer allocation ops are guaranteed to
      *not* be the same allocation as block arguments of containing blocks.
      This fact can be used to allow for more aggressive simplification of
      `bufferization.dealloc` ops.
      43198b0a
    • Henrik G. Olsson's avatar
      [clang] Add missing canonicalization in int literal profile (#67822) · 53179129
      Henrik G. Olsson authored
      The addition of the type kind to the profile ID of IntegerLiterals
      results in e.g. size_t and unsigned long literals mismatch even on
      platforms where they are canonically the same type. This patch checks
      the Canonical field to determine whether to canonicalize the type first.
      
      rdar://116063468
      53179129
    • Timm Bäder's avatar
      8245ca99
    • jeanPerier's avatar
      [flang] Zero initialize uninitialized components in saved default init (#67777) · 87e25210
      jeanPerier authored
      Follow up up of https://github.com/llvm/llvm-project/pull/67693
      
      - Zero initialize uninitialized components of saved derived type entity
      with a default initial value.
      - Zero initialize uninitialized storage of common blocks with a member
      with an initial value.
      - Zero initialized uninitialized saved equivalence
      
      This removes all the cases where fir.global are created with an initial
      value that results in an undef in LLVM for part of the global, leading
      in surprising LLVM optimizations at -O2 for Fortran folks that expects
      there saved variables to be zero initialized if there is no explicit or
      default initial value.
      87e25210
    • Owen Pan's avatar
      [clang-format] Fix a bug in mis-annotating arrows (#67780) · c83d64f1
      Owen Pan authored
      Fixed #66923.
      c83d64f1
    • Owen Pan's avatar
      [clang-format] Fix a bug in RemoveParentheses: ReturnStatement (#67911) · 75441a68
      Owen Pan authored
      Don't remove the outermost parentheses surrounding a return statement
      expression when inside a function/lambda that has the decltype(auto)
      return type.
      
      Fixed #67892.
      75441a68
    • David Green's avatar
      aacefaf1
    • Timm Bäder's avatar
      dcb946a1
    • Kai Sasaki's avatar
      [mlir][affine] Check the input vector sizes to be greater than 0 (#65293) · 97829935
      Kai Sasaki authored
      In the process of vectorization of the affine loop, the 0 vector size
      causes the crash with building the invalid AffineForOp. We can catch the
      case beforehand propagating to the assertion.
      
      See: https://github.com/llvm/llvm-project/issues/64262
      97829935
    • Philip Reames's avatar
      [RISCV] Form vredsum from explode_vector + scalar (left) reduce (#67821) · f0505c3d
      Philip Reames authored
      This change adds two related DAG combines which together will take a
      left-reduce scalar add tree of an explode_vector, and will incrementally
      form a vector reduction of the vector prefix. If the entire vector is
      reduced, the result will be a reduction over the entire vector.
      
      Profitability wise, this relies on vredsum being cheaper than a pair of
      extracts and scalar add. Given vredsum is linear in LMUL, and the
      vslidedown required for the extract is *also* linear in LMUL, this is
      clearly true at higher index values. At N=2, it's a bit questionable,
      but I think the vredsum form is probably a better canonical form
      anyways.
      
      Note that this only matches left reduces. This happens to be the
      motivating example I have (from spec2017 x264). This approach could be
      generalized to handle right reduces without much effort, and could be
      generalized to handle any reduce whose tree starts with adjacent
      elements if desired. The approach fails for a reduce such as (A+C)+(B+D)
      because we can't find a root to start the reduce with without scanning
      the entire associative add expression. We could maybe explore using
      masked reduces for the root node, but that seems of questionable
      profitability. (As in, worth questioning - I haven't explored in any
      detail.)
      
      This is covering up a deficiency in SLP. If SLP encounters the scalar
      form of reduce_or(A) + reduce_sum(a) where a is some common
      vectorizeable tree, SLP will sometimes fail to revisit one of the
      reductions after vectorizing the other. Fixing this in SLP is hard, and
      there's no good reason not to handle the easy cases in the backend.
      
      Another option here would be to do this in VectorCombine or generic DAG.
      I chose not to as the profitability of the non-legal typed prefix cases
      is very target dependent. I think this makes sense as a starting point,
      even if we move it elsewhere later.
      
      This is currently restructed only to add reduces, but obviously makes
      sense for any associative reduction operator. Once this is approved, I
      plan to extend it in this manner. I'm simply staging work in case we
      decide to go in another direction.
      f0505c3d
    • Mehdi Amini's avatar
    • Kazu Hirata's avatar
      [BOLT] Fix the initialization of DWARFDataExtractor · a7517e12
      Kazu Hirata authored
      Without this patch, we pass Endian as one of the parameters to the
      constructor of DWARFDataExtractor.  The problem is that Endian is of:
      
        enum endianness {big, little, native};
      
      whereas the constructor is expecting "bool IsLittleEndian".  That is,
      we are relying on an implicit conversion to convert big and little to
      false and true, respectively.
      
      When we migrate llvm::support::endianness to std::endian in future, we
      can no longer rely on an implicit conversion because std::endian is
      declared with "enum class".  Even if we could, the conversion would
      not be guaranteed to work because, for example, libcxx defines:
      
        enum class endian {
          little = 0xDEAD,
          big = 0xFACE,
          :
      
      where big and little are not boolean values.
      
      This patch fixes the problem by properly converting Endian to a
      boolean value.
      a7517e12