1. Mar 31, 2020
    • Sameer Sahasrabuddhe's avatar
      Introduce unify-loop-exits pass. · 3cbbded6
      Sameer Sahasrabuddhe authored
      For each natural loop with multiple exit blocks, this pass creates a
      new block N such that all exiting blocks now branch to N, and then
      control flow is redistributed to all the original exit blocks.
      
      The bulk of the tranformation is a new function introduced in
      BasicBlockUtils that an redirect control flow from a set of incoming
      blocks to a set of outgoing blocks via a common "hub".
      
      This is a useful workaround for a limitation in the structurizer which
      incorrectly orders blocks when processing a nest of loops. This pass
      bypasses that issue by ensuring that each natural loop is recognized
      as a separate region. Since the structurizer is a region pass, it no
      longer sees a nest of loops in a single region, and instead processes
      each "level" in the nesting as a separate region.
      
      The AMDGPU backend provides a new option to enable this pass before
      the structurizer, which may eventually be enabled by default.
      
      Reviewers: madhur13490, arsenm, nhaehnle
      
      Reviewed By: nhaehnle
      
      Differential Revision: https://reviews.llvm.org/D75865
      3cbbded6
    • Vedant Kumar's avatar
      [LoopVectorize] Fix crash on "getNoopOrZeroExtend cannot truncate!" (PR45259) · dcc410b5
      Vedant Kumar authored
      In InnerLoopVectorizer::getOrCreateTripCount, when the backedge taken
      count is a SCEV add expression, its type is defined by the type of the
      last operand of the add expression.
      
      In the test case from PR45259, this last operand happens to be a
      pointer, which (according to llvm::Type) does not have a primitive size
      in bits. In this case, LoopVectorize fails to truncate the SCEV and
      crashes as a result.
      
      Uing ScalarEvolution::getTypeSizeInBits makes the truncation work as expected.
      
      https://bugs.llvm.org/show_bug.cgi?id=45259
      
      Differential Revision: https://reviews.llvm.org/D76669
      dcc410b5
    • Fangrui Song's avatar
      [ELF] Allow SHF_LINK_ORDER and non-SHF_LINK_ORDER to be mixed · 673e81ee
      Fangrui Song authored
      Currently, `error: incompatible section flags for .rodata` is reported
      when we mix SHF_LINK_ORDER and non-SHF_LINK_ORDER sections in an output section.
      
      This is overconstrained. This patch allows mixed flags with the
      requirement that SHF_LINK_ORDER sections must be contiguous. Mixing
      flags is used by Linux aarch64 (https://github.com/ClangBuiltLinux/linux/issues/953)
      
        .init.data : { ... KEEP(*(__patchable_function_entries)) ... }
      
      When the integrated assembler is enabled, clang's -fpatchable-function-entry=N[,M]
      implementation sets the SHF_LINK_ORDER flag (D72215) to fix a number of
      garbage collection issues.
      
      Strictly speaking, the ELF specification does not require contiguous
      SHF_LINK_ORDER sections but for many current uses of SHF_LINK_ORDER like
      .ARM.exidx/__patchable_function_entries there has been a requirement for
      the sections to be contiguous on top of the requirements of the ELF
      specification.
      
      This patch also imposes one restriction: SHF_LINK_ORDER sections cannot
      be separated by a symbol assignment or a BYTE command. Not allowing BYTE
      is a natural extension that a non-SHF_LINK_ORDER cannot be a separator.
      Symbol assignments can delimiter the contents of SHF_LINK_ORDER
      sections.  Allowing SHF_LINK_ORDER sections across symbol assignments
      (especially __start_/__stop_) can make things hard to explain. The
      restriction should not be a problem for practical use cases.
      
      Reviewed By: psmith
      
      Differential Revision: https://reviews.llvm.org/D77007
      673e81ee
    • Raul Tambre's avatar
      [libc++] Fix wrong default value for LIBCXX_ENABLE_ASSERTIONS in documentation · 094b11c3
      Raul Tambre authored
      It's set to OFF by default at libcxx/CMakeLists.txt:73.
      
      Differential Revision: https://reviews.llvm.org/D76905
      094b11c3
    • Louis Dionne's avatar
      [libc++] Add support for a new keyword ADDITIONAL_COMPILE_FLAGS · 32c9efb4
      Louis Dionne authored
      This allows adding compilation flags for a single test, which can help
      eliminate some .sh.cpp tests and some custom handling in the libc++
      test format.
      
      It also works around the issue that .sh.cpp substitutions are _not_
      equivalent to the actual compiler command lines used to compile tests,
      since the compiler flags can be modified in local lit configurations,
      and substitutions are frozen at that point. For example using %{compile}
      in a .sh.cpp test in the coroutines subdirectory will not include the
      -fcoroutines-ts flag, which is added in the local lit config, because
      the %{compile} substitution is created long before we add -fcoroutines-ts
      to the compiler flags (in the lit.local.cfg for coroutines).
      32c9efb4
    • Fangrui Song's avatar
      2d19270e
    • Yuanfang Chen's avatar
      [X86] make sure POP has implicit def/use of stack pointer when materializing... · ece79f47
      Yuanfang Chen authored
      [X86] make sure POP has implicit def/use of stack pointer when materializing 8-bit immediates for minsize
      
      Summary:
      Otherwise PostRA list scheduler may reorder instruction, such as
      
      schedule this
      '''
      pushq  $0x8
      pop    %rbx
      lea    0x2a0(%rsp),%r15
      '''
      to
      '''
      pushq  $0x8
      lea    0x2a0(%rsp),%r15
      pop    %rbx
      '''
      by mistake. The patch is to prevent this to happen by making sure POP has
      implicit use of SP.
      
      Reviewers: craig.topper
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D77031
      ece79f47
  2. Mar 30, 2020