1. Mar 31, 2020
    • Jonas Devlieghere's avatar
      [lldb/CMake] Make check-lldb-* work for the standalone build. · 63aaecd5
      Jonas Devlieghere authored
      In order to run check-lldb-* we need the correct map_config directives
      in llvm-lit. For the standalone build, LLVM doesn't know about LLDB, and
      the lldb mappings are missing. In that case we build our own llvm-lit,
      and tell LLVM to use the llvm-lit in the lldb build directory.
      
      Differential revision: https://reviews.llvm.org/D76945
      63aaecd5
    • Matt Arsenault's avatar
      GlobalISel: Add accessor to known bits to CombinerHelper · a87ca9e4
      Matt Arsenault authored
      I need to pass known bits to a target combine matcher (which for some
      reason aren't methods in a subclass of CombinerHelper?)
      a87ca9e4
    • Matt Arsenault's avatar
      GlobalISel: Translate llvm.fshl/llvm.fshr · 23da702d
      Matt Arsenault authored
      23da702d
    • Nico Weber's avatar
      lld: Reduce number of references to undefined printed from 10 to 3. · 20eb719f
      Nico Weber authored
      As of a while ago, lld groups all undefined references to a single
      symbol in a single diagnostic. Back then, I made it so that we
      print up to 10 references to each undefined symbol.
      
      Having used this for a while, I never wished there were more
      references, but I sometimes found that this can print a lot of
      output. lld prints up to 10 diagnostics by default, and if
      each has 10 references (which I've seen in practice), and each
      undefined symbol produces 2 (possibly very long) lines of output,
      that's over 200 lines of error output.
      
      Let's try it with just 3 references for a while and see how
      that feels in practice.
      
      Differential Revision: https://reviews.llvm.org/D77017
      20eb719f
    • Thomas Raoux's avatar
      [ConstantFold][NFC] Compile time optimization for large vectors · 3ea0774b
      Thomas Raoux authored
      Optimize the common case of splat vector constant. For large vector
      going through all elements is expensive. For splatr/broadcast cases we
      can skip going through all elements.
      
      Differential Revision: https://reviews.llvm.org/D76664
      3ea0774b
    • LLVM GN Syncbot's avatar
      [gn build] Port 3cbbded6 · 8242509a
      LLVM GN Syncbot authored
      8242509a
    • Nico Weber's avatar
      Move CLANG_SYSTEMZ_DEFAULT_ARCH to config.h. · c506adcd
      Nico Weber authored
      Instead of using a global define; see comments on D75914.
      
      While here, port 9c9d88d8 to the GN build.
      c506adcd
    • Uday Bondhugula's avatar
      [MLIR] Fix permuteLoops utility · f273e5c5
      Uday Bondhugula authored
      Rewrite mlir::permuteLoops (affine loop permutation utility) to fix
      incorrect approach. Avoiding using sinkLoops entirely - use single move
      approach. Add test pass.
      
      This fixes https://bugs.llvm.org/show_bug.cgi?id=45328
      
      
      
      Depends on D77003.
      
      Signed-off-by: default avatarUday Bondhugula <uday@polymagelabs.com>
      
      Differential Revision: https://reviews.llvm.org/D77004
      f273e5c5
    • Jakub Kuderski's avatar
      [AMDGPU] Add Relocation Constant Support · 77ce2e21
      Jakub Kuderski authored
      Summary:
      This change adds amdgcn.reloc.constant intrinsic to the amdgpu backend, which will compile into a relocation entry in the resulting elf.
      
      The intrinsics takes a MetadataNode (String) as its only argument, which specifies the symbol name of the relocation entry.
      
      `SelectionDAGBuilder::getValueImpl` is changed to allow metadata operands passed through to ISel.
      
      Author: csyonghe <yonghe@google.com>
      
      Reviewers: tpr, nhaehnle
      
      Reviewed By: nhaehnle
      
      Subscribers: arsenm, kzhuravl, jvesely, wdng, yaxunl, dstuttard, t-tye, hiraditya, kerbowa, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D76440
      77ce2e21
    • Alexey Bataev's avatar
      [OPENMP50]Add codegen support for array shaping expression in depend · 7842e7eb
      Alexey Bataev authored
      clauses.
      
      Implemented codegen for array shaping operation in depend clauses. The
      begin of the expression is the pointer itself, while the size of the
      dependence data is the mukltiplacation of all dimensions in the array
      shaping expression.
      7842e7eb
    • Sid Manning's avatar
      [Hexagon] MaxAtomicPromoteWidth and MaxAtomicInlineWidth are not getting set. · 81194bfe
      Sid Manning authored
      Noticed when building llvm's c++ library.
      
      Differential Revision: https://reviews.llvm.org/D76546
      81194bfe
    • Sameer Sahasrabuddhe's avatar
      Introduce unify-loop-exits pass. · 3cbbded6
      Sameer Sahasrabuddhe authored
      For each natural loop with multiple exit blocks, this pass creates a
      new block N such that all exiting blocks now branch to N, and then
      control flow is redistributed to all the original exit blocks.
      
      The bulk of the tranformation is a new function introduced in
      BasicBlockUtils that an redirect control flow from a set of incoming
      blocks to a set of outgoing blocks via a common "hub".
      
      This is a useful workaround for a limitation in the structurizer which
      incorrectly orders blocks when processing a nest of loops. This pass
      bypasses that issue by ensuring that each natural loop is recognized
      as a separate region. Since the structurizer is a region pass, it no
      longer sees a nest of loops in a single region, and instead processes
      each "level" in the nesting as a separate region.
      
      The AMDGPU backend provides a new option to enable this pass before
      the structurizer, which may eventually be enabled by default.
      
      Reviewers: madhur13490, arsenm, nhaehnle
      
      Reviewed By: nhaehnle
      
      Differential Revision: https://reviews.llvm.org/D75865
      3cbbded6
    • Vedant Kumar's avatar
      [LoopVectorize] Fix crash on "getNoopOrZeroExtend cannot truncate!" (PR45259) · dcc410b5
      Vedant Kumar authored
      In InnerLoopVectorizer::getOrCreateTripCount, when the backedge taken
      count is a SCEV add expression, its type is defined by the type of the
      last operand of the add expression.
      
      In the test case from PR45259, this last operand happens to be a
      pointer, which (according to llvm::Type) does not have a primitive size
      in bits. In this case, LoopVectorize fails to truncate the SCEV and
      crashes as a result.
      
      Uing ScalarEvolution::getTypeSizeInBits makes the truncation work as expected.
      
      https://bugs.llvm.org/show_bug.cgi?id=45259
      
      Differential Revision: https://reviews.llvm.org/D76669
      dcc410b5
    • Fangrui Song's avatar
      [ELF] Allow SHF_LINK_ORDER and non-SHF_LINK_ORDER to be mixed · 673e81ee
      Fangrui Song authored
      Currently, `error: incompatible section flags for .rodata` is reported
      when we mix SHF_LINK_ORDER and non-SHF_LINK_ORDER sections in an output section.
      
      This is overconstrained. This patch allows mixed flags with the
      requirement that SHF_LINK_ORDER sections must be contiguous. Mixing
      flags is used by Linux aarch64 (https://github.com/ClangBuiltLinux/linux/issues/953)
      
        .init.data : { ... KEEP(*(__patchable_function_entries)) ... }
      
      When the integrated assembler is enabled, clang's -fpatchable-function-entry=N[,M]
      implementation sets the SHF_LINK_ORDER flag (D72215) to fix a number of
      garbage collection issues.
      
      Strictly speaking, the ELF specification does not require contiguous
      SHF_LINK_ORDER sections but for many current uses of SHF_LINK_ORDER like
      .ARM.exidx/__patchable_function_entries there has been a requirement for
      the sections to be contiguous on top of the requirements of the ELF
      specification.
      
      This patch also imposes one restriction: SHF_LINK_ORDER sections cannot
      be separated by a symbol assignment or a BYTE command. Not allowing BYTE
      is a natural extension that a non-SHF_LINK_ORDER cannot be a separator.
      Symbol assignments can delimiter the contents of SHF_LINK_ORDER
      sections.  Allowing SHF_LINK_ORDER sections across symbol assignments
      (especially __start_/__stop_) can make things hard to explain. The
      restriction should not be a problem for practical use cases.
      
      Reviewed By: psmith
      
      Differential Revision: https://reviews.llvm.org/D77007
      673e81ee
    • Raul Tambre's avatar
      [libc++] Fix wrong default value for LIBCXX_ENABLE_ASSERTIONS in documentation · 094b11c3
      Raul Tambre authored
      It's set to OFF by default at libcxx/CMakeLists.txt:73.
      
      Differential Revision: https://reviews.llvm.org/D76905
      094b11c3
    • Louis Dionne's avatar
      [libc++] Add support for a new keyword ADDITIONAL_COMPILE_FLAGS · 32c9efb4
      Louis Dionne authored
      This allows adding compilation flags for a single test, which can help
      eliminate some .sh.cpp tests and some custom handling in the libc++
      test format.
      
      It also works around the issue that .sh.cpp substitutions are _not_
      equivalent to the actual compiler command lines used to compile tests,
      since the compiler flags can be modified in local lit configurations,
      and substitutions are frozen at that point. For example using %{compile}
      in a .sh.cpp test in the coroutines subdirectory will not include the
      -fcoroutines-ts flag, which is added in the local lit config, because
      the %{compile} substitution is created long before we add -fcoroutines-ts
      to the compiler flags (in the lit.local.cfg for coroutines).
      32c9efb4
    • Fangrui Song's avatar
      2d19270e
    • Yuanfang Chen's avatar
      [X86] make sure POP has implicit def/use of stack pointer when materializing... · ece79f47
      Yuanfang Chen authored
      [X86] make sure POP has implicit def/use of stack pointer when materializing 8-bit immediates for minsize
      
      Summary:
      Otherwise PostRA list scheduler may reorder instruction, such as
      
      schedule this
      '''
      pushq  $0x8
      pop    %rbx
      lea    0x2a0(%rsp),%r15
      '''
      to
      '''
      pushq  $0x8
      lea    0x2a0(%rsp),%r15
      pop    %rbx
      '''
      by mistake. The patch is to prevent this to happen by making sure POP has
      implicit use of SP.
      
      Reviewers: craig.topper
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D77031
      ece79f47
  2. Mar 30, 2020