1. Jan 20, 2020
    • dfukalov's avatar
      [SCEV] Swap guards estimation sequence. NFC · de34b54e
      dfukalov authored
      Summary:
      Loop unroll spends a lot of time in SCEVs processing in case when a function
      contains hundreds of simple 'for' loops with a quite complex arrays indexes like
      
        for (int i = 0; i < 8; ++i) {
          for (int j = 0; j < 32; ++j) {
            C[j*8+i] = B[j*32+i+128] + A[i*64+128];
          }
        }
        for (int i = 0; i < 8; ++i) {
          for (int j = 0; j < 8; ++j) {
            for (int k = 0; k < 32; ++k) {
              D[k*64+i*8+j] = D[k*64+i*8+j] + E[i+16] * C[k*8+j+256];
            }
          }
        }
      
      The patch improves loop unroll speed since isLoopBackedgeGuardedByCond takes
      much less time than isLoopEntryGuardedByCond in the edge case.
      
      Reviewers: skatkov, sanjoy, mkazantsev
      
      Reviewed By: sanjoy
      
      Subscribers: fhahn, hiraditya, javed.absar, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72929
      de34b54e
    • Raphael Isemann's avatar
      [lldb] Mark the implicit copy constructor as deleted when a move constructor is provided. · 22447a61
      Raphael Isemann authored
      Summary:
      CXXRecordDecls that have a move constructor but no copy constructor need to
      have their implicit copy constructor marked as deleted (see C++11 [class.copy]p7, p18)
      Currently we don't do that when building an AST with ClangASTContext which causes
      Sema to realise that the AST is malformed and asserting when trying to create an implicit
      copy constructor for us in the expression:
      ```
      Assertion failed: ((data().DefaultedCopyConstructorIsDeleted || needsOverloadResolutionForCopyConstructor())
          && "Copy constructor should not be deleted"), function setImplicitCopyConstructorIsDeleted, file include/clang/AST/DeclCXX.h, line 828.
      ```
      
      In the test case there is a class `NoCopyCstr` that should have its copy constructor marked as
      deleted (as it has a move constructor). When we end up trying to tab complete in the
      `IndirectlyDeletedCopyCstr` constructor, Sema realises that the `IndirectlyDeletedCopyCstr`
      has no implicit copy constructor and tries to create one for us. It then realises that
      `NoCopyCstr` also has no copy constructor it could find via lookup. However because we
      haven't marked the FieldDecl as having a deleted copy constructor the
      `needsOverloadResolutionForCopyConstructor()` returns false and the assert fails.
      `needsOverloadResolutionForCopyConstructor()` would return true if during the time we
      added the `NoCopyCstr` FieldDecl to `IndirectlyDeletedCopyCstr` we would have actually marked
      it as having a deleted copy constructor (which would then mark the copy constructor of
      `IndirectlyDeletedCopyCstr ` as needing overload resolution and Sema is happy).
      
      This patch sets the correct mark when we complete our CXXRecordDecls (which is the time when
      we know whether a copy constructor has been declared). In theory we don't have to do this if
      we had a Sema around when building our debug info AST but at the moment we don't have this
      so this has to do the job for now.
      
      Reviewers: shafik
      
      Reviewed By: shafik
      
      Subscribers: aprantl, JDevlieghere, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D72694
      22447a61
    • Alex Zinenko's avatar
      [mlir] clarify LangRef wording around control flow in regions · f63f5a22
      Alex Zinenko authored
      It was unclear what "exiting a region" meant in the existing formulation.
      Phrase it in terms of control flow transfer to the operation enclosing the
      region.
      
      Discussion: https://groups.google.com/a/tensorflow.org/d/msg/mlir/73d2O8gjTuA/xVj1KoCTBAAJ
      f63f5a22
    • Simon Tatham's avatar
      [ARM,MVE] Fix confusing MC names for MVE VMINA/VMAXA insns. · f3e73e88
      Simon Tatham authored
      Summary:
      A recent commit accidentally defined names like `MVE_VMAXAs8` as
      instances of the multiclass `MVE_VMINA`, and vice versa. This has no
      effect on the test suite, because nothing directly refers to those
      instruction names (the isel patterns are generated in Tablegen using
      `!cast<Instruction>(NAME)` inside a lower-level multiclass). But it
      means that `llvm-mc -show-inst` was listing VMAXA as VMINA, and it
      would also affect any further draft code gen patches that use those
      instruction ids.
      
      Reviewers: MarkMurrayARM, dmgreen, miyuki, ostannard
      
      Reviewed By: dmgreen
      
      Subscribers: kristof.beyls, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73034
      f3e73e88
    • Yi Kong's avatar
      [llvm-profdata] Fix hint message since argument format has changed · 01bfb366
      Yi Kong authored
      "-sample" option is now changed to "--sample".
      01bfb366
    • Christian Sigg's avatar
      [mlir] Add in-dialect lowering of gpu.all_reduce. · 8b2eb7c4
      Christian Sigg authored
      Reviewers: ftynse, nicolasvasilache, herhut
      
      Reviewed By: ftynse, herhut
      
      Subscribers: liufengdb, aartbik, herhut, merge_guards_bot, mgorny, mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, nicolasvasilache, arpith-jacob, mgester, lucyrfox, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72129
      8b2eb7c4
    • Andrzej Warzynski's avatar
      [AArch64][SVE] Extend int_aarch64_sve_ld1_gather_imm · 7e717b39
      Andrzej Warzynski authored
      The ACLE distinguishes between the following addressing modes for gather
      loads:
        * "scalar base, vector offset", and
        * "vector base, scalar offset".
      For the "vector base, scalar offset" case, the
      `int_aarch64_sve_ld1_gather_imm` intrinsic was added in 79f2422d.
      Currently, that intrinsic assumes that the scalar offset is passed as an
      immediate.  As a result, it does not cater for cases where scalar offset
      is stored in a register.
      
      In this patch `int_aarch64_sve_ld1_gather_imm` is extended so that all
      cases are covered:
      * `int_aarch64_sve_ld1_gather_imm` is renamed as
        `int_aarch64_sve_ld1_gather_scalar_offset`
      * new DAG combine rules are added for GLD1_IMM for scenarios where the
        offset is a non-immediate scalar or an out-of-range immediate
      * sve-intrinsics-gather-loads-vector-base.ll is renamed as
        sve-intrinsics-gather-loads-vector-base-imm-offset.ll
      * sve-intrinsics-gather-loads-vector-base-scalar-offset.ll is added to test
        file for non-immediate offsets
      
      Similar changes are made for scatter store intrinsics.
      
      Reviewed By: sdesmalen, efriedma
      
      Differential Revision: https://reviews.llvm.org/D71773
      7e717b39
    • Pavel Labath's avatar
      [lldb] Allow loading of minidumps with no process id · 468ca490
      Pavel Labath authored
      Summary:
      Normally, on linux we retrieve the process ID from the LinuxProcStatus
      stream (which is just the contents of /proc/%d/status pseudo-file).
      
      However, this stream is not strictly required (it's a breakpad
      extension), and we are encountering a fair amount of minidumps which do
      not have it present. It's not clear whether this is the case with all
      these minidumps, but the two known situations where this stream can be
      missing are:
      - /proc filesystem not mounted (or something to that effect)
      - process crashing after exhausting (almost) all file descriptors (so
        the minidump writer may not be able to open the /proc file)
      
      Since this is a corner case which will become less and less relevant
      (crashpad-generated minidumps should not suffer from this problem), I
      work around this problem by hardcoding the PID to 1 in these cases.
      The same thing is done by the gdb plugin when talking to a stub which
      does not report a process id (e.g. a hardware probe).
      
      Reviewers: jingham, clayborg
      
      Subscribers: markmentovai, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D70238
      468ca490
    • Pavel Labath's avatar
      [lldb] Don't process symlinks deep inside DWARFUnit · 27df2d9f
      Pavel Labath authored
      Summary:
      This code is handling debug info paths starting with /proc/self/cwd,
      which is one of the mechanisms people use to obtain "relocatable" debug
      info (the idea being that one starts the debugger with an appropriate
      cwd and things "just work").
      
      Instead of resolving the symlinks inside DWARFUnit, we can do the same
      thing more elegantly by hooking into the existing Module path remapping
      code. Since llvm::DWARFUnit does not support any similar functionality,
      doing things this way is also a step towards unifying llvm and lldb
      dwarf parsers.
      
      Reviewers: JDevlieghere, aprantl, clayborg, jdoerfert
      
      Subscribers: lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D71770
      27df2d9f
    • Stephen Kelly's avatar
      9a3ff478
    • Unnar Freyr Erlendsson's avatar
      Make SymbolFileDWARF::ParseLineTable use std::sort instead of insertion sort · 39f13354
      Unnar Freyr Erlendsson authored
      Summary:
      Motivation: When setting breakpoints in certain projects line sequences are frequently being inserted out of order.
      
      Rather than inserting sequences one at a time into a sorted line table, store all the line sequences as we're building them up and sort and flatten afterwards.
      
      Reviewers: jdoerfert, labath
      
      Reviewed By: labath
      
      Subscribers: teemperor, labath, mgrang, JDevlieghere, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D72909
      39f13354
    • Pavel Labath's avatar
      [lldb/DWARF] Simplify DWARFDebugInfoEntry::LookupAddress · b7af1bfa
      Pavel Labath authored
      Summary:
      This method was doing a lot more than it's only caller needed
      (DWARFDIE::LookupDeepestBlock) needed, so I inline it into the caller,
      and remove any code which is not actually used. This includes code for
      searching for the deepest function, and the code for working around
      incomplete DW_AT_low_pc/high_pc attributes on a compile unit DIE (modern
      compiler get this right, and this method is called on function DIEs
      anyway).
      
      This also improves our llvm consistency, as llvm::DWARFDebugInfoEntry is
      just a very simple struct with no nontrivial logic.
      
      Reviewers: JDevlieghere, aprantl
      
      Subscribers: lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D72920
      b7af1bfa
    • Stephen Kelly's avatar
      Fix clang-formatting for recent commits · 8248190a
      Stephen Kelly authored
      8248190a
    • Evgeniy Brevnov's avatar
      [LV] Vectorizer should adjust trip count in profile information · af7e1588
      Evgeniy Brevnov authored
      Summary: Vectorized loop processes VFxUF number of elements in one iteration thus total number of iterations decreases proportionally. In addition epilog loop may not have more than VFxUF - 1 iterations. This patch updates profile information accordingly.
      
      Reviewers: hsaito, Ayal, fhahn, reames, silvas, dcaballe, SjoerdMeijer, mkuper, DaniilSuchkov
      
      Reviewed By: Ayal, DaniilSuchkov
      
      Subscribers: fedor.sergeev, hiraditya, rkruppe, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D67905
      af7e1588
    • Kadir Cetinkaya's avatar
      [clang][CodeComplete] Propogate printing policy to FunctionDecl · 1f946ee2
      Kadir Cetinkaya authored
      Summary:
      Printing policy was not propogated to functiondecls when creating a
      completion string which resulted in canonical template parameters like
      `foo<type-parameter-0-0>`. This patch propogates printing policy to those as
      well.
      
      Fixes https://github.com/clangd/clangd/issues/76
      
      Reviewers: ilya-biryukov
      
      Subscribers: jkorous, arphaman, usaxena95, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72715
      1f946ee2
    • Stephen Kelly's avatar
    • Stephen Kelly's avatar
      Add missing tests for parent traversal · 514e3c36
      Stephen Kelly authored
      514e3c36
    • Haojian Wu's avatar
      [clangd] Remove a stale FIXME, NFC. · 61b56340
      Haojian Wu authored
      61b56340
    • Simon Pilgrim's avatar
      [X86][SSE] Add PACKSS SimplifyMultipleUseDemandedBits 'sign bit' handling. · eaa45484
      Simon Pilgrim authored
      Attempt to use SimplifyMultipleUseDemandedBits to simplify PACKSS if we're only after the sign bit.
      eaa45484
    • Pavel Labath's avatar
      [lldb/DWARF] Change how we construct a llvm::DWARFContext · 06e73f07
      Pavel Labath authored
      Summary:
      The goal of this patch is two-fold. First, it fixes a use-after-free in
      the construction of the llvm DWARFContext. This happened because the
      construction code was throwing away the lldb DataExtractors it got while
      reading the sections (unlike their llvm counterparts, these are also
      responsible for memory ownership). In most cases this did not matter,
      because the sections are just slices of the mmapped file data. But this
      isn't the case for compressed elf sections, in which case the section is
      decompressed into a heap buffer. A similar thing also happen with object
      files which are loaded from process memory.
      
      The second goal is to make it explicit which sections go into the llvm
      DWARFContext -- any access to the sections through both DWARF parsers
      carries a risk of parsing things twice, so it's better if this is a
      conscious decision. Also, this avoids loading completely irrelevant
      sections (e.g. .text). At present, the only section that needs to be
      present in the llvm DWARFContext is the debug_line_str. Using it through
      both APIs is not a problem, as there is no parsing involved.
      
      The first goal is achieved by loading the sections through the existing
      lldb DWARFContext APIs, which already do the caching. The second by
      explicitly enumerating the sections we wish to load.
      
      Reviewers: JDevlieghere, aprantl
      
      Subscribers: lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D72917
      06e73f07
    • Sjoerd Meijer's avatar
      [ARM][MVE] Tail-Predication: rematerialise iteration count in exit blocks · 8cba99e2
      Sjoerd Meijer authored
      This patch uses helper function rewriteLoopExitValues that is refactored in
      D72602 to rematerialise the iteration count in exit blocks, so that we can
      clean-up loop update expressions inside the hardware-loops later in
      ARMLowOverheadLoops, which is necessary to get actual performance gains for
      tail-predicated loops.
      
      Differential Revision: https://reviews.llvm.org/D72714
      8cba99e2
    • Evgeniy Brevnov's avatar
    • David Spickett's avatar
      [test] On Mac, don't try to use result of sysctl command if calling it failed. · 952a540b
      David Spickett authored
      If sysctl is not found at all, let the usual exception propogate
      so that the user can fix their env. If it fails because of the
      permissions required to read the property then print a warning
      and continue.
      
      Differential Revision: https://reviews.llvm.org/D72278
      952a540b
    • Evgeniy Brevnov's avatar
      [LoopUtils] Better accuracy for getLoopEstimatedTripCount. · 10357e1c
      Evgeniy Brevnov authored
      Summary: Current implementation of getLoopEstimatedTripCount returns 1 iteration less than it should. The reason is that in bottom tested loop first iteration is executed before first back branch is taken. For example for loop with !{!"branch_weights", i32 1 // taken, i32 1 // exit} metadata getLoopEstimatedTripCount gives 1 while actual number of iterations is 2.
      
      Reviewers: Ayal, fhahn
      
      Reviewed By: Ayal
      
      Subscribers: mgorny, hiraditya, zzheng, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D71990
      10357e1c
    • Awanish Pandey's avatar
      Recommit "[DWARF5][DebugInfo]: Added support for DebugInfo generation for auto... · 84c4c87e
      Awanish Pandey authored
      Recommit "[DWARF5][DebugInfo]: Added support for DebugInfo generation for auto return type for C++ member functions."
      
      Summary:
      This was reverted in 328e0f3d due to
      chromium bot failure. This revision addresses that case.
      
      Original commit message:
      Summary:
          This patch will provide support for auto return type for the C++ member
          functions. Before this return type of the member function is deduced and
          stored in the DIE.
          This patch includes llvm side implementation of this feature.
      
          Patch by: Awanish Pandey <Awanish.Pandey@amd.com>
      
          Reviewers: dblaikie, aprantl, shafik, alok, SouraVX, jini.susan.george
      
          Reviewed by: dblaikie
      
          Differential Revision: https://reviews.llvm.org/D70524
      84c4c87e
    • Georgii Rymar's avatar
      [llvm-objdump] - Fix the indentation when printing dynamic tags. · 547530cc
      Georgii Rymar authored
      We have a bug currently: printed tag names might overlap the
      value column. It happens for MIPS now.
      
      This patch adds a logic to calculate the size of indentation on fly
      to fix such issues.
      
      Differential revision: https://reviews.llvm.org/D72838
      547530cc
    • Sjoerd Meijer's avatar
      [IndVarSimplify][LoopUtils] rewriteLoopExitValues. NFCI · 93175a5c
      Sjoerd Meijer authored
      This moves `rewriteLoopExitValues()` from IndVarSimplify to LoopUtils thus
      making it a generic loop utility function.  This allows to rewrite loop exit
      values by just calling this function without running the whole IndVarSimplify
      pass.
      
      We use this in D72714 to rematerialise the iteration count in exit blocks, so
      that we can clean-up loop update expressions inside the hardware-loops later.
      
      Differential Revision: https://reviews.llvm.org/D72602
      93175a5c
    • Fangrui Song's avatar
      854f7be2
    • Georgii Rymar's avatar
      [llvm-mc] - Produce R_X86_64_PLT32 relocation for branches with JCC opcodes too. · 11e8e324
      Georgii Rymar authored
      The idea is to produce R_X86_64_PLT32 instead of
      R_X86_64_PC32 for branches.
      
      It fixes https://bugs.llvm.org/show_bug.cgi?id=44397.
      
      This patch teaches MC to do that for JCC (jump if condition is met)
      instructions. The new behavior matches modern GNU as.
      It is similar to D43383, which did the same for "call/jmp foo",
      but missed JCC cases.
      
      Differential revision: https://reviews.llvm.org/D72831
      11e8e324
    • Nathan Chancellor's avatar
      345e8ed4
    • David Green's avatar
      [ARM] MVE VLDn postinc · ff2e67a4
      David Green authored
      This adds Post inc variants of the VLD2/4 and VST2/4 instructions in
      MVE. It uses the same mechanism/nodes as Neon, transforming the
      intrinsic+add pair into a ARMISD::VLD2_UPD, which gets selected to a
      post-inc instruction. The code to do that is mostly taken from the
      existing Neon code, but simplified as less variants are needed.
      
      It also fills in some getTgtMemIntrinsic for the arm.mve.vld2/4
      instrinsics, which allow the nodes to have MMO's, calculated as the full
      length to the memory being loaded/stored.
      
      Differential Revision: https://reviews.llvm.org/D71194
      ff2e67a4
    • David Green's avatar
      [ARM] MVE VLDn post inc tests. NFC · d6075726
      David Green authored
      d6075726
    • David Green's avatar
      [ARM] Favour post inc for MVE loops · 5e51f755
      David Green authored
      We were previously not necessarily favouring postinc for the MVE loads
      and stores, leading to extra code prior to the loop to set up the
      preinc. MVE in general can benefit from postinc (as we don't have
      unrolled loops), and certain instructions like the VLD2's only post-inc
      versions are available.
      
      Differential Revision: https://reviews.llvm.org/D70790
      5e51f755
    • Fangrui Song's avatar
      [StackColoring] Remap FixedStackPseudoSourceValue frame index referenced by MachineMemOperand · eaab1bf2
      Fangrui Song authored
      StackColoring::remapInstructions() remaps MachineOperand frame index (e.g. %stack.1 -> %stack.0)
      but does not remap FixedStackPseudoSourceValue frame index (e.g. store 4 into %stack.1.ap2.i.i)
      referenced by MachineMemoryOperand.
      
      This can cause an assertion failure when LiveDebugValues references a dead stack object.
      
      It is difficult to craft a test case. -g, va_copy and stack-coloring are required.
      I can only reproduce it on ppc32.
      eaab1bf2
    • Kazuaki Ishizaki's avatar
      [mlir] NFC: Fix trivial typos in comments · fc817b09
      Kazuaki Ishizaki authored
      Differential Revision: https://reviews.llvm.org/D73012
      fc817b09
    • Eric Fiselier's avatar
      [libc++][libc++abi] Fix or suppress failing tests in single-threaded · d15fad26
      Eric Fiselier authored
      builds.
      
      Fix a libc++abi test that was incorrectly checking for threading
      primitives even when threading was disabled.
      
      Additionally, temporarily XFAIL some module tests that fail because
      the <atomic> header is unsupported but still built as a part of the
      std module.
      
      To properly address this libc++ would either need to produce a different
      module.modulemap for single-threaded configurations, or it would need
      to make the <atomic> header not hard-error and instead be empty
      for single-threaded configurations
      d15fad26
    • Richard Smith's avatar
    • Richard Smith's avatar
      List implicit operator== after implicit destructors in a vtable. · add2b7e4
      Richard Smith authored
      Summary:
      We previously listed first declared members, then implicit operator=,
      then implicit operator==, then implicit destructors. Per discussion on
      https://github.com/itanium-cxx-abi/cxx-abi/issues/88, put the implicit
      equality comparison operators at the very end, after all special member
      functions.
      
      Reviewers: rjmccall
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72897
      add2b7e4
    • Richard Smith's avatar
      PR42108 Consistently diagnose binding a reference template parameter to · 13fa4e2e
      Richard Smith authored
      a temporary.
      
      We previously failed to materialize a temporary when performing an
      implicit conversion to a reference type, resulting in our thinking the
      argument was a value rather than a reference in some cases.
      13fa4e2e
    • Michael Liao's avatar
      Reorder targets in alphabetical order. NFC. · 81942174
      Michael Liao authored
      81942174