1. Aug 28, 2020
  2. Aug 27, 2020
    • Benjamin Kramer's avatar
      fddf543e
    • Teresa Johnson's avatar
      [HeapProf] Clang and LLVM support for heap profiling instrumentation · 7ed8124d
      Teresa Johnson authored
      See RFC for background:
      http://lists.llvm.org/pipermail/llvm-dev/2020-June/142744.html
      
      Note that the runtime changes will be sent separately (hopefully this
      week, need to add some tests).
      
      This patch includes the LLVM pass to instrument memory accesses with
      either inline sequences to increment the access count in the shadow
      location, or alternatively to call into the runtime. It also changes
      calls to memset/memcpy/memmove to the equivalent runtime version.
      The pass is modeled on the address sanitizer pass.
      
      The clang changes add the driver option to invoke the new pass, and to
      link with the upcoming heap profiling runtime libraries.
      
      Currently there is no attempt to optimize the instrumentation, e.g. to
      aggregate updates to the same memory allocation. That will be
      implemented as follow on work.
      
      Differential Revision: https://reviews.llvm.org/D85948
      7ed8124d
    • Mikhail Maltsev's avatar
      Revert "[libcxx] Fix compile for BUILD_EXTERNAL_THREAD_LIBRARY" · a19fd1aa
      Mikhail Maltsev authored
      This reverts commit 3b71f915.
      
      The commit is breaking some build bots.
      a19fd1aa
    • Roman Lebedev's avatar
      [InstSimplify][EarlyCSE] Try to CSE PHI nodes in the same basic block · 6102310d
      Roman Lebedev authored
      Apparently, we don't do this, neither in EarlyCSE, nor in InstSimplify,
      nor in (old) GVN, but do in NewGVN and SimplifyCFG of all places..
      
      While i could teach EarlyCSE how to hash PHI nodes,
      we can't really do much (anything?) even if we find two identical
      PHI nodes in different basic blocks, same-BB case is the interesting one,
      and if we teach InstSimplify about it (which is what i wanted originally,
      https://reviews.llvm.org/D86530), we get EarlyCSE support for free.
      
      So i would think this is pretty uncontroversial.
      
      On vanilla llvm test-suite + RawSpeed, this has the following effects:
      ```
      | statistic name                                     | baseline  | proposed  |      Δ |        % |    \|%\| |
      |----------------------------------------------------|-----------|-----------|-------:|---------:|---------:|
      | instsimplify.NumPHICSE                             | 0         | 23779     |  23779 |    0.00% |    0.00% |
      | asm-printer.EmittedInsts                           | 7942328   | 7942392   |     64 |    0.00% |    0.00% |
      | assembler.ObjectBytes                              | 273069192 | 273084704 |  15512 |    0.01% |    0.01% |
      | correlated-value-propagation.NumPhis               | 18412     | 18539     |    127 |    0.69% |    0.69% |
      | early-cse.NumCSE                                   | 2183283   | 2183227   |    -56 |    0.00% |    0.00% |
      | early-cse.NumSimplify                              | 550105    | 542090    |  -8015 |   -1.46% |    1.46% |
      | instcombine.NumAggregateReconstructionsSimplified  | 73        | 4506      |   4433 | 6072.60% | 6072.60% |
      | instcombine.NumCombined                            | 3640264   | 3664769   |  24505 |    0.67% |    0.67% |
      | instcombine.NumDeadInst                            | 1778193   | 1783183   |   4990 |    0.28% |    0.28% |
      | instcount.NumCallInst                              | 1758401   | 1758799   |    398 |    0.02% |    0.02% |
      | instcount.NumInvokeInst                            | 59478     | 59502     |     24 |    0.04% |    0.04% |
      | instcount.NumPHIInst                               | 330557    | 330533    |    -24 |   -0.01% |    0.01% |
      | instcount.TotalInsts                               | 8831952   | 8832286   |    334 |    0.00% |    0.00% |
      | simplifycfg.NumInvokes                             | 4300      | 4410      |    110 |    2.56% |    2.56% |
      | simplifycfg.NumSimpl                               | 1019808   | 999607    | -20201 |   -1.98% |    1.98% |
      ```
      I.e. it fires ~24k times, causes +110 (+2.56%) more `invoke` -> `call`
      transforms, and counter-intuitively results in *more* instructions total.
      
      That being said, the PHI count doesn't decrease that much,
      and looking at some examples, it seems at least some of them
      were previously getting PHI CSE'd in SimplifyCFG of all places..
      
      I'm adjusting `Instruction::isIdenticalToWhenDefined()` at the same time.
      As a comment in `InstCombinerImpl::visitPHINode()` already stated,
      there are no guarantees on the ordering of the operands of a PHI node,
      so if we just naively compare them, we may false-negatively say that
      the nodes are not equal when the only difference is operand order,
      which is especially important since the fold is in InstSimplify,
      so we can't rely on InstCombine sorting them beforehand.
      
      Fixing this for the general case is costly (geomean +0.02%),
      and does not appear to catch anything in test-suite, but for
      the same-BB case, it's trivial, so let's fix at least that.
      
      As per http://llvm-compile-time-tracker.com/compare.php?from=04879086b44348cad600a0a1ccbe1f7776cc3cf9&to=82bdedb888b945df1e9f130dd3ac4dd3c96e2925&stat=instructions
      this appears to cause geomean +0.03% compile time increase (regression),
      but geomean -0.01%..-0.04% code size decrease (improvement).
      6102310d
    • Roman Lebedev's avatar
      [NFC][EarlyCSE][InstSimplify] Add tests for CSE of PHI nodes · 94d3dd8b
      Roman Lebedev authored
      PHI nodes depend on the block they're in,
      so we can only deal with the most basic case of same-BB PHI's.
      94d3dd8b
    • Russell Gallop's avatar
      [Test] Tidy up loose ends from LLVM_HAS_GLOBAL_ISEL · c9455d3c
      Russell Gallop authored
      This hasn't been allowed as a build option since r309990
      
      Remove leftover REQUIRES: global-isel
      
      Differential Revision: https://reviews.llvm.org/D86714
      c9455d3c
    • Louis Dionne's avatar
      49644cd9
    • David Nicuesa's avatar
      [libcxx] Fix compile for BUILD_EXTERNAL_THREAD_LIBRARY · 3b71f915
      David Nicuesa authored
      Fix compilation with -DLIBCXX_BUILD_EXTERNAL_THREAD_LIBRARY when using clang. Now linking target  'cxx_external_threads' with 'cxx-headers'. Fix mismatching visibility for `libcpp_timed_backoff_policy` function in file <__threading_support>.
      
      Reviewed By: #libc, ldionne
      
      Differential Revision: https://reviews.llvm.org/D86598
      3b71f915
    • Cullen Rhodes's avatar
      [CodeGen][AArch64] Support arm_sve_vector_bits attribute · 42587345
      Cullen Rhodes authored
      This patch implements codegen for the 'arm_sve_vector_bits' type
      attribute, defined by the Arm C Language Extensions (ACLE) for SVE [1].
      The purpose of this attribute is to define vector-length-specific (VLS)
      versions of existing vector-length-agnostic (VLA) types.
      
      VLSTs are represented as VectorType in the AST and fixed-length vectors
      in the IR everywhere except in function args/return. Implemented in this
      patch is codegen support for the following:
      
        * Implicit casting between VLA <-> VLS types.
        * Coercion of VLS types in function args/return.
        * Mangling of VLS types.
      
      Casting is handled by the CK_BitCast operation, which has been extended
      to support the two new vector kinds for fixed-length SVE predicate and
      data vectors, where the cast is implemented through memory rather than a
      bitcast which is unsupported. Implementing this as a normal bitcast
      would require relaxing checks in LLVM to allow bitcasting between
      scalable and fixed types. Ano...
      42587345
    • Alexandre Ganea's avatar
      [Support] On Windows, add optional support for {rpmalloc|snmalloc|mimalloc} · a6a37a2f
      Alexandre Ganea authored
      This patch optionally replaces the CRT allocator (i.e., malloc and free) with rpmalloc (mixed public domain licence/MIT licence) or snmalloc (MIT licence) or mimalloc (MIT licence). Please note that the source code for these allocators must be available outside of LLVM's tree.
      
      To enable, use `cmake ... -DLLVM_INTEGRATED_CRT_ALLOC=D:/git/rpmalloc -DLLVM_USE_CRT_RELEASE=MT` where `D:/git/rpmalloc` has already been git clone'd from `https://github.com/mjansson/rpmalloc`. The same applies to snmalloc and mimalloc.
      
      When enabled, the allocator will be embeded (statically linked) into the LLVM tools & libraries. This currently only works with the static CRT (/MT), although using the dynamic CRT (/MD) could potentially work as well in the future.
      
      When enabled, this changes the memory stack from:
        new/delete -> MS VC++ CRT malloc/free -> HeapAlloc -> VirtualAlloc
      to:
        new/delete -> {rpmalloc|snmalloc|mimalloc} -> VirtualAlloc
      
      The goal of this patch is to bypass the application's global heap - which is thread-safe thus inducing locking - and instead take advantage of a modern lock-free, thread cache, allocator. On a 6-core Xeon Skylake we observe a 2.5x decrease in execution time when linking a large scale application with LLD and ThinLTO (12 min 20 sec -> 5 min 34 sec), when all hardware threads are being used (using LLD's flag /opt:lldltojobs=all). On a dual 36-core Xeon Skylake with all hardware threads used, we observe a 24x decrease in execution time (1 h 2 min -> 2 min 38 sec) when linking a large application with LLD and ThinLTO. Clang build times also see a decrease in the range 5-10% depending on the configuration.
      
      Differential Revision: https://reviews.llvm.org/D71786
      a6a37a2f
    • diggerlin's avatar
      Revert "[AIX][XCOFF] emit symbol visibility for xcoff object file." · 6923b0a7
      diggerlin authored
      This reverts commit a0818689.
      
      Based on the Hubert Tong'comment  https://reviews.llvm.org/D84265#inline-799085
      6923b0a7
    • Alexandre E. Eichenberger's avatar
      [MLIR] MemRef Normalization for Dialects · a14a2805
      Alexandre E. Eichenberger authored
      When dealing with dialects that will results in function calls to
      external libraries, it is important to be able to handle maps as some
      dialects may require mapped data.  Before this patch, the detection of
      whether normalization can apply or not, operations are compared to an
      explicit list of operations (`alloc`, `dealloc`, `return`) or to the
      presence of specific operation interfaces (`AffineReadOpInterface`,
      `AffineWriteOpInterface`, `AffineDMAStartOp`, or `AffineDMAWaitOp`).
      
      This patch add a trait, `MemRefsNormalizable` to determine if an
      operation can have its `memrefs` normalized.
      
      This trait can be used in turn by dialects to assert that such
      operations are compatible with normalization of `memrefs` with
      nontrivial memory layout specification. An example is given in the
      literal tests.
      
      Differential Revision: https://reviews.llvm.org/D86236
      a14a2805
    • Benjamin Kramer's avatar
    • Benjamin Kramer's avatar
      2b7df270
    • Pavel Labath's avatar
      [lldb/cmake] Fix linking of lldbSymbolHelpers for 9cb222e7 · dd635062
      Pavel Labath authored
      I didn't find this locally because I have a /usr/include/gtest which is
      similar enough to the bundled one to make things appear to work.
      dd635062
    • Matt Arsenault's avatar
      AMDGPU: Hoist subtarget lookup · 6c770a09
      Matt Arsenault authored
      6c770a09