1. May 04, 2020
    • Raul Tambre's avatar
      [AArch64] Add NVIDIA Carmel support · 0863e94e
      Raul Tambre authored
      Summary:
      NVIDIA's Carmel ARM64 cores are used in Tegra194 chips found in Jetson AGX Xavier, DRIVE AGX Xavier and DRIVE AGX Pegasus.
      
      References:
      * https://devblogs.nvidia.com/nvidia-jetson-agx-xavier-32-teraops-ai-robotics/#h.huq9xtg75a5e
      * NVIDIA Xavier Series System-on-Chip Technical Reference Manual 1.3 (https://developer.nvidia.com/embedded/downloads#?search=Xavier%20Series%20SoC%20Technical%20Reference%20Manual)
      
      Reviewers: sdesmalen, paquette
      
      Reviewed By: sdesmalen
      
      Subscribers: llvm-commits, ianshmean, kristof.beyls, hiraditya, jfb, danielkiss, cfe-commits, t.p.northover
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D77940
      0863e94e
    • Melanie Blower's avatar
      Reapply "Add support for #pragma float_control" with buildbot fixes · f5360d4b
      Melanie Blower authored
      Add support for #pragma float_control
      
      Reviewers: rjmccall, erichkeane, sepavloff
      
      Differential Revision: https://reviews.llvm.org/D72841
      
      This reverts commit fce82c0e.
      f5360d4b
    • Marcel Koester's avatar
      [mlir] Removed tight coupling of BufferPlacement pass to Alloc and Dealloc. · 67b466de
      Marcel Koester authored
      The current BufferPlacement implementation tries to find Alloc and Dealloc
      operations in order to move them. However, this is a tight coupling to
      standard-dialect ops which has been removed in this CL.
      
      Differential Revision: https://reviews.llvm.org/D78993
      67b466de
    • Kerry McLaughlin's avatar
      [SVE][Codegen] Lower legal min & max operations · 19f5da9c
      Kerry McLaughlin authored
      Summary:
      This patch adds AArch64ISD nodes for [S|U]MIN_PRED
      and [S|U]MAX_PRED, and lowers both SVE intrinsics and
      IR operations for min and max to these nodes.
      
      There are two forms of these instructions for SVE: a predicated
      form and an immediate (unpredicated) form. The patterns
      which existed for the latter have been updated to match a
      predicated node with an immediate and map this
      to the immediate instruction.
      
      Reviewers: sdesmalen, efriedma, dancgr, rengolin
      
      Reviewed By: efriedma
      
      Subscribers: huihuiz, tschuett, kristof.beyls, hiraditya, rkruppe, psnobl, cfe-commits, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D79087
      19f5da9c
    • Jay Foad's avatar
      [SLC] Allow llvm.pow(x,2.0) -> x*x etc even if no pow() lib func · e737847b
      Jay Foad authored
      optimizePow does not create any new calls to pow, so it should work
      regardless of whether the pow library function is available. This allows
      it to optimize the llvm.pow intrinsic on targets with no math library.
      
      Based on a patch by Tim Renouf.
      
      Differential Revision: https://reviews.llvm.org/D68231
      e737847b
    • Simon Pilgrim's avatar
      [InstCombine] Add tests showing failure to fold mul(abs(x),abs(x)) -> mul(x,x) (PR39476) · 8e9a8dc1
      Simon Pilgrim authored
      Includes abs() and nabs() variants
      8e9a8dc1
    • Florian Hahn's avatar
      [SCCP] Re-use pushToWorkList in pushToWorkListMsg (NFC). · 935685f4
      Florian Hahn authored
      There's no need to duplicate the logic to push to the different
      work-lists.
      935685f4
    • Hans Wennborg's avatar
      Fix building with GCC5 after e64f99c5 · 3c2c7760
      Hans Wennborg authored
      It was failing with:
      
        /work/llvm.monorepo/clang-tools-extra/clangd/ClangdServer.cpp: In lambda function:
        /work/llvm.monorepo/clang-tools-extra/clangd/ClangdServer.cpp:374:75:
        error: could not convert ‘(const char*)""’ from ‘const char*’ to ‘llvm::StringLiteral’
                                                        trace::Metric::Distribution);
                                                                                   ^
      3c2c7760
    • Jay Foad's avatar
      Precommit test updates for D68231. · 6c42814a
      Jay Foad authored
      6c42814a
    • Wen-Heng (Jack) Chung's avatar
      [mlir][rocdl] add rocdl.barier op. · bc23c1d8
      Wen-Heng (Jack) Chung authored
      - Add rocdl.barrier op.
      - Lower gpu.barier to rocdl.barrier in -convert-gpu-to-rocdl.
      
      Differential Revision: https://reviews.llvm.org/D79126
      bc23c1d8
    • Wen-Heng (Jack) Chung's avatar
      [mlir][vector] add tests for type_cast taking non-zero addrspace · a581c6f8
      Wen-Heng (Jack) Chung authored
      Add tests for vector.type_cast that takes memrefs on non-zero
      addrspaces.
      
      Differential Revision: https://reviews.llvm.org/D79099
      a581c6f8
    • Simon Moll's avatar
      [VE][NFC] formatting VEISD enum · 1e89f36c
      Simon Moll authored
      1e89f36c
    • Djordje Todorovic's avatar
      [llvm-dwarfdump][Stats] Clean up · 0a4defe8
      Djordje Todorovic authored
      This addresses:
        -Clean up the source code
        -Refactor the JSON fields
        -Fix the test cases
        -Improve the docs for the stats output
      
      Differential Revision: https://reviews.llvm.org/D77789
      0a4defe8
    • Craig Topper's avatar
      [X86] Simplify some code in combineTruncatedArithmetic. NFC · 243ffc0e
      Craig Topper authored
      We haven't promoted AND/OR/XOR to vXi64 types for a while. So
      there's no reason to use isOperationLegalOrPromote. So we can
      just use isOperationLegal by merging with ADD handling.
      243ffc0e
    • Craig Topper's avatar
      [X86] Custom legalize v16i64->v16i8 truncate with avx512. · 8b53fdd3
      Craig Topper authored
      Default legalization will create two v8i64 truncs to v8i32, concat
      them to v16i32, and then truncate the rest of the way to v16i8.
      
      Instead we can truncate directly from v8i64 to v8i8 in the lower
      half of an xmm. Then concat the two halves to use vpunpcklqdq.
      This is the same number of uops, but the dependency chain through
      the uops is better since the halves are merged at the end.
      
      I had to had SimplifyDemandedBits support for VTRUNC to prevent
      a regression on vector-trunc-math.ll. combineTruncatedArithmetic
      no longer gets a chance to shrink vXi64 mul so we were producing
      the v8i64 multiply sequence using multiple PMULUDQs. With the
      demanded bits fix we are able to prune out the extra ops leaving
      just two PMULUDQs, one for each v8i64 half. This is twice the
      width of the 2 v8i32 PMULLDs we had before, but PMULUDQ is 1
      uop and PMULLD is 2. We also save some truncates. It's probably
      worth using PMULUDQ even when PMULLQ is available since the latter
      is 3 uops, but that will require a different change.
      
      Differential Revision: https://reviews.llvm.org/D79231
      8b53fdd3
    • serge-sans-paille's avatar
      Make Polly tests dependencies explicit · 8ceee08d
      serge-sans-paille authored
      Due to libPolly now using the component infrastructure, it no longer carries all
      dependencies as it used to do.
      
      Differential Revision: https://reviews.llvm.org/D79295
      8ceee08d
    • Fangrui Song's avatar
      [llvm-objcopy] Avoid invalid Sec.Offset after D79229 · 762fb1c4
      Fangrui Song authored
      To avoid undefined behavior caught by -fsanitize=undefined on binary-paddr.test
      
        void SectionWriter::visit(const Section &Sec) {
          if (Sec.Type != SHT_NOBITS)
            // Sec.Contents is empty while Sec.Offset may be out of bound
            llvm::copy(Sec.Contents, Out.getBufferStart() + Sec.Offset);
        }
      762fb1c4
    • Johannes Doerfert's avatar
      [Attributor][NFC] Replace the nested AAMap with a key pair · 14cb0bdf
      Johannes Doerfert authored
      No functional change is intended.
      
      ---
      
      Single run of the Attributor module and then CGSCC pass (oldPM)
      for SPASS/clause.c (~10k LLVM-IR loc):
      
      Before:
      ```
      calls to allocation functions: 512375 (362871/s)
      temporary memory allocations: 98746 (69933/s)
      peak heap memory consumption: 22.54MB
      peak RSS (including heaptrack overhead): 106.78MB
      total memory leaked: 269.10KB
      ```
      
      After:
      ```
      calls to allocation functions: 509833 (338534/s)
      temporary memory allocations: 98902 (65671/s)
      peak heap memory consumption: 18.71MB
      peak RSS (including heaptrack overhead): 103.00MB
      total memory leaked: 269.10KB
      ```
      
      Difference:
      ```
      calls to allocation functions: -2542 (-27042/s)
      temporary memory allocations: 156 (1659/s)
      peak heap memory consumption: -3.83MB
      peak RSS (including heaptrack overhead): 0B
      total memory leaked: 0B
      ```
      14cb0bdf
    • Johannes Doerfert's avatar
      [Attributor] Remember only necessary dependences · 95e0d28b
      Johannes Doerfert authored
      Before we eagerly put dependences into the QueryMap as soon as we
      encountered them (via `Attributor::getAAFor<>` or
      `Attributor::recordDependence`). Now we will wait to see if the
      dependence is useful, that is if the target is not already in a fixpoint
      state at the end of the update. If so, there is no need to record the
      dependence at all.
      
      Due to the abstraction via `Attributor::updateAA` we will now also treat
      the very first update (during attribute creation) as we do subsequent
      updates.
      
      Finally this resolves the problematic usage of QueriedNonFixAA.
      
      ---
      
      Single run of the Attributor module and then CGSCC pass (oldPM)
      for SPASS/clause.c (~10k LLVM-IR loc):
      
      Before:
      ```
      calls to allocation functions: 554675 (389245/s)
      temporary memory allocations: 101574 (71280/s)
      peak heap memory consumption: 28.46MB
      peak RSS (including heaptrack overhead): 116.26MB
      total memory leaked: 269.10KB
      ```
      
      After:
      ```
      calls to allocation functions: 512465 (345559/s)
      temporary memory allocations: 98832 (66643/s)
      peak heap memory consumption: 22.54MB
      peak RSS (including heaptrack overhead): 106.58MB
      total memory leaked: 269.10KB
      ```
      
      Difference:
      ```
      calls to allocation functions: -42210 (-727758/s)
      temporary memory allocations: -2742 (-47275/s)
      peak heap memory consumption: -5.92MB
      peak RSS (including heaptrack overhead): 0B
      total memory leaked: 0B
      ```
      95e0d28b
    • Johannes Doerfert's avatar
      [Attributor] Inititialize "value attributes" w/ must-be-executed-context info · 231026a5
      Johannes Doerfert authored
      Attributes that only depend on the value (=bit pattern) can be
      initialized from uses in the must-be-executed-context (MBEC). We did use
      `AAComposeTwoGenericDeduction` and `AAFromMustBeExecutedContext` before
      to do this for some positions of these attributes but not for all. This
      was fairly complicated and also problematic as we did run it in every
      `updateImpl` call even though we only use known information. The new
      implementation removes `AAComposeTwoGenericDeduction`* and
      `AAFromMustBeExecutedContext` in favor of a simple interface
      `AddInformation::fromMBEContext(...)` which we call from the
      `initialize` methods of the "value attribute" `Impl` classes, e.g.
      `AANonNullImpl:initialize`.
      
      There can be two types of test changes:
        1) Artifacts were we miss some information that was known before a
           global fixpoint was reached and therefore available in an update
           but not at the beginning.
        2) Deduction for values we did not derive via the MBEC before or which
           were not found as the `AAFromMustBeExecutedContext::updateImpl` was
           never invoked.
      
      * An improved version of AAComposeTwoGenericDeduction can be found in
        D78718. Once we find a new use case that implementation will be able
        to handle "generic" AAs better.
      
      ---
      
      Single run of the Attributor module and then CGSCC pass (oldPM)
      for SPASS/clause.c (~10k LLVM-IR loc):
      
      Before:
      ```
      calls to allocation functions: 468428 (328952/s)
      temporary memory allocations: 77480 (54410/s)
      peak heap memory consumption: 32.71MB
      peak RSS (including heaptrack overhead): 122.46MB
      total memory leaked: 269.10KB
      ```
      
      After:
      ```
      calls to allocation functions: 554720 (351310/s)
      temporary memory allocations: 101650 (64376/s)
      peak heap memory consumption: 28.46MB
      peak RSS (including heaptrack overhead): 116.75MB
      total memory leaked: 269.10KB
      ```
      
      Difference:
      ```
      calls to allocation functions: 86292 (556722/s)
      temporary memory allocations: 24170 (155935/s)
      peak heap memory consumption: -4.25MB
      peak RSS (including heaptrack overhead): 0B
      total memory leaked: 0B
      ```
      
      Reviewed By: uenoku
      
      Differential Revision: https://reviews.llvm.org/D78719
      231026a5
    • Johannes Doerfert's avatar
      87f1e939
    • Johannes Doerfert's avatar
      [Attributor][NFC] Proactively ask for `nocapure` on call site arguments · 2f97b8b8
      Johannes Doerfert authored
      This minimizes test noise later on and is in line with other attributes
      we derive proactively.
      2f97b8b8
    • Sam McCall's avatar
      6fe20a44
    • Shilei Tian's avatar
      [OpenMP] Fix an issue of wrong return type of DeviceRTLTy::getNumOfDevices · cb038927
      Shilei Tian authored
      Summary: There is a typo in DeviceRTLTy::getNumOfDevices that the type of its return value is bool. It will lead to a problem of wrong device number returned from omp_get_num_devices.
      
      Reviewers: jdoerfert
      
      Reviewed By: jdoerfert
      
      Subscribers: yaxunl, guansong, openmp-commits
      
      Tags: #openmp
      
      Differential Revision: https://reviews.llvm.org/D79255
      cb038927
    • Kadir Cetinkaya's avatar
      [clangd] Reland LSP latency test · 81e48ae2
      Kadir Cetinkaya authored
      81e48ae2
    • Sergey Dmitriev's avatar
      [Attributor] Bitcast constant to the returned value type if it has different type · 0f70f733
      Sergey Dmitriev authored
      Reviewers: jdoerfert, sstefan1, uenoku
      
      Reviewed By: jdoerfert
      
      Subscribers: hiraditya, uenoku, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D79277
      0f70f733
    • Nikita Popov's avatar
      Revert "[InstSimplify] Remove known bits constant folding" · 46ee652c
      Nikita Popov authored
      This reverts commit 08556afc.
      
      This breaks some AMDGPU tests.
      46ee652c
    • Nikita Popov's avatar
      [InstSimplify] Remove known bits constant folding · 08556afc
      Nikita Popov authored
      If SimplifyInstruction() does not succeed in simplifying the
      instruction, it will compute the known bits of the instruction
      in the hope that all bits are known and the instruction can be
      folded to a constant. I have removed a similar optimization
      from InstCombine in D75801, and would like to drop this one as well.
      
      On average, we spend ~1% of total compile-time performing this
      known bits calculation. However, if we introduce some additional
      statistics for known bits computations and how many of them succeed
      in simplifying the instruction we get (on test-suite):
      
          instsimplify.NumKnownBits: 216
          instsimplify.NumKnownBitsComputed: 13828375
          valuetracking.NumKnownBitsComputed: 45860806
      
      Out of ~14M known bits calculations (accounting for approximately
      one third of all known bits calculations), only 0.0015% succeed in
      producing a constant. Those cases where we do succeed to compute
      all known bits will get folded by other passes like InstCombine
      later. On test-suite, only lencod.test and GCC-C-execute-pr44858.test
      show a hash difference after this change. On lencod we see an
      improvement (a loop phi is optimized away), on the GCC torture
      test a regression (a function return value is determined only
      after IPSCCP, preventing propagation from a noinline function.)
      
      There are various regressions in InstSimplify tests. However, all
      of these cases are already handled by InstCombine, and corresponding
      tests have already been added there.
      
      Differential Revision: https://reviews.llvm.org/D79294
      08556afc
    • Casey Carter's avatar
      [libc++][test] Use a non-narrowing conversion in assign_pair.pass.cpp · 7e3ef299
      Casey Carter authored
      ...to avoid warnings, e.g., from MSVC.
      7e3ef299
    • Hongtao Yu's avatar
      [ICP] Handling must tail calls in indirect call promotion · 911e06f5
      Hongtao Yu authored
      Per the IR convention, a musttail call must precede a ret with an optional bitcast. This was violated by the indirect call promotion optimization which could result an IR like:
      
          ; <label>:2192:
            br i1 %2198, label %2199, label %2201, !dbg !226012, !prof !229483
      
          ; <label>:2199:                                   ; preds = %2192
            musttail call fastcc void @foo(i8* %2195), !dbg !226012
            br label %2202, !dbg !226012
      
          ; <label>:2201:                                   ; preds = %2192
            musttail call fastcc void %2197(i8* %2195), !dbg !226012
            br label %2202, !dbg !226012
      
          ; <label>:2202:                                   ; preds = %605, %2201, %2199
            ret void, !dbg !229485
      
      This is being fixed in this change where the return statement goes together with the promoted indirect call. The code generated is like:
      
          ; <label>:2192:
            br i1 %2198, label %2199, label %2201, !dbg !226012, !prof !229483
      
          ; <label>:2199:                                   ; preds = %2192
            musttail call fastcc void @foo(i8* %2195), !dbg !226012
            ret void, !dbg !229485
      
          ; <label>:2201:                                   ; preds = %2192
            musttail call fastcc void %2197(i8* %2195), !dbg !226012
            ret void, !dbg !229485
      
      Differential Revision: https://reviews.llvm.org/D79258
      911e06f5
    • Mircea Trofin's avatar
      [llvm][NFC] Inliner: factor cost and reporting out of inlining process · bec4ab95
      Mircea Trofin authored
      Summary:
      This factors cost and reporting out of the inlining workflow, thus
      making it easier to reuse when driving inlining from the upcoming
      InliningAdvisor.
      
      Depends on: D79215
      
      Reviewers: davidxl, echristo
      
      Subscribers: eraman, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D79275
      bec4ab95
    • Florian Hahn's avatar
      bbdfcf8f
    • Johannes Doerfert's avatar
      [Attributor][NFC] Encode IRPositions in the bits of a single pointer · 8228153f
      Johannes Doerfert authored
      This reduces memory consumption for IRPositions by eliminating the
      vtable pointer and the `KindOrArgNo` integer. Since each abstract
      attribute has an associated IRPosition, the 12-16 bytes we save add up
      quickly.
      
      No functional change is intended.
      
      ---
      
      Single run of the Attributor module and then CGSCC pass (oldPM)
      for SPASS/clause.c (~10k LLVM-IR loc):
      
      Before:
      ```
      calls to allocation functions: 469545 (260135/s)
      temporary memory allocations: 77137 (42735/s)
      peak heap memory consumption: 30.50MB
      peak RSS (including heaptrack overhead): 119.50MB
      total memory leaked: 269.07KB
      ```
      
      After:
      ```
      calls to allocation functions: 468999 (274108/s)
      temporary memory allocations: 77002 (45004/s)
      peak heap memory consumption: 28.83MB
      peak RSS (including heaptrack overhead): 118.05MB
      total memory leaked: 269.07KB
      ```
      
      Difference:
      ```
      calls to allocation functions: -546 (5808/s)
      temporary memory allocations: -135 (1436/s)
      peak heap memory consumption: -1.67MB
      peak RSS (including heaptrack overhead): 0B
      total memory leaked: 0B
      ```
      
      ---
      
      CTMark 15 runs
      
      Metric: compile_time
      
      Program                                        lhs    rhs    diff
       test-suite...:: CTMark/sqlite3/sqlite3.test    25.07  24.09 -3.9%
       test-suite...Mark/mafft/pairlocalalign.test    14.58  14.14 -3.0%
       test-suite...-typeset/consumer-typeset.test    21.78  21.58 -0.9%
       test-suite :: CTMark/SPASS/SPASS.test          21.95  22.03  0.4%
       test-suite :: CTMark/lencod/lencod.test        25.43  25.50  0.3%
       test-suite...ark/tramp3d-v4/tramp3d-v4.test    23.88  23.83 -0.2%
       test-suite...TMark/7zip/7zip-benchmark.test    60.24  60.11 -0.2%
       test-suite :: CTMark/kimwitu++/kc.test         15.69  15.69 -0.0%
       test-suite...:: CTMark/ClamAV/clamscan.test    25.43  25.42 -0.0%
       test-suite :: CTMark/Bullet/bullet.test        37.63  37.62 -0.0%
       Geomean difference                                          -0.8%
      
      ---
      
      Reviewed By: lebedev.ri
      
      Differential Revision: https://reviews.llvm.org/D78722
      8228153f
    • Johannes Doerfert's avatar
      [Attributor][NFC] Let AbstractAttribute be an IRPosition · 6bf16ee4
      Johannes Doerfert authored
      Since every AbstractAttribute so far, and for the foreseeable future,
      corresponds to a single IRPosition we can simplify the class structure.
      We already did this for IRAttribute but there is no reason to stop
      there.
      6bf16ee4
    • Nico Weber's avatar
      Revert "Optimize path::remove_dots" · fb5fd746
      Nico Weber authored
      This reverts commit 53913a65.
      Breaks VFSFromYAMLTest.DirectoryIterationSameDirMultipleEntries
      in SupportTests on non-Windows.
      fb5fd746
    • Simon Pilgrim's avatar
    • Mircea Trofin's avatar
    • Kadir Cetinkaya's avatar
  2. May 03, 2020
    • Reid Kleckner's avatar
      Optimize path::remove_dots · 53913a65
      Reid Kleckner authored
      LLD calls this on every source file string in every object file when
      writing PDBs, so it is somewhat hot.
      
      Avoid rewriting paths that do not contain path traversal components
      (./..). Use find_first_not_of(separators) directly instead of using the
      path iterators. The path component iterators appear to be slow, and
      directly searching for slashes makes it easier to find double separators
      that need to be canonicalized.
      
      I discovered that the VFS relies on remote_dots to not canonicalize
      early slashes (/foo or C:/foo) on Windows, so I had to leave that
      behavior behind with unit tests for it. This is undesirable, but I claim
      that my change is NFC.
      53913a65
    • Reid Kleckner's avatar
      9b7f6146