1. Jul 14, 2021
    • Simon Pilgrim's avatar
      [InstCombine] Fold (select C, (gep Ptr, Idx), Ptr) -> (gep Ptr, (select C,... · d561b6fb
      Simon Pilgrim authored
      [InstCombine] Fold (select C, (gep Ptr, Idx), Ptr) -> (gep Ptr, (select C, Idx, 0)) (PR50183) (REAPPLIED)
      
      As discussed on PR50183, we already fold to prefer 'select-of-idx' vs 'select-of-gep':
      
      define <4 x i32>* @select0a(<4 x i32>* %a0, i64 %a1, i1 %a2, i64 %a3) {
        %gep0 = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a1
        %gep1 = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a3
        %sel = select i1 %a2, <4 x i32>* %gep0, <4 x i32>* %gep1
        ret <4 x i32>* %sel
      }
      -->
      define <4 x i32>* @select1a(<4 x i32>* %a0, i64 %a1, i1 %a2, i64 %a3) {
        %sel = select i1 %a2, i64 %a1, i64 %a3
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %sel
        ret <4 x i32>* %gep
      }
      
      This patch adds basic handling for the 'fallthrough' cases where the gep idx == 0 has been folded away to the base address:
      
      define <4 x i32>* @select0(<4 x i32>* %a0, i64 %a1, i1 %a2) {
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a1
        %sel = select i1 %a2, <4 x i32>* %a0, <4 x i32>* %gep
        ret <4 x i32>* %sel
      }
      -->
      define <4 x i32>* @select1(<4 x i32>* %a0, i64 %a1, i1 %a2) {
        %sel = select i1 %a2, i64 0, i64 %a1
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %sel
        ret <4 x i32>* %gep
      }
      
      Reapplied with a fix for the bpf "-bpf-disable-avoid-speculation" tests
      
      Differential Revision: https://reviews.llvm.org/D105901
      d561b6fb
    • Chuanqi Xu's avatar
      [NFC] [Coroutines] Remove unused CoroFree · 12d04ce9
      Chuanqi Xu authored
      12d04ce9
    • Bruce Mitchener's avatar
      [lldb][docs] Remove mention of subversion. NFC. · f7d931ac
      Bruce Mitchener authored
      Reviewed By: DavidSpickett
      
      Differential Revision: https://reviews.llvm.org/D103744
      f7d931ac
    • Simon Pilgrim's avatar
      [X86] Implement smarter instruction lowering for FP_TO_UINT from f32/f64 to... · ee71c1bb
      Simon Pilgrim authored
      [X86] Implement smarter instruction lowering for FP_TO_UINT from f32/f64 to i32/i64 and vXf32/vXf64 to vXi32 for SSE2 and AVX2 by using the exact semantic of the CVTTPS2SI instruction.
      
      We know that "CVTTPS2SI" returns 0x80000000 for out of range inputs (and for FP_TO_UINT, negative float values are undefined). We can use this to make unsigned conversions from vXf32 to vXi32 more efficient, particularly on targets without blend using the following logic:
      
      small := CVTTPS2SI(x);
      fp_to_ui(x) := small | (CVTTPS2SI(x - 2^31) & ARITHMETIC_RIGHT_SHIFT(small, 31))
      
      Even on targets where "PBLENDVPS"/"PBLENDVB" exists, it is often a latency 2, low throughput instruction so this logic is applied there too (in particular for AVX2 also). It furthermore gets rid of one high latency floating point comparison in the previous lowering.
      
      @TomHender checked the correctness of this for all possible floats between -1 and 2^32 (both ends excluded).
      
      Original Patch by @TomHender (Tom Hender)
      
      Differential Revision: https://reviews.llvm.org/D89697
      ee71c1bb
    • LLVM GN Syncbot's avatar
      [gn build] Port c08dabb0 · 90e7f5d2
      LLVM GN Syncbot authored
      90e7f5d2
    • Simon Pilgrim's avatar
      Revert rGb803294c : "[InstCombine] Fold... · 0722f3d0
      Simon Pilgrim authored
      Revert rGb803294c : "[InstCombine] Fold (select C, (gep Ptr, Idx), Ptr) -> (gep Ptr, (select C, Idx, 0)) (PR50183)"
      
      Missed some BPF test changes that need addressing
      0722f3d0
    • Nico Weber's avatar
      [gn build] (manually) merge 462d4de3 · aff09545
      Nico Weber authored
      aff09545
    • Stefan Pintilie's avatar
      [NFC][PowerPC] Added test to check regsiter allocation for ACC registers · cf0aa0b6
      Stefan Pintilie authored
      ACC regsiters are a combination of 4 consecutive vector regsiters and therefore
      somtimes require special treatment for register allocation. This patch only
      adds a test.
      cf0aa0b6
    • Stephen Tozer's avatar
      [DebugInfo] Correctly update dbg.values with duplicated location ops · 810e4c3c
      Stephen Tozer authored
      This patch fixes code that incorrectly handled dbg.values with duplicate
      location operands, i.e. !DIArgList(i32 %a, i32 %a). The errors in
      question were caused by either applying an update to dbg.value multiple
      times when the update is only valid once, or by updating the
      DIExpression for only the first instance of a value that appears
      multiple times.
      
      Differential Revision: https://reviews.llvm.org/D105831
      810e4c3c
    • Simon Pilgrim's avatar
      [InstCombine] Fold (select C, (gep Ptr, Idx), Ptr) -> (gep Ptr, (select C, Idx, 0)) (PR50183) · b803294c
      Simon Pilgrim authored
      As discussed on PR50183, we already fold to prefer 'select-of-idx' vs 'select-of-gep':
      
      define <4 x i32>* @select0a(<4 x i32>* %a0, i64 %a1, i1 %a2, i64 %a3) {
        %gep0 = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a1
        %gep1 = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a3
        %sel = select i1 %a2, <4 x i32>* %gep0, <4 x i32>* %gep1
        ret <4 x i32>* %sel
      }
      -->
      define <4 x i32>* @select1a(<4 x i32>* %a0, i64 %a1, i1 %a2, i64 %a3) {
        %sel = select i1 %a2, i64 %a1, i64 %a3
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %sel
        ret <4 x i32>* %gep
      }
      
      This patch adds basic handling for the 'fallthrough' cases where the gep idx == 0 has been folded away to the base address:
      
      define <4 x i32>* @select0(<4 x i32>* %a0, i64 %a1, i1 %a2) {
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %a1
        %sel = select i1 %a2, <4 x i32>* %a0, <4 x i32>* %gep
        ret <4 x i32>* %sel
      }
      -->
      define <4 x i32>* @select1(<4 x i32>* %a0, i64 %a1, i1 %a2) {
        %sel = select i1 %a2, i64 0, i64 %a1
        %gep = getelementptr inbounds <4 x i32>, <4 x i32>* %a0, i64 %sel
        ret <4 x i32>* %gep
      }
      
      Differential Revision: https://reviews.llvm.org/D105901
      b803294c
    • Butygin's avatar
    • Fraser Cormack's avatar
      [RISCV] Fix the neutral element in vector 'fadd' reductions · 03a4702c
      Fraser Cormack authored
      Using positive zero as the neutral element in 'fadd' reductions, while
      it generates better code, is incorrect. The correct neutral element is
      negative zero: 0.0 + -0.0 = 0.0, whereas -0.0 + -0.0 = -0.0.
      
      There are perhaps more optimal lowerings of negative zero avoiding
      constant-pool loads which could be left as future work.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D105902
      03a4702c
    • Sebastian Neubauer's avatar
      [AMDGPU] Init scratch only if necessary · 4359b870
      Sebastian Neubauer authored
      If no scratch or flat instructions are used, we do not need to
      initialize the flat scratch hardware register.
      
      Differential Revision: https://reviews.llvm.org/D105920
      4359b870
    • Sebastian Neubauer's avatar
      a12e5518
    • Cullen Rhodes's avatar
      [AArch64][SME] Add matrix register definitions and parsing support · c08dabb0
      Cullen Rhodes authored
      SME introduces the ZA array, a new piece of architectural register state
      consisting of a matrix of [SVLb x SVLb] bytes, where SVL is the
      implementation defined Streaming SVE vector length and SVLb is the
      number of 8-bit elements in a vector of SVL bits.
      
      SME instructions consist of three types of matrix operands:
      
        * Tiles: a ZA tile is a square, two-dimensional sub-array of elements
        within the ZA array. These tiles make up the larger accumulator array
        and the granularity varies based on the element size, i.e.
          - ZAQ0..ZAQ15 (smallest tile granule)
          - ZAD0..ZAD7
          - ZAS0..ZAS3
          - ZAH0..ZAH1
          or ZAB0       (largest tile granule, single tile)
        * Tile vectors: similar to regular tiles, but have an extra 'h' or 'v'
        to tell how the vector at [reg+offset] is layed out in the tile,
        horizontally or vertically. E.g. za1h.h or za15v.q, which corresponds
        to vectors in registers ZAH1 and ZAQ15, respectively.
        * Accumulator matrix: this is the entire accumulator array ZA.
      
      This patch adds the register classes and related operands and parsing
      for SME instructions operating on the accumulator array.
      
      The ADDHA and ADDVA instructions which operate on tiles are also added
      in this patch to make some use of the code added, later patches will
      make use of the other operands introduced here.
      
      The reference can be found here:
      https://developer.arm.com/documentation/ddi0602/2021-06
      
      Co-authored by: Sander de Smalen (@sdesmalen)
      
      Reviewed By: david-arm
      
      Differential Revision: https://reviews.llvm.org/D105570
      c08dabb0
    • Sam McCall's avatar
      [clangd] Add CMake option to (not) link in clang-tidy checks · 462d4de3
      Sam McCall authored
      This reduces the size of the dependency graph and makes incremental
      development a little more pleasant (less rebuilding).
      
      This introduces a bit of complexity/fragility as some tests verify
      clang-tidy behavior. I attempted to isolate these and build/run as much
      of the tests as possible in both configs to prevent rot.
      
      Expectation is that (some) developers will use this locally, but
      buildbots etc will keep testing clang-tidy.
      
      Fixes https://github.com/clangd/clangd/issues/233
      
      Differential Revision: https://reviews.llvm.org/D105679
      462d4de3
    • Ruiling Song's avatar
      [AMDGPU] Don't handle export done when unify exit nodes · d9b9fdd9
      Ruiling Song authored
      This patch aims to revert the changes introduced by D70781 D71192 D76364
      
      D70781 was introduced to fix hardware hang where we do not insert exp-
      null-done for a kill inside infinit loop. At that time we have not added
      exp-null-done for kill early termination, but I believe as for now, we will
      always add the exp-null-done for early termination case in LaterBranchLowering.
      
      D71192 was introduced to handle the only_kill case, which is also been
      handled by the kill early termination work.
      
      D76364 was used to fix a regression by D71192, where we cleared the done
      bit of the export in the existing program and not let the normal return
      block branching to the new unified return block.
      
      With this change, we just trust frontends have setup exp-done correctly
      which is true for all existing frontends. The backend only inserts
      exp-null-done for the kill cases which is handled in SILateBranchLowering.cpp.
      
      Reviewed by: critson
      
      Differential Revision: https://reviews.llvm.org/D105610
      d9b9fdd9
    • Ruiling Song's avatar
      [NFC][AMDGPU] autogenerate kill-infinite-loop.ll checks · 1d9585c8
      Ruiling Song authored
      This would help us to track the assembly changes to these tests.
      
      Reviewed by: foad
      
      Differential Revision: https://reviews.llvm.org/D105609
      1d9585c8
    • Ruiling Song's avatar
      [RegisterCoalescer] Resolve conflict based on liveness of subregister · 40e3df2a
      Ruiling Song authored
      Currently we are resolving lane/subregister conflict by visiting
      instructions sequentially in current block to see whether there is any
      use of the tainted lanes. To save compile time, we are not doing further
      check in successor blocks. This sounds reasonable without subgregister liveness.
      
      But since we have added subregister liveness tracking capability to
      register coalescer, we can easily determine whether we have subregister
      liveness conflict by checking subranges. This would help coalescing more
      COPYs for target that enables subregister liveness tracking.
      
      Reviewed by: arsenm, qcolombet
      
      Differential Revision: https://reviews.llvm.org/D104509
      40e3df2a
    • Kito Cheng's avatar
      [RISCV] Pass -u to linker correctly. · 5635d2a5
      Kito Cheng authored
      `-u` is a linker option used to pretend a symbol is undefined,
      this option are common used for forcing archive member extraction.
      
      This option should pass to `ld`, and many other toolchain in Clang
      like `tools::gnutools` has pass that too.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D105091
      5635d2a5
    • Martin Storsjö's avatar
      [libcxx] [test] Clarify weak_ptr_ret on Windows, remove a LIBCXX-WINDOWS-FIXME · 2c425c17
      Martin Storsjö authored
      On Windows, structs with a destructor are always returned indirectly;
      add this to the list of known exceptions in the test where the class
      isn't returned in registers as expected.
      
      Differential Revision: https://reviews.llvm.org/D105906
      2c425c17
    • Dmitry Vyukov's avatar
      sanitizer_common: add simpler ThreadRegistry ctor · dfd9808b
      Dmitry Vyukov authored
      Currently ThreadRegistry is overcomplicated because of tsan,
      it needs tid quarantine and reuse counters. Other sanitizers
      don't need that. It also seems that no other sanitizer now
      needs max number of threads. Asan used to need 2^24 limit,
      but it does not seem to be needed now. Other sanitizers blindly
      copy-pasted that without reasons. Lsan also uses quarantine,
      but I don't see why that may be potentially needed.
      
      Add a ThreadRegistry ctor that does not require any sizes
      and use it in all sanitizers except for tsan.
      In preparation for new tsan runtime, which won't need
      any of these parameters as well.
      
      Reviewed By: vitalybuka
      
      Differential Revision: https://reviews.llvm.org/D105713
      dfd9808b
    • Yuichi Yoshida's avatar
      Reformulate OrcJIT tutorial doc to make it more clear. · 8ae31b08
      Yuichi Yoshida authored
      Fixed a minor writing error. The text was hard to understand.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D105899
      8ae31b08
    • Zakk Chen's avatar
      [RISCV] Support overloading for RVV miscellaneous functions. · 08cf69c3
      Zakk Chen authored
      Based on this update to the intrinsic doc
      https://github.com/riscv/rvv-intrinsic-doc/pull/103
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D105611
      08cf69c3
    • Vitaly Buka's avatar
      [sanitizer] Fix type error in python 3 · 16f8207d
      Vitaly Buka authored
      16f8207d
    • Vitaly Buka's avatar
      94210b12
    • David Green's avatar
      Revert "[clang] Refactor AST printing tests to share more infrastructure" · 40ce58d0
      David Green authored
      This reverts commit 20176bc7 as some
      versions of GCC do not seem to handle the new code very well. They
      complain about:
      
      /tmp/ccqUQZyw.s: Assembler messages:
      /tmp/ccqUQZyw.s:1151: Error: symbol `_ZNSt14_Function_base13_Base_managerIN5clangUlPKNS1_4StmtEE2_EE10_M_managerERSt9_Any_dataRKS7_St18_Manager_operation' is already defined
      /tmp/ccqUQZyw.s:11963: Error: symbol `_ZNSt17_Function_handlerIFbPKN5clang4StmtEENS0_UlS3_E2_EE9_M_invokeERKSt9_Any_dataOS3_' is already defined
      
      This seems like it is some GCC issue, but multiple buildbots (and my
      local machine) are all failing because of it.
      40ce58d0
    • Vitaly Buka's avatar
      [sanitizer] Convert script to python 3 · ba127a45
      Vitaly Buka authored
      ba127a45
    • Michael Kruse's avatar
      [Polly] Fix typo. NFC. · d5c0b010
      Michael Kruse authored
      Thanks to Mugerwa Martin for reporting.
      d5c0b010
    • Jinsong Ji's avatar
      [AIX] Update testcase to use aix triple · 64785ac1
      Jinsong Ji authored
      We have implemented the basic MCAsmParser now, we can use the triple
      directly now.
      64785ac1
    • Hongtao Yu's avatar
      [CSSPGO][llvm-profgen] Fix a missing initalization · 6b04ecaa
      Hongtao Yu authored
      Fixing a missing initalization that accidentaly caused by https://reviews.llvm.org/D103178 .
      6b04ecaa
    • Hongtao Yu's avatar
      Revert "[CSSPGO][llvm-profgen] Fix a missing initalization" · 597e9c61
      Hongtao Yu authored
      This reverts commit fef5f445.
      597e9c61
    • Hongtao Yu's avatar
      [CSSPGO][llvm-profgen] Fix a missing initalization · fef5f445
      Hongtao Yu authored
      Fixing a missing initalization that accidentaly caused by https://reviews.llvm.org/D103178 .
      fef5f445
    • Shilei Tian's avatar
      [AbstractAttributor] Fold function calls to `__kmpc_is_spmd_exec_mode` if possible · 1100e4aa
      Shilei Tian authored
      In the device runtime there are many function calls to `__kmpc_is_spmd_exec_mode`
      to query the execution mode of current kernels. In many cases, user programs
      only contain target region executing in one mode. As a consequence, those runtime
      function calls will only return one value. If we can get rid of these function
      calls during compliation, it can potentially improve performance.
      
      In this patch, we use `AAKernelInfo` to analyze kernel execution. Basically, for
      each kernel (device) function `F`, we collect all kernel entries `K` that can
      reach `F`. A new AA, `AAFoldRuntimeCall`, is created for each call site. In each
      iteration, it will check all reaching kernel entries, and update the folded value
      accordingly.
      
      In the future we will support more function.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D105787
      1100e4aa
    • Philip Reames's avatar
      [SCEV] Handle zero stride correctly in howManyLessThans · 205ed009
      Philip Reames authored
      This is split from D105216, but the code is hoisted much earlier into
      the path where we can actually get a zero stride flowing through. Some
      fairly simple proofs handle the cases which show up in practice. The
      only test changes are the cases where we really do need a non-zero
      divider to produce the right result.
      
      Recommitting with isLoopInvariant() check.
      
      Differential Revision: https://reviews.llvm.org/D105921
      205ed009
    • Richard Smith's avatar
      Fix test trying to write a spurious output file into the source · 8a0f1163
      Richard Smith authored
      directory.
      
      This causes test failures if the source directory is read-only.
      8a0f1163
    • Hongtao Yu's avatar
      cda2394d
    • Hongtao Yu's avatar
      [CSSPGO] Do not import pseudo probe desc in thinLTO · 74b99b5c
      Hongtao Yu authored
      Previously we reliedy on pseudo probe descriptors to look up precomputed GUID during probe emission for inlined probes. Since we are moving to always using unique linkage names, GUID for functions can be computed in place from dwarf names. This eliminates the need of importing pseudo probe descs in thinlto, since those descs should be emitted by the original modules.
      
      This significantly reduces thinlto memory footprint in some extreme case where the number of imported modules for a single module is massive.
      
      Test Plan:
      
      Reviewed By: wenlei
      
      Differential Revision: https://reviews.llvm.org/D105248
      74b99b5c
    • Hongtao Yu's avatar
      [CSSPGO][llvm-profgen] Allow multiple executable load segments. · 07120384
      Hongtao Yu authored
      The linker or post-link optimizer can create an ELF image with multiple executable segments each of which will be loaded separately at run time. This breaks the assumption of llvm-profgen that currently only supports one base load address. What it ends up with is that the subsequent mmap events will be treated as an overwrite of the first mmap event which will in turn screw up address mapping. While it is non-trivial to support multiple separate load addresses and given that on x64 those segments will always be loaded at consecutive addresses (though via separate mmap
      sys calls), I'm adding an error checking logic to bail out if that's violated and keep using a single load address which is the address of the first executable segment.
      
      Also changing the disassembly output from printing section offset to printing the virtual address instead, which matches the behavior of objdump.
      
      Differential Revision: https://reviews.llvm.org/D103178
      07120384
    • Vitaly Buka's avatar
      35ce6633