1. Nov 13, 2019
    • Sjoerd Meijer's avatar
      [ARM][MVE] canTailPredicateLoop · d90804d2
      Sjoerd Meijer authored
      This implements TTI hook 'preferPredicateOverEpilogue' for MVE.  This is a
      first version and it operates on single block loops only. With this change, the
      vectoriser will now determine if tail-folding scalar remainder loops is
      possible/desired, which is the first step to generate MVE tail-predicated
      vector loops.
      
      This is disabled by default for now. I.e,, this is depends on option
      -disable-mve-tail-predication, which is off by default.
      
      I will follow up on this soon with a patch for the vectoriser to respect loop
      hint 'vectorize.predicate.enable'. I.e., with this loop hint set to Disabled,
      we don't want to tail-fold and we shouldn't query this TTI hook, which is
      done in D70125.
      
      Differential Revision: https://reviews.llvm.org/D69845
      d90804d2
    • Luís Marques's avatar
      [RISCV] Fix wrong CFI directives · a5ce8bd7
      Luís Marques authored
      Summary: Removes CFI CFA directives that could incorrectly propagate
      beyond the basic block they were inteded for. Specifically it removes
      the epilogue CFI directives. See the branch_and_tail_call test for an
      example of the issue. Should fix the stack unwinding issues caused by
      the incorrect directives.
      
      Reviewers: asb, lenary, shiva0217
      Reviewed By: lenary
      Tags: #llvm
      Differential Revision: https://reviews.llvm.org/D69723
      a5ce8bd7
    • Simon Tatham's avatar
      [ARM,MVE] Add intrinsics for contiguous load/stores. · a12f588e
      Simon Tatham authored
      This patch adds the ACLE intrinsics for all the MVE load and store
      instructions not already handled by D69791. These ones don't need new
      IR intrinsics, because they can be implemented in terms of standard
      LLVM IR constructions.
      
      Some of the load and store instructions access less than 128 bits of
      memory, sign/zero extending each value to a wider vector lane on load
      or truncating it on store. These are represented in IR by a load of a
      shorter vector followed by a zext/sext, and conversely, a trunc
      followed by a short store. Existing ISel patterns already recognize
      those combinations and turn them into the right MVE instructions.
      
      The predicated forms of all these instructions are represented in the
      same way, except that the ordinary load/store operation is replaced
      with the existing intrinsics @llvm.masked.{load,store}. These are
      currently only code-generated as predicated MVE load/store
      instructions if you give LLVM the `-enable-arm-maskedldst` option; so
      I've done that in the LLVM codegen test. When we make that the
      default, that option can be removed.
      
      In the Tablegen backend, I've had to add a handful of extra support
      features:
      
      * We need to be able to make clang::Address objects out of a
        pointer and an alignment (previously we only needed these when the
        user passed us an existing one).
      
      * We can now specify vector types that aren't 128 bits wide (for use
        in those intermediate values in IR), the parametrized type system
        can make one starting from two existing vector types (using the lane
        count of one and the element type of the other).
      
      * I've added support for code generation of pointer casts, and for
        specifying LLVM types as operands to IRBuilder operations (for zext
        and sext, though I think they'll come in useful again).
      
      * Now not all IR construction operations need to be specified as
        Builder.CreateFoo; some don't involve a Builder at all, and one
        passes it as a parameter to a tiny static helper function in
        CGBuiltin.cpp.
      
      Reviewers: ostannard, MarkMurrayARM, dmgreen
      
      Subscribers: kristof.beyls, cfe-commits, llvm-commits
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D70088
      a12f588e
    • Simon Pilgrim's avatar
      [X86][AVX] Add plausible schedule classes to MASKPAIR/VP2INTERSECT/VDPBF16PS instructions · 4d0e7b62
      Simon Pilgrim authored
      These are really just placeholders that use approximately the right resources - once we have CPUs scheduler models that support these instructions they will need revisiting.
      
      In the meantime this means that all instructions have a class of some kind., meaning models can be more easily flagged as complete.
      4d0e7b62
    • JonChesterfield's avatar
      [libomptarget] Move supporti.h to support.cu · fd9fa999
      JonChesterfield authored
      Summary:
      [libomptarget] Move supporti.h to support.cu
      Reimplementation of D69652, without the unity build and refactors.
      Will need a clean build of libomptarget as the cmakelists changed.
      
      Reviewers: ABataev, jdoerfert
      
      Reviewed By: jdoerfert
      
      Subscribers: mgorny, jfb, openmp-commits
      
      Tags: #openmp
      
      Differential Revision: https://reviews.llvm.org/D70131
      fd9fa999
    • Hans Wennborg's avatar
      Revert 57dd4b03 "[ValueTracking] Allow context-sensitive nullness check for non-pointers" · 6ea47759
      Hans Wennborg authored
      This caused miscompiles of Chromium (https://crbug.com/1023818). The reduced
      repro is small enough to fit here:
      
        $ cat /tmp/a.c
        unsigned char f(unsigned char *p) {
          unsigned char result = 0;
          for (int shift = 0; shift < 1; ++shift)
            result |= p[0] << (shift * 8);
          return result;
        }
        $ bin/clang -O2 -S -o - /tmp/a.c | grep -A4 f:
        f:                                      # @f
                .cfi_startproc
        # %bb.0:                                # %entry
                xorl    %eax, %eax
                retq
      
      That's nicely optimized, but I don't think it's the right result :-)
      
      > Same as D60846 but with a fix for the problem encountered there which
      > was a missing context adjustment in the handling of PHI nodes.
      >
      > The test that caused D60846 to be reverted was added in e15ab8f2.
      >
      > Reviewers: nikic, nlopes, mkazantsev,spatel, dlrobertson, uabelho, hakzsam
      >
      > Subscribers: hiraditya, bollu, llvm-commits
      >
      > Tags: #llvm
      >
      > Differential Revision: https://reviews.llvm.org/D69571
      
      This reverts commit 57dd4b03.
      6ea47759
    • Mirko Brkusanin's avatar
      [Mips] Add rematerialization support for ldi.fmt · fed17867
      Mirko Brkusanin authored
      Instruction ldi.fmt can be considered cheap enough to avoid spill and restore
      of value that it produces since it's loaded from immediate.
      
      Differential Revision: https://reviews.llvm.org/D69898
      fed17867
    • Simon Atanasyan's avatar
      [mips] Show an error if 64-bit target triple provided with 32-bit CPU · 068db2ed
      Simon Atanasyan authored
      When a 64-bit triple is used emit an error if the CPU only supports
      32-bit code.
      
      Patch by Miloš Stojanović.
      
      Differential Revision: https://reviews.llvm.org/D70018
      068db2ed
    • Simon Atanasyan's avatar
      [mips][test] Add Mips CPU tests. NFC · b3853d85
      Simon Atanasyan authored
      Adding tests check all available CPUs on Mips.
      
      Patch by Miloš Stojanović.
      
      Differential Revision: https://reviews.llvm.org/D70017
      b3853d85
    • Sven van Haastregt's avatar
      [OpenCL] Add remaining vector data builtin functions · 2fe674ba
      Sven van Haastregt authored
      Add the remaining half (fp16) vector data load and store builtin
      functions from the OpenCL C specification.
      
      Patch by Pierre Gondois and Sven van Haastregt.
      2fe674ba
    • Daniil Suchkov's avatar
      Temporarily revert "[InstCombine] Fold PHIs with equal incoming pointers" · cba4a277
      Daniil Suchkov authored
      Revert due to sanitizer-windows buildbot failure.
      
      This reverts commit bbb29738.
      cba4a277
    • David Stenberg's avatar
      [DebugInfo] Avoid creating entry values for clobbered registers · 5e646ff5
      David Stenberg authored
      Summary:
      Entry values are considered for parameters that have register-described
      DBG_VALUEs in the entry block (along with other conditions).
      
      If a parameter's value has been propagated from the caller to the
      callee, then the parameter's DBG_VALUE in the entry block may be
      described using a register defined by some instruction, and entry values
      should not be emitted for the parameter, which can currently occur.
      One such case was seen in the attached test case, in which the second
      parameter, which is described by a redefinition of the first parameter's
      register, would incorrectly get an entry value using the first
      parameter's register. This commit intends to solve such cases by keeping
      track of register defines, and ignoring DBG_VALUEs in the entry block
      that are described by such registers.
      
      In a RelWithDebInfo build of clang-8, the average size of the set was
      27, and in a RelWithDebInfo+ASan build it was 30.
      
      Reviewers: djtodoro, NikolaPrica, aprantl, vsk
      
      Reviewed By: djtodoro, vsk
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #debug-info, #llvm
      
      Differential Revision: https://reviews.llvm.org/D69889
      5e646ff5
    • David Stenberg's avatar
      [DebugInfo] Add helper for finding entry value candidates [NFC] · 4fec44cd
      David Stenberg authored
      Summary:
      The conditions that are used to determine if entry values should be
      emitted for a parameter are quite many, and will grow slightly
      in a follow-up commit, so move those to a helper function, as was
      suggested in the code review for D69889.
      
      Reviewers: djtodoro, NikolaPrica
      
      Reviewed By: djtodoro
      
      Subscribers: probinson, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D69955
      4fec44cd
    • Sander de Smalen's avatar
      [AArch64] Extend storeRegToStackSlot to spill SVE registers. · 3367686b
      Sander de Smalen authored
      This patch allows the register allocator to spill SVE registers to the stack.
      
      Reviewers: ostannard, efriedma, rengolin, cameron.mcinally
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D70082
      3367686b
    • Daniil Suchkov's avatar
      [InstCombine] Fold PHIs with equal incoming pointers · bbb29738
      Daniil Suchkov authored
      In case when all incoming values of a PHI are equal pointers, this
      transformation inserts a definition of such a pointer right after
      definition of the base pointer and replaces with this value both PHI and
      all it's incoming pointers. Primary goal of this transformation is
      canonicalization of this pattern in order to enable optimizations that
      can't handle PHIs. Non-inbounds pointers aren't currently supported.
      
      Reviewers: spatel, RKSimon, lebedev.ri, apilipenko
      
      Reviewed By: apilipenko
      
      Tags: #llvm
      
      Subscribers: hiraditya, llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D68128
      bbb29738
    • Sander de Smalen's avatar
      [AArch64][SVE] Allocate locals that are scalable vectors. · 9a1c243a
      Sander de Smalen authored
      This patch adds a target interface to set the StackID for a given type,
      which allows scalable vectors (e.g. `<vscale x 16 x i8>`) to be assigned a
      'sve-vec' StackID, so it is allocated in the SVE area of the stack frame.
      
      Reviewers: ostannard, efriedma, rengolin, cameron.mcinally
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D70080
      9a1c243a
    • Simon Tatham's avatar
      [ARM,MVE] Use VMOV.{S8,S16} for sign-extended extractelement. · 5b9e4dae
      Simon Tatham authored
      MVE includes instructions that extract an 8- or 16-bit lane from a
      vector and sign-extend it into the output 32-bit GPR. `ARMInstrMVE.td`
      already included isel patterns to select those instructions in
      response to the `ARMISD::VGETLANEs` selection-DAG node type. But
      `ARMISD::VGETLANEs` was never actually generated, because the code
      that creates it was conditioned on NEON only.
      
      It's an easy fix to enable the same code for integer MVE, and now IR
      that sign-extends the result of an extractelement (whether explicitly
      or as part of the function call ABI) will use `vmov.s8` instead of
      `vmov.u8` followed by `sxtb`.
      
      Reviewers: SjoerdMeijer, dmgreen, ostannard
      
      Subscribers: kristof.beyls, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70132
      5b9e4dae
    • David Zarzycki's avatar
      1d55c9e5
    • joanlluch's avatar
      [TargetLowering][DAGCombine][MSP430] Shift Amount Threshold in DAGCombine (4) · d384ad6b
      joanlluch authored
      Summary:
      Replaces
      ```
      unsigned getShiftAmountThreshold(EVT VT)
      ```
      by
      
      ```
      bool shouldAvoidTransformToShift(EVT VT, unsigned amount)
      ```
      thus giving more flexibility for targets to decide whether particular shift amounts must be considered expensive or not.
      
      Updates the MSP430 target with a custom implementation.
      
      This continues  D69116, D69120, D69326 and updates them, so all of them must be committed before this.
      
      Existing tests apply, a few more have been added.
      
      Reviewers: asl, spatel
      
      Reviewed By: spatel
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70042
      d384ad6b
    • Craig Topper's avatar
      [X86] Remove setOperationAction for FP_TO_SINT v8i16. · a4b7613a
      Craig Topper authored
      This is no longer needed after widening legalization as we
      custom legalize v8i8 ourselves.
      
      Added entries to the cost model, but bumped the cost slightly
      to account for the truncate shuffle that wasn't costed before.
      a4b7613a
    • Michael Kruse's avatar
      [GPGPU] Fix regression test after 395124. · 7be6ec5f
      Michael Kruse authored
      Commit 395124 "NVPTX: Don't insert an extra empty line at the end of the last section"
      changed the length of the kernel payload. Update the regression test to the new binary size.
      7be6ec5f
    • Jonas Devlieghere's avatar
      [Reproducer] Discard reproducer directory if not generated. · 7ba28644
      Jonas Devlieghere authored
      If lldb was run in capture mode, but no reproducer was generated, make
      sure we clean up the reproducer directory.
      7ba28644
    • Francesco Petrogalli's avatar
      [VFABI] Add LLVM internal mangling for vector functions. · d8b6b111
      Francesco Petrogalli authored
      Summary:
      This patch adds a custom ISA for vector functions for internal use
      in LLVM. The <isa> token is set to "_LLVM_", and it is not attached
      to any specific instruction Vector ISA, or Vector Function ABI.
      
      The ISA is used as a token for handling Vector Function ABI-style
      vectorization for those vector functions that are not directly
      associated to any existing Vector Function ABI (for example, some of
      the vector functions exposed by TargetLibraryInfo). The demangling
      function for this ISA in a Vector Function ABI context is set to be
      the same as the common one shared between X86 and AArch64.
      
      Reviewers: jdoerfert, sdesmalen, simoll
      
      Subscribers: kristof.beyls, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70089
      d8b6b111
    • Justin Hibbits's avatar
      Add 8548 CPU definition and attributes · bc4bc5aa
      Justin Hibbits authored
      8548 CPU is GCC's name for the e500v2, so accept this in clang.  The
      e500v2 doesn't support lwsync, so define __NO_LWSYNC__ for this as well,
      as GCC does.
      
      Differential Revision:  https://reviews.llvm.org/D67787
      bc4bc5aa
    • Matt Arsenault's avatar
      AMDGPU: Extend add x, (ext setcc) combine to sub · 9d7bccab
      Matt Arsenault authored
      This is the same as the add case, but inverts the operation type.
      
      This avoids regressions in a future patch.
      9d7bccab
    • Matt Arsenault's avatar
      AMDGPU: Switch backend default max workgroup size to 1024 · 4b472139
      Matt Arsenault authored
      Previously this would default to 256, not the maximum supported size
      of 1024. Using a maximum lower than the hardware maximum requires
      language runtimes to enforce this limit for correctness, which no
      language has correctly done. Switch the default to the conservatively
      correct maximum, and force frontends to opt-in to the more optimal 256
      default maximum.
      
      I don't really understand why the changes in occupancy-levels.ll
      increased the computed occupancy, which I expected to decrease. I'm
      not sure if these tests should be forcing the old maximum.
      4b472139
    • Matt Arsenault's avatar
      AMDGPU Reduce reported maximum group size to 1024 · 25c5da5a
      Matt Arsenault authored
      While some targets allow encoding 2048, this was never tested or
      supported.
      25c5da5a
    • Alina Sbirlea's avatar
      [GlobalsAA] Reenable test. · 793b42a4
      Alina Sbirlea authored
      793b42a4
    • Muhammad Omair Javaid's avatar
      [LLDB] Add core definition for armv8l and armv7l · 9b958356
      Muhammad Omair Javaid authored
      
      
      This patch adds core definitions in lldb ArchSpecs for armv8l and armv7l cores.
      
      This was needed because on Linux running on 32-bit Arm v8 we are returned
      armv8l in case we are running 32-bit sysroot on 64bit kernel. In case of 32-bit
      kernel and 32-bit sysroot running on arm v8 hardware we are returned armv7l.
      This is quite common when we run 32 bit arm using docker container.
      
      Signed-off-by: default avatarMuhammad Omair Javaid <omair.javaid@linaro.org>
      
      Differential Revision: https://reviews.llvm.org/D69904
      9b958356
    • Richard Smith's avatar
      Don't assume that the clang binary's resolved name includes the string · 5ad6f279
      Richard Smith authored
      'clang'.
      
      This is not true in practice in some content-addressed file systems.
      5ad6f279
    • Leonard Chan's avatar
      [Sema] Add MacroQualified case for FunctionTypeUnwrapper · e278c138
      Leonard Chan authored
      This is a fix for PR43315. An assertion error is hit for this minimal example:
      
      ```
      //clang -cc1 -triple x86_64-- -S tstVMStructRC-min.cpp
      int (a b)();  // Assertion `Chunk.Kind == DeclaratorChunk::Function' failed.
      ```
      
      This is because we do not cover the case in the FunctionTypeUnwrapper where it
      receives a MacroQualifiedType. We have not run into this earlier because this
      is a unique case where the __attribute__ contains both __cdecl__ and
      __regparm__ (in that order), and we are compiling for x86_64. Changing the
      architecture or the order of __cdecl__ and __regparm__ does not raise the
      assertion.
      
      Differential Revision: https://reviews.llvm.org/D67992
      e278c138
    • Alina Sbirlea's avatar
      Temporarily disable test. · 92611da5
      Alina Sbirlea authored
      92611da5
    • Eric Christopher's avatar
      Temporarily Revert "Reapply [LVI] Normalize pointer behavior" as it's broken python 3.6. · 7a3ad48d
      Eric Christopher authored
      Reverting to figure out if it's a problem in python or the compiler for now.
      
      This reverts commit 885a05f4.
      7a3ad48d
    • Jonas Devlieghere's avatar
    • Jonas Devlieghere's avatar
      34ca6e1f
    • Douglas Yung's avatar
      Add a shim for setenv on PS4 since it does not exist. · 7ebde1bf
      Douglas Yung authored
      A few years back a similar change was made for getenv since neither function is supported on the PS4 platform.
      
      Recently, commit d889d1ef added a call to setenv in compiler-rt which was causing linking errors because the symbol was not found. This fixes that issue by putting in a shim similar to how we previously dealt with the lack of getenv.
      
      Differential Revision: https://reviews.llvm.org/D70033
      7ebde1bf
    • Craig Topper's avatar
      [X86] Don't consider v64i1 as a legal type unless v64i8 is also a legal type. · 3e1aee2b
      Craig Topper authored
      This avoids some nasty issues with argument passing and lowering of
      arbitrary v64i8 shuffles.
      3e1aee2b
    • Craig Topper's avatar
      [X86] Only pass v64i8/v32i16 as v16i32 on non-avx512bw targets if the v16i32... · 0f04ffc0
      Craig Topper authored
      [X86] Only pass v64i8/v32i16 as v16i32 on non-avx512bw targets if the v16i32 type won't be split by prefer-vector-width=256
      
      Otherwise just let the v64i8/v32i16 types be split to v32i8/v16i16.
      
      In reality this shouldn't happen because it means we have a 512-bit
      vector argument, but min-legal-vector-width says a value less than
      512. But a 512-bit argument should have been factored into the
      preferred vector width.
      0f04ffc0
    • Sterling Augustine's avatar
      Fix include guard and properly order __deregister_frame_info. · 38c35617
      Sterling Augustine authored
      Summary:
      This patch fixes two problems with the crtbegin.c as written:
      
      1. In do_init, register_frame_info is not guarded by a #define, but in
      do_fini, deregister_frame_info is guarded by #ifndef
      CRT_HAS_INITFINI_ARRAY. Thus when CRT_HAS_INITFINI_ARRAY is not
      defined, frames are registered but then never deregistered.
      
      The frame registry mechanism builds a linked-list from the .so's
      static variable do_init.object, and when the .so is unloaded, this
      memory becomes invalid and should be deregistered.
      
      Further, libgcc's crtbegin treats the frame registry as independent
      from the initfini array mechanism.
      
      This patch fixes this by adding a new #define,
      "EH_USE_FRAME_INFO_REGISTRY", which is set by the cmake option
      COMPILER_RT_CRT_USE_EH_FRAME_REGISTRY Currently, do_init calls
      register_frame_info, and then calls the binary's constructors. This
      allows constructors to safely use libunwind. However, do_fini calls
      deregister_frame_info and then calls the binary's destructors. This
      prevents destructors from safely using libunwind.
      
      This patch also switches that ordering, so that destructors can safely
      use libunwind. As it happens, this is a fairly common scenario for
      thread sanitizer.
      38c35617
    • Weverything's avatar
      Add -Wtautological-compare to -Wall · 9740f9f0
      Weverything authored
      Some warnings in -Wtautological-compare subgroups are DefaultIgnore.
      Adding this group to -Wmost, which is part of -Wall, will aid in their
      discoverability.
      
      Differential Revision: https://reviews.llvm.org/D69292
      9740f9f0