1. Feb 24, 2021
  2. Feb 23, 2021
    • Nicolai Hähnle's avatar
      [AMDGPU][SelectionDAG] Don't combine uniform multiplies to MUL_[UI]24 · 52bc2e75
      Nicolai Hähnle authored
      Prefer to keep uniform (non-divergent) multiplies on the scalar ALU when
      possible. This significantly improves some game cases by eliminating
      v_readfirstlane instructions when the result feeds into a scalar
      operation, like the address calculation for a scalar load or store.
      
      Since isDivergent is only an approximation of whether a value is in
      SGPRs, it can potentially regress some situations where a uniform value
      ends up in a VGPR. These should be rare in real code, although the test
      changes do contain a number of examples.
      
      Most of the test changes are just using s_mul instead of v_mul/mad which
      is generally better for both register pressure and latency (at least on
      GFX10 where sgpr pressure doesn't affect occupancy and vector ALU
      instructions have significantly longer latency than scalar ALU). Some
      R600 tests now use MULLO_INT instead of MUL_UINT24.
      
      GlobalISel appears to handle more scenarios in the desirable way,
      although it can also be thrown off and fails to select the 24-bit
      multiplies in some cases.
      
      Alternative solution considered and rejected was to allow selecting
      MUL_[UI]24 to S_MUL_I32. I've rejected this because the definition of
      those SD operations works is don't-care on the most significant 8 bits,
      and this fact is used in some combines via SimplifyDemandedBits.
      
      Based on a patch by Nicolai Hähnle.
      
      Differential Revision: https://reviews.llvm.org/D97063
      52bc2e75
    • Juneyoung Lee's avatar
      [JumpThreading] Update computeValueKnownInPredecessors to recognize logical and/or patterns · 19c2e129
      Juneyoung Lee authored
      This allows JumpThreading's computeValueKnownInPredecessors to
      recognize select form of and/or patterns as well.
      19c2e129
    • Jay Foad's avatar
      [AMDGPU] Rename a prefix for sanity. NFC. · 64831fb0
      Jay Foad authored
      64831fb0
    • Nate Chandler's avatar
      Add @llvm.coro.async.size.replace intrinsic. · 01b4890e
      Nate Chandler authored
      The new intrinsic replaces the size in one specified AsyncFunctionPointer with
      the size in another.  This ability is necessary for functions which merely
      forward to async functions such as those defined for partial applications.
      
      Reviewed By: aschwaighofer
      
      Differential Revision: https://reviews.llvm.org/D97229
      01b4890e
    • Jessica Clarke's avatar
      22215e49
    • Martin Storsjö's avatar
      [libcxx] [test] Define _CRT_STDIO_ISO_WIDE_SPECIFIERS while building tests · f97ea0d5
      Martin Storsjö authored
      This matches how libc++ itself is built. This avoids errors due to
      mismatch if linking libc++ statically.
      
      Differential Revision: https://reviews.llvm.org/D97169
      f97ea0d5
    • Nathan James's avatar
      [clang-tidy] Remove IncludeInserter from MoveConstructorInit check. · e96f9cca
      Nathan James authored
      This check registers an IncludeInserter, however the check itself doesn't actually emit any fixes or includes, so the inserter is redundant.
      
      From what I can tell the fixes were removed in D26453(rL290051) but the inserter was left in, probably an oversight.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D97243
      e96f9cca
    • Joe Ellis's avatar
      [clang][SVE] Don't warn on vector to sizeless builtin implicit conversion · 1b1b30cf
      Joe Ellis authored
      This commit prevents warnings from -Wconversion when a clang vector type
      is implicitly converted to a sizeless builtin type -- for example, when
      implicitly converting a fixed-predicate to a scalable predicate.
      
      The code below:
      
           1    #include <arm_sve.h>
           2
           3    #define N __ARM_FEATURE_SVE_BITS
           4    #define FIXED_ATTR __attribute__((arm_sve_vector_bits (N)))
           5    typedef svbool_t fixed_svbool_t FIXED_ATTR;
           6
           7    inline fixed_svbool_t foo(fixed_svbool_t p) {
           8      return svnot_z(svptrue_b64(), p);
           9    }
      
      would previously raise this warning:
      
          warning: implicit conversion turns vector to scalar: \
          'fixed_svbool_t' (vector of 8 'unsigned char' values) to 'svbool_t' \
          (aka '__SVBool_t') [-Wconversion]
      
      Note that many cases of these implicit conversions were already
      permitted because many functions inside arm_sve.h are spawned via
      preprocessor macros, and the call to isInSystemMacro would cover us in
      this case. This commit fixes the remaining cases.
      
      Differential Revision: https://reviews.llvm.org/D97053
      1b1b30cf
    • Balázs Kéri's avatar
      [clang-tidy] Extending bugprone-signal-handler with POSIX functions. · 2c54b293
      Balázs Kéri authored
      An option is added to the check to select wich set of functions is
      defined as asynchronous-safe functions.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D90851
      2c54b293
    • Michał Górny's avatar
    • Michał Górny's avatar
    • Simon Pilgrim's avatar
      [X86] Cleanup overflow test check prefixes. NFCI. · 2315410f
      Simon Pilgrim authored
      Tidy up the check prefixes to improve reuse.
      2315410f
    • Jay Foad's avatar
      [AMDGPU] Use divergent addresses for vector loads · fdaa2d02
      Jay Foad authored
      Change some test cases to use divergent addresses for vector loads,
      which should be the common case in real world code. Using uniform
      addresses causes poor instruction selection for the surrounding
      code which has to be fixed up post-register-allocation, and this causes
      a lot of testsuite churn for a forthcoming patch to stop selecting
      24-bit vector multiply instructions for uniform multiplies.
      
      This shows up some problems in the idot tests where we fail to select
      v_dot instructions because the patterns only match MUL_[UI]24 ISD nodes,
      but the DAG contains i16 mul nodes instead.
      
      Differential Revision: https://reviews.llvm.org/D97062
      fdaa2d02
    • Sjoerd Meijer's avatar
      [ARM] do not consider sp as deprecated for ldm/stm · e1c3bf6a
      Sjoerd Meijer authored
      Early versions of the ARMv7 reference manuals considered the sp register
      as a deprecated register for ldm/stm familiy of instructions. However,
      later versions such as ARM DDI 0406C.d added a note to the Appendix:
      
      D9.3 Use of the SP as a general-purpose register
      Most ARM instructions, unlike Thumb instructions, provide exactly the
      same access to the SP as to R0-R12. This means that it is possible to
      use the SP as a general-purpose register.  Earlier issues of this manual
      deprecated the use of SP in an ARM instruction, in any way that is
      deprecated, not permitted, or not possible in the corresponding
      Thumb instruction. However, user feedback indicates a number of cases
      where these instructions are useful. Therefore, ARM no longer deprecates
      these instruction uses.
      Also Armv8 manuals no longer consider SP as deprecated register for ldm/
      stm A32 instructions.
      
      Furthermore, GNU as also does not print a deprecated warning when using
      SP with those instructions.
      
      Drop deprecation warning for pop/ldm/push/stm instructions.
      
      Patch by: Stefan Agner.
      
      Differential Revision: https://reviews.llvm.org/D82692
      e1c3bf6a
    • David Green's avatar
      [TTI] Change getOperandsScalarizationOverhead to take Type args · dd2dbf7e
      David Green authored
      As a followup to D95291, getOperandsScalarizationOverhead was still
      using a VF as a vector factor if the arguments were scalar, and would
      assert on certain matrix intrinsics with differently sized vector
      arguments. This patch removes the VF arg, instead passing the Types
      through directly. This should allow it to more accurately compute the
      cost without having to guess at which operands will be vectorized,
      something difficult with more complex intrinsics.
      
      This adjusts one SVE test as it is now calling the wrong intrinsic vs
      veccall. Without invalid InstructCosts the cost of the scalarized
      intrinsic is too low. This should get fixed when the cost of
      scalarization is accounted for with scalable types.
      
      Differential Revision: https://reviews.llvm.org/D96287
      dd2dbf7e
    • David Green's avatar
      [CostModel] Remove VF from IntrinsicCostAttributes · bd4b61ef
      David Green authored
      getIntrinsicInstrCost takes a IntrinsicCostAttributes holding various
      parameters of the intrinsic being costed. It can either be called with a
      scalar intrinsic (RetTy==Scalar, VF==1), with a vector instruction
      (RetTy==Vector, VF==1) or from the vectorizer with a scalar type and
      vector width (RetTy==Scalar, VF>1). A RetTy==Vector, VF>1 is considered
      an error. Both of the vector modes are expected to be treated the same,
      but because this is confusing many backends end up getting it wrong.
      
      Instead of trying work with those two values separately this removes the
      VF parameter, widening the RetTy/ArgTys by VF used called from the
      vectorizer. This keeps things simpler, but does require some other
      modifications to keep things consistent.
      
      Most backends look like this will be an improvement (or were not using
      getIntrinsicInstrCost). AMDGPU needed the most changes to keep the code
      from c230965c working. ARM removed the fix in
      dfac521d, webassembly happens to get a fixup for an SLP cost
      issue and both X86 and AArch64 seem to now be using better costs from
      the vectorizer.
      
      Differential Revision: https://reviews.llvm.org/D95291
      bd4b61ef