1. Apr 29, 2019
  2. Apr 28, 2019
    • Eric Fiselier's avatar
      Attempt to switch to auto-scaling bots · 2f5f9a15
      Eric Fiselier authored
      llvm-svn: 359403
      2f5f9a15
    • Nikita Popov's avatar
      [ConstantRange] Add makeExactNoWrapRegion() · 7a94795b
      Nikita Popov authored
      I got confused on the terminology, and the change in D60598 was not
      correct. I was thinking of "exact" in terms of the result being
      non-approximate. However, the relevant distinction here is whether
      the result is
      
       * Largest range such that:
         Forall Y in Other: Forall X in Result: X BinOp Y does not wrap.
         (makeGuaranteedNoWrapRegion)
       * Smallest range such that:
         Forall Y in Other: Forall X not in Result: X BinOp Y wraps.
         (A hypothetical makeAllowedNoWrapRegion)
       * Both. (makeExactNoWrapRegion)
      
      I'm adding a separate makeExactNoWrapRegion method accepting a
      single APInt (same as makeExactICmpRegion) and using it in the
      places where the guarantee is relevant.
      
      Differential Revision: https://reviews.llvm.org/D60960
      
      llvm-svn: 359402
      7a94795b
    • Simon Pilgrim's avatar
      [X86][AVX] Enabled AVX512F tests and add PR40815 test case · d3941952
      Simon Pilgrim authored
      llvm-svn: 359401
      d3941952
    • Simon Pilgrim's avatar
      [X86][AVX] Combine non-lane crossing binary shuffles using X86ISD::VPERMV3 · 22d1476b
      Simon Pilgrim authored
      Some of the combines might be further improved if we lower more shuffles with X86ISD::VPERMV3 directly, instead of waiting to combine the results.
      
      llvm-svn: 359400
      22d1476b
    • Sanjay Patel's avatar
      [SelectionDAG] include FP min/max variants as binary operators · ce8cfe96
      Sanjay Patel authored
      The x86 test diffs don't look great because of extra move ops,
      but FP min/max should clearly be included in the list.
      
      llvm-svn: 359399
      ce8cfe96
    • Sanjay Patel's avatar
      [DAGCombiner] try repeated fdiv divisor transform before building estimate · fb9a5307
      Sanjay Patel authored
      This was originally part of D61028, but it's an independent diff.
      
      If we try the repeated divisor reciprocal transform before producing an estimate sequence,
      then we have an opportunity to use scalar fdiv. On x86, the trade-off is 1 divss vs. 5
      vector FP ops in the default estimate sequence. On recent chips (Skylake, Ryzen), the
      full-precision division is only 3 cycle throughput, so that's probably the better perf
      default option and avoids problems from x86's inaccurate estimates.
      
      The last 2 tests show that users still have the option to override the defaults by using
      the function attributes for reciprocal estimates, but those patterns are potentially made
      faster by converting the vector ops (including ymm ops) to scalar math.
      
      Differential Revision: https://reviews.llvm.org/D61149
      
      llvm-svn: 359398
      fb9a5307
    • Andrea Di Biagio's avatar
      [MCA] Fix typo in AVX2 gather tests. NFC · 43003f0f
      Andrea Di Biagio authored
      llvm-svn: 359397
      43003f0f
    • Simon Pilgrim's avatar
      [X86][SSE] Optimize llvm.experimental.vector.reduce.xor.vXi1 parity reduction (PR38840) · 93ad4821
      Simon Pilgrim authored
      An xor reduction of a bool vector can be optimized to a parity check of the MOVMSK/BITCAST'd integer - if the population count is odd return 1, else return 0.
      
      Differential Revision: https://reviews.llvm.org/D61230
      
      llvm-svn: 359396
      93ad4821
    • Simon Pilgrim's avatar
      fed302ae
    • Dan Liew's avatar
      [CMake] Don't modify `FUZZER_SUPPORTED_ARCH` is place. · 8651edf8
      Dan Liew authored
      On a Darwin host we were modifying the `FUZZER_SUPPORTED_ARCH` in place
      which would strip out non-x86 architectures. This unhelpful if we
      want to use `FUZZER_SUPPORTED_ARCH` later.
      
      To fix this we introduce `FUZZER_TEST_ARCH` which is similar to what we
      have for for the other sanitizers. For non-Darwin host platforms
      `FUZZER_TEST_ARCH` is the same as `FUZZER_SUPPORTED_ARCH` but for Darwin
      host platforms we use `darwin_filter_host_archs(...)` as the previous
      code did.
      
      llvm-svn: 359394
      8651edf8
    • Qiu Chaofan's avatar
      [PowerPC][Clang] Add tests for PowerPC MMX intrinsics · 8eeb3349
      Qiu Chaofan authored
      Add the rest of test cases covering functions defined in mmintrin.h on PowerPC.
      
      Reviewed By: Jinsong Ji
      
      llvm-svn: 359393
      8eeb3349
    • Craig Topper's avatar
      [X86] Remove (V)MOV64toSDrr/m and (V)MOVDI2SSrr/m. Use 128-bit result... · bd35a309
      Craig Topper authored
      [X86] Remove (V)MOV64toSDrr/m and (V)MOVDI2SSrr/m. Use 128-bit result MOVD/MOVQ and COPY_TO_REGCLASS instead
      
      Summary:
      The register form of these instructions are CodeGenOnly instructions that cover
      GR32->FR32 and GR64->FR64 bitcasts. There is a similar set of instructions for
      the opposite bitcast. Due to the patterns using bitcasts these instructions get
      marked as "bitcast" machine instructions as well. The peephole pass is able to
      look through these as well as other copies to try to avoid register bank copies.
      
      Because FR32/FR64/VR128 are all coalescable to each other we can end up in a
      situation where a GR32->FR32->VR128->FR64->GR64 sequence can be reduced to
      GR32->GR64 which the copyPhysReg code can't handle.
      
      To prevent this, this patch removes one set of the 'bitcast' instructions. So
      now we can only go GR32->VR128->FR32 or GR64->VR128->FR64. The instruction that
      converts from GR32/GR64->VR128 has no special significance to the peephole pass
      and won't be looked through.
      
      I guess the other option would be to add support to copyPhysReg to just promote
      the GR32->GR64 to a GR64->GR64 copy. The upper bits were basically undefined
      anyway. But removing the CodeGenOnly instruction in favor of one that won't be
      optimized seemed safer.
      
      I deleted the peephole test because it couldn't be made to work with the bitcast
      instructions removed.
      
      The load version of the instructions were unnecessary as the pattern that selects
      them contains a bitcasted load which should never happen.
      
      Fixes PR41619.
      
      Reviewers: RKSimon, spatel
      
      Reviewed By: RKSimon
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D61223
      
      llvm-svn: 359392
      bd35a309
    • Simon Pilgrim's avatar
      Revert rL359389: [X86][SSE] Add support for <64 x i1> bool reduction · 03c4e266
      Simon Pilgrim authored
      Minor generalization of the existing <32 x i1> pre-AVX2 split code.
      ........
      Causing irregular buildbot failures.
      
      llvm-svn: 359391
      03c4e266
    • Simon Pilgrim's avatar
      1a4a4325
    • Simon Pilgrim's avatar
      [X86][SSE] Add support for <64 x i1> bool reduction · 4118be3a
      Simon Pilgrim authored
      Minor generalization of the existing <32 x i1> pre-AVX2 split code.
      
      llvm-svn: 359389
      4118be3a
    • Simon Pilgrim's avatar
      [X86][AVX] Cleanup and add additional expandload and compressstore tests · 399746ea
      Simon Pilgrim authored
      sort order by types and add vXi32/vXi16/vXi8 test coverage
      
      llvm-svn: 359388
      399746ea
    • Raphael Isemann's avatar
      Fix UNPREDICTABLE check in EmulateInstructionARM::EmulateADDRegShift · e2849a03
      Raphael Isemann authored
      Summary:
      As reported in LLVM bug 41487, the check in this function is wrong and should be
      the same as the described check in the comment (which is correctly copied from the
      ARM ISA reference).
      
      Reviewers: #lldb, davide, JDevlieghere
      
      Reviewed By: #lldb, davide, JDevlieghere
      
      Subscribers: davide, javed.absar, kristof.beyls, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D60654
      
      llvm-svn: 359387
      e2849a03
    • Simon Pilgrim's avatar
      [X86][AVX512] Improve vector bool reductions · 2a2d4224
      Simon Pilgrim authored
      As predicate masks are legal on AVX512 targets, we avoid MOVMSK in these cases, but we can just bitcast the bool vector to the integer equivalent directly - avoiding expansion of the reduction to a shuffle pattern.
      
      llvm-svn: 359386
      2a2d4224
    • Simon Pilgrim's avatar
      [X86] Add vector boolean reduction tests (PR38840) · 913bfd33
      Simon Pilgrim authored
      AND/OR/XOR tests for the @llvm.experimental.vector.reduce intrinsics
      
      AND/OR are pretty good (pre-AVX512), XOR (not so common but used for parity reduction) is still pretty bad.
      
      llvm-svn: 359385
      913bfd33