1. Sep 13, 2020
  2. Sep 12, 2020
    • Simon Pilgrim's avatar
      [InstCombine][X86] Covert masked load/stores with (sign extended) bool vector... · 3170d548
      Simon Pilgrim authored
      [InstCombine][X86] Covert masked load/stores with (sign extended) bool vector masks to generic intrinsics.
      
      As detailed on PR11210, if the mask is known to come from a (sign extended) bool vector (e.g. comparisons) then we can represent with a generic masked load/store without losing anything.
      
      We already do something similar for BLENDV -> SELECT conversion.
      3170d548
    • Florian Hahn's avatar
      [Clang] Add option to allow marking pass-by-value args as noalias. · a874d633
      Florian Hahn authored
      After the recent discussion on cfe-dev 'Can indirect class parameters be
      noalias?' [1], it seems like using using noalias is problematic for
      current C++, but should be allowed for C-only code.
      
      This patch introduces a new option to let the user indicate that it is
      safe to mark indirect class parameters as noalias. Note that this also
      applies to external callers, e.g. it might not be safe to use this flag
      for C functions that are called by C++ functions.
      
      In targets that allocate indirect arguments in the called function, this
      enables more agressive optimizations with respect to memory operations
      and brings a ~1% - 2% codesize reduction for some programs.
      
      [1] : http://lists.llvm.org/pipermail/cfe-dev/2020-July/066353.html
      
      Reviewed By: rjmccall
      
      Differential Revision: https://reviews.llvm.org/D85473
      a874d633
    • Evgeny Leviant's avatar
    • Tyker's avatar
      Reland [AssumeBundles] Use operand bundles to encode alignment assumptions · 78de7297
      Tyker authored
      NOTE: There is a mailing list discussion on this: http://lists.llvm.org/pipermail/llvm-dev/2019-December/137632.html
      
      Complemantary to the assumption outliner prototype in D71692, this patch
      shows how we could simplify the code emitted for an alignemnt
      assumption. The generated code is smaller, less fragile, and it makes it
      easier to recognize the additional use as a "assumption use".
      
      As mentioned in D71692 and on the mailing list, we could adopt this
      scheme, and similar schemes for other patterns, without adopting the
      assumption outlining.
      78de7297
    • Simon Pilgrim's avatar
      [InstCombine][X86] Add tests for masked load/stores with comparisons. · d030aad7
      Simon Pilgrim authored
      As detailed on PR11210, if the mask is known to come from a (sign extended) bool vector (e.g. comparisons) then we can represent with a generic masked load/store without losing anything.
      d030aad7
    • David Green's avatar
      [ARM] Fixup single source mla reductions. · 6cfd38d0
      David Green authored
      This fixes a complication on top of D87276. If we are sign extending
      around a mul with the two operands that are the same, instcombine will
      helpfully convert one of the sext to a zext. Reverse that so that we
      again generate a reduction.
      
      Differnetial Revision: https://reviews.llvm.org/D87287
      6cfd38d0
    • Sanjay Patel's avatar
      [Intrinsics] define semantics for experimental fmax/fmin vector reductions · 3a8ea860
      Sanjay Patel authored
      As discussed on llvm-dev:
      http://lists.llvm.org/pipermail/llvm-dev/2020-April/140729.html
      
      This is hopefully the final remaining showstopper before we can remove
      the 'experimental' from the reduction intrinsics.
      
      No behavior was specified for the FP min/max reductions, so we have a
      mess of different interpretations.
      
      There are a few potential options for the semantics of these max/min ops.
      I think this is the simplest based on current behavior/implementation:
      make the reductions inherit from the existing llvm.maxnum/minnum intrinsics.
      These correspond to libm fmax/fmin, and those are similar to the (now
      deprecated?) IEEE-754 maxNum/minNum functions (NaNs are treated as missing
      data). So the default expansion creates calls to libm functions.
      
      Another option would be to inherit from llvm.maximum/minimum (NaNs propagate),
      but most targets just crash in codegen when given those nodes because no
      default expansion was ever implemented AFAICT.
      
      We could also just assume 'nnan' semantics by default (we are already
      assuming 'nsz' semantics in the maxnum/minnum intrinsics), but some targets
      (AArch64, PowerPC) support the more defined behavior, so it doesn't make much
      sense to not allow a tighter spec. Fast-math-flags (nnan) can be used to
      loosen the semantics.
      
      (Note that D67507 was proposed to update the LangRef to acknowledge the more
      recent IEEE-754 2019 standard, but that patch seems to have stalled. If we do
      update based on the new standard, the reduction instructions can seamlessly
      inherit from whatever updates are made to the max/min intrinsics.)
      
      x86 sees a regression here on 'nnan' tests because we have underlying,
      longstanding bugs in FMF creation/propagation. Those need to be fixed apart
      from this change (for example: https://llvm.org/PR35538). The expansion
      sequence before this patch may not have been correct.
      
      Differential Revision: https://reviews.llvm.org/D87391
      3a8ea860
    • Simon Pilgrim's avatar
      [InstCombine][X86] getNegativeIsTrueBoolVec - use ConstantExpr evaluators. NFCI. · 50ee0b99
      Simon Pilgrim authored
      Don't do this manually, we can just use the ConstantExpr evaluators to do it more tidily for us.
      50ee0b99
    • David Green's avatar
      [ARM] Recognize "double extend" reduction patterns · c437446d
      David Green authored
      We can sometimes get code that does:
        xe = zext i16 x to i32
        ye = zext i16 y to i32
        m = mul i32 xe, ye
        me = zext i32 m to i64
        r = vecreduce.add(me)
      This "double extend" can trip up the reduction identification, but
      should give identical results.
      
      This extends the pattern matching to handle them.
      
      Differential Revision: https://reviews.llvm.org/D87276
      c437446d
    • Nikita Popov's avatar
      [InstCombine] Fix incorrect SimplifyWithOpReplaced transform (PR47322) · 36e2e2e1
      Nikita Popov authored
      This is a followup to D86834, which partially fixed this issue in
      InstSimplify. However, InstCombine repeats the same transform while
      dropping poison flags -- which does not cover cases where poison is
      introduced in some other way.
      
      The fix here is a bit more comprehensive, because things are quite
      entangled, and it's hard to only partially address it without
      regressing optimization. There are really two changes here:
      
       * Export the SimplifyWithOpReplaced API from InstSimplify, with an
         added AllowRefinement flag. For replacements inside the TrueVal
         we don't actually care whether refinement occurs or not, the
         replacement is always legal. This part of the transform is now
         done in InstSimplify only. (It should be noted that the current
         AllowRefinement check is not sufficient -- that's an issue we
         need to address separately.)
       * Change the InstCombine fold to work by temporarily dropping
         poison generating flags, running the fold and then restoring the
         flags if it didn't work out. This will ensure that the InstCombine
         fold is correct as long as the InstSimplify fold is correct.
      
      Differential Revision: https://reviews.llvm.org/D87445
      36e2e2e1
    • Simon Pilgrim's avatar
      [X86][SSE] lowerShuffleAsDecomposedShuffleBlend - support decomposed unpacks... · 35dc91ae
      Simon Pilgrim authored
      [X86][SSE] lowerShuffleAsDecomposedShuffleBlend - support decomposed unpacks for some vXi8/vXi16 cases
      
      Follow up to D86429 to handle the remaining regressions.
      
      This patch generalizes lowerShuffleAsDecomposedShuffleBlend to lowerShuffleAsDecomposedShuffleMerge, and attempts to use an UNPCKL shuffle mask instead of a blend for the cases where the inputs are coming from alternating vXi8/vXi16 sources. Technically they don't have to be alternating (just as long as they can fit into a lower lane half for the unpack) but I didn't find as many general cases and it needed a lot more of the function to be altered.
      
      For vXi32/vXi64 cases this could still be beneficial but in most cases the existing permute+blend approach was better.
      
      Differential Revision: https://reviews.llvm.org/D87405
      35dc91ae