1. Sep 13, 2020
  2. Sep 12, 2020
    • Simon Pilgrim's avatar
      [InstCombine][X86] Covert masked load/stores with (sign extended) bool vector... · 3170d548
      Simon Pilgrim authored
      [InstCombine][X86] Covert masked load/stores with (sign extended) bool vector masks to generic intrinsics.
      
      As detailed on PR11210, if the mask is known to come from a (sign extended) bool vector (e.g. comparisons) then we can represent with a generic masked load/store without losing anything.
      
      We already do something similar for BLENDV -> SELECT conversion.
      3170d548
    • Florian Hahn's avatar
      [Clang] Add option to allow marking pass-by-value args as noalias. · a874d633
      Florian Hahn authored
      After the recent discussion on cfe-dev 'Can indirect class parameters be
      noalias?' [1], it seems like using using noalias is problematic for
      current C++, but should be allowed for C-only code.
      
      This patch introduces a new option to let the user indicate that it is
      safe to mark indirect class parameters as noalias. Note that this also
      applies to external callers, e.g. it might not be safe to use this flag
      for C functions that are called by C++ functions.
      
      In targets that allocate indirect arguments in the called function, this
      enables more agressive optimizations with respect to memory operations
      and brings a ~1% - 2% codesize reduction for some programs.
      
      [1] : http://lists.llvm.org/pipermail/cfe-dev/2020-July/066353.html
      
      Reviewed By: rjmccall
      
      Differential Revision: https://reviews.llvm.org/D85473
      a874d633
    • Evgeny Leviant's avatar
    • Tyker's avatar
      Reland [AssumeBundles] Use operand bundles to encode alignment assumptions · 78de7297
      Tyker authored
      NOTE: There is a mailing list discussion on this: http://lists.llvm.org/pipermail/llvm-dev/2019-December/137632.html
      
      Complemantary to the assumption outliner prototype in D71692, this patch
      shows how we could simplify the code emitted for an alignemnt
      assumption. The generated code is smaller, less fragile, and it makes it
      easier to recognize the additional use as a "assumption use".
      
      As mentioned in D71692 and on the mailing list, we could adopt this
      scheme, and similar schemes for other patterns, without adopting the
      assumption outlining.
      78de7297
    • Simon Pilgrim's avatar
      [InstCombine][X86] Add tests for masked load/stores with comparisons. · d030aad7
      Simon Pilgrim authored
      As detailed on PR11210, if the mask is known to come from a (sign extended) bool vector (e.g. comparisons) then we can represent with a generic masked load/store without losing anything.
      d030aad7
    • David Green's avatar
      [ARM] Fixup single source mla reductions. · 6cfd38d0
      David Green authored
      This fixes a complication on top of D87276. If we are sign extending
      around a mul with the two operands that are the same, instcombine will
      helpfully convert one of the sext to a zext. Reverse that so that we
      again generate a reduction.
      
      Differnetial Revision: https://reviews.llvm.org/D87287
      6cfd38d0
    • Sanjay Patel's avatar
      [Intrinsics] define semantics for experimental fmax/fmin vector reductions · 3a8ea860
      Sanjay Patel authored
      As discussed on llvm-dev:
      http://lists.llvm.org/pipermail/llvm-dev/2020-April/140729.html
      
      This is hopefully the final remaining showstopper before we can remove
      the 'experimental' from the reduction intrinsics.
      
      No behavior was specified for the FP min/max reductions, so we have a
      mess of different interpretations.
      
      There are a few potential options for the semantics of these max/min ops.
      I think this is the simplest based on current behavior/implementation:
      make the reductions inherit from the existing llvm.maxnum/minnum intrinsics.
      These correspond to libm fmax/fmin, and those are similar to the (now
      deprecated?) IEEE-754 maxNum/minNum functions (NaNs are treated as missing
      data). So the default expansion creates calls to libm functions.
      
      Another option would be to inherit from llvm.maximum/minimum (NaNs propagate),
      but most targets just crash in codegen when given those nodes because no
      default expansion was ever implemented AFAICT.
      
      We could also just assume 'nnan' semantics by default (we are already
      assuming 'nsz' semantics in the maxnum/minnum intrinsics), but some targets
      (AArch64, PowerPC) support the more defined behavior, so it doesn't make much
      sense to not allow a tighter spec. Fast-math-flags (nnan) can be used to
      loosen the semantics.
      
      (Note that D67507 was proposed to update the LangRef to acknowledge the more
      recent IEEE-754 2019 standard, but that patch seems to have stalled. If we do
      update based on the new standard, the reduction instructions can seamlessly
      inherit from whatever updates are made to the max/min intrinsics.)
      
      x86 sees a regression here on 'nnan' tests because we have underlying,
      longstanding bugs in FMF creation/propagation. Those need to be fixed apart
      from this change (for example: https://llvm.org/PR35538). The expansion
      sequence before this patch may not have been correct.
      
      Differential Revision: https://reviews.llvm.org/D87391
      3a8ea860