1. May 06, 2020
    • Sanjay Patel's avatar
      [SLP] add another bailout for load-combine patterns · 86dfbc67
      Sanjay Patel authored
      This builds on the or-reduction bailout that was added with D67841.
      We still do not have IR-level load combining, although that could
      be a target-specific enhancement for -vector-combiner.
      
      The heuristic is narrowly defined to catch the motivating case from
      PR39538:
      https://bugs.llvm.org/show_bug.cgi?id=39538
      ...while preserving existing functionality.
      
      That is, there's an unmodified test of pure load/zext/store that is
      not seen in this patch at llvm/test/Transforms/SLPVectorizer/X86/cast.ll.
      That's the reason for the logic difference to require the 'or'
      instructions. The chances that vectorization would actually help a
      memory-bound sequence like that seem small, but it looks nicer with:
      
        vpmovzxwd	(%rsi), %xmm0
        vmovdqu	%xmm0, (%rdi)
      
      rather than:
      
        movzwl	(%rsi), %eax
        movl	%eax, (%rdi)
        ...
      
      In the motivating test, we avoid creating a vector mess that is
      unrecoverable in the backend, and SDAG forms the expected bswap
      instructions after load combining:
      
        movzbl (%rdi), %eax
        vmovd %eax, %xmm0
        movzbl 1(%rdi), %eax
        vmovd %eax, %xmm1
        movzbl 2(%rdi), %eax
        vpinsrb $4, 4(%rdi), %xmm0, %xmm0
        vpinsrb $8, 8(%rdi), %xmm0, %xmm0
        vpinsrb $12, 12(%rdi), %xmm0, %xmm0
        vmovd %eax, %xmm2
        movzbl 3(%rdi), %eax
        vpinsrb $1, 5(%rdi), %xmm1, %xmm1
        vpinsrb $2, 9(%rdi), %xmm1, %xmm1
        vpinsrb $3, 13(%rdi), %xmm1, %xmm1
        vpslld $24, %xmm0, %xmm0
        vpmovzxbd %xmm1, %xmm1 # xmm1 = xmm1[0],zero,zero,zero,xmm1[1],zero,zero,zero,xmm1[2],zero,zero,zero,xmm1[3],zero,zero,zero
        vpslld $16, %xmm1, %xmm1
        vpor %xmm0, %xmm1, %xmm0
        vpinsrb $1, 6(%rdi), %xmm2, %xmm1
        vmovd %eax, %xmm2
        vpinsrb $2, 10(%rdi), %xmm1, %xmm1
        vpinsrb $3, 14(%rdi), %xmm1, %xmm1
        vpinsrb $1, 7(%rdi), %xmm2, %xmm2
        vpinsrb $2, 11(%rdi), %xmm2, %xmm2
        vpmovzxbd %xmm1, %xmm1 # xmm1 = xmm1[0],zero,zero,zero,xmm1[1],zero,zero,zero,xmm1[2],zero,zero,zero,xmm1[3],zero,zero,zero
        vpinsrb $3, 15(%rdi), %xmm2, %xmm2
        vpslld $8, %xmm1, %xmm1
        vpmovzxbd %xmm2, %xmm2 # xmm2 = xmm2[0],zero,zero,zero,xmm2[1],zero,zero,zero,xmm2[2],zero,zero,zero,xmm2[3],zero,zero,zero
        vpor %xmm2, %xmm1, %xmm1
        vpor %xmm1, %xmm0, %xmm0
        vmovdqu %xmm0, (%rsi)
      
        movl	(%rdi), %eax
        movl	4(%rdi), %ecx
        movl	8(%rdi), %edx
        movbel	%eax, (%rsi)
        movbel	%ecx, 4(%rsi)
        movl	12(%rdi), %ecx
        movbel	%edx, 8(%rsi)
        movbel	%ecx, 12(%rsi)
      
      Differential Revision: https://reviews.llvm.org/D78997
      86dfbc67
    • Pete Steinfeld's avatar
      [flang] New implementation for checks for constraints C741 through C750 · 8d0c3c05
      Pete Steinfeld authored
      Summary:
      Most of these checks were already implemented, and I just added references to
      them to the code and tests. Also, much of this code was already
      reviewed in the old flang/f18 GitHub repository, but I didn't get to
      merge it before we switched repositories.
      
      I implemented the check for C747 to not allow coarray components in derived
      types that are of type C_PTR, C_FUNPTR, or type TEAM_TYPE.
      
      I implemented the check for C748 that requires a data component whose type has
      a coarray ultimate component to be a nonpointer, nonallocatable scalar and not
      be a coarray.
      
      I implemented the check for C750 that adds additional restrictions to the
      bounds expressions of a derived type component that's an array.
      These bounds expressions are sepcification expressions as defined in
      10.1.11.  There was already code in lib/Evaluate/check-expression.cpp to
      check semantics for specification expressions, but it did not check for
      the extra requirements of C750.
      
      C750 prohibits specification functions, the intrinsic functions
      ALLOCATED, ASSOCIATED, EXTENDS_TYPE_OF, PRESENT, and SAME_TYPE_AS. It
      also requires every specification inquiry reference to be a constant
      expression, and requires that the value of the bound not depend on the
      value of a variable.
      
      To implement these additional checks, I added code to the intrinsic proc
      table to get the intrinsic class of a procedure.  I also added an
      enumeration to distinguish between specification expressions for
      derived type component bounds versus for type parameters.  I then
      changed the code to pass an enumeration value to
      "CheckSpecificationExpr()" to indicate that the expression was a bounds
      expression and used this value to determine whether to emit an error
      message when violations of C750 are found.
      
      I changed the implementation of IsPureProcedure() to handle statement
      functions and changed some references in the code that tested for the
      PURE attribute to call IsPureProcedure().
      
      I also fixed some unrelated tests that got new errors when I implemented these
      new checks.
      
      Reviewers: tskeith, DavidTruby, sscalpone
      
      Subscribers: jfb, llvm-commits
      
      Tags: #llvm, #flang
      
      Differential Revision: https://reviews.llvm.org/D79263
      8d0c3c05
    • Francesco Petrogalli's avatar
      [clang][OpenMP] Fix getNDSWDS for aarch64. · 4fa13a3d
      Francesco Petrogalli authored
      Summary:
      This change fixes an aarch64-specific bug in the generation of the NDS and WDS values used to compute the signature of the vector functions out of OpenMP directives like `declare simd`. When the directive is used in conjunction with the `linear` clause, the size of the pointee must be used instead of the size of the pointer to compute NDS and WDS.
      
      The code-fix is strictly related to the behavior for `linear`, but given that the only way we have to test the NDS and WDS values is to check the resulting `<vlen>` token in the mangled name of the vector function, the tests have been extended to cover all the possible values of WDS and NDS as defined in the ABI at https://github.com/ARM-software/abi-aa/tree/master/vfabia64.
      
      Reviewers: ABataev, jdoerfert, andwar
      
      Reviewed By: jdoerfert
      
      Subscribers: yaxunl, kristof.beyls, guansong, danielkiss, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D78969
      4fa13a3d
    • Stephen Neuendorffer's avatar
      [MLIR] Add a tests for out of tree dialect example. · 175a3df9
      Stephen Neuendorffer authored
      This attempts to ensure that out of tree usage remains stable.
      
      Differential Revision: https://reviews.llvm.org/D78656
      175a3df9
    • Vedant Kumar's avatar
      c05f3544
    • Jinsong Ji's avatar
      [MachinePipeliner] Add ORE for MachinePipeliner · 80b78a47
      Jinsong Ji authored
      This patch adds ORE for MachinePipeliner, so that people can anaylyze
      their code using opt-viewer or other tools, then optimize the code to
      catch more piplining opportunities.
      
      Reviewed By: bcahoon
      
      Differential Revision: https://reviews.llvm.org/D79368
      80b78a47
  2. May 05, 2020