1. Jan 21, 2020
  2. Jan 20, 2020
    • Sid Manning's avatar
      Add support for Linux/Musl ABI · 7fee4fed
      Sid Manning authored
      Differential revision: https://reviews.llvm.org/D72701
      
      The patch adds a new option ABI for Hexagon. It primary deals with
      the way variable arguments are passed and is use in the Hexagon Linux Musl
      environment.
      
      If a callee function has a variable argument list, it must perform the
      following operations to set up its function prologue:
      
        1. Determine the number of registers which could have been used for passing
           unnamed arguments. This can be calculated by counting the number of
           registers used for passing named arguments. For example, if the callee
           function is as follows:
      
               int foo(int a, ...){ ... }
      
           ... then register R0 is used to access the argument ' a '. The registers
           available for passing unnamed arguments are R1, R2, R3, R4, and R5.
      
        2. Determine the number and size of the named arguments on the stack.
      
        3. If the callee has named arguments on the stack, it should copy all of these
           arguments to a location below the current position on the stack, and the
           difference should be the size of the register-saved area plus padding
           (if any is necessary).
      
           The register-saved area constitutes all the registers that could have
           been used to pass unnamed arguments. If the number of registers forming
           the register-saved area is odd, it requires 4 bytes of padding; if the
           number is even, no padding is required. This is done to ensure an 8-byte
           alignment on the stack.  For example, if the callee is as follows:
      
             int foo(int a, ...){ ... }
      
           ... then the named arguments should be copied to the following location:
      
             current_position - 5 (for R1-R5) * 4 (bytes) - 4 (bytes of padding)
      
           If the callee is as follows:
      
              int foo(int a, int b, ...){ ... }
      
           ... then the named arguments should be copied to the following location:
      
              current_position - 4 (for R2-R5) * 4 (bytes) - 0 (bytes of padding)
      
        4. After any named arguments have been copied, copy all the registers that
           could have been used to pass unnamed arguments on the stack. If the number
           of registers is odd, leave 4 bytes of padding and then start copying them
           on the stack; if the number is even, no padding is required. This
           constitutes the register-saved area. If padding is required, ensure
           that the start location of padding is 8-byte aligned.  If no padding is
           required, ensure that the start location of the on-stack copy of the
           first register which might have a variable argument is 8-byte aligned.
      
        5. Decrement the stack pointer by the size of register saved area plus the
           padding.  For example, if the callee is as follows:
      
              int foo(int a, ...){ ... } ;
      
           ... then the decrement value should be the following:
      
              5 (for R1-R5) * 4 (bytes) + 4 (bytes of padding) = 24 bytes
      
           The decrement should be performed before the allocframe instruction.
           Increment the stack-pointer back by the same amount before returning
           from the function.
      7fee4fed
    • Thomas Preud'homme's avatar
      [FileCheck] Clean and improve unit tests · abd0ab38
      Thomas Preud'homme authored
      Summary:
      Clean redundant unit test checks (codepath already tested elsewhere) and
      add a few missing checks for existing numeric substitution and match
      logic.
      
      Reviewers: jhenderson, jdenny, probinson, grimar, arichardson, rnk
      
      Reviewed By: jhenderson
      
      Subscribers: llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72912
      abd0ab38
    • Sanjay Patel's avatar
      [InstCombine] form copysign from select of FP constants (PR44153) · 7bee9441
      Sanjay Patel authored
      This should be the last step needed to solve the problem in the
      description of PR44153:
      https://bugs.llvm.org/show_bug.cgi?id=44153
      
      If we're casting an FP value to int, testing its signbit, and then
      choosing between a value and its negated value, that's a
      complicated way of saying "copysign":
      
      (bitcast X) <  0 ? -TC :  TC --> copysign(TC,  X)
      
      Differential Revision: https://reviews.llvm.org/D72643
      7bee9441
    • LLVM GN Syncbot's avatar
      [gn build] Port 24b7b99b · 9ecfaad7
      LLVM GN Syncbot authored
      9ecfaad7
    • Miloš Stojanović's avatar
      [llvm-exegesis][NFC] Disassociate snippet generators from benchmark runners · 24b7b99b
      Miloš Stojanović authored
      The addition of `inverse_throughput` mode highlighted the disjointedness
      of snippet generators and benchmark runners because it used the
      `UopsSnippetGenerator` with the  `LatencyBenchmarkRunner`.
      To keep the code consistent tie the snippet generators to
      parallelization/serialization rather than their benchmark runners.
      
      Renaming `LatencySnippetGenerator` -> `SerialSnippetGenerator`.
      Renaming `UopsSnippetGenerator` -> `ParallelSnippetGenerator`.
      
      Differential Revision: https://reviews.llvm.org/D72928
      24b7b99b
    • Eric Astor's avatar
      Fix build - removing legacy target reference. · 6ccebe00
      Eric Astor authored
      6ccebe00
    • Jon Chesterfield's avatar
      [libomptarget] Implement smid for amdgcn · 03c2a59c
      Jon Chesterfield authored
      Summary:
      [libomptarget] Implement smid for amdgcn
      
      Implementation is in a new file as it uses an intrinsic with
      complicated encoding that warranted substantial comments.
      
      Reviewers: jdoerfert, grokos, ABataev, ronlieb
      
      Reviewed By: jdoerfert
      
      Subscribers: jvesely, mgorny, openmp-commits
      
      Tags: #openmp
      
      Differential Revision: https://reviews.llvm.org/D72956
      03c2a59c
    • Guillaume Chatelet's avatar
      [Alignment][NFC] Use Align with CreateElementUnorderedAtomicMemCpy · 46b9563c
      Guillaume Chatelet authored
      Summary:
      This is patch is part of a series to introduce an Alignment type.
      See this thread for context: http://lists.llvm.org/pipermail/llvm-dev/2019-July/133851.html
      See this patch for the introduction of the type: https://reviews.llvm.org/D64790
      
      Reviewers: courbet, nicolasvasilache
      
      Subscribers: hiraditya, jfb, mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, csigg, arpith-jacob, mgester, lucyrfox, herhut, liufengdb, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73041
      46b9563c
    • Mark Murray's avatar
      [ARM][MVE][Intrinsics] Take abs() of VMINNMAQ, VMAXNMAQ intrinsics' first arguments. · b10a0eb0
      Mark Murray authored
      Summary: Fix VMINNMAQ, VMAXNMAQ intrinsics; BOTH arguments have the absolute values taken.
      
      Reviewers: dmgreen, simon_tatham
      
      Subscribers: kristof.beyls, hiraditya, cfe-commits, llvm-commits
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D72830
      b10a0eb0
    • Eric Astor's avatar
      [ms] [llvm-ml] Add placeholder for llvm-ml, based on llvm-mc · 5f6dfa80
      Eric Astor authored
      As discussed on the mailing list, I plan to introduce an ml-compatible MASM assembler as part of providing more of the Windows build tools. This will be similar to llvm-mc, but with different command-line parameters.
      
      This placeholder is purely a stripped-down version of llvm-mc; we'll eventually add support for the Microsoft-style command-line flags, and back it with a MASM parser.
      
      Relanding this revision after fixing ARM-compatibility issues.
      
      Reviewers: rnk, thakis, RKSimon
      
      Reviewed By: thakis, RKSimon
      
      Differential Revision: https://reviews.llvm.org/D72679
      5f6dfa80
    • Raphael Isemann's avatar
    • Sanjay Patel's avatar
      [InstSimplify] fold select of vector constants that include undef elements · da9c93f3
      Sanjay Patel authored
      As mentioned in D72643, we'd like to be able to assert that any select
      of equivalent constants has been removed before we're deep into InstCombine.
      
      But there's a loophole in that assertion for vectors with undef elements
      that don't match exactly.
      
      This patch should close that gap. If we have undefs, we can't safely
      propagate those unless both constants elements for that lane are undef.
      
      Differential Revision: https://reviews.llvm.org/D72958
      da9c93f3
    • dfukalov's avatar
      [SCEV] Swap guards estimation sequence. NFC · de34b54e
      dfukalov authored
      Summary:
      Loop unroll spends a lot of time in SCEVs processing in case when a function
      contains hundreds of simple 'for' loops with a quite complex arrays indexes like
      
        for (int i = 0; i < 8; ++i) {
          for (int j = 0; j < 32; ++j) {
            C[j*8+i] = B[j*32+i+128] + A[i*64+128];
          }
        }
        for (int i = 0; i < 8; ++i) {
          for (int j = 0; j < 8; ++j) {
            for (int k = 0; k < 32; ++k) {
              D[k*64+i*8+j] = D[k*64+i*8+j] + E[i+16] * C[k*8+j+256];
            }
          }
        }
      
      The patch improves loop unroll speed since isLoopBackedgeGuardedByCond takes
      much less time than isLoopEntryGuardedByCond in the edge case.
      
      Reviewers: skatkov, sanjoy, mkazantsev
      
      Reviewed By: sanjoy
      
      Subscribers: fhahn, hiraditya, javed.absar, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72929
      de34b54e