1. Feb 20, 2021
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Run redundant_sext_inreg in the post-legalizer combiner · 8d3442ed
      Jessica Paquette authored
      This is to ensure that we can eliminate G_ASSERT_SEXT.
      
      In a follow-up patch, I'm going to make CallLowering emit G_ASSERT_SEXT for
      signext parameters.
      
      Differential Revision: https://reviews.llvm.org/D96913
      8d3442ed
    • Nicolas Vasilache's avatar
      0ee4bf15
    • Geoffrey Martin-Noble's avatar
    • Benjamin Kramer's avatar
      59f442e6
    • Nikita Popov's avatar
      [MemCopyOpt] Enable MemorySSA by default · 71a8e4e7
      Nikita Popov authored
      This enables use of MemorySSA instead of MemDep in MemCpyOpt. To
      allow this without significant compile-time impact, the MemCpyOpt
      pass is moved directly before DSE (in the cases where this was not
      already the case), which allows us to reuse the existing MemorySSA
      analysis.
      
      Unlike the MemDep-based implementation, the MemorySSA-based MemCpyOpt
      can also perform simple optimizations across basic blocks.
      
      Differential Revision: https://reviews.llvm.org/D94376
      71a8e4e7
    • Matthew Malcomson's avatar
      Hwasan InitPrctl check for error using internal_iserror · c1653b8c
      Matthew Malcomson authored
      When adding this function in https://reviews.llvm.org/D68794 I did not
      notice that internal_prctl has the API of the syscall to prctl rather
      than the API of the glibc (posix) wrapper.
      
      This means that the error return value is not necessarily -1 and that
      errno is not set by the call.
      
      For InitPrctl this means that the checks do not catch running on a
      kernel *without* the required ABI (not caught since I only tested this
      function correctly enables the ABI when it exists).
      This commit updates the two calls which check for an error condition to
      use internal_iserror. That function sets a provided integer to an
      equivalent errno value and returns a boolean to indicate success or not.
      
      Tested by running on a kernel that has this ABI and on one that does
      not. Verified that running on the kernel without this ABI the current
      code prints the provided error message and does not attempt to run the
      program. Verified that running on the kernel with this ABI the current
      code does not print an error message and turns on the ABI.
      This done on an x86 kernel (where the ABI does not exist), an AArch64
      kernel without this ABI, and an AArch64 kernel with this ABI.
      
      In order to keep running the testsuite on kernels that do not provide
      this new ABI we add another option to the HWASAN_OPTIONS environment
      variable, this option determines whether the library kills the process
      if it fails to enable the relaxed syscall ABI or not.
      This new flag is `fail_without_syscall_abi`.
      The check-hwasan testsuite results do not change with this patch on
      either x86, AArch64 without a kernel supporting this ABI, and AArch64
      with a kernel supporting this ABI.
      
      Differential Revision: https://reviews.llvm.org/D96964
      c1653b8c
    • Philip Reames's avatar
      [SCEV] Use both known bits and sign bits when computing range of SCEV unknowns · 4a5edea1
      Philip Reames authored
      When computing a range for a SCEVUnknown, today we use computeKnownBits for unsigned ranges, and computeNumSignBots for signed ranges. This means we miss opportunities to improve range results.
      
      One common missed pattern is that we have a signed range of a value which CKB can determine is positive, but CNSB doesn't convey that information. The current range includes the negative part, and is thus double the size.
      
      Per the removed comment, the original concern which delayed using both (after some code merging years back) was a compile time concern. CTMark results (provided by Nikita, thanks!) showed a geomean impact of about 0.1%. This doesn't seem large enough to avoid higher quality results.
      
      Differential Revision: https://reviews.llvm.org/D96534
      4a5edea1
    • Marek Kurdej's avatar
    • Joel E. Denny's avatar
      [OpenMP] Fix nvptx CUDA_VERSION conversion · ef8b3b5f
      Joel E. Denny authored
      As mentioned in PR#49250, without this patch, ptxas for CUDA 9.1 fails
      in the following two tests:
      
      - openmp/libomptarget/test/mapping/lambda_mapping.cpp
      - openmp/libomptarget/test/offloading/bug49021.cpp
      
      The error looks like:
      
      ```
      ptxas /tmp/lambda_mapping-081ea9.s, line 828; error   : Not a name of any known instruction: 'activemask'
      ```
      
      The problem is that our cmake script converts CUDA version strings
      incorrectly: 9.1 becomes 9100, but it should be 9010, as shown in
      `getCudaVersion` in `clang/lib/Driver/ToolChains/Cuda.cpp`.  Thus,
      `openmp/libomptarget/deviceRTLs/nvptx/src/target_impl.cu`
      inadvertently enables `activemask` because it apparently becomes
      available in 9.2.  This patch fixes the conversion.
      
      This patch does not fix the other two tests in PR#49250.
      
      Reviewed By: tianshilei1992
      
      Differential Revision: https://reviews.llvm.org/D97012
      ef8b3b5f
    • Joel E. Denny's avatar
      [OpenMP] Fix always,from and delete for data absent at exit · d2147b1a
      Joel E. Denny authored
      Without this patch, there's a runtime error for those map types at
      exit from an "omp target data" or at "omp target exit data", but the
      spec says the list item should be ignored.
      
      This patch tests that fix in data_absent_at_exit.c, and it also
      improves other testing for data that is not fully present at exit.
      
      Reviewed By: grokos, RaviNarayanaswamy
      
      Differential Revision: https://reviews.llvm.org/D96999
      d2147b1a
  2. Feb 19, 2021