1. Jun 08, 2021
    • Vignesh Balasubramanian's avatar
      [OpenMP][OMPD] Implementation of OMPD debugging library - libompd. · f61602b0
      Vignesh Balasubramanian authored
      This is the first of seven patches that implements OMPD, a debugging interface to support debugging of OpenMP programs.
      It contains support code required in "openmp/runtime" for OMPD implementation.
      
      Reviewed By: @hbae
      Differential Revision: https://reviews.llvm.org/D100181
      f61602b0
    • Kerry McLaughlin's avatar
      [CostModel] Return an invalid cost for memory ops with unsupported types · 5db52751
      Kerry McLaughlin authored
      Fixes getTypeConversion to return `TypeScalarizeScalableVector` when a scalable vector
      type cannot be legalized by widening/splitting. When this is the method of legalization
      found, getTypeLegalizationCost will return an Invalid cost.
      
      The getMemoryOpCost, getMaskedMemoryOpCost & getGatherScatterOpCost functions already call
      getTypeLegalizationCost and will now also return an Invalid cost for unsupported types.
      
      Reviewed By: sdesmalen, david-arm
      
      Differential Revision: https://reviews.llvm.org/D102515
      5db52751
    • Sven van Haastregt's avatar
      [OpenCL] Add memory_scope_all_devices · d54e7b73
      Sven van Haastregt authored
      Add the `memory_scope_all_devices` enum value, which is restricted to
      OpenCL 3.0 or newer and the `__opencl_c_atomic_scope_all_devices`
      feature.  Also guard `memory_scope_all_svm_devices` accordingly, which
      is already available in OpenCL 2.0.
      
      The `__opencl_c_atomic_scope_all_devices` feature is header-only, so
      set its define to 1 in `opencl-c-base.h`.  This is done
      unconditionally at the moment, as the mechanism for disabling
      header-only options hasn't been decided yet.
      
      This patch only adds a negative test for now.  Ideally adding a CL3.0
      run line to atomic-ops.cl should suffice as a positive test, but we
      cannot do that yet until (at least) generic address spaces and program
      scope variables are supported in OpenCL 3.0 mode.
      
      Differential Revision: https://reviews.llvm.org/D103241
      d54e7b73
    • Fraser Cormack's avatar
    • Caroline Concatto's avatar
      [InstCombine] Add instcombine fold for extractelement + splat for scalable vectors · 6fd1604d
      Caroline Concatto authored
      This patch allows that scalable vector can also use the fold that already
      exists for fixed vector, only when the lane index is lower than the minimum
      number of elements of the vector.
      
      Differential Revision: https://reviews.llvm.org/D102404
      6fd1604d
    • Simon Pilgrim's avatar
      OptBisect.cpp - remove unused include. NFCI. · f96b5e80
      Simon Pilgrim authored
      StringRef.h is included in OptBisect.h and we have no uses of std::string.
      f96b5e80
    • Simon Pilgrim's avatar
      [CostModel][X86] Improve AVX1/AVX2 truncation costs · 49d3a367
      Simon Pilgrim authored
      Based off the worse case numbers generated by D103695, we were overestimating the cost of a number of vector truncations:
      
      AVX2: v2i32->v2i8, v2i64->v2i16 + v4i64->v4i32
      AVX1: v2i32->v2i8, v4i64->v4i16 + v16i16->v16i8
      
      Once we have a working set of conversion costs, the intention is to cleanup the tables and use legalized types a lot more to reduce the number of entries we currently have.
      49d3a367
    • Simon Pilgrim's avatar
    • Simon Pilgrim's avatar
    • Simon Pilgrim's avatar
    • Kerry McLaughlin's avatar
      [LoopVectorize] Don't use strict reductions when reordering is allowed · 14eeccfe
      Kerry McLaughlin authored
      If the `-enable-strict-reductions` flag is set to true, then currently we will
      always choose to vectorize the loop with strict in-order reductions. This is
      not necessary where we allow the reordering of FP operations, such as
      when loop hints are passed via metadata.
      
      This patch moves useOrderedReductions so that we can also check whether
      loop hints allow reordering, in which case we should use the default
      behaviour of vectorizing with unordered reductions.
      
      Reviewed By: sdesmalen
      
      Differential Revision: https://reviews.llvm.org/D103814
      14eeccfe
    • Alex Zinenko's avatar
      [mlir] fix shared-libs build · 7116468c
      Alex Zinenko authored
      7116468c
    • David Green's avatar
      [DAG] Allow isNullOrNullSplat to see truncated zeroes · b889c6ee
      David Green authored
      This sets the AllowTruncation flag on isConstOrConstSplat in
      isNullOrNullSplat, allowing it to see truncated constant zeroes on
      architectures such as AArch64, where only a i32.i64 are legal. As a
      truncation of 0 is always 0, this should always be valid, allowing some
      extra folding to happen including some of the cases from D103755.
      
      Differential Revision: https://reviews.llvm.org/D103756
      b889c6ee
    • Martin Storsjö's avatar
      [clang] Apply MS ABI details on __builtin_ms_va_list on non-windows platforms on x86_64 · b34da6ff
      Martin Storsjö authored
      This fixes inconsistencies in the ms_abi.c testcase.
      
      Also add a couple cases of missing double pointers in the windows part
      of the testcase; the outcome of building that testcase on windows hasn't
      changed, but the previous form of the test was imprecise (checking
      for "%[[STRUCT_FOO]]*" when clang actually generates "%[[STRUCT_FOO]]**"),
      which still used to match.
      
      Ideally this would share code with the native Windows case, but
      X86_64ABIInfo and WinX86_64ABIInfo aren't superclasses/subclasses of
      each other so it's impractical, and the code to share currently only
      consists of a couple lines.
      
      Differential Revision: https://reviews.llvm.org/D103837
      b34da6ff
    • Alex Zinenko's avatar
      [mlir] support memref of memref in standard-to-llvm conversion · c59ce1f6
      Alex Zinenko authored
      Now that memref supports arbitrary element types, add support for memref of
      memref and make sure it is properly converted to the LLVM dialect. The type
      support itself avoids adding the interface to the memref type itself similarly
      to other built-in types. This allows the shape, and therefore byte size, of the
      memref descriptor to remain a lowering aspect that is easier to customize and
      evolve as opposed to sanctifying it in the data layout specification for the
      memref type itself.
      
      Factor out the code previously in a testing pass to live in a dedicated data
      layout analysis and use that analysis in the conversion to compute the
      allocation size for memref of memref. Other conversions will be ported
      separately.
      
      Depends On D103827
      
      Reviewed By: rriddle
      
      Differential Revision: https://reviews.llvm.org/D103828
      c59ce1f6
    • Alex Zinenko's avatar
      [mlir] Make MemRef element type extensible · ada9aa5a
      Alex Zinenko authored
      Historically, MemRef only supported a restricted list of element types that
      were known to be storable in memory. This is unnecessarily restrictive given
      the open nature of MLIR's type system. Allow types to opt into being used as
      MemRef elements by implementing a type interface. For now, the interface is
      merely a declaration with no methods. Later, methods to query, e.g., the type
      size or whether a type can alias elements of another type may be added.
      
      Harden the "standard"-to-LLVM conversion against memrefs with non-builtin
      types.
      
      See https://llvm.discourse.group/t/rfc-memref-of-custom-types/3558.
      
      Depends On D103826
      
      Reviewed By: rriddle
      
      Differential Revision: https://reviews.llvm.org/D103827
      ada9aa5a
    • Alex Zinenko's avatar
      [mlir] fix integer type mismatch in alloc conversion to LLVM · 3c70a82e
      Alex Zinenko authored
      Some places in the alloc-like op conversion use the converted index type
      whereas other places use the pointer-sized integer type, which may not be the
      same. Consistently use the converted index type, similarly to other address
      calculations.
      
      Reviewed By: pifon2a
      
      Differential Revision: https://reviews.llvm.org/D103826
      3c70a82e
    • Javier Setoain's avatar
      Revert "[mlir][ArmSVE] Add basic mask generation operations" · 57546f5b
      Javier Setoain authored
      This reverts commit 392af6a7
      57546f5b
    • Lang Hames's avatar
    • David Spickett's avatar
      [lldb] Set return status to failed when adding a command error · e05b03cf
      David Spickett authored
      There is a common pattern:
      result.AppendError(...);
      result.SetStatus(eReturnStatusFailed);
      
      I found that some commands don't actually "fail" but only
      print "error: ..." because the second line got missed.
      
      This can cause you to miss a failed command when you're
      using the Python interface during testing.
      (and produce some confusing script results)
      
      I did not find any place where you would want to add
      an error without setting the return status, so just
      set eReturnStatusFailed whenever you add an error to
      a command result.
      
      This change does not remove any of the now redundant
      SetStatus. This should allow us to see if there are any
      tests that have commands unexpectedly fail with this change.
      (the test suite passes for me but I don't have access to all
      the systems we cover so there could be some corner cases)
      
      Some tests that failed on x86 and AArch64 have been modified
      to work with the new behaviour.
      
      Differential Revision: https://reviews.llvm.org/D103701
      e05b03cf
    • Tomasz Miąsko's avatar
      [Demangle][Rust] Parse const backreferences · f9a79356
      Tomasz Miąsko authored
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D103848
      f9a79356
    • Tomasz Miąsko's avatar
      [Demangle][Rust] Parse type backreferences · 44d63c57
      Tomasz Miąsko authored
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D103847
      44d63c57
    • Tomasz Miąsko's avatar
      [Demangle][Rust] Parse path backreferences · 82b7e822
      Tomasz Miąsko authored
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D103459
      82b7e822
    • Javier Setoain's avatar
      [mlir][ArmSVE] Add basic mask generation operations · 392af6a7
      Javier Setoain authored
      These `arm_sve.cmp` functions are needed to generate scalable vector
      masks as long as scalable vectors are not part of the standard types.
      Once in standard, these can be removed and `std.cmp` can be used
      instead.
      
      Differential Revision: https://reviews.llvm.org/D103473
      392af6a7
    • Denys Petrov's avatar
      [analyzer] [NFC] Implement a wrapper SValBuilder::getCastedMemRegionVal for... · d3a6181e
      Denys Petrov authored
      [analyzer]  [NFC] Implement a wrapper SValBuilder::getCastedMemRegionVal for similar functionality on region cast
      
      Summary: Replaced code on region cast with a function-wrapper SValBuilder::getCastedMemRegionVal. This is a next step of code refining due to suggestions in D103319.
      
      Differential Revision: https://reviews.llvm.org/D103803
      d3a6181e
    • Petr Hosek's avatar
      [Driver] Support libc++ in MSVC · 9625d61e
      Petr Hosek authored
      This implements support for using libc++ headers and library in the MSVC
      toolchain.  We only support libc++ that is a part of the toolchain, and
      not headers installed elsewhere on the system.
      
      Differential Revision: https://reviews.llvm.org/D101479
      9625d61e
    • Craig Topper's avatar
      7c4e9a68
    • Craig Topper's avatar
      [RISCV] Masked compares should use a tail agnostic policy. · ae3ab4f0
      Craig Topper authored
      Writes of a mask result are always tail agnostic.
      
      Unfortunately, this seems to have made codegen worse. I can only
      think this must be because the vsetvli was acting as some sort
      of barrier that prevented some code movement in the scheduler.
      
      Reviewed By: arcbbb
      
      Differential Revision: https://reviews.llvm.org/D103331
      ae3ab4f0
    • Craig Topper's avatar
      [RISCV] Use AVL Operand instead of GPR for tied mask pseudo for vwadd.wv and similar. · 7a105b57
      Craig Topper authored
      I mistakenly copied this from an older version of our internal
      repo.
      7a105b57
    • Yonghong Song's avatar
      BPF: fix relocation types in lib/Object/RelocationResolver.cpp · 8ce45f97
      Yonghong Song authored
      Commit 6a2ea846 ("BPF: Add more relocation kinds")
      added new relocations R_BPF_64_ABS64 and R_BPF_64_ABS32
      for normal 64-bit and 32-bit data relocations.
      This is to replace some of functionalities with
      R_BPF_64_64 and R_BPF_64_32 so that new R_BPF_64_64
      and R_BPF_64_32 semantics are for ld_imm64 and
      call instructions only.
      
      The BPF support in lib/Object/RelocationResolver.cpp
      is used to perform normal data relocations for
      the case like DWARFObjInMemory with an object file
      (search function getRelocationResolver() in file
      DebugInfo/DWARF/DWARFContext.cpp) or llvm-readobj
      to dump ".stack_sizes" section data.
      In all these casees, normal 64-bit and 32-bit relocations
      are performed and such resolution resolution
      is exactly what implemented in RelocationResolver.cpp.
      
      But Commit 6a2ea846 missed to change
      R_BPF_64_64/R_BPF_64_32 to R_BPF_64_ABS64/R_BPF_64_ABS32.
      This patch fixed the issue and added a test for it
      with llvm-readobj dumping ".stack_sizes" section.
      
      Differential Revision: https://reviews.llvm.org/D103864
      8ce45f97
    • Jez Ng's avatar
      [lld-macho] Implement -force_load_swift_libs · 447dfbe0
      Jez Ng authored
      It causes libraries whose names start with "swift" to be force-loaded.
      Note that unlike the more general `-force_load`, this flag only applies
      to libraries specified via LC_LINKER_OPTIONS, and not those passed on
      the command-line. This is what ld64 does.
      
      Reviewed By: #lld-macho, thakis
      
      Differential Revision: https://reviews.llvm.org/D103709
      447dfbe0
    • Jez Ng's avatar
      [lld-macho] Implement cstring deduplication · 04259cde
      Jez Ng authored
      Our implementation draws heavily from LLD-ELF's, which in turn delegates
      its string deduplication to llvm-mc's StringTableBuilder. The messiness of
      this diff is largely due to the fact that we've previously assumed that
      all InputSections get concatenated together to form the output. This is
      no longer true with CStringInputSections, which split their contents into
      StringPieces. StringPieces are much more lightweight than InputSections,
      which is important as we create a lot of them. They may also overlap in
      the output, which makes it possible for strings to be tail-merged. In
      fact, the initial version of this diff implemented tail merging, but
      I've dropped it for reasons I'll explain later.
      
      **Alignment Issues**
      
      Mergeable cstring literals are found under the `__TEXT,__cstring`
      section. In contrast to ELF, which puts strings that need different
      alignments into different sections, clang's Mach-O backend puts them all
      in one section. Strings that need to be aligned have the `.p2align`
      directive emitted before them, which simply translates into zero padding
      in the object file.
      
      I *think* ld64 extracts the desired per-string alignment from this data
      by preserving each string's offset from the last section-aligned
      address. I'm not entirely certain since it doesn't seem consistent about
      doing this; but perhaps this can be chalked up to cases where ld64 has
      to deduplicate strings with different offset/alignment combos -- it
      seems to pick one of their alignments to preserve. This doesn't seem
      correct in general; we can in fact can induce ld64 to produce a crashing
      binary just by linking in an additional object file that only contains
      cstrings and no code. See PR50563 for details.
      
      Moreover, this scheme seems rather inefficient: since unaligned and
      aligned strings are all put in the same section, which has a single
      alignment value, it doesn't seem possible to tell whether a given string
      doesn't have any alignment requirements. Preserving offset+alignments
      for strings that don't need it is wasteful.
      
      In practice, the crashes seen so far seem to stem from x86_64 SIMD
      operations on cstrings. X86_64 requires SIMD accesses to be
      16-byte-aligned. So for now, I'm thinking of just aligning all strings
      to 16 bytes on x86_64. This is indeed wasteful, but implementation-wise
      it's simpler than preserving per-string alignment+offsets. It also
      avoids the aforementioned crash after deduplication of
      differently-aligned strings. Finally, the overhead is not huge: using
      16-byte alignment (vs no alignment) is only a 0.5% size overhead when
      linking chromium_framework.
      
      With these alignment requirements, it doesn't make sense to attempt tail
      merging -- most strings will not be eligible since their overlaps aren't
      likely to start at a 16-byte boundary. Tail-merging (with alignment) for
      chromium_framework only improves size by 0.3%.
      
      It's worth noting that LLD-ELF only does tail merging at `-O2`. By
      default (at `-O1`), it just deduplicates w/o tail merging. @thakis has
      also mentioned that they saw it regress compressed size in some cases
      and therefore turned it off. `ld64` does not seem to do tail merging at
      all.
      
      **Performance Numbers**
      
      CString deduplication reduces chromium_framework from 250MB to 242MB, or
      about a 3.2% reduction.
      
      Numbers for linking chromium_framework on my 3.2 GHz 16-Core Intel Xeon W:
      
            N           Min           Max        Median           Avg        Stddev
        x  20          3.91          4.03         3.935          3.95   0.034641016
        +  20          3.99          4.14         4.015        4.0365     0.0492336
        Difference at 95.0% confidence
                0.0865 +/- 0.027245
                2.18987% +/- 0.689746%
                (Student's t, pooled s = 0.0425673)
      
      As expected, cstring merging incurs some non-trivial overhead.
      
      When passing `--no-literal-merge`, it seems that performance is the
      same, i.e. the refactoring in this diff didn't cost us.
      
            N           Min           Max        Median           Avg        Stddev
        x  20          3.91          4.03         3.935          3.95   0.034641016
        +  20          3.89          4.02         3.935        3.9435   0.043197831
        No difference proven at 95.0% confidence
      
      Reviewed By: #lld-macho, gkm
      
      Differential Revision: https://reviews.llvm.org/D102964
      04259cde
    • Esme-Yi's avatar
      [yaml2obj] Fix buildbot-issue-4886 · 310d2b49
      Esme-Yi authored
      XCOFFEmitter.cpp:67:16: runtime error: null pointer passed as argument 2,
      which is declared to never be null
      310d2b49
    • Carl Ritson's avatar
      [AMDGPU] Allow oversize vaddr in GFX10 MIMG assembly · c8bbfb8c
      Carl Ritson authored
      As a follow up to D103672, we should allow vaddr to be larger than
      required when assembling GFX10 MIMG instructions.
      
      Reviewed By: dp
      
      Differential Revision: https://reviews.llvm.org/D103733
      c8bbfb8c
    • Jake.Egan's avatar
      [AIX] Define __STDC_NO_ATOMICS__ and __STDC_NO_THREADS__ · f38eff77
      Jake.Egan authored
      Revert/reapply to fix Git authorship metadata
      
      Differential Revision: https://reviews.llvm.org/D103707
      f38eff77
    • Chris Bowler's avatar
    • Carl Ritson's avatar
      [AMDGPU] Add v5f32/VReg_160 support for MIMG instructions · f8816c74
      Carl Ritson authored
      Avoid having to round up to v8f32/VReg_256 when only 5 VGPRs are
      required for a MIMG address operand.
      
      Maintain _V8 instruction variants of pseudo instructions allowing
      assembly prior to GFX10 to work as-is.  Currently the validator
      can tell for GFX10 what the correct size is, so will disallow
      oversize address registers.
      
      Reviewed By: rampitec
      
      Differential Revision: https://reviews.llvm.org/D103672
      f8816c74
    • =Jake Egan's avatar
    • Vitaly Buka's avatar
      [NFC][scudo] Print errno of fork failure · b41b76b3
      Vitaly Buka authored
      This fork fails sometime on sanitizer-x86_64-linux-qemu bot.
      b41b76b3
    • Craig Topper's avatar
      [RISCV] Use bitfields to shrink the size of the vector load/store intrinsics... · 0aa94165
      Craig Topper authored
      [RISCV] Use bitfields to shrink the size of the vector load/store intrinsics to pseudo instruction lookup tables.
      0aa94165