1. Sep 03, 2020
  2. Sep 02, 2020
    • Simon Pilgrim's avatar
      [X86][SSE] Fold vselect(pshufb,pshufb) -> or(pshufb,pshufb) · 888049b9
      Simon Pilgrim authored
      If the PSHUFBs have no other uses, then we can force the unselected elements to zero to OR them instead, avoiding both an extra mask load and a costly variable blend.
      
      Eventually we should try to bring this into shuffle combining, once we can more easily convert between shuffles + select patterns.
      888049b9
    • Ehsan Toosi's avatar
      [mlir] Extend BufferAssignmentTypeConverter with result conversion callbacks · 39cf83cc
      Ehsan Toosi authored
      In this PR, the users of BufferPlacement can configure
      BufferAssginmentTypeConverter. These new configurations would give the user more
      freedom in the process of converting function signature, and return and call
      operation conversions.
      
      These are the new features:
          - Accepting callback functions for decomposing types (i.e. 1 to N type
          conversion such as unpacking tuple types).
          - Defining ResultConversionKind for specifying whether a function result
          with a certain type should be appended to the function arguments list or
          should be kept as function result. (Usage:
          converter.setResultConversionKind<MemRefType>(AppendToArgumentList))
          - Accepting callback functions for composing or decomposing values (i.e. N
          to 1 and 1 to N value conversion).
      
      Differential Revision: https://reviews.llvm.org/D85133
      39cf83cc
    • Jordan Rupprecht's avatar
      [lldb/Host] Add missing proc states · c5aa63dd
      Jordan Rupprecht authored
      The /proc/<pid>/status parsing is missing a few cases:
      - Idle
      - Parked
      - Dead
      
      If we encounter an unknown proc state, this leads to an msan warning. In reality, we only check that the state != Zombie, so it doesn't really matter that we handle all cases, but handle them anyway (current list: [1]). Also explicitly set it to unknown if we encounter an unknown state. There will still be an msan warning if the proc entry has no `State:` line, but that should not happen.
      
      Use a StringSwitch to make the handling of proc states a little more compact.
      
      [1] https://github.com/torvalds/linux/blob/master/fs/proc/array.c
      
      Reviewed By: labath
      
      Differential Revision: https://reviews.llvm.org/D86818
      c5aa63dd
    • Congzhe Cao's avatar
      [IPSCCP] Fix a bug that the "returned" attribute is not cleared when function... · ec489ae0
      Congzhe Cao authored
      [IPSCCP] Fix a bug that the "returned" attribute is not cleared when function is optimized to return undef
      
      In IPSCCP when a function is optimized to return undef, it should clear the returned attribute for all its input arguments
      and its corresponding call sites.
      
      The bug is exposed when the value of an input argument of the function is assigned to a physical register and
      because of the argument having a returned attribute, the value of this physical register will continue to be used
      as the function return value right after the call instruction returns, even if the value that this register holds may
      be clobbered during the function call. This potentially results in incorrect values being used afterwards.
      
      Reviewed By: jdoerfert, fhahn
      
      Differential Revision: https://reviews.llvm.org/D84220
      ec489ae0
    • Qiu Chaofan's avatar
      b6b63684
    • Med Ismail Bennani's avatar
      [lldb/Target] Add custom interpreter option to `platform shell` · addb5148
      Med Ismail Bennani authored
      This patch adds the ability to use a custom interpreter with the
      `platform shell` command. If the user set the `-s|--shell` option
      with the path to a binary, lldb passes it down to the platform's
      `RunShellProcess` method and set it as the shell to use in
      `ProcessLaunchInfo to run commands.
      
      Note that not all the Platforms support running shell commands with
      custom interpreters (i.e. RemoteGDBServer is only expected to use the
      default shell).
      
      This patch also makes some refactoring and cleanups, like swapping
      CString for StringRef when possible and updating `SBPlatformShellCommand`
      with new methods and a new constructor.
      
      rdar://67759256
      
      Differential Revision: https://reviews.llvm.org/D86667
      
      
      
      Signed-off-by: default avatarMed Ismail Bennani <medismail.bennani@gmail.com>
      addb5148
    • Anna Thomas's avatar
      [ImplicitNullChecks] NFC: Refactor dependence safety check · 425573a2
      Anna Thomas authored
      After computing dependence, we check if it is safe to hoist by
      identifying if it clobbers any liveIns in the sibling block (NullSucc).
      This check is moved to its own function which will be used in the
      soon-to-be modified dependence checking algorithm for implicit null
      checks pass.
      
      Tests-Run: lit tests on X86/implicit-*
      425573a2
    • Anna Thomas's avatar
      [ImplicitNullChecks] NFC: Separated out checks and added comments · 6f7737c4
      Anna Thomas authored
      Separated out some checks in isSuitableMemoryOp and added comments
      explaining why some of those checks are done.
      
      Tests-Run:X86 implicit null checks tests.
      6f7737c4
    • Louis Dionne's avatar
      [libc++] Make some testing utilities constexpr · 255a60cd
      Louis Dionne authored
      This will be needed in order to test constexpr std::vector.
      255a60cd
    • Lei Zhang's avatar
      Revert "[mlir] Extend BufferAssignmentTypeConverter with result conversion callbacks" · 1b88bbf5
      Lei Zhang authored
      This reverts commit 94f5d248 because
      of failing the following tests:
      
      MLIR :: Dialect/Linalg/tensors-to-buffers.mlir
      MLIR :: Transforms/buffer-placement-preparation-allowed-memref-results.mlir
      MLIR :: Transforms/buffer-placement-preparation.mlir
      1b88bbf5
    • David Stenberg's avatar
      [GlobalOpt] Fix an incorrect Modified status · 6d36b22b
      David Stenberg authored
      When marking a global variable constant, and simplifying users using
      CleanupConstantGlobalUsers(), the pass could incorrectly return false if
      there were still some uses left, and no further optimizations was done.
      
      This was caught using the check introduced by D80916.
      
      This fixes PR46749.
      
      Reviewed By: fhahn
      
      Differential Revision: https://reviews.llvm.org/D85837
      6d36b22b
    • Jakub Lichman's avatar
      [mlir][VectorToSCF] 128 byte alignment of alloc ops · f5ed22f0
      Jakub Lichman authored
      Added 128 byte alignment to alloc ops created in VectorToSCF pass.
      128b alignment was already introduced to this pass but not to all alloc
      ops. This commit changes that by adding 128b alignment to the remaining ops.
      The point of specifying alignment is to prevent possible memory alignment errors
      on weakly tested architectures.
      
      Differential Revision: https://reviews.llvm.org/D86454
      f5ed22f0
    • Venkataramanan Kumar's avatar
      [InstCombine] Transform 1.0/sqrt(X) * X to X/sqrt(X) · 626c3738
      Venkataramanan Kumar authored
      These transforms will now be performed irrespective of the number of uses for the expression "1.0/sqrt(X)":
      1.0/sqrt(X) * X => X/sqrt(X)
      X * 1.0/sqrt(X) => X/sqrt(X)
      
      We already handle more general cases, and we are intentionally not creating extra (and likely expensive)
      fdiv ops in IR. This pattern is the exception to the rule because we always expect the Backend to reduce
      X/sqrt(X) to sqrt(X), if it has the necessary (reassoc) fast-math-flags.
      
      Ref: DagCombiner optimizes the X/sqrt(X) to sqrt(X).
      
      Differential Revision: https://reviews.llvm.org/D86726
      626c3738
    • Sanjay Patel's avatar
      [VectorCombine] allow vector loads with mismatched insert type · 8fb05593
      Sanjay Patel authored
      This is an enhancement to D81766 to allow loading the minimum target
      vector type into an IR vector with a different number of elements.
      
      In one of the motivating tests from PR16739, SLP creates <2 x float>
      load ops mixed with <4 x float> insert ops, so we want to handle that
      pattern in addition to potential oversized vectors created by the
      vectorizers.
      
      For now, we are assuming the insert/extract subvector with undef is
      free because there is no exact corresponding TTI modeling for that.
      
      Differential Revision: https://reviews.llvm.org/D86160
      8fb05593
    • Daniel Grumberg's avatar
      Move all fields of '-cc1' option related classes into def file databases · c4a2a130
      Daniel Grumberg authored
      Once the new option parsing system is committed, this will allow to generate a
      check to ensure that correct command line generation happens
      
      Differential Revision: https://reviews.llvm.org/D86290
      c4a2a130
    • Max Kazantsev's avatar
      8a3907cd
    • Ehsan Toosi's avatar
      [mlir] Extend BufferAssignmentTypeConverter with result conversion callbacks · 94f5d248
      Ehsan Toosi authored
      In this PR, the users of BufferPlacement can configure
      BufferAssginmentTypeConverter. These new configurations would give the user more
      freedom in the process of converting function signature, and return and call
      operation conversions.
      
      These are the new features:
          - Accepting callback functions for decomposing types (i.e. 1 to N type
          conversion such as unpacking tuple types).
          - Defining ResultConversionKind for specifying whether a function result
          with a certain type should be appended to the function arguments list or
          should be kept as function result. (Usage:
          converter.setResultConversionKind<MemRefType>(AppendToArgumentList))
          - Accepting callback functions for composing or decomposing values (i.e. N
          to 1 and 1 to N value conversion).
      
      Differential Revision: https://reviews.llvm.org/D85133
      94f5d248
    • Paul Walker's avatar
      [SVE] Don't reorder subvector/binop sequences when the resulting binop is not legal. · f7212125
      Paul Walker authored
      When lowering fixed length vector operations for SVE the subvector
      operations are used extensively to marshall data between scalable
      and fixed-length vectors. This means that sequences like:
      
        extract_subvec(binop(insert_subvec(a), insert_subvec(b)))
      
      are very common. DAGCombine only checks if the resulting binop is
      legal or can be custom lowered when undoing such sequences. When
      it's custom lowering that is introducing them the result is an
      infinite legalise->combine->legalise loop.
      
      This patch extends the isOperationLegalOr... functions to include
      a "LegalOnly" parameter to restrict the check to legal operations
      only. Although isOperationLegal could be used it's common for
      the affected code paths to be visited pre and post legalisation,
      so the extra parameter keeps the code tidy.
      
      Differential Revision: https://reviews.llvm.org/D86450
      f7212125
    • Jay Foad's avatar
      [AMDGPU] Fix offset for REL32_HI relocs · 4bdab2e8
      Jay Foad authored
      The addend in a REL32 reloc needs to be adjusted to account for the
      offset from the PC value returned by the s_getpc instruction to the
      point where the reloc is applied. This was being done correctly for
      (GOTPC)REL32_LO but not for (GOTPC)REL32_HI. This will only make a
      difference if the target symbol happens to get loaded almost exactly
      a multiple of 4G away from the relocated instructions.
      
      Differential Revision: https://reviews.llvm.org/D86938
      4bdab2e8
    • Sander de Smalen's avatar
      [AArch64][SVE] Preserve full vector regs over EH edge. · f13beac5
      Sander de Smalen authored
      Unwinders may only preserve the lower 64bits of Neon and SVE registers,
      as only the registers in the base ABI are guaranteed to be preserved
      over the exception edge. The caller will need to preserve additional
      registers for when the call throws an exception and the unwinder has
      tried to recover state.
      
      For  e.g.
      
          svint32_t bar(svint32_t);
          svint32_t foo(svint32_t x, bool *err) {
            try { bar(x); } catch (...) { *err = true; }
            return x;
          }
      
      `z0` needs to be spilled before the call to `bar(x)` and reloaded before
      returning from foo, as the exception handler may have clobbered z0.
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D84737
      f13beac5
    • Igor Kudrin's avatar
      [DebugInfo] Emit a 1-byte value as a terminator of entries list in the name index. · 3445ec9b
      Igor Kudrin authored
      As stated in section 6.1.1.2, DWARFv5, p. 142,
      | The last entry for each name is followed by a zero byte that
      | terminates the list. There may be gaps between the lists.
      
      The patch changes emitting a 4-byte zero value to a 1-byte one, which
      effectively removes the gap between entry lists, and thus saves
      approximately 3 bytes per name; the calculation is not exact because
      the total size of the table is aligned to 4.
      
      Differential Revision: https://reviews.llvm.org/D86927
      3445ec9b
    • Igor Kudrin's avatar
      [DebugInfo] Remove Dwarf5AccelTableWriter::Header::UnitLength. NFC. · 71eed480
      Igor Kudrin authored
      The member is not in use; the unit length for the table is emitted as
      a difference between two labels. Moreover, the type of the member might
      be misleading, because for DWARF64 the field should be 64 bit long.
      
      Differential Revision: https://reviews.llvm.org/D86912
      71eed480
    • Martin Storsjö's avatar
    • Benjamin Kramer's avatar
      [mlir][VectorOps] Fail fast when a strided memref is passed to vector_transfer · 2bf491c7
      Benjamin Kramer authored
      Otherwise we'll silently miscompile things.
      
      Differential Revision: https://reviews.llvm.org/D86951
      2bf491c7
    • Simon Pilgrim's avatar
      [X86][SSE] SimplifyDemandedVectorEltsForTargetNode - add general shuffle combining support · 21d02dc5
      Simon Pilgrim authored
      This patch uses partial DemandedElts masks to further simplify target shuffle chains and finally starts making target shuffle combining part of SimplifyDemandedBits/SimplifyDemandedVectorElts.
      
      We already manage this for Depth == 0 cases, where combineX86ShuffleChain would early-out if the shuffle combined to the same op, but the patch generalizes this by manipulating the depth handling of combineX86ShufflesRecursively - calling with a new Depth = 0 and reducing the maximum shuffle combine depth accordingly.
      
      Differential Revision: https://reviews.llvm.org/D66004
      21d02dc5
    • Raphael Isemann's avatar
      Revert "[libc++] Workaround timespec_get not always being available in Apple SDKs" · 81424257
      Raphael Isemann authored
      This reverts commit 99f3b231. It breaks
      libcxx/modules/stds_include.sh.cpp on macOS as the new include to sys/cdefs.h
      causes a dependency from __config to the Darwin module (which already has
      a dependency on __config). This cyclic dependency breaks compiling the std
      module which breaks compiling pretty much every program with ToT libc++ and
      enabled modules.
      
      I'll revert for now to get the bots green again. Sorry for the inconvenience.
      81424257
    • Shinji Okumura's avatar
      [Attributor] Make use of AANoUndef in AAUndefinedBehavior · 5d134795
      Shinji Okumura authored
      This patch makes it possible for AAUB to use information from AANoUndef.
      This is the next patch of D86983
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D86984
      5d134795
    • Shinji Okumura's avatar
      [Attributor] Fix AANoUndef initialization · 7558e9e5
      Shinji Okumura authored
      When the associated value is undef, we immediately forced to indicate a pessimistic fixpoint so far.
      This patch changes the initialization to check the attribute given in IR at first and to indicate an optimistic fixpoint when it is given.
      This change will enable us to catch , for example, the following case in AAUB.
      ```
      call void @foo(i32 noundef undef)
      ```
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D86983
      7558e9e5
    • ZHANG Hongbin's avatar
      [mlir] Add Complex Type, Vector Type and Tuple Type subclasses to python bindings · 1d994728
      ZHANG Hongbin authored
      Based on the PyType and PyConcreteType classes, this patch implements the bindings of Complex Type, Vector Type and Tuple Type subclasses.
      For the convenience of type checking, this patch defines a `mlirTypeIsAIntegerOrFloat` function to check whether the given type is an integer or float type.
      These three subclasses in this patch have similar binding strategy:
      - The function pointer `isaFunction` points to `mlirTypeIsA***`.
      - The `mlir***TypeGet` C API is bound with the `get_***` method in the python side.
      - The Complex Type and Vector Type check whether the given type is an integer or float type.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D86785
      1d994728