1. Feb 25, 2020
    • Andrzej Warzynski's avatar
      [AArch64][SVE] Update names and comments for gathers/scatters (NFC) · cff90c93
      Andrzej Warzynski authored
      Summary:
      This patch renames functions and TableGen classes for SVE gathers and
      scatters. The original names implied that the corresponding
      methods/classes are only suited for regular gathers/scatters (i.e. LD1
      and ST1), which is not the case. Indeed, we will be re-using them for
      non-temporal and first-faulting gathers/scatters in the forthcoming
      patches. The new names also highlight the split into Vector-Scalar (VS)
      and Scalar-Vector (SV) cases.
      
      List of changes:
      * `performLD1GatherCombine` and `performST1ScatterCombine` are renamed
        as `performGatherLoadCombine` and `performScatterStoreCombine`,
        respectively.
      * Selection DAG types for scatters and gathers from
        AArch64SVEInstrInfo.td are renamed. For example, `SDT_AArch64_GLD1` is
        renamed as `SDT_AArch64_GATHER_SV`. SV stands for Scalar-Vector, as
        opposed to Vector-Scalar (VS).
      * The intrinsic classes from IntrinsicsAArch64.td are renamed. For
        example, `AdvSIMD_GatherLoad_64bitOffset_Intrinsic` is renamed as
        `AdvSIMD_GatherLoad_SV_64b_Offsets_Intrinsic`.
      * Updated comments in `performGatherLoadCombine` and
        `performScatterStoreCombine`.
      
      Reviewers: sdesmalen, rengolin, efriedma
      
      Reviewed By: sdesmalen
      
      Subscribers: tschuett, kristof.beyls, hiraditya, rkruppe, psnobl, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D75035
      cff90c93
    • Raphael Isemann's avatar
    • Alex Zinenko's avatar
      [mlir] simplify affine maps and operands in affine.min/max · 5f9b543e
      Alex Zinenko authored
      Affine dialect already has a map+operand simplification infrastructure in
      place. Plug the recently added affine.min/max operations into this
      infrastructure and add a simple test. More complex behavior of the simplifier
      is already tested by other ops.
      
      Addresses https://bugs.llvm.org/show_bug.cgi?id=45008.
      
      Differential Revision: https://reviews.llvm.org/D75058
      5f9b543e
    • Alex Zinenko's avatar
      [mlir] Intrinsics generator: use TableGen-defined builder function · 3a1b34ff
      Alex Zinenko authored
      Originally, intrinsics generator for the LLVM dialect has been producing
      customized code fragments for the translation of MLIR operations to LLVM IR
      intrinsics. LLVM dialect ODS now provides a generalized version of the
      translation code, parameterizable with the properties of the operation.
      Generate ODS that uses this version of the translation code instead of
      generating a new version of it for each intrinsic.
      
      Differential Revision: https://reviews.llvm.org/D74893
      3a1b34ff
    • Alex Zinenko's avatar
      [mlir] Generalize intrinsic builders in the LLVM dialect definition · 00d4814f
      Alex Zinenko authored
      All LLVM IR intrinsics are constructed in a similar way. The ODS definition of
      the LLVM dialect in MLIR also lists multiple intrinsics, many of which
      reproduce the same (or similar enough) code stanza to translate the MLIR
      operation into the LLVM IR intrinsic. Provide a single base class containing
      parameterizable code to build LLVM IR intrinsics given their name and the lists
      of overloadable operands and results. Use this class to remove (almost)
      duplicate translations for intrinsics defined in LLVMOps.td.
      
      Differential Revision: https://reviews.llvm.org/D74889
      00d4814f
    • Hans Wennborg's avatar
      Don't generate libcalls for wide shift on Windows ARM (PR42711) · decd021f
      Hans Wennborg authored
      The previous patch (cff90f07) didn't
      cover ARM.
      decd021f
    • Stephan Herhut's avatar
      [MLIR][GPU] Implement a simple greedy loop mapper. · 7a7eacc7
      Stephan Herhut authored
      Summary:
      The mapper assigns annotations to loop.parallel operations that
      are compatible with the loop to gpu mapping pass. The outermost
      loop uses the grid dimensions, followed by block dimensions. All
      remaining loops are mapped to sequential loops.
      
      Differential Revision: https://reviews.llvm.org/D74963
      7a7eacc7
    • Georgii Rymar's avatar
      [yaml2obj] - Address post commit comments for D74764 · 157b3d50
      Georgii Rymar authored
      It removes a stale comment and fixes the comment in the test
      and section names related accordingly.
      157b3d50
    • Cullen Rhodes's avatar
      [AArch64][SVE] Add predicate reinterpret intrinsics · 72848f26
      Cullen Rhodes authored
      Summary:
      Implements the following intrinsics:
      
          * llvm.aarch64.sve.convert.to.svbool
          * llvm.aarch64.sve.convert.from.svbool
      
      For converting the ACLE svbool_t type (<n x 16 x i1>) to and from the
      other predicate types: <n x 8 x i1>, <n x 4 x i1> and <n x 2 x i1>.
      
      Reviewers: sdesmalen, kmclaughlin, efriedma, dancgr, rengolin
      
      Reviewed By: sdesmalen, efriedma
      
      Subscribers: tschuett, kristof.beyls, hiraditya, rkruppe, psnobl, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D74471
      72848f26
    • Kristóf Umann's avatar
      [analyzer][MallocChecker][NFC] Communicate the allocation family to auxiliary... · 9fd7ce7f
      Kristóf Umann authored
      [analyzer][MallocChecker][NFC] Communicate the allocation family to auxiliary functions with parameters
      
      The following series of refactoring patches aim to fix the horrible mess that MallocChecker.cpp is.
      
      I genuinely hate this file. It goes completely against how most of the checkers
      are implemented, its by far the biggest headache regarding checker dependencies,
      checker options, or anything you can imagine. On top of all that, its just bad
      code. Its seriously everything that you shouldn't do in C++, or any other
      language really. Bad variable/class names, in/out parameters... Apologies, rant
      over.
      
      So: there are a variety of memory manipulating function this checker models. One
      aspect of these functions is their AllocationFamily, which we use to distinguish
      between allocation kinds, like using free() on an object allocated by operator
      new. However, since we always know which function we're actually modeling, in
      fact we know it compile time, there is no need to use tricks to retrieve this
      information out of thin air n+1 function calls down the line. This patch changes
      many methods of MallocChecker to take a non-optional AllocationFamily template
      parameter (which also makes stack dumps a bit nicer!), and removes some no
      longer needed auxiliary functions.
      
      Differential Revision: https://reviews.llvm.org/D68162
      9fd7ce7f
    • Igor Kudrin's avatar
      [DebugInfo] Fix printing CIE offsets in EH FDEs. · bd2df13e
      Igor Kudrin authored
      While the value of the CIE pointer field in a DWARF FDE record is
      an offset to the corresponding CIE record from the beginning of
      the section, for EH FDE records it is relative to the current offset.
      Previously, we did not make that distinction when dumped both kinds
      of FDE records and just printed the same value for the CIE pointer
      field and the CIE offset; that was acceptable for DWARF FDEs but was
      wrong for EH FDEs.
      
      This patch fixes the issue by explicitly printing the offset of the
      linked CIE object.
      
      Differential Revision: https://reviews.llvm.org/D74613
      bd2df13e
    • Hans Wennborg's avatar
    • Calixte Denizet's avatar
      [profile] gcov_mutex must be static · 62c7d840
      Calixte Denizet authored
      Summary: Forget static keyword for gcov_mutex in https://reviews.llvm.org/D74953 and that causes test failure on mac.
      
      Reviewers: erik.pilkington, vsk
      
      Reviewed By: vsk
      
      Subscribers: vsk, dexonsmith, #sanitizers, llvm-commits
      
      Tags: #sanitizers, #llvm
      
      Differential Revision: https://reviews.llvm.org/D75080
      62c7d840
    • Jay Foad's avatar
    • Jay Foad's avatar
      dc781908
    • Jan Vesely's avatar
      libclc: cmake configure should depend on file list · 814fb658
      Jan Vesely authored
      This makes sure targets are rebuilt if a file is added or removed.
      Reviewer: tstellar
      Differential Revision: https://reviews.llvm.org/D74662
      814fb658
    • Raphael Isemann's avatar
      [lldb][NFC] Move namespace lookup in ClangASTSource to own function. · 05d174d3
      Raphael Isemann authored
      Beside being cleaner we can probably reuse that logic elsewhere.
      05d174d3
    • Kang Zhang's avatar
      27c89ced
    • Pavel Labath's avatar
      [lldb] s/CHECK-NEXT/CHECK-DAG in dwp-debug-types.s · eefbff00
      Pavel Labath authored
      These can come out nondeterministically for two reasons:
      - sorting based on ConstStringified pointer values
      - different relative speeds of the indexing threads
      
      Making these nondeterministic without incurring performance penalties is
      hard, so I just make the test expect them in any order (the order is not
      important in this test anyway.
      eefbff00
    • Raphael Isemann's avatar
      [lldb][NFC] Make ArrayRef initialization more obvious in lldb-test.cpp · ea6b95dc
      Raphael Isemann authored
      Seems like this code raised some alarm bells as it looks like an ArrayRef
      to a temporary initializer list, but it's actually just calling the ArrayRef(T*, T*)
      constructor. Let's clarify this and directly call the right ArrayRef constructor here.
      
      Fixes rdar://problem/59176052
      ea6b95dc
    • Alex Brachet's avatar
      [libc] [UnitTest] Give UnitTest gtest like colors · 29e2cb87
      Alex Brachet authored
      Summary:
      This is a quality of life change to make it a little nicer to look at, NFC.
      This patch makes the RUN and OK lines green and FAILED lines red to match gtest's output.
      
      Reviewers: sivachandra, gchatelet, PaulkaToast
      
      Reviewed By: gchatelet
      
      Subscribers: MaskRay, tschuett, libc-commits
      
      Differential Revision: https://reviews.llvm.org/D75103
      29e2cb87
    • Craig Topper's avatar
      [X86] Pass parameters into selectVectorAddr to remove dependency on X86MaskedGatherScatterSDNode. · 89ba4aca
      Craig Topper authored
      Might be able to get rid of X86ISD::SCATTER and some uses of
      X86ISD::GATHER. Which require isel to use ISD::SCATTER and
      ISD::GATHER as well.
      89ba4aca
    • Craig Topper's avatar
      [X86] Remove mask output from X86 gather/scatter ISD opcodes. · 9238dfb4
      Craig Topper authored
      Instead add it when we make the machine nodes during instruction
      selections.
      
      This makes this ISD node closer to ISD::MGATHER. Trying to see
      if we remove the X86 specific ones.
      9238dfb4
    • Nathan James's avatar
      [ASTMatchers] Adds a matcher called `hasAnyOperatorName` · 6a0c066c
      Nathan James authored
      Summary:
      Acts on `BinaryOperator` and `UnaryOperator` and functions the same as `anyOf(hasOperatorName(...), hasOperatorName(...), ...)`
      
      Documentation generation isn't perfect but I feel that the python doc script needs updating for that
      
      Reviewers: aaron.ballman, gribozavr2
      
      Reviewed By: aaron.ballman, gribozavr2
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D75040
      6a0c066c
    • Nathan James's avatar
      [ASTMatchers] Matcher macros with params move params instead of copying · 3e9a7b2b
      Nathan James authored
      Summary: Use move semantics instead of copying for AST Matchers with parameters
      
      Reviewers: aaron.ballman, gribozavr2
      
      Reviewed By: gribozavr2
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D75096
      3e9a7b2b
    • Raphael Isemann's avatar
      [lldb] Fix that a crashing test is marked as unsupported when it prints UNSUPPORTED before crashing · 55d4b0d7
      Raphael Isemann authored
      Summary:
      I added an `abort()` call to some code and noticed that the test suite was still passing and it just marked my test as "UNSUPPORTED".
      
      It seems the reason for that is that we expect failing tests to print "FAIL:" which doesn't happen when we crash. If we then also
      have an unsupported because we skipped some debug information in the output, we just mark the test passing because it is unsupported
      on the current platform.
      
      This patch marks any test that has a non-zero exit code as failing even if it doesn't print "FAIL:" (e.g., because it crashed).
      
      Reviewers: labath, JDevlieghere
      
      Reviewed By: labath, JDevlieghere
      
      Subscribers: aprantl, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D75031
      55d4b0d7
    • Pavel Labath's avatar
      [lldb] Mark ObjectFileBreakpad test inputs as non-text · c08a1c70
      Pavel Labath authored
      These are technically text files, but the object file layer treats them
      as binary, and the relevant tests verify the parsed contents byte for
      byte. Git's crlf conversion can make those tests fail. Marking the files
      as non-text disables that.
      c08a1c70
    • Jim Lin's avatar
      [Sparc][NFC] Remove trailing space · 84c3d3f3
      Jim Lin authored
      84c3d3f3
    • Jonas Devlieghere's avatar
      [lldb/Utility] Fix unspecified behavior. · 35a06145
      Jonas Devlieghere authored
      Order of evaluation of the operands of any C++ operator [...] is
      unspecified. This patch fixes the issue in Stream::Indent by calling the
      function consecutively.
      
      On my Windows setup, TestSettings.py fails because the function prints
      the value first, followed by the indentation.
      
      Expected result:
        MY_FILE=this is a file name with spaces.txt
      
      Actual result:
      MY_FILE  =this is a file name with spaces.txt
      35a06145
    • Hideto Ueno's avatar
    • Matt Arsenault's avatar
      GlobalISel: Remove unneeded initialiation · 1612d382
      Matt Arsenault authored
      Removes implicit unsigned->Register conversion.
      1612d382
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Introduce post-legalize combiner · fee41517
      Matt Arsenault authored
      The current set of custom combines are only really useful after
      legalization, so move them there. There is a lot of overlap in the
      boilerplate here, but I think we do want a pretty different set of
      combines before and after legalize. I think we will want a lot of
      overlap between the post-legalize and a post-regbankselect combiner.
      fee41517
    • Jason Molenda's avatar
      Revert "Unwind past an interrupt handler correctly on arm or at pc==0" · 4fdd2edb
      Jason Molenda authored
      The aarcht64-ubuntu bot is showing a test failure in TestHandleAbort.py
      with this patch.  Adding some logging to that file, it looks like
      the saved register context above the trap handler does not have
      save state for $pc, but it does have it for $lr on that platform.
      I need to fall back to looking for $lr if the $pc cannot be retrieved.
      I'll update the patch and re-commit once that's fixed.
      
      This reverts commit edc4f4c9.
      4fdd2edb
    • Jason Molenda's avatar
      d5a4fa05
    • Roman Tereshin's avatar
      [MachineVerifier] Doing ::calcRegsPassed over faster sets: ~15-20% faster MV, NFC · b3bce6a3
      Roman Tereshin authored
      MachineVerifier still takes 45-50% of total compile time with
      -verify-machineinstrs, with calcRegsPassed dataflow taking ~50-60% of
      MachineVerifier.
      
      The majority of that time is spent in BBInfo::addPassed, mostly within
      DenseSet implementing the sets the dataflow is operating over.
      
      In particular, 1/4 of that DenseSet time is spent just iterating over it
      (operator++), 40-50% on insertions, and most of the rest in ::count.
      
      Given that, we're implementing custom sets just for this analysis here,
      focusing on cheap insertions and O(n) iteration time (as opposed to
      O(U), where U is the universe).
      
      As it's based _mostly_ on BitVector for sparse and SmallVector for
      dense, it may remotely resemble SparseSet. The difference is, our
      solution is a lot less clever, doesn't have constant time `clear` that
      we won't use anyway as reusing these sets across analyses is cumbersome,
      and thus more space efficient and safer (got a resizable Universe and a
      fallback to DenseSet for sparse if it gets too big).
      
      With this patch MachineVerifier gets ~15-20% faster, its contribution to
      total compile time drops from 45-50% to ~35%, while contribution of
      calcRegsPassed to MachineVerifier drops from 50-60% to ~35% as well.
      
      calcRegsPassed itself gets another 2x faster here.
      
      All measured on a large suite of shaders targeting a number of GPUs.
      
      Reviewers: bogner, stoklund, rudkx, qcolombet
      
      Reviewed By: rudkx
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D75033
      b3bce6a3
    • Bill Wendling's avatar
      Support output constraints on "asm goto" · 50cac248
      Bill Wendling authored
      Summary:
      Clang's "asm goto" feature didn't initially support outputs constraints. That
      was the same behavior as gcc's implementation. The decision by gcc not to
      support outputs was based on a restriction in their IR regarding terminators.
      LLVM doesn't restrict terminators from returning values (e.g. 'invoke'), so
      it made sense to support this feature.
      
      Output values are valid only on the 'fallthrough' path. If an output value's used
      on an indirect branch, then it's 'poisoned'.
      
      In theory, outputs *could* be valid on the 'indirect' paths, but it's very
      difficult to guarantee that the original semantics would be retained. E.g.
      because indirect labels could be used as data, we wouldn't be able to split
      critical edges in situations where two 'callbr' instructions have the same
      indirect label, because the indirect branch's destination would no longer be
      the same.
      
      Reviewers: jyknight, nickdesaulniers, hfinkel
      
      Reviewed By: jyknight, nickdesaulniers
      
      Subscribers: MaskRay, rsmith, hiraditya, llvm-commits, cfe-commits, craig.topper, rnk
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D69876
      50cac248
    • Bill Wendling's avatar
      Allow "callbr" to return non-void values · 23c2a5ce
      Bill Wendling authored
      Summary:
      Terminators in LLVM aren't prohibited from returning values. This means that
      the "callbr" instruction, which is used for "asm goto", can support "asm goto
      with outputs."
      
      This patch removes all restrictions against "callbr" returning values. The
      heavy lifting is done by the code generator. The "INLINEASM_BR" instruction's
      a terminator, and the code generator doesn't allow non-terminator instructions
      after a terminator. In order to correctly model the feature, we need to copy
      outputs from "INLINEASM_BR" into virtual registers. Of course, those copies
      aren't terminators.
      
      To get around this issue, we split the block containing the "INLINEASM_BR"
      right before the "COPY" instructions. This results in two cheats:
      
        - Any physical registers defined by "INLINEASM_BR" need to be marked as
          live-in into the block with the "COPY" instructions. This violates an
          assumption that physical registers aren't marked as "live-in" until after
          register allocation. But it seems as if the live-in information only
          needs to be correct after register allocation. So we're able to get away
          with this.
      
        - The indirect branches from the "INLINEASM_BR" are moved to the "COPY"
          block. This is to satisfy PHI nodes.
      
      I've been told that MLIR can support this handily, but until we're able to
      use it, we'll have to stick with the above.
      
      Reviewers: jyknight, nickdesaulniers, hfinkel, MaskRay, lattner
      
      Reviewed By: nickdesaulniers, MaskRay, lattner
      
      Subscribers: rriddle, qcolombet, jdoerfert, MatzeB, echristo, MaskRay, xbolva00, aaron.ballman, cfe-commits, JonChesterfield, hiraditya, llvm-commits, rnk, craig.topper
      
      Tags: #llvm, #clang
      
      Differential Revision: https://reviews.llvm.org/D69868
      23c2a5ce
    • Sourabh Singh Tomar's avatar
      [DebugInfo]: Refactored Macinfo section consumption part to allow future · 226bddce
      Sourabh Singh Tomar authored
      macro section dumping.
      
      Summary: Previously macinfo infrastructure was using functions
      names that were ambiguous i.e `getMacro/getMacroDWO` in a sense
      of conveying stated intentions. This patch refactored them into more
      reasonable `getDebugMacinfo/getDebugMacinfoDWO` names thus making
      room for macro implementation.
      
      Reviewers: aprantl, probinson, jini.susan.george, dblaikie
      
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D75037
      226bddce
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Fix incorrect VOP3P fneg folding · 0b46b078
      Matt Arsenault authored
      We use some s32 values in VOP3P operands, and won't see any
      intervening casts from a 32-bit fneg. Make sure it's really a packed
      fneg before folding.
      0b46b078
    • Matt Arsenault's avatar
      GlobalISel: Reimplement fewerElementsVectorBasic · 11e3dde6
      Matt Arsenault authored
      Changes the handling of odd breakdowns, and avoids using
      G_EXTRACT/G_INSERT. Pad with undef to a wider size, and unmerge. Also
      avoid introducing instructions for the fully undef components.
      11e3dde6