1. Dec 19, 2020
    • Craig Topper's avatar
      [RISCV] Add intrinsics for vmv.x.s and vmv.s.x · 86d282ba
      Craig Topper authored
      This adds intrinsics for vmv.x.s and vmv.s.x.
      
      I've used stricter type constraints on these intrinsics than what we've been doing on the arithmetic intrinsics so far. This will allow us to not need to pass the scalar type to the Intrinsic::getDeclaration call when creating these intrinsics.
      
      A custom ISD is used for vmv.x.s in order to implement the change in computeNumSignBitsForTargetNode which can remove sign extends on the result.
      
      I also modified the MC layer description of these instructions to show the tied source/dest operand. This is different than what we do for masked instructions where we drop the tied source operand when converting to MC. But it is a more accurate description of the instruction. We can't do this for masked instructions since we use the same MC instruction for masked and unmasked. Tools like llvm-mca operate in the MC layer and rely on ins/outs and Uses/Defs for analysis so I don't know if we'll be able to maintain the current behavior for masked instructions. So I went with the accurate description here since it was easy.
      
      Reviewed By: frasercrmck
      
      Differential Revision: https://reviews.llvm.org/D93365
      86d282ba
    • Kazu Hirata's avatar
      [GVNHoist] Remove successorDominate (NFC) · 5ac37725
      Kazu Hirata authored
      The function was introduced on Aug 25, 2016 in commit
      5f0d0e60.
      
      Its last use was removed on Sep 13, 2017 in commit
      dfa8741c.
      5ac37725
    • Roman Lebedev's avatar
      [InstCombine] Canonicalize SPF to abs intrinsic · 897c985e
      Roman Lebedev authored
      This patch enables canonicalization of SPF_ABS and SPF_ABS
      to the abs intrinsic.
      
      This is a recommit, the original try was
      05d4c4eb,
      but it was reverted due to an apparent miscompile,
      which since then has just been fixed by the previous commit.
      
      Differential Revision: https://reviews.llvm.org/D87188
      897c985e
    • Roman Lebedev's avatar
      [InstSimplify] Don't miscompile `X == 0 ? abs(X) : -abs(X) --> -abs(X)` xform · e9289dc2
      Roman Lebedev authored
      The transform wasn't checking that the LHS of the comparison
      *is* the `X` in question...
      This is the miscompile that was holding up D87188.
      
      Thanks to Dave Green for producing an actionable reproducer!
      e9289dc2
    • Roman Lebedev's avatar
      [NFC][InstSimplify] Add miscompiled testcase from D87188/D87197 · 9b183a14
      Roman Lebedev authored
      Thanks to Dave Green for producing an actionable reproducer!
      It is (obviously) a miscompile:
      ```
      ----------------------------------------
      define i32 @select_abs_of_abs_eq_wrong(i32 %x, i32 %y) {
      %0:
        %abs = abs i32 %x, 0
        %neg = sub i32 0, %abs
        %cmp = icmp eq i32 %y, 0
        %sel = select i1 %cmp, i32 %neg, i32 %abs
        ret i32 %sel
      }
      =>
      define i32 @select_abs_of_abs_eq_wrong(i32 %x, i32 %y) {
      %0:
        %abs = abs i32 %x, 0
        ret i32 %abs
      }
      Transformation doesn't verify!
      ERROR: Value mismatch
      
      Example:
      i32 %x = #xe0000000 (3758096384, -536870912)
      i32 %y = #x00000000 (0)
      
      Source:
      i32 %abs = #x20000000 (536870912)
      i32 %neg = #xe0000000 (3758096384, -536870912)
      i1 %cmp = #x1 (1)
      i32 %sel = #xe0000000 (3758096384, -536870912)
      
      Target:
      i32 %abs = #x20000000 (536870912)
      Source value: #xe0000000 (3758096384, -536870912)
      Target value: #x20000000 (536870912)
      
      Alive2: Transform doesn't verify!
      
      ```
      9b183a14
    • Chih-Ping Chen's avatar
      [DebugInfo] Support Fortran 'use <external module>' statement. · 5f75dcf5
      Chih-Ping Chen authored
      The main change is to add a 'IsDecl' field to DIModule so
      that when IsDecl is set to true, the debug info entry generated
      for the module would be marked as a declaration. That way, the debugger
      would look up the definition of the module in the gloabl scope.
      
      Please see the comments in llvm/test/DebugInfo/X86/dimodule.ll
      for what the debug info entries would look like.
      
      Differential Revision: https://reviews.llvm.org/D93462
      5f75dcf5
    • Björn Schäpers's avatar
      5e5ef535
    • diggerlin's avatar
      d551e40f
    • Florian Hahn's avatar
      Revert "[BasicAA] Handle two unknown sizes for GEPs" · a74941da
      Florian Hahn authored
      Temporarily revert commit 8b1c4e31.
      
      After 8b1c4e31 the compile-time for `MultiSource/Benchmarks/MiBench/consumer-lame`
      dramatically increases with -O3 & LTO, causing issues for builders with
      that configuration.
      
      I filed PR48553 with a smallish reproducer that shows a 10-100x compile
      time increase.
      a74941da
    • Craig Topper's avatar
      [RISCV] Add intrinsics for vmv.v.v, vmv.v.x, and vmv.x.i · fc7b7fc0
      Craig Topper authored
      
      
      We work with @rogfer01 from BSC to come out this patch.
      
      Authored-by: default avatarRoger Ferrer Ibanez <rofirrim@gmail.com>
      Co-Authored-by: default avatarCraig Topper <craig.topper@sifive.com>
      
      Differential Revision: https://reviews.llvm.org/D93514
      fc7b7fc0
    • Kevin P. Neal's avatar
      Revert "Revert "[FPEnv] Teach the IRBuilder about invoke's correct use of the strictfp attribute."" · 7fef551c
      Kevin P. Neal authored
      Similar to D69312, and documented in D69839, the IRBuilder needs to add
      the strictfp attribute to invoke instructions when constrained floating
      point is enabled.
      
      This is try 2, with the test corrected.
      
      Differential Revision: https://reviews.llvm.org/D93134
      7fef551c
    • Whitney Tsang's avatar
      Ensure SplitEdge to return the new block between the two given blocks · 2a814cd9
      Whitney Tsang authored
      This PR implements the function splitBasicBlockBefore to address an
      issue
      that occurred during SplitEdge(BB, Succ, ...), inside splitBlockBefore.
      The issue occurs in SplitEdge when the Succ has a single predecessor
      and the edge between the BB and Succ is not critical. This produces
      the result ‘BB->Succ->New’. The new function splitBasicBlockBefore
      was added to splitBlockBefore to handle the issue and now produces
      the correct result ‘BB->New->Succ’.
      
      Below is an example of splitting the block bb1 at its first instruction.
      
      /// Original IR
      bb0:
      	br bb1
      bb1:
              %0 = mul i32 1, 2
      	br bb2
      bb2:
      /// IR after splitEdge(bb0, bb1) using splitBasicBlock
      bb0:
      	br bb1
      bb1:
      	br bb1.split
      bb1.split:
              %0 = mul i32 1, 2
      	br bb2
      bb2:
      /// IR after splitEdge(bb0, bb1) using splitBasicBlockBefore
      bb0:
      	br bb1.split
      bb1.split
      	br bb1
      bb1:
              %0 = mul i32 1, 2
      	br bb2
      bb2:
      
      Differential Revision: https://reviews.llvm.org/D92200
      2a814cd9
    • Kazu Hirata's avatar
    • Craig Blackmore's avatar
      [RegisterScavenging] Fix assert in scavengeRegisterBackwards · 698ae90f
      Craig Blackmore authored
      According to the documentation, if a spill is required to make a
      register available and AllowSpill is false, then NoRegister should be
      returned, however, this scenario was actually triggering an assertion
      failure.
      
      This patch moves the assertion after the handling of AllowSpill.
      
      Authored by: Lewis Revill
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D92104
      698ae90f
    • Arnamoy Bhattacharyya's avatar
      [SROA] Remove Dead Instructions while creating speculative instructions · 06d5b1c9
      Arnamoy Bhattacharyya authored
      The SROA pass tries to be lazy for removing dead instructions that are collected during iterative run of the pass in the DeadInsts list.  However it does not remove instructions from the dead list while running eraseFromParent() on those instructions.
      
      This causes (rare) null pointer dereferences.  For example, in the speculatePHINodeLoads() instruction, in the following code snippet:
      
      ```
         while (!PN.use_empty()) {
           LoadInst *LI = cast<LoadInst>(PN.user_back());
           LI->replaceAllUsesWith(NewPN);
           LI->eraseFromParent();
         }
      ```
      
      If the Load instruction LI belongs to the DeadInsts list, it should be removed when eraseFromParent() is called.  However, the bug does not show up in most cases, because immediately in the same function, a new LoadInst is created in the following line:
      
      ```
      LoadInst *Load = PredBuilder.CreateAlignedLoad(
               LoadTy, InVal, Alignment,
               (PN.getName() + ".sroa.speculate.load." + Pred->getName()));
      ```
      
      This new LoadInst object takes the same memory address of the just deleted LI using eraseFromParent(), therefore the bug does not materialize.  In very rare cases, the addresses differ and therefore, a dangling pointer is created, causing a crash.
      
      Reviewed By: lebedev.ri
      
      Differential Revision: https://reviews.llvm.org/D92431
      06d5b1c9
    • Fangrui Song's avatar
      [ELF] Rename R_TLS to R_TPREL and R_NEG_TLS to R_TPREL_NEG. NFC · 22c1bd57
      Fangrui Song authored
      The scope of R_TLS (TP offset relocation types (TPREL/TPOFF) used for the
      local-exec TLS model) is actually narrower than its name may imply. R_TLS_NEG
      is only used by Solaris R_386_TLS_LE_32.
      
      Rename them so that they will be less confusing.
      
      Reviewed By: grimar, psmith, rprichard
      
      Differential Revision: https://reviews.llvm.org/D93467
      22c1bd57
    • Nicolas Vasilache's avatar
      [mlir][Linlag] Reflow Linalg.md - NFC · b88ed4ec
      Nicolas Vasilache authored
      Markdown formatting seems to now be available, reflowing the doc without changing any content.
      b88ed4ec
    • David Green's avatar
      [ARM] Match dual lane vmovs from insert_vector_elt · e1c1adf9
      David Green authored
      MVE has a dual lane vector move instruction, capable of moving two
      general purpose registers into lanes of a vector register. They look
      like one of:
        vmov q0[2], q0[0], r2, r0
        vmov q0[3], q0[1], r3, r1
      They only accept these lane indices though (and only insert into an
      i32), either moving lanes 1 and 3, or 0 and 2.
      
      This patch adds some tablegen patterns for them, selecting from vector
      inserts elements. Because the insert_elements are know to be
      canonicalized to ascending order there are several patterns that we need
      to select. These lane indices are:
      
      3 2 1 0    -> vmovqrr 31; vmovqrr 20
      3 2 1      -> vmovqrr 31; vmov 2
      3 1        -> vmovqrr 31
      2 1 0      -> vmovqrr 20; vmov 1
      2 0        -> vmovqrr 20
      
      With the top one being the most common. All other potential patterns of
      lane indices will be matched by a combination of these and the
      individual vmov pattern already present. This does mean that we are
      selecting several machine instructions at once due to the need to
      re-arrange the inserts, but in this case there is nothing else that will
      attempt to match an insert_vector_elt node.
      
      This is a recommit of 6cc3d80a after
      fixing the backward instruction definitions.
      e1c1adf9
    • Xun Li's avatar
      Cleanup coro-inline.ll · 4652718e
      Xun Li authored
      Following up with the comments in D92706.
      - Use -passes instead of -enable-new-pm
      - CoroEarly should happen before AlwaysInliner, adjust it.
      - Remove some unnecessary barriers (still kept one)
      - Cleanup unnecessary debug info
      
      Differential Revision: https://reviews.llvm.org/D93342
      4652718e
    • Matt Arsenault's avatar
      PEI: Only call updateLiveness once per function · fd0f5fb8
      Matt Arsenault authored
      This only needs to be called once for the function, and it visits all
      the necessary blocks in the function. It looks like
      631f6b88 accidentally moved this into
      the loop over all save blocks.
      fd0f5fb8
    • Simon Pilgrim's avatar
      [X86] Avoid std::string creation in RecognizableInstr constructor. NFCI. · 94da2cf6
      Simon Pilgrim authored
      The value names in byteFromRec calls are compile time constants - just create StringRef directly instead of via std::string.
      94da2cf6
  2. Dec 18, 2020