1. Feb 06, 2020
    • Shu-Chun Weng's avatar
      [GlobalISel][AArch64] Fix contract cross-bank copies with SIMD instructions · ce963363
      Shu-Chun Weng authored
      contractCrossBankCopyIntoStore() finds the instruction defines the
      source register and uses its output to replace the register. There are,
      however, instructions that have multiple outputs, e.g. G_UNMERGE_VALUES.
      Current implementation hardcodes to operand 0 and has no way of knowing
      which output should be used.
      
      This change adds another function to directly return the register that
      is the source of the register and use that for folding.
      
      This fixes https://bugs.llvm.org/show_bug.cgi?id=44783
      
      Differential Revision: https://reviews.llvm.org/D74005
      ce963363
    • David Green's avatar
      f64b3466
    • River Riddle's avatar
      [mlir][ODS] Add documentation for the declarative assembly format. · c1bcdb93
      River Riddle authored
      Summary: This details the structure of the format, it's requirements, and gives a few examples.
      
      Differential Revision: https://reviews.llvm.org/D73983
      c1bcdb93
    • Fangrui Song's avatar
      [test] yaml2obj -docnum => --docnum= · 77519b60
      Fangrui Song authored
      77519b60
    • Matt Arsenault's avatar
      GlobalISel: Assume G_INTRINSIC* are convergent · ccc11a9f
      Matt Arsenault authored
      This is safer in case anyone tries to run MI optimization passes on
      pre-selected MIR. If there turns out to be a real reason to do this,
      we might need to add separate convergent intrinsic opcodes.
      ccc11a9f
    • LLVM GN Syncbot's avatar
      [gn build] Port fc62b36a · d2182d6c
      LLVM GN Syncbot authored
      d2182d6c
    • Nick Desaulniers's avatar
      [llvm-reduce] add ReduceAttribute delta pass · fc62b36a
      Nick Desaulniers authored
      Summary:
      The output from llvm-reduce still has significantly more attributes than
      bugpoint does.  Teach llvm-reduce to remove attributes.
      
      Reviewers: diegotf, dblaikie, george.burgess.iv
      
      Subscribers: mgorny, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73853
      fc62b36a
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Fold G_ASHR into TB(N)Z bit calculation · 292f7257
      Jessica Paquette authored
      This implements walking over G_ASHR in the same way as `getTestBitOperand` in
      AArch64ISelLowering.
      
      ```
      (tbz (ashr x, c), b) -> (tbz x, b+c) or (tbz x, msb) if b+c is > # bits in x
      ```
      
      Differential Revision: https://reviews.llvm.org/D73933
      292f7257
    • Christopher Tetreault's avatar
      Reapply: [SVE] Fix bug in simplification of scalable vector instructions · b03f3fbd
      Christopher Tetreault authored
      This reverts commit a0544103, reapplying
      commit 31574d38
      b03f3fbd
    • Petr Hosek's avatar
      [CMake] Filter libc++abi and libunwind from runtimes build in MSVC · 9986b88e
      Petr Hosek authored
      These don't build on MSVC at the moment, so filter these out altogether
      from the list of runtimes and print a warning.
      
      Differential Revision: https://reviews.llvm.org/D73812
      9986b88e
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Prefer merge/unmerge ops to legalize TFE · 7bffa972
      Matt Arsenault authored
      These have a better chance of combining with other operations and are
      currently much better supported than G_EXTRACT.
      7bffa972
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Fix one use check in getTestBitReg · a82a28ae
      Jessica Paquette authored
      (1) The check needs to be on the 0th operand of whatever we're folding
      (2) Checks for validity should happen before we change the bit
      
      Fixes a bug which caused MultiSource/Applications/JM/lencod to fail at -O3.
      
      Differential Revision: https://reviews.llvm.org/D74002
      a82a28ae
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Legalize TFE image result loads · e65e6d05
      Matt Arsenault authored
      Rewrite the result register pair into the expected sinigle register
      format in the legalizer.
      
      I'm also operating under the assumption that TFE doesn't apply to
      stores or atomics, but don't know if this is true or not.
      e65e6d05
    • Jonathan Coe's avatar
      [clang-format] Do not merge short C# class definitions into one line · f40a7972
      Jonathan Coe authored
      Summary: Skip access specifiers before record definitions when deciding whether
      or not to wrap lines so that C# class definitions do not get wrapped into a
      single line.
      
      Reviewers: krasimir, MyDeveloperDay
      
      Reviewed By: krasimir
      
      Tags: #clang-format
      
      Differential Revision: https://reviews.llvm.org/D74050
      f40a7972
    • Hiroshi Yamauchi's avatar
      [PGO][PGSO] Tune flags for profile guided size optimization. · b70f23f5
      Hiroshi Yamauchi authored
      Summary:
      Tune the profile threshold flag value for instrumentation PGO based on internal
      benchmarks.
      
      Also, add flags to allow profile guided size optimizations for non-cold code
      to be enabled separately for instrumentation and sample PGSO.
      
      Neither changes the default behavior (yet) as it's disabled for non-cold code.
      
      Reviewers: davidxl
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D72937
      b70f23f5
    • Michał Górny's avatar
      [lldb] [test] Pass LLVM_LIBS_DIR from CMake for linking liblldb · dcab9736
      Michał Górny authored
      Pass the correct library directory from CMake to dotest.py when linking
      liblldb, instead of trying to reconstruct the path from executable path.
      This fixes link failures on platforms having non-null
      LLVM_LIBDIR_SUFFIX.
      
      Differential Revision: https://reviews.llvm.org/D73767
      dcab9736
    • Matt Arsenault's avatar
      AMDGPU: Fix divergence analysis of control flow intrinsics · 096cd991
      Matt Arsenault authored
      The mask results of these should be uniform. The trickier part is the
      dummy booleans used as IR glue need to be treated as divergent. This
      should make the divergence analysis results correct for the IR the DAG
      is constructed from.
      
      This should allow us to eliminate requiresUniformRegister, which has
      an expensive, recursive scan over all users looking for control flow
      intrinsics. This should avoid recent compile time regressions.
      096cd991
    • Jonathan Coe's avatar
      [clang-format] Do not treat C# attribute targets as labels · ca1fd460
      Jonathan Coe authored
      Summary: Merge '[', 'target' , ':' into a single token for C# attributes to
      prevent the target from being seen as a label.
      
      Reviewers: MyDeveloperDay, krasimir
      
      Reviewed By: krasimir
      
      Tags: #clang-format
      
      Differential Revision: https://reviews.llvm.org/D74043
      ca1fd460
    • Jordan Rupprecht's avatar
      9f507bfd
    • Kazu Hirata's avatar
      Resubmit^2: [JumpThreading] Thread jumps through two basic blocks · 4698bf14
      Kazu Hirata authored
      This reverts commit 41784bed.
      
      Since the original revision ead81592,
      this revision fixes three issues:
      
      - This revision fixes the Windows build.  My original patch improperly
        copied EH pads on Windows.  This patch disregards jump threading
        opportunities having to do with EH pads.
      
      - This revision fixes jump threading to a wrong destination.
        Specifically, my original patch treated any Constant other than 0 as 1
        while evaluating the branch condition.  This bug led to treating
        constant expressions like:
      
          icmp ugt i8* null, inttoptr (i64 4 to i8*)
      
        to "true".  This patch fixes the bug by calling isOneValue.
      
      - This revision fixes the cost calculation of two basic blocks being
        threaded through.  Note that getJumpThreadDuplicationCost returns
        "(unsigned)~0" for those basic blocks that cannot be duplicated.  If
        we sum of two return values from getJumpThreadDuplicationCost, we
        could have an unsigned overflow like:
      
          (unsigned)~0 + 5 = 4
      
        and mistakenly determine that it's safe and profitable to proceed
        with the jump threading opportunity.  The patch fixes the bug by
        checking each return value before summing them up.
      
      [JumpThreading] Thread jumps through two basic blocks
      
      Summary:
      This patch teaches JumpThreading.cpp to thread through two basic
      blocks like:
      
        bb3:
          %var = phi i32* [ null, %bb1 ], [ @a, %bb2 ]
          %tobool = icmp eq i32 %cond, 0
          br i1 %tobool, label %bb4, label ...
      
        bb4:
          %cmp = icmp eq i32* %var, null
          br i1 %cmp, label bb5, label bb6
      
      by duplicating basic blocks like bb3 above.  Once we duplicate bb3 as
      bb3.dup and redirect edge bb2->bb3 to bb2->bb3.dup, we have:
      
        bb3:
          %var = phi i32* [ @a, %bb2 ]
          %tobool = icmp eq i32 %cond, 0
          br i1 %tobool, label %bb4, label ...
      
        bb3.dup:
          %var = phi i32* [ null, %bb1 ]
          %tobool = icmp eq i32 %cond, 0
          br i1 %tobool, label %bb4, label ...
      
        bb4:
          %cmp = icmp eq i32* %var, null
          br i1 %cmp, label bb5, label bb6
      
      Then the existing code in JumpThreading.cpp can thread edge
      bb3.dup->bb4 through bb4 and eventually create bb3.dup->bb5.
      
      Reviewers: wmi
      
      Subscribers: hiraditya, jfb, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D70247
      4698bf14
    • Alina Sbirlea's avatar
      [IRCE] Make IRCE a Function pass. · 67904db2
      Alina Sbirlea authored
      Summary: Make InductiveRangeCheckElimination a FunctionPass.
      
      Reviewers: reames, mkazantsev
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73592
      67904db2
    • Stephan Herhut's avatar
      [MLIR][GPU] Fix build files for mlir-opt. · 921d4e7c
      Stephan Herhut authored
      The recent refactoring of build files broke building with the MIR CUDA
      integration enabled. This fixes it by adding some additional
      dependencies to mlir-opt.
      
      Differential Revision: https://reviews.llvm.org/D74041
      921d4e7c
    • LLVM GN Syncbot's avatar
      [gn build] Port b198f16e · 622ef91b
      LLVM GN Syncbot authored
      622ef91b
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Legalize llvm.amdgcn.s.buffer.load · 69cc9f30
      Matt Arsenault authored
      The 96-bit results need to be widened.
      
      I find the interaction between LegalizerHelper and MIRBuilder somewhat
      awkward. The custom legalization is called by the LegalizerHelper, but
      then does not have access to the helper. You have to construct a new
      helper, which then does not own the MachineIRBuilder, but does modify
      it. Maybe custom legalization should be passed the helper?
      69cc9f30
    • Teresa Johnson's avatar
      [WPD/LowerTypeTests] Delay lowering/removal of type tests until after ICP · 748bb5a0
      Teresa Johnson authored
      Summary:
      Currently type test assume sequences inserted for devirtualization are
      removed during WPD. This patch delays their removal until later in the
      optimization pipeline. This is an enabler for upcoming enhancements to
      indirect call promotion, for example streamlined promotion guard
      sequences that compare against vtable address instead of the target
      function, when there are small number of possible vtables (either
      determined via WPD or by in-progress type profiling). We need the type
      tests to correlate the callsites with the address point offset needed in
      the compare sequence, and optionally to associated type summary info
      computed during WPD.
      
      This depends on work in D71913 to enable invocation of LowerTypeTests to
      drop type test assume sequences, which will now be invoked following ICP
      in the ThinLTO post-LTO link pipelines, and also after the existing
      export phase LowerTypeTests invocation in regular LTO (which is already
      after ICP). We cannot simply move the existing import phase
      LowerTypeTests pass later in the ThinLTO post link pipelines, as the
      comment in PassBuilder.cpp notes (it must run early because when
      performing CFI other passes may disturb the sequences it looks for).
      
      This necessitated adding a new type test resolution "Unknown" that we
      can use on the type test assume sequences previously removed by WPD,
      that we now want LTT to ignore.
      
      Depends on D71913.
      
      Reviewers: pcc, evgeny777
      
      Subscribers: mehdi_amini, Prazek, hiraditya, steven_wu, dexonsmith, arphaman, davidxl, cfe-commits, llvm-commits
      
      Tags: #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D73242
      748bb5a0
    • Adam Balogh's avatar
      [Analyzer] Model STL Algoirthms to improve the iterator checkers · b198f16e
      Adam Balogh authored
      STL Algorithms are usually implemented in a tricky for performance
      reasons which is too complicated for the analyzer. Furthermore inlining
      them is costly. Instead of inlining we should model their behavior
      according to the specifications.
      
      This patch is the first step towards STL Algorithm modeling. It models
      all the `find()`-like functions in a simple way: the result is either
      found or not. In the future it can be extended to only return success if
      container modeling is also extended in a way the it keeps track of
      trivial insertions and deletions.
      
      Differential Revision: https://reviews.llvm.org/D70818
      b198f16e
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Fix processing new phi in waterfall loop · 307e0d54
      Matt Arsenault authored
      The adjusted iterator range included the last we just inserted, and
      don't want to process. Figure out the new iterator range before
      inserting phis. This was a harmless problem, but added an unnecessary
      complication for a future patch.
      307e0d54
    • Matt Arsenault's avatar
      GlobalISel: Make LegalizerHelper primitives public · cc1cffbe
      Matt Arsenault authored
      I want to re-use widenScalarDst/moreElementsVectorDst directly.
      cc1cffbe
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Don't use legal v2s16 G_BUILD_VECTOR · dfa9420f
      Matt Arsenault authored
      If we have s_pack_* instructions, legalize this to
      G_BUILD_VECTOR_TRUNC from s32 elements. This is closer to how how the
      s_pack_* instructions really behave.
      
      If we don't have s_pack_ instructions, expand this by creating a merge
      to s32 and bitcasting. This expands to the expected bit operations. I
      think this eventually should go in a new bitcast legalize action type
      in LegalizerHelper.
      
      We already directly emit the shift operations in RegBankSelect for the
      vector case. This could possibly be cleaned up, but I also may want to
      defer doing this expansion to selection anyway. I'll see about that
      when I try to actually match VOP3P instructions.
      
      This breaks the selection of the build_vector since tablegen doesn't
      know how to match G_BUILD_VECTOR_TRUNC yet, so just xfail it for now.
      dfa9420f
    • Med Ismail Bennani's avatar
      [lldb/Target] Add Assert StackFrame Recognizer · 2b7f3289
      Med Ismail Bennani authored
      When a thread stops, this checks depending on the platform if the top frame is
      an abort stack frame. If so, it looks for an assert stack frame in the upper
      frames and set it as the most relavant frame when found.
      
      To do so, the StackFrameRecognizer class holds a "Most Relevant Frame" and a
      "cooked" stop reason description. When the thread is about to stop, it checks
      if the current frame is recognized, and if so, it fetches the recognized frame's
      attributes and applies them.
      
      rdar://58528686
      
      Differential Revision: https://reviews.llvm.org/D73303
      
      
      
      Signed-off-by: default avatarMed Ismail Bennani <medismail.bennani@gmail.com>
      2b7f3289
    • Momchil Velikov's avatar
      [ARM][TargetParser] Improve handling of dependencies between target features · 3627c91e
      Momchil Velikov authored
      The patch at https://reviews.llvm.org/D64048 added "negative"
      dependency handling in `ARM::appendArchExtFeatures`: feature "noX"
      removes all features, which imply "X".
      
      This patch adds the "positive" handling: feature "X" adds all the
      feature strings implied by "X".
      
      (This patch also comes from the suggestion here
      https://reviews.llvm.org/D72633#inline-658582)
      
      Differential Revision: https://reviews.llvm.org/D72762
      3627c91e
    • Sven van Haastregt's avatar
      [OpenCL] Fix tblgen support for cl_khr_mipmap_image_writes · 91b3083a
      Sven van Haastregt authored
      Apply the fix of f780e15c ("[OpenCL] Fix support for
      cl_khr_mipmap_image_writes", 2020-01-27) also to the TableGen OpenCL
      builtin function definitions.
      91b3083a
  2. Feb 05, 2020