1. Oct 02, 2019
    • Adrian Prantl's avatar
      Fix a syntax error. · dffe5dfa
      Adrian Prantl authored
      llvm-svn: 373355
      dffe5dfa
    • Adrian Prantl's avatar
      Fix a condition-flip regression introduced in r373344. · ad08a5f0
      Adrian Prantl authored
      llvm-svn: 373354
      ad08a5f0
    • Adrian Prantl's avatar
      Typo (NFC) · c7f19caa
      Adrian Prantl authored
      llvm-svn: 373353
      c7f19caa
    • Adrian Prantl's avatar
      Simplify condition (NFC) · 771d464f
      Adrian Prantl authored
      llvm-svn: 373352
      771d464f
    • Philip Reames's avatar
      [IndVars] An implementation of loop predication without a need for speculation · 0200626f
      Philip Reames authored
      This patch implements a variation of a well known techniques for JIT compilers - we have an implementation in tree as LoopPredication - but with an interesting twist. This version does not assume the ability to execute a path which wasn't taken in the original program (such as a guard or widenable.condition intrinsic). The benefit is that this works for arbitrary IR from any frontend (including C/C++/Fortran). The tradeoff is that it's restricted to read only loops without implicit exits.
      
      This builds on SCEV, and can thus eliminate the loop varying portion of the any early exit where all exits are understandable by SCEV. A key advantage is that fixing deficiency exposed in SCEV - already found one while writing test cases - will also benefit all of full redundancy elimination (and most other loop transforms).
      
      I haven't seen anything in the literature which quite matches this. Given that, I'm not entirely sure that keeping the name "loop predication" is helpful. Anyone have suggestions for a better name? This is analogous to partial redundancy elimination - since we remove the condition flowing around the backedge - and has some parallels to our existing transforms which try to make conditions invariant in loops.
      
      Factoring wise, I chose to put this in IndVarSimplify since it's a generally applicable to all workloads. I could split this off into it's own pass, but we'd then probably want to add that new pass every place we use IndVars.  One solid argument for splitting it off into it's own pass is that this transform is "too good". It breaks a huge number of existing IndVars test cases as they tend to be simple read only loops.  At the moment, I've opted it off by default, but if we add this to IndVars and enable, we'll have to update around 20 test files to add side effects or disable this transform.
      
      Near term plan is to fuzz this extensively while off by default, reflect and discuss on the factoring issue mentioned just above, and then enable by default.  I also need to give some though to supporting widenable conditions in this framing.
      
      Differential Revision: https://reviews.llvm.org/D67408
      
      llvm-svn: 373351
      0200626f
    • Matt Arsenault's avatar
      AMDGPU/GlobalISel: Increase max legal size to 1024 · 9dba6037
      Matt Arsenault authored
      There are 1024 bit register classes defined for AGPRs. Additionally
      OpenCL defines vectors up to 16 x i64, and this helps those tests
      legalize.
      
      llvm-svn: 373350
      9dba6037
    • Craig Topper's avatar
      [X86] Add a VBROADCAST_LOAD ISD opcode representing a scalar load broadcasted to a vector. · 105e82ed
      Craig Topper authored
      Summary:
      This adds the ISD opcode and a DAG combine to create it. There are
      probably some places where we can directly create it, but I'll
      leave that for future work.
      
      This updates all of the isel patterns to look for this new node.
      I had to add a few additional isel patterns for aligned extloads
      which we should probably fix with a DAG combine or something. This
      does mean that the broadcast load folding for avx512 can no
      longer match a broadcasted aligned extload.
      
      There's still some work to do here for combining a broadcast of
      a broadcast_load. We also need to improve extractelement or
      demanded vector elements of a broadcast_load. I'll try to get
      those done before I submit this patch.
      
      Reviewers: RKSimon, spatel
      
      Reviewed By: RKSimon
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D68198
      
      llvm-svn: 373349
      105e82ed
    • Alexey Bataev's avatar
      [OPENMP]Fix PR43516: Compiler crash with collapse(2) on non-rectangular · 658ad4d4
      Alexey Bataev authored
      loop.
      
      Missed check if the condition is also dependent when building final
      expressions for the collapsed loop directives.
      
      llvm-svn: 373348
      658ad4d4
    • Peter Collingbourne's avatar
      ELF: Add .interp synthetic sections first in createSyntheticSections(). · 0bb825d2
      Peter Collingbourne authored
      Our .interp section is not a SyntheticSection. As a result, it terminates the
      loop in removeUnusedSyntheticSections(). This has at least two consequences:
      
      - The synthetic .bss and .bss.rel.ro sections are always present in
        dynamically linked executables, even when they are not needed.
      - The synthetic .ARM.exidx (and possibly other) sections are always present
        in partitions other than the last one, even when not needed.
        .ARM.exidx in particular is problematic because it assumes that its
        list of code sections is non-empty in getLinkOrderDep(), which can
        lead to a crash if the partition does not have any code sections.
      
      Fix these problems by moving the creation of the .interp sections to the
      top of createSyntheticSections(). While here, make the code a little less
      error-prone by changing the add() lambdas to take a SyntheticSection instead
      of an InputSectionBase.
      
      Differential Revision: https://reviews.llvm.org/D68256
      
      llvm-svn: 373347
      0bb825d2
  2. Oct 01, 2019