1. Feb 26, 2020
    • Jim Lin's avatar
      [ARC][NFC] Remove trailing space · f6603aed
      Jim Lin authored
      f6603aed
    • Greg Clayton's avatar
      Add a llvm-gsymutil tool that can convert object files to GSYM and perform lookups. · 2f6cc21f
      Greg Clayton authored
      Summary:
      This patch creates the llvm-gsymutil binary that can convert object files to GSYM using the --convert <path> option. It can also dump and lookup addresses within GSYM files that have been saved to disk.
      
      To dump a file:
      
      llvm-gsymutil /path/to/a.gsym
      
      To perform address lookups, like with atos, on GSYM files:
      
      llvm-gsymutil --address 0x1000 --address 0x1100 /path/to/a.gsym
      
      To convert a mach-o or ELF file, including any DWARF debug info contained within the object files:
      
      llvm-gsymutil --convert /path/to/a.out --out-file /path/to/a.out.gsym
      
      Conversion highlights:
      - convert DWARF debug info in mach-o or ELF files to GSYM
      - convert symbols in symbol table to GSYM and don't convert symbols that overlap with DWARF debug info
      - extract UUID from object files
      - extract .text (read + execute) section address ranges and filter out any DWARF or symbols that don't fall in those ranges.
      - if .text sections are extracted, and if the last gsym::FunctionInfo object has no size, cap the size to the end of the section the function was contained in
      
      Dumping GSYM files will dump all sections of the GSYM file in textual format.
      
      Reviewers: labath, aadsm, serhiy.redko, jankratochvil, xiaobai, wallace, aprantl, JDevlieghere, jdoerfert
      
      Subscribers: mgorny, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D74883
      2f6cc21f
    • Juneyoung Lee's avatar
      [SimpleLoopUnswitch] Fix introduction of UB when hoisted condition may be undef or poison · 181628b5
      Juneyoung Lee authored
      Summary:
      Loop unswitch hoists branches on loop-invariant conditions. However, if this
      condition is poison/undef and the branch wasn't originally reachable, loop
      unswitch introduces UB (since the optimized code will branch on poison/undef and
      the original one didn't)).
      We fix this problem by freezing the condition to ensure we don't introduce UB.
      
      We will now transform the following:
        while (...) {
          if (C) { A }
          else   { B }
        }
      
      Into:
        C' = freeze(C)
        if (C') {
          while (...) { A }
        } else {
          while (...) { B }
        }
      
      This patch fixes the root cause of the following bug reports (which use the old loop unswitch, but can be reproduced with minor changes in the code and -enable-nontrivial-unswitch):
      - https://llvm.org/bugs/show_bug.cgi?id=27506
      - https://llvm.org/bugs/show_bug.cgi?id=31652
      
      Reviewers: reames, majnemer, chenli, sanjoy, hfinkel
      
      Reviewed By: reames
      
      Subscribers: hiraditya, jvesely, nhaehnle, filcab, regehr, trentxintong, nlopes, llvm-commits, mzolotukhin
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D29015
      181628b5
    • Kang Zhang's avatar
      [PowerPC] Fix the unexpected modification caused by D62993 in LowerSELECT_CC for power9 · b083d7a3
      Kang Zhang authored
      Summary:
      The patch D62993 : `[PowerPC] Emit scalar min/max instructions with unsafe fp math`
      has modified the functionality when `Subtarget.hasP9Vector() && (!HasNoInfs || !HasNoNaNs)`,
       this modification is not expected.
      
      Reviewed By: nemanjai
      
      Differential Revision: https://reviews.llvm.org/D74701
      b083d7a3
    • Fangrui Song's avatar
      [MC] Default MCContext::UseNamesOnTempLabels to false and only set it to true for MCAsmStreamer · b61a4aac
      Fangrui Song authored
      Only MCAsmStreamer (assembly output) needs to keep names of temporary labels created by
      MCContext::createTempSymbol().
      
      This change made the rL236642 optimization available for cc2as and
      probably some other users.
      
      This eliminates a behavior difference between llvm-mc -filetype=obj and cc1as, which caused
      https://reviews.llvm.org/D74006#1890487
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D75097
      b61a4aac
    • Fangrui Song's avatar
      [MC][ARM] Don't create multiple .ARM.exidx associated to one .text · d0c4277d
      Fangrui Song authored
      Fixed an issue exposed by D74006.
      
      In clang cc1as, MCContext::UseNamesOnTempLabels is true.
      When parsing a .fnstart directive, FnStart gets redefined to a temporary symbol of a different name (.Ltmp0, .Ltmp1, ...).
      MCContext::getELFSection() called by SwitchToEHSection() will create a different .ARM.exidx each time.
      
      llvm-mc uses `Ctx.setUseNamesOnTempLabels(false);` and FnStart is unnamed.
      MCContext::getELFSection() called by SwitchToEHSection() will reuse the same .ARM.exidx .
      
      Reviewed By: efriedma
      
      Differential Revision: https://reviews.llvm.org/D75095
      d0c4277d
    • Fangrui Song's avatar
      6fb70c87
    • Nathan James's avatar
      [docs] dump-ast-matchers removes const from Matcher args and handles template... · b653ab0e
      Nathan James authored
      [docs] dump-ast-matchers removes const from Matcher args and handles template functions slightly better
      
      Reviewers: aaron.ballman, gribozavr2, joerg
      
      Reviewed By: aaron.ballman
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D75113
      b653ab0e
    • Reid Kleckner's avatar
      Remove namespace lld { namespace coff { from COFF LLD cpp files · 8a310f40
      Reid Kleckner authored
      Instead, use `using namespace lld(::coff)`, and fully qualify the names
      of free functions where they are defined in cpp files.
      
      This effectively reverts d79c3be6 to follow the new style guide added
      in 236fcbc2.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D74882
      8a310f40
    • Max Moroz's avatar
    • Craig Topper's avatar
      [SelectionDAG][PowerPC][AArch64][X86][ARM] Add chain input and output the ISD::FLT_ROUNDS_ · 735d27dc
      Craig Topper authored
      This node reads the rounding control which means it needs to be ordered properly with operations that change the rounding control. So it needs to be chained to maintain order.
      
      This patch adds a chain input and output to the node and connects it to the chain in SelectionDAGBuilder. I've update all in-tree targets to connect their chain through their lowering code.
      
      Differential Revision: https://reviews.llvm.org/D75132
      735d27dc
    • zoecarver's avatar
      Remove std::shared_ptr::allocate_shared · 28d38a25
      zoecarver authored
      std::shared_ptr::allocate_shared isn't in the standard. This commit removes it from libc++. It updates std::allocate_shared to use __create_with_cntrl_block.
      
      Differential Revision: https://reviews.llvm.org/D66178
      28d38a25
    • Lang Hames's avatar
      [ORC] Remove the JITDylib::SymbolTableEntry::isInMaterializingState() method. · b7aa1cc3
      Lang Hames authored
      It was being used inconsistently. Uses have been replaced with direct checks
      on the symbol state.
      b7aa1cc3
    • Adrian Prantl's avatar
      828fb0c5
    • Nico Weber's avatar
    • Quentin Colombet's avatar
      [GISel][KnownBits] Update a comment regarding the effect of cache on PHIs · 5bf0023b
      Quentin Colombet authored
      Unlike what I claimed in my previous commit. The caching is
      actually not NFC on PHIs.
      
      When we put a big enough max depth, we end up simulating loops.
      The cache is effectively cutting the simulation short and we
      get less information as a result.
      E.g.,
      ```
      v0 = G_CONSTANT i8 0xC0
      jump
      v1 = G_PHI i8 v0, v2
      v2 = G_LSHR i8 v1, 1
      ```
      
      Let say we want the known bits of v1.
      - With cache:
      Set v1 cache to we know nothing
      v1 is v0 & v2
      v0 gives us 0xC0
      v2 gives us known bits of v1 >> 1
      v1 is in the cache
      => v1 is 0, thus v2 is 0x80
      Finally v1 is v0 & v2 => 0x80
      
      - Without cache and enough depth to do two iteration of the loop:
      v1 is v0 & v2
      v0 gives us 0xC0
      v2 gives us known bits of v1 >> 1
      v1 is v0 & v2
      v0 is 0xC0
      v2 is v1 >> 1
      Reach the max depth for v1...
      unwinding
      v1 is know nothing
      v2 is 0x80
      v0 is 0xC0
      v1 is 0x80
      v2 is 0xC0
      v0 is 0xC0
      v1 is 0xC0
      
      Thus now v1 is 0xC0 instead of 0x80.
      
      I've added a unittest demonstrating that.
      
      NFC
      5bf0023b
    • aartbik's avatar
      [mlir] [VectorOps] Add vector.print to EDSC · 3cefebc3
      aartbik authored
      Summary: This prepares using the operation in model builder runner.
      
      Reviewers: nicolasvasilache, andydavis1
      
      Reviewed By: nicolasvasilache
      
      Subscribers: mehdi_amini, rriddle, jpienaar, burmako, shauheen, antiagainst, nicolasvasilache, arpith-jacob, mgester, lucyrfox, liufengdb, Joonsoo, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D75147
      3cefebc3
    • Louis Dionne's avatar
      [NFC][libc++] Refactor some future tests to reduce code duplication · b051cc93
      Louis Dionne authored
      The same test was being repeated over and over again.
      That's what functions are for.
      b051cc93
    • River Riddle's avatar
      [mlir][DenseElementsAttr] Fix storage size for bfloat16 when parsing from hex. · b3e6487f
      River Riddle authored
      Summary: bfloat16 is stored internally as a double, so we can't direct use Type::getIntOrFloatBitWidth.
      
      Differential Revision: https://reviews.llvm.org/D75133
      b3e6487f
    • Jason Molenda's avatar
      Re-land Unwind past an interrupt handler correctly on arm or at pc==0 · 4b2b8b96
      Jason Molenda authored
      Updated the patch to only fetch $pc on a Return Address-using
      target only if we're in a trap frame *and* if there is a saved
      location for $pc in the trap frame's unwind rules.  If not,
      we fall back to fetching the Return Address register (eg $lr).
      
      Original commit msg:
      
          Unwind past an interrupt handler correctly on arm or at pc==0
      
          Fix RegisterContextLLDB::InitializeNonZerothFrame so that it
          will fetch a FullUnwindPlan instead of falling back to the
          architectural default unwind plan -- GetFullUnwindPlan knows
          how to spot a jmp 0x0 that results in a fault, which may be
          the case when we see a trap handler on the stack.
      
          Fix RegisterContextLLDB::SavedLocationForRegister so that when
          the pc value is requested from a trap handler frame, where we
          have a complete register context available to us, don't provide
          the Return Address register (lr) instead of the pc.  We have
          an actual pc value here, and it's pointing to the instruction
          that faulted.
      
          Differential revision: https://reviews.llvm.org/D75007
          <rdar://problem/59416588>
      4b2b8b96
    • Louis Dionne's avatar
      [libc++] Avoid including <semaphore.h> on Apple · 3b5530cf
      Louis Dionne authored
      It turns out that <semaphore.h> is not well-behaved, as it transitively
      includes <sys/param.h>, and that one defines several non-reserved macros
      that clash with some downstream projects in modular builds. For the time
      being, using <sys/semaphore.h> instead gives us the declarations we need
      without the macros.
      
      rdar://59744472
      3b5530cf
    • Vedant Kumar's avatar
      Revert "[X86MCTargetDesc.h] Speculative fix for macro collision with sys/param.h" · 8594f3d8
      Vedant Kumar authored
      This reverts commit eee22ec3.
      
      This is not the correct fix, the root cause seems to be a bug in the
      stage1 host clang compiler. See https://reviews.llvm.org/D75091 for more
      discussion.
      8594f3d8
    • Roman Lebedev's avatar
      [clang] Annotating C++'s `operator new` with more attributes · 3dd5a298
      Roman Lebedev authored
      Summary:
      Right now we annotate C++'s `operator new` with `noalias` attribute,
      which very much is healthy for optimizations.
      
      However as per [[ http://eel.is/c++draft/basic.stc.dynamic.allocation | `[basic.stc.dynamic.allocation]` ]],
      there are more promises on global `operator new`, namely:
      * non-`std::nothrow_t` `operator new` *never* returns `nullptr`
      * If `std::align_val_t align` parameter is taken, the pointer will also be `align`-aligned
      * ~~global `operator new`-returned pointer is `__STDCPP_DEFAULT_NEW_ALIGNMENT__`-aligned ~~ It's more caveated than that.
      
      Supplying this information may not cause immediate landslide effects
      on any specific benchmarks, but it for sure will be healthy for optimizer
      in the sense that the IR will better reflect the guarantees provided in the source code.
      
      The caveat is `-fno-assume-sane-operator-new`, which currently prevents emitting `noalias`
      attribute, and is automatically passed by Sanitizers ([[ https://bugs.llvm.org/show_bug.cgi?id=16386 | PR16386 ]]) - should it also cover these attributes?
      The problem is that the flag is back-end-specific, as seen in `test/Modules/explicit-build-flags.cpp`.
      But while it is okay to add `noalias` metadata in backend, we really should be adding at least
      the alignment metadata to the AST, since that allows us to perform sema checks on it.
      
      Reviewers: erichkeane, rjmccall, jdoerfert, eugenis, rsmith
      
      Reviewed By: rsmith
      
      Subscribers: xbolva00, jrtc27, atanasyan, nlopes, cfe-commits
      
      Tags: #llvm, #clang
      
      Differential Revision: https://reviews.llvm.org/D73380
      3dd5a298
    • Roman Lebedev's avatar
      [Sema] Perform call checking when building CXXNewExpr · b8fdafe6
      Roman Lebedev authored
      Summary:
      There was even a TODO for this.
      The main motivation is to make use of call-site based
      `__attribute__((alloc_align(param_idx)))` validation (D72996).
      
      Reviewers: rsmith, erichkeane, aaron.ballman, jdoerfert
      
      Reviewed By: rsmith
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D73020
      b8fdafe6
    • Cyndy Ishida's avatar
      [llvm][TextAPI] rename test vars, NFC · 6d2372ce
      Cyndy Ishida authored
      * Conforms to clang tidy
      6d2372ce
    • Johannes Doerfert's avatar
      [OpenMP][Opt] Combine `struct ident_t*` during deduplication · 396b7253
      Johannes Doerfert authored
      If we deduplicate OpenMP runtime calls we have multiple `ident_t*` that
      represent information like source location. So far, we simply kept the
      one used by the replacement call. However, as exposed by PR44893, that
      can cause problems if we have stack allocated `ident_t` objects. While
      we need to revisit the use of these as well, it is clear that we
      eventually want to merge source location information in some way. With
      this patch we add the infrastructure to do so but without doing the
      actual merge. Instead we pick a global `ident_t` from the replaced
      calls, if possible, or create a new one with an unknown location
      instead.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D74925
      396b7253
    • Thomas Lively's avatar
      [WebAssembly] Simplify extract_vector lowering · 0906dca4
      Thomas Lively authored
      Summary:
      Removes patterns that were not doing useful work, changes the
      default extract instructions to be the unsigned versions now that
      they are enabled by default, fixes PR44988, and adds tests for
      sext_inreg lowering.
      
      Reviewers: aheejin
      
      Reviewed By: aheejin
      
      Subscribers: dschuff, sbc100, jgravelle-google, hiraditya, sunfish, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D75005
      0906dca4
    • Yitzhak Mandelbaum's avatar
      [libTooling] Add function to determine associated text of a declaration. · 9c54f615
      Yitzhak Mandelbaum authored
      Summary:
      This patch adds `getAssociatedRange` which, for a given decl, computes preceding
      and trailing text that would conceptually be associated with the decl by the
      reader. This includes comments, whitespace, and separators like ';'.
      
      Reviewers: gribozavr
      
      Subscribers: cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D72153
      9c54f615
    • Akira Hatanaka's avatar
      [ObjC][ARC] Don't move a retain call living outside a loop into the loop · 430512ed
      Akira Hatanaka authored
      body
      
      We started seeing cases where ARC optimizer would move retain calls into
      loop bodies, causing imbalance in the number of retain and release
      calls, after changes were made to delete inert ARC calls since the inert
      calls that used to block code motion are gone.
      
      Fix the bug by setting the CFG hazard flag when visiting a loop header.
      
      rdar://problem/56908836
      430512ed
    • Alexey Bataev's avatar
      [LIBOMPTARGET]Fix PR44933: fix crash because of the too early deinitialization of libomptarget. · 63cef621
      Alexey Bataev authored
      Summary:
      Instead of using global variables with unpredicted time of
      deinitialization, use dynamically allocated variables with functions
      explicitly marked as global constructor/destructor and priority. This
      allows to prevent the crash because of the incorrect order of dynamic
      libraries deinitialization.
      
      Reviewers: grokos, hfinkel
      
      Subscribers: caomhin, kkwli0, openmp-commits
      
      Tags: #openmp
      
      Differential Revision: https://reviews.llvm.org/D74837
      63cef621
    • Craig Topper's avatar
      [X86] Add test to show incorrect ordering of flt.rounds intrinsic relative to calls to fesetround. · c5ce6d8b
      Craig Topper authored
      We don't order flt.rounds intrinsics relative to side effecting
      operations in SelectionDAG. And we CSE multiple calls because of
      this.
      c5ce6d8b
    • zoecarver's avatar
      Check args passed to __builtin_frame_address and __builtin_return_address. · 6201f660
      zoecarver authored
      Verifies that an argument passed to __builtin_frame_address or __builtin_return_address is within the range [0, 0xFFFF]
      
      Differential revision: https://reviews.llvm.org/D66839
      
      Re-committed after fixed: c93112dc
      6201f660
    • Bill Wendling's avatar
      Use "nop" to avoid size warnings. · 6d0d1a63
      Bill Wendling authored
      6d0d1a63
    • Roman Lebedev's avatar
      [SCEV][IndVars] Always provide insertion point to the SCEVExpander::isHighCostExpansion() · 400ceda4
      Roman Lebedev authored
      Summary: This addresses the `llvm/test/Transforms/IndVarSimplify/elim-extend.ll` `@nestedIV` regression from D73728
      
      Reviewers: reames, mkazantsev, wmi, sanjoy
      
      Reviewed By: mkazantsev
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73777
      400ceda4
    • Roman Lebedev's avatar
      [SCEV] rewriteLoopExitValues(): even if have hard uses, still rewrite if cheap (PR44668) · 44edc6fd
      Roman Lebedev authored
      Summary:
      Replacing uses of IV outside of the loop is likely generally useful,
      but `rewriteLoopExitValues()` is cautious, and if it isn't told to always
      perform the replacement, and there are hard uses of IV in loop,
      it doesn't replace.
      
      In [[ https://bugs.llvm.org/show_bug.cgi?id=44668 | PR44668 ]],
      that prevents `-indvars` from replacing uses of induction variable
      after the loop, which might be one of the optimization failures
      preventing that code from being vectorized.
      
      Instead, now that the cost model is fixed, i believe we should be
      a little bit more optimistic, and also perform replacement
      if we believe it is within our budget.
      
      Fixes [[ https://bugs.llvm.org/show_bug.cgi?id=44668 | PR44668 ]].
      
      Reviewers: reames, mkazantsev, asbirlea, fhahn, skatkov
      
      Reviewed By: mkazantsev
      
      Subscribers: nikic, hiraditya, zzheng, javed.absar, dmgreen, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73501
      44edc6fd
    • Roman Lebedev's avatar
      [SCEV] SCEVExpander::isHighCostExpansionHelper(): cost-model min/max (PR44668) · d6f47aeb
      Roman Lebedev authored
      Summary:
      Previosly we simply always said that `SCEVMinMaxExpr` is too costly to expand.
      But this isn't really true, it expands into just a comparison+swap pair.
      And again much like with add/mul, there will be one less such pair
      than the number of operands. And we need to count the cost of operands themselves.
      
      This does change a number of testcases, and as far as i can tell,
      all of these changes are improvements, in the sense that
      we fixed up more latches to do the [in]equality comparison.
      
      This concludes cost-modelling changes, no other SCEV expressions exist as of now.
      
      This is a part of addressing [[ https://bugs.llvm.org/show_bug.cgi?id=44668 | PR44668 ]].
      
      Reviewers: reames, mkazantsev, wmi, sanjoy
      
      Reviewed By: mkazantsev
      
      Subscribers: hiraditya, javed.absar, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73744
      d6f47aeb
    • Roman Lebedev's avatar
      [SCEV] SCEVExpander::isHighCostExpansionHelper(): cost-model polynomial recurrence · 0f3c9b54
      Roman Lebedev authored
      Summary:
      So, i wouldn't call this *obviously* correct,
      but i think i got it right this time :)
      
      Roughly, we have
      ```
      Op0*x^0 + Op1*x^1 + Op2*x^2 ...
      ```
      where `Op_{n} * x^{n}` is called term, and `n` the degree of term.
      
      Due to the way they are stored internally in `SCEVAddRecExpr`,
      i believe we can have `Op_{n}` to be `0`, so we should not charge for those.
      
      I think it is most straight-forward to count the cost in 4 steps:
      1. First, count it the same way we counted `scAddExpr`, but be sure to skip terms with zero constants.
         Much like with `add` expr we will have one less addition than number of terms.
      2. Each non-constant term (term degree >= 1) requires a multiplication between the `Op_{n}` and `x^{n}`.
         But again, only charge for it if it is required - `Op_{n}` must not be 0 (no term) or 1 (no multiplication needed),
         and obviously don't charge constant terms (`x^0 == 1`).
      3. We must charge for all the `x^0`..`x^{poly_degree}` themselves.
         Since `x^{poly_degree}` is `x * x * ...  * x`, i.e. `poly_degree` `x`'es multiplied,
         for final `poly_degree` term we again require `poly_degree-1` multiplications.
         Note that all the `x^{0}`..`x^{poly_degree-1}` will be computed for the free along the way there.
      4. And finally, the operands themselves.
      
      Here, much like with add/mul exprs, we really don't look for preexisting instructions..
      
      Reviewers: reames, mkazantsev, wmi, sanjoy
      
      Reviewed By: mkazantsev
      
      Subscribers: hiraditya, javed.absar, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73741
      0f3c9b54
    • Roman Lebedev's avatar
      [SCEV] SCEVExpander::isHighCostExpansionHelper(): cost-model add/mul · 756af2f8
      Roman Lebedev authored
      Summary:
      While this resolves the regression from D73722 in `llvm/test/Transforms/IndVarSimplify/exit_value_test2.ll`,
      this now regresses `llvm/test/Transforms/IndVarSimplify/elim-extend.ll` `@nestedIV` test,
      we no longer can perform that expansion within default budget of `4`, but require budget of `6`.
      That regression is being addressed by D73777.
      
      The basic idea here is simple.
      ```
      Op0,  Op1, Op2 ...
       |     |    |
       \--+--/    |
          |       |
          \---+---/
      ```
      I.e. given N operands, we will have N-1 operations,
      so we have to add cost of an add (mul) for **every** Op processed,
      **except** the first one, plus we need to recurse into *every* Op.
      
      I'm guessing there's already canonicalization that ensures we won't
      have `1` operand in `scMulExpr`, and no `0` in `scAddExpr`/`scMulExpr`.
      
      Reviewers: reames, mkazantsev, wmi, sanjoy
      
      Reviewed By: mkazantsev
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73728
      756af2f8
    • Roman Lebedev's avatar
      [SCEV] SCEVExpander::isHighCostExpansionHelper(): cost-model plain UDiv · cc29600b
      Roman Lebedev authored
      Summary:
      If we don't believe this UDiv is actually a LShr in disguise, things are much worse.
      First, we try to see if this UDiv actually originates from user code,
      by looking for `S + 1`, and if found considering this UDiv to be free.
      But otherwise, we always considered this UDiv to be high-cost.
      
      However that is no longer the case with TTI-driven cost model:
      our default budget is 4, which matches the default cost of UDiv,
      so now we allow a single UDiv to not be counted as high-cost.
      
      While that is the case, it is evident this is actually a regression
      due to the fact that cost-modelling is incomplete - we did not account
      for the `add`, `mul` costs yet. That is being addressed in D73728.
      
      Cost-modelling for UDiv also seems pretty straight-forward:
      subtract cost of the UDiv itself, and recurse into both the LHS and RHS.
      
      Reviewers: reames, mkazantsev, wmi, sanjoy
      
      Reviewed By: mkazantsev
      
      Subscribers: hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73722
      cc29600b
    • Roman Lebedev's avatar
      [NFC][IndVarSimplify] Adjust value names in IndVarSimplify/exit_value_test2.ll · b8abdf9a
      Roman Lebedev authored
      %tmp prefix confuses auto-update scripts
      b8abdf9a