1. May 20, 2021
    • Joseph Huber's avatar
      [Diagnostics] Allow emitting analysis and missed remarks on functions · 2db182ff
      Joseph Huber authored
      Summary:
      Currently, only `OptimizationRemarks` can be emitted using a Function.
      Add constructors to allow this for `OptimizationRemarksAnalysis` and
      `OptimizationRemarkMissed` as well.
      
      Reviewed By: jdoerfert thegameg
      
      Differential Revision: https://reviews.llvm.org/D102784
      2db182ff
    • Sanjay Patel's avatar
      [x86] add tests for fma folds with fast-math-flags; NFC · 9b59a61c
      Sanjay Patel authored
      Part of prep work for D90901
      9b59a61c
    • Sanjay Patel's avatar
      [x86] propagate FMF from x86-specific intrinsic nodes to others during combining · f12f9beb
      Sanjay Patel authored
      This is another FMF gap exposed by D90901, but I don't see a way
      to show the difference in a regression test as with:
      f66ba4cf
      60256635
      
      We will see an asm difference if we add a test as part of D90901.
      f12f9beb
    • Nico Weber's avatar
      [lld/mac] Remove dead declaration · fd09a764
      Nico Weber authored
      fd09a764
    • Christopher Di Bella's avatar
      [libcxx][ranges] adds concept `sized_range` and cleans up `ranges::size` · d8fad661
      Christopher Di Bella authored
      * adds `sized_range` and conformance tests
      * moves `disable_sized_range` into namespace `std::ranges`
      * removes explicit type parameter
      
      Implements part of P0896 'The One Ranges Proposal'.
      
      Differential Revision: https://reviews.llvm.org/D102434
      d8fad661
    • Andrea Di Biagio's avatar
      [MCA] Unbreak the buildbots by passing flag -mcpu=generic to the new test... · 9acabe8b
      Andrea Di Biagio authored
      [MCA] Unbreak the buildbots by passing flag -mcpu=generic to the new test added by commit e5d59db4.
      
      This should unbreak buildbot clang-ppc64le-linux-lnt.
      9acabe8b
    • Christopher Di Bella's avatar
    • Sanjay Patel's avatar
      [x86] update fma test with deprecated intrinsics; NFC · 333c968d
      Sanjay Patel authored
      Similar to 8854b27b -
      
      All of the CHECK lines should be identical to before,
      but without any of the x86-specific calls that were
      replaced with generic FMA long ago.
      
      The file still has value because it shows a miscompile
      as demonstrated in D90901, but we probably need to
      add tests with FMF to make that explicit without
      losing coverage.
      333c968d
    • Stephen Neuendorffer's avatar
      [MLIR] Update Vector To LLVM conversion to be aware of assume_alignment · 29a50c58
      Stephen Neuendorffer authored
      vector.transfer_read and vector.transfer_write operations are converted
      to llvm intrinsics with specific alignment information, however there
      doesn't seem to be a way in llvm to take information from llvm.assume
      intrinsics and change this alignment information.  In any
      event, due the to the structure of the llvm.assume instrinsic, applying
      this information at the llvm level is more cumbersome.  Instead, let's
      generate the masked vector load and store instrinsic with the right
      alignment information from MLIR in the first place.  Since
      we're bothering to do this, lets just emit the proper alignment for
      loads, stores, scatter, and gather ops too.
      
      Differential Revision: https://reviews.llvm.org/D100444
      29a50c58
    • Pirama Arumuga Nainar's avatar
      [CoverageMapping] Handle gaps in counter IDs for source-based coverage · e4274cfe
      Pirama Arumuga Nainar authored
      For source-based coverage, the frontend sets the counter IDs and the
      constraints of counter IDs is not defined.  For e.g., the Rust frontend
      until recently had a reserved counter #0
      (https://github.com/rust-lang/rust/pull/83774).  Rust coverage
      instrumentation also creates counters on edges in addition to basic
      blocks.  Some functions may have more counters than regions.
      
      This breaks an assumption in CoverageMapping.cpp where the number of
      counters in a function is assumed to be bounded by the number of
      regions:
        Counts.assign(Record.MappingRegions.size(), 0);
      
      This assumption causes CounterMappingContext::evaluate() to fail since
      there are not enough counter values created in the above call to
      `Counts.assign`.  Consequently, some uncovered functions are not
      reported in coverage reports.
      
      This change walks a Function's CoverageMappingRecord to find the maximum
      counter ID, and uses it to initialize the counter array when instrprof
      records are missing for a function in sparse profiles.
      
      Differential Revision: https://reviews.llvm.org/D101780
      e4274cfe
    • Roman Lebedev's avatar
    • Roman Lebedev's avatar
    • Roman Lebedev's avatar
    • Patrick Holland's avatar
      [MCA] llvm-mca MCTargetStreamer segfault fix · e5d59db4
      Patrick Holland authored
      In order to create the code regions for llvm-mca to analyze, llvm-mca creates an
      AsmCodeRegionGenerator and calls AsmCodeRegionGenerator::parseCodeRegions().
      Within this function, both an MCAsmParser and MCTargetAsmParser are created so
      that MCAsmParser::Run() can be used to create the code regions for us.
      
      These parser classes were created for llvm-mc so they are designed to emit code
      with an MCStreamer and MCTargetStreamer that are expected to be setup and passed
      into the MCAsmParser constructor. Because llvm-mca doesn’t want to emit any
      code, an MCStreamerWrapper class gets created instead and passed into the
      MCAsmParser constructor. This wrapper inherits from MCStreamer and overrides
      many of the emit methods to just do nothing. The exception is the
      emitInstruction() method which calls Regions.addInstruction(Inst).
      
      This works well and allows llvm-mca to utilize llvm-mc’s MCAsmParser to build
      our code regions, however there are a few directives which rely on the
      MCTargetStreamer. llvm-mc assumes that the MCStreamer that gets passed into the
      MCAsmParser’s constructor has a valid pointer to an MCTargetStreamer. Because
      llvm-mca doesn’t setup an MCTargetStreamer, when the parser encounters one of
      those directives, a segfault will occur.
      
      In x86, each one of these 7 directives will cause this segfault if they exist in
      the input assembly to llvm-mca:
      
      .cv_fpo_proc
      .cv_fpo_setframe
      .cv_fpo_pushreg
      .cv_fpo_stackalloc
      .cv_fpo_stackalign
      .cv_fpo_endprologue
      .cv_fpo_endproc
      I haven’t looked at other targets, but I wouldn’t be surprised if some of the
      other ones also have certain directives which could result in this same
      segfault.
      
      My proposed solution is to simply initialize an MCTargetStreamer after we
      initialize the MCStreamerWrapper. The MCTargetStreamer requires an ostream
      object, but we don’t actually want any of these directives to be emitted
      anywhere, so I use an ostream created with the nulls() function. Since this
      needs to happen after the MCStreamerWrapper has been initialized, it needs to
      happen within the AsmCodeRegionGenerator::parseCodeRegions() function. The
      MCTargetStreamer also needs an MCInstPrinter which is easiest to initialize
      within the main() function of llvm-mca. So this MCInstPrinter gets constructed
      within main() then passed into the parseCodeRegions() function as a parameter.
      (If you feel like it would be appropriate and possible to create the
      MCInstPrinter within the parseCodeRegions() function, then feel free to modify
      my solution. That would stop us from having to pass it into the function and
      would limit its scope / lifetime.)
      
      My solution stops the segfault from happening and still passes all of the
      current (expected) llvm-mca tests. I also added a new test for x86 that checks
      for this segfault on an input that includes one of the .cv_fpo directives (this
      test fails without my solution, but passes with it).
      
      As far as I can tell, all of the functions that I modified are only called from
      within llvm-mca so there shouldn’t be any worries about breaking other tools.
      
      Differential Revision: https://reviews.llvm.org/D102709
      e5d59db4
    • Philip Reames's avatar
      Do actual DCE in LoopUnroll (try 4) · 449d14eb
      Philip Reames authored
      Turns out simplifyLoopIVs sometimes returns a non-dead instruction in it's DeadInsts out param.  I had done a bit of NFC cleanup which was only NFC if simplifyLoopIVs obeyed it's documentation.  I'm simplfy dropping that part of the change.
      
      Commit message from try 3:
      
      Recommitting after fixing a bug found post commit. Amusingly, try 1 had been correct, and by reverting to incorporate last minute review feedback, I introduce the bug. Oops. :)
      
      Original commit message:
      
      The problem was that recursively deleting an instruction can delete instructions beyond the current iterator (via a dead phi), thus invalidating iteration. Test case added in LoopUnroll/dce.ll to cover this case.
      
      LoopUnroll does a limited DCE pass after unrolling, but if you have a chain of dead instructions, it only deletes the last one. Improve the code to recursively delete all trivially dead instructions.
      
      Differential Revision: https://reviews.llvm.org/D102511
      449d14eb
    • Frederik Gossen's avatar
      Revert "Reapply "[clang][deps] Support inferred modules"" · 76b8754d
      Frederik Gossen authored
      This reverts commit c98833cd.
      The test `ClangScanDeps/modules-inferred-explicit-build.m` creates files
      in the current directory.
      76b8754d
    • Sanjay Patel's avatar
      [x86] propagate FMF from x86-specific intrinsic nodes to others during lowering · f66ba4cf
      Sanjay Patel authored
      This is another fast-math-flags failure exposed by D90901.
      f66ba4cf
    • Sanjay Patel's avatar
    • Nikita Popov's avatar
      [ScalarEvolution] Remove unused ExitLimit::hasOperand() method (NFC) · b661a55a
      Nikita Popov authored
      We only use BackedgeTakenInfo::hasOperand().
      b661a55a
    • Vedant Kumar's avatar
      [profile] Skip mmap() if there are no counters · 7014a101
      Vedant Kumar authored
      If there are no counters, an mmap() of the counters section would fail
      due to the size argument being too small (EINVAL).
      
      rdar://78175925
      
      Differential Revision: https://reviews.llvm.org/D102735
      7014a101
    • Jessica Paquette's avatar
      Recommit "[GlobalISel] Simplify G_ICMP to true/false when the result is known" · 84ae1cf8
      Jessica Paquette authored
      Add missing REQUIRES line to
      prelegalizer-combiner-icmp-to-true-false-known-bits.
      84ae1cf8
    • Hongtao Yu's avatar
      [CSSPGO] Overwrite branch weight annotated in previous pass. · 4ca6e37b
      Hongtao Yu authored
      Sample profile loader can be run in both LTO prelink and postlink. Currently the counts annoation in postilnk doesn't fully overwrite what's done in prelink. I'm adding a switch (`-overwrite-existing-weights=1`) to enable a full overwrite, which includes:
      
      1. Clear old metadata for calls when their parent block has a zero count. This could be caused by prelink code duplication.
      
      2. Clear indirect call metadata if somehow all the rest targets have a sum of zero count.
      
      3. Overwrite branch weight for basic blocks.
      
      With a CS profile, I was seeing #1 and #2 help reduce code size by preventing post-sample ICP and CGSCC inliner working on obsolete metadata, which come from a partial global inlining in prelink.  It's not expected to work well for non-CS case with a less-accurate post-inline count quality.
      
      It's worth calling out that some prelink optimizations can damage counts quality in an irreversible way. One example is the loop rotate optimization. Due to lack of exact loop entry count (profiling can only give loop iteration count and loop exit count), moving one iteration out of the loop body leaves the rest iteration count unknown. We had to turn off prelink loop rotate to achieve a better postlink counts quality. A even better postlink counts quality can be archived by turning off prelink CGSCC inlining which is not context-sensitive.
      
      Reviewed By: wenlei, wmi
      
      Differential Revision: https://reviews.llvm.org/D102537
      4ca6e37b
  2. May 19, 2021