1. Feb 23, 2021
  2. Feb 20, 2021
    • Johannes Doerfert's avatar
      Avoid use of stack allocations in asynchronous calls · 76d5d54f
      Johannes Doerfert authored
      NOTE: This is an adaption of the original patch to be applicable to the
            LLVM 12 release branch. Logic is the same though.
      
      As reported by Guilherme Valarini [0], we used to pass stack allocations
      to calls that can nowadays be asynchronous. This is arguably a problem
      and it will inevitably result in UB. To remedy the situation we allocate
      the locations as part of the AsyncInfoTy object. The lifetime of that
      object matches what we need for now. If the synchronization is not tied
      to the AsyncInfoTy object anymore we might need to have a different
      buffer construct in global space.
      
      This should be back-ported to LLVM 12 but needs slight modifications as
      it is based on refactoring patches we do not need to backport.
      
      [0] https://lists.llvm.org/pipermail/openmp-dev/2021-February/003867.html
      
      Differential Revision: https://reviews.llvm.org/D96667
      76d5d54f
    • Fangrui Song's avatar
      [llvm-objdump] --source: drop the warning when there is no debug info · ee7eaf86
      Fangrui Song authored
      Warnings have been added for three cases (PR41905): (1) missing debug info, (2)
      the source file cannot be found, (3) the debug info points at a line beyond the
      end of the file.
      
      (1) is probably less useful. This was brought up once on
      http://lists.llvm.org/pipermail/llvm-dev/2020-April/141264.html and two
      internal users mentioned it to me that it was annoying. (I personally
      find the warning confusing, too.)
      
      Users specify --source to get additional information if sources happen to be
      available.  If sources are not available, it should be obvious as the output
      will have no interleaved source lines. The warning can be especially annoying
      when using llvm-objdump -S on a bunch of files.
      
      This patch drops the warning when there is no debug info.
      (If LLVMSymbolizer::symbolizeCode returns an `Error`, there will still be
      an error. There is currently no test for an `Error` return value.
      The only code path is probably a broken symbol table, but we probably already emit a warning
      in that case)
      
      `source-interleave-prefix.test` has an inappropriate "malformed" test - the test simply has no
      .debug_* because new llc does not produce debug info when the filename is empty (invalid).
      I have tried tampering the header of .debug_info/.debug_line but llvm-symbolizer does not warn.
      This patch does not intend to add the missing test coverage.
      
      Differential Revision: https://reviews.llvm.org/D88715
      
      (cherry picked from commit eecbb1c7)
      ee7eaf86
    • William S. Moses's avatar
      [SROA] Amend failing test from D95826 · bdafd284
      William S. Moses authored
      (cherry picked from commit 892d2822)
      bdafd284
    • Florian Hahn's avatar
      [clang] Add -ffinite-loops & -fno-finite-loops options. · b5b31112
      Florian Hahn authored
      This cherry-picks the following patches on the release branch:
      
      6280bb4c [clang] Remove redundant condition (NFC).
      51bf4c0e [clang] Add -ffinite-loops & -fno-finite-loops options.
      fb4d8fe8 [clang] Update mustprogress tests
      
      This patch adds 2 new options to control when Clang adds `mustprogress`:
      
        1. -ffinite-loops: assume all loops are finite; mustprogress is added
           to all loops, regardless of the selected language standard.
        2. -fno-finite-loops: assume no loop is finite; mustprogress is not
           added to any loop or function. We could add mustprogress to
           functions without loops, but we would have to detect that in Clang,
           which is probably not worth it.
      
      Differential Revision: https://reviews.llvm.org/D96850
      b5b31112
    • wlei's avatar
      [CSSPGO][llvm-profgen] Filter out the instructions without location info for symbolizer · 610b51c0
      wlei authored
      It appears some instructions doesn't have the debug location info and the symbolizer will return an empty call stack for them which will cause some crash later in profile unwinding. Actually we do not record the sample info for them, so this change just filter out those instruction.
      
      As those instruction would appears at the begin and end of the instruction list, without them we need to add the boundary check for IP `advance` and `backward`.
      
      Also for pseudo probe based profile, we actually don't need the symbolized location info, so here just change to use an empty stack for it. This could save half of the binary loading time.
      
      Differential Revision: https://reviews.llvm.org/D96434
      610b51c0
    • wlei's avatar
      [CSSPGO][llvm-profgen] Renovate perfscript check and command line input validation · 66873fb6
      wlei authored
      This include some changes related with PerfReader's the input check and command line change:
      
      1) It appears there might be thousands of leading MMAP-Event line in the perfscript for large workload. For this case, the 4k threshold is not eligible to determine it's a hybrid sample. This change renovated the `isHybridPerfScript` by going through the script without threshold limitation checking whether there is a non-empty call stack immediately followed by a LBR sample. It will stop once it find a valid one.
      
      2) Added several input validations for the command line switches in PerfReader.
      
      3) Changed the command line `show-disassembly` to `show-disassembly-only`, it will print to stdout and exit early which leave an empty output profile.
      
      Reviewed By: hoy, wenlei
      
      Differential Revision: https://reviews.llvm.org/D96387
      66873fb6
    • wlei's avatar
      [CSSPGO][llvm-profgen] Add brackets for context id to support extended binary format · beb80ffe
      wlei authored
      To align with https://reviews.llvm.org/D95547, we need to add brackets for context id before initializing the `SampleContext`.
      
      Also added test cases for extended binary format from llvm-profgen side.
      
      Differential Revision: https://reviews.llvm.org/D95929
      beb80ffe
    • Hongtao Yu's avatar
      Remove test code that cause MSAN failure. · 989b5c95
      Hongtao Yu authored
      Summary:
      The negative test (with the feature being added disabled) caused MSAN failure and that's the added feature is supposed to fix. Therefore the negative test code is being removed.
      989b5c95
    • Hongtao Yu's avatar
      [CSSPGO] Process functions in a top-down order on a dynamic call graph. · 1f5e2016
      Hongtao Yu authored
      Functions are currently processed by the sample profiler loader in a top-down order defined by the static call graph. The order is being adjusted to be a top-down order based on the input context-sensitive profile. One benefit is that the processing order of caller and callee in one SCC would follow the context order in the profile to favor more inlining. Another benefit is that the processing order of caller and callee through an indirect call (which is not on the static call graph) can be honored which in turn allows for more inlining.
      
      The profile top-down order for SCC is also extended to support non-CS profiles.
      
      Two switches `-mllvm -use-profile-indirect-call-edges` and `-mllvm -use-profile-top-down-order` are being introduced.
      
      Reviewed By: wmi
      
      Differential Revision: https://reviews.llvm.org/D95988
      1f5e2016
    • Hongtao Yu's avatar
      1a5bb1e4
    • Hongtao Yu's avatar
      [CSSPGO] Unblock optimizations with pseudo probe instrumentation. · e8e45f52
      Hongtao Yu authored
      The IR/MIR pseudo probe intrinsics don't get materialized into real machine instructions and therefore they don't incur runtime cost directly. However, they come with indirect cost by blocking certain optimizations. Some of the blocking are intentional (such as blocking code merge) for better counts quality while the others are accidental. This change unblocks perf-critical optimizations that do not affect counts quality. They include:
      
      1. IR InstCombine, sinking load operation to shorten lifetimes.
      2. MIR LiveRangeShrink, similar to #1
      3. MIR TwoAddressInstructionPass, i.e, opeq transform
      4. MIR function argument copy elision
      5. IR stack protection. (though not perf-critical but nice to have).
      
      Reviewed By: wmi
      
      Differential Revision: https://reviews.llvm.org/D95982
      e8e45f52
    • Wenlei He's avatar
      [CSSPGO] Use merged base profile for hot threshold calculation · 10712791
      Wenlei He authored
      Context-sensitive profile effectively split a function profile into many copies each representing the CFG profile of a particular calling context. That makes the count distribution looks more flat as we now have more function profiles each with lower counts, which in turn leads to lower hot thresholds. Now we tells threshold computation to merge context profile first before calculating percentile based cutoffs to compensate for seemingly flat context profile. This can be controlled by swtich `sample-profile-contextless-threshold`.
      
      Earlier measurement showed ~0.4% perf boost with this tuning on spec2k6 for CSSPGO (with pseudo-probe and new inliner).
      
      Differential Revision: https://reviews.llvm.org/D95980
      10712791
    • wlei's avatar
      [CSSPGO][llvm-profgen] Fix bug with parsing hybrid sample trace line · db88d922
      wlei authored
      when we skip the call stack starting with an external address, we should also skip the bottom LBR entry, otherwise it will cause a truncated context issue.
      
      Reviewed By: hoy, wenlei
      
      Differential Revision: https://reviews.llvm.org/D95480
      db88d922
    • wlei's avatar
      [CSSPGO][llvm-profgen] Merge and trim profile for cold context to reduce profile size · 87c27020
      wlei authored
      This change allows merging and trimming cold context profile in llvm-profgen to solve profile size bloat problem. Currently when the profile's total sample is below threshold(supported by a switch), it will be considered cold and merged into a base context-less profile, which will at least keep the profile quality as good as the baseline(non-cs).
      
      For example, two input profiles:
       [main @ foo @ bar]:60
       [main @ bar]:50
      Under threshold = 100, the two profiles will be merge into one with the base context, get result:
       [bar]:110
      
      Added two switches:
      `--csprof-cold-thres=<value>`: Specified the total samples threshold for a context profile to be considered cold, with 100 being the default. Any cold context profiles will be merged into context-less base profile by default.
      `--csprof-keep-cold`: Force profile generation to keep cold context profiles instead of dropping them. By default, any cold context will not be written to output profile.
      
      Results:
      Though not yet evaluating it with the latest CSSPGO, our internal branch shows neutral on performance but significantly reduce the profile size. Detailed evaluation on llvm-profgen with CSSPGO will come later.
      
      Differential Revision: https://reviews.llvm.org/D94111
      87c27020
    • wlei's avatar
      [CSSPGO][llvm-profgen] Aggregate samples on call frame trie to speed up profile generation · e562ff08
      wlei authored
      For CS profile generation, the process of call stack unwinding is time-consuming since for each LBR entry we need linear time to generate the context( hash, compression, string concatenation). This change speeds up this by grouping all the call frame within one LBR sample into a trie and aggregating the result(sample counter) on it, deferring the context compression and string generation to the end of unwinding.
      
      Specifically, it uses `StackLeaf` as the top frame on the stack and manipulates(pop or push a trie node) it dynamically during virtual unwinding so that the raw sample can just be recoded on the leaf node, the path(root to leaf) will represent its calling context. In the end, it traverses the trie and generates the context on the fly.
      
      Results:
      Our internal branch shows about 5X speed-up on some large workloads in SPEC06 benchmark.
      
      Differential Revision: https://reviews.llvm.org/D94110
      e562ff08
    • wlei's avatar
      [CSSPGO][llvm-profgen] Compress recursive cycles in calling context · 6209b075
      wlei authored
      This change compresses the context string by removing cycles due to recursive function for CS profile generation. Removing recursion cycles is a way to normalize the calling context which will be better for the sample aggregation and also make the context promoting deterministic.
      Specifically for implementation, we recognize adjacent repeated frames as cycles and deduplicated them through multiple round of iteration.
      For example:
      Considering a input context string stack:
      [“a”, “a”, “b”, “c”, “a”, “b”, “c”, “b”, “c”, “d”]
      For first iteration,, it removed all adjacent repeated frames of size 1:
      [“a”, “b”, “c”, “a”, “b”, “c”, “b”, “c”, “d”]
      For second iteration, it removed all adjacent repeated frames of size 2:
      [“a”, “b”, “c”, “a”, “b”, “c”, “d”]
      So in the end, we get compressed output:
      [“a”, “b”, “c”, “d”]
      
      Compression will be called in two place: one for sample's context key right after unwinding, one is for the eventual context string id in the ProfileGenerator.
      Added a switch `compress-recursion` to control the size of duplicated frames, default -1 means no size limit.
      Added unit tests and regression test for this.
      
      Differential Revision: https://reviews.llvm.org/D93556
      6209b075
    • wlei's avatar
      [CSSPGO][llvm-profgen] Pseudo probe based CS profile generation · 78b35e27
      wlei authored
      This change implements profile generation infra for pseudo probe in llvm-profgen. During virtual unwinding, the raw profile is extracted into range counter and branch counter and aggregated to sample counter map indexed by the call stack context. This change introduces the last step and produces the eventual profile. Specifically, the body of function sample is recorded by going through each probe among the range and callsite target sample is recorded by extracting the callsite probe from branch's source.
      
      Please refer https://groups.google.com/g/llvm-dev/c/1p1rdYbL93s and https://reviews.llvm.org/D89707 for more context about CSSPGO and llvm-profgen.
      
      **Implementation**
      
      - Extended `PseudoProbeProfileGenerator` for pseudo probe based profile generation.
      - `populateBodySamplesWithProbes` reading range counter is responsible for recording function body samples and inferring caller's body samples.
      - `populateBoundarySamplesWithProbes` reading branch counter is responsible for recording call site target samples.
      - Each sample is recorded with its calling context(named `ContextId`). Remind that the probe based context key doesn't include the leaf frame probe info, so the `ContextId` string is created from two part: one from the probe stack strings' concatenation and other one from the leaf frame probe.
      - Added regression test
      
      Test Plan:
      
      ninja & ninja check-llvm
      
      Differential Revision: https://reviews.llvm.org/D92998
      78b35e27
    • Yang Fan's avatar
      [CSSPGO] Fix MSVC initializing truncation warning (NFC) · a7629a22
      Yang Fan authored
      MSVC warning:
      ```
      \llvm-project\llvm\include\llvm\Transforms\IPO\SampleProfileProbe.h(65): warning C4305: 'initializing': truncation from 'double' to 'const float'
      ```
      a7629a22
    • William S. Moses's avatar
      [SROA] Propagate correct TBAA/TBAA Struct offsets · d3f9f512
      William S. Moses authored
      SROA does not correctly account for offsets in TBAA/TBAA struct metadata.
      This patch creates functionality for generating new MD with the corresponding
      offset and updates SROA to use this functionality.
      
      Differential Revision: https://reviews.llvm.org/D95826
      
      (cherry picked from commit 40862b1a)
      d3f9f512
    • Georgii Rymar's avatar
      [llvm-symbolizer] - Fix the crash in GNU output style with --no-inlines and missing input file. · 0d4f8a3f
      Georgii Rymar authored
      Fixes https://bugs.llvm.org/show_bug.cgi?id=48882.
      
      If the input file does not exist (or has a reading error), the
      following code will crash if there are two or more input addresses.
      
      ```
      auto ResOrErr = Symbolizer.symbolizeInlinedCode(
        ModuleName, {Offset, object::SectionedAddress::UndefSection});
      Printer << (error(ResOrErr) ? DILineInfo() : ResOrErr.get().getFrame(0));
      ```
      
      For the first address, `symbolizeInlinedCode` returns an error.
      For the second address, `symbolizeInlinedCode` returns an empty result
      (not an error) and `.getFrame(0)` will crash.
      
      Differential revision: https://reviews.llvm.org/D95609
      
      (cherry picked from commit d2214068)
      0d4f8a3f
    • Simonas Kazlauskas's avatar
      [llvm-dwp] Join dwo paths correctly when DWOPath is absolute · b1106a5b
      Simonas Kazlauskas authored
      When the `DWOPath` is absolute, we want to use `DWOPath` as is, without prepending any other
      components to the path. The `sys::path::append` does not join, but rather unconditionally appends
      the paths, so something like `sys::path::append("/tmp", "/tmp/banana")` will result in
      `/tmp/tmp/banana` rather than the desired `/tmp/banana`.
      
      This then causes `llvm-dwp` to fail in a following situation:
      
      ```
      $ clang -gsplit-dwarf /tmp/banana/test.c -c -o /tmp/outdir/foo.o
      $ clang outdir/foo.o -o outdir/hm
      $ llvm-dwarfdump outdir/hm | grep -C2 foo.dwo
                        DW_AT_comp_dir    ("/tmp")
                        DW_AT_GNU_pubnames  (true)
                        DW_AT_GNU_dwo_name    ("/tmp/outdir/foo.dwo")
                                      DW_AT_GNU_dwo_id    (0xde4d396f3bf0e257)
                        DW_AT_low_pc  (0x0000000000401100)
      $ strace -o trace llvm-dwp -e outdir/hm -o outdir/hm.dwp
      error: No such file or directory
      $ cat trace | grep foo.dwo
      openat(AT_FDCWD, "/tmp/tmp/outdir/foo.dwo", O_RDONLY|O_CLOEXEC) = -1 ENOENT (No such file or directory)
      ```
      
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D96678
      
      (cherry picked from commit 6ffcb293)
      b1106a5b
    • Kadir Cetinkaya's avatar
      [clangd] Treat "null" optional fields as missing · 34e8fd50
      Kadir Cetinkaya authored
      Clangd currently throws away any protocol messages whenever an optional
      field has an unexpected type. This patch changes the behaviour to treat
      `null` fields as missing.
      
      This enables clangd to be more tolerant against small violations to the
      LSP spec.
      
      Fixes https://github.com/clangd/vscode-clangd/issues/134
      
      Differential Revision: https://reviews.llvm.org/D95229
      
      (cherry picked from commit af20232b)
      34e8fd50
    • Shilei Tian's avatar
      [OpenMP][NVPTX] Add the support for CUDA 11.2 and CUDA 11.1 · 2f74c220
      Shilei Tian authored
      CUDA 11.2 and CUDA 11.1 are all available now.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D97004
      
      (cherry picked from commit 89827fd4)
      2f74c220
    • Jeroen Dobbelaere's avatar
      [clang] functions with the 'const' or 'pure' attribute must always return. · a338d577
      Jeroen Dobbelaere authored
      As described in
      * https://gcc.gnu.org/onlinedocs/gcc/Common-Function-Attributes.html#index-pure-function-attribute
      * https://gcc.gnu.org/onlinedocs/gcc/Common-Function-Attributes.html#index-const-function-attribute
      
      An `__attribute__((pure))` function must always return, as well as an `__attribute__((const))` function.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D96960
      
      (cherry picked from commit 46757ccb)
      a338d577
    • Nikita Popov's avatar
      [LLD] Fix tests after D96993 · 17daef8b
      Nikita Popov authored
      We now need mustprogress to eliminate these calls. The code doesn't
      really make sense, but that's not the point of the test...
      
      (cherry picked from commit ac065b7a)
      17daef8b
    • Nikita Popov's avatar
      [DCE] Don't remove non-willreturn calls · 8e9c2ad9
      Nikita Popov authored
      In both ADCE and BDCE (via DemandedBits) we should not remove
      instructions that are not guaranteed to return. This issue was
      pointed out by fhahn in the recent llvm-dev thread.
      
      Differential Revision: https://reviews.llvm.org/D96993
      
      (cherry picked from commit 2f17ed29)
      8e9c2ad9
    • Nikita Popov's avatar
      [IR] Move willReturn() to Instruction · d1d7dc77
      Nikita Popov authored
      This moves the willReturn() helper from CallBase to Instruction,
      so that it can be used in a more generic manner. This will make
      it easier to fix additional passes (ADCE and BDCE), and will give
      us one place to change if additional instructions should become
      non-willreturn (e.g. there has been talk about handling volatile
      operations this way).
      
      I have also included the IntrinsicInst workaround directly in
      here, so that it gets applied consistently. (As such this change
      is not entirely NFC -- FuncAttrs will now use this as well.)
      
      Differential Revision: https://reviews.llvm.org/D96992
      
      (cherry picked from commit 370addb9)
      d1d7dc77
    • Nikita Popov's avatar
      [DCE] Add tests for non-willreturn function being removed (NFC) · c2a0b081
      Nikita Popov authored
      (cherry picked from commit 4045ad6b)
      c2a0b081
    • Lei Huang's avatar
  3. Feb 18, 2021