1. Feb 13, 2021
  2. Feb 10, 2021
  3. Feb 09, 2021
  4. Feb 06, 2021
  5. Feb 05, 2021
  6. Feb 04, 2021
    • Shilei Tian's avatar
      [OpenMP] Disabled profiling in `libomp` by default to unblock link errors · 92a5106e
      Shilei Tian authored
      Link error occurred when time profiling in libomp is enabled by default
      because `libomp` is assumed to be a C library but the dependence on
      `libLLVMSupport` for profiling is a C++ library. Currently the issue blocks all
      OpenMP tests in Phabricator.
      
      This patch set a new CMake option `OPENMP_ENABLE_LIBOMP_PROFILING` to
      enable/disable the feature. By default it is disabled. Note that once time
      profiling is enabled for `libomp`, it becomes a C++ library.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95585
      
      (cherry picked from commit c571b168)
      92a5106e
    • Giorgis Georgakoudis's avatar
      [OpenMP] Fix building using LLVM_ENABLE_RUNTIMES · 66c7b449
      Giorgis Georgakoudis authored
      Fix when time profiling is enabled.
      
      Related to: D94855
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D95398
      
      (cherry picked from commit bb40e673)
      66c7b449
    • Peter Waller's avatar
      [clang][aarch64][WOA64][docs] Release note for longjmp crash with /guard:cf · bc2dad16
      Peter Waller authored
      Add a release note workaround for PR47463.
      
      Bug: https://bugs.llvm.org/show_bug.cgi?id=47463
      
      Differential Revision: https://reviews.llvm.org/D95435
      bc2dad16
    • Shilei Tian's avatar
      7fad20ec
    • Craig Topper's avatar
      [X86] Accept 64-bit GPRs for vextractps when using a register that requires EVEX. · e8cdcaea
      Craig Topper authored
      This is consistent with the VEX version. It also fixes a sorting
      issue in the matching table that caused the EVEX version to be
      prioritized over VEX in intel syntax.
      
      Fixes issue [2] from PR48991.
      
      (cherry picked from commit c691fe14)
      e8cdcaea
    • Shilei Tian's avatar
      [OpenMP][NVPTX] Take functions in `deviceRTLs` as `convergent` · ad208665
      Shilei Tian authored
      OpenMP device compiler (similar to other SPMD compilers) assumes that
      functions are convergent by default to avoid invalid transformations, such as
      the bug (https://bugs.llvm.org/show_bug.cgi?id=49021).
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95971
      
      (cherry picked from commit 0f0ce3c1)
      ad208665
    • Hongtao Yu's avatar
      [CSSPGO] Introducing distribution factor for pseudo probe. · a9157c56
      Hongtao Yu authored
      Sample re-annotation is required in LTO time to achieve a reasonable post-inline profile quality. However, we have seen that such LTO-time re-annotation degrades profile quality. This is mainly caused by preLTO code duplication that is done by passes such as loop unrolling, jump threading, indirect call promotion etc, where samples corresponding to a source location are aggregated multiple times due to the duplicates. In this change we are introducing a concept of distribution factor for pseudo probes so that samples can be distributed for duplicated probes scaled by a factor. We hope that optimizations duplicating code well-maintain the branch frequency information (BFI) based on which probe distribution factors are calculated. Distribution factors are updated at the end of preLTO pipeline to reflect an estimated portion of the real execution count.
      
      This change also introduces a pseudo probe verifier that can be run after each IR passes to detect duplicated pseudo probes.
      
      A saturated distribution factor stands for 1.0. A pesudo probe will carry a factor with the value ranged from 0.0 to 1.0. A 64-bit integral distribution factor field that represents [0.0, 1.0] is associated to each block probe. Unfortunately this cannot be done for callsite probes due to the size limitation of a 32-bit Dwarf discriminator. A 7-bit distribution factor is used instead.
      
      Changes are also needed to the sample profile inliner to deal with prorated callsite counts. Call sites duplicated by PreLTO passes, when later on inlined in LTO time, should have the callees’s probe prorated based on the Prelink-computed distribution factors. The distribution factors should also be taken into account when computing hotness for inline candidates. Also, Indirect call promotion results in multiple callisites. The original samples should be distributed across them. This is fixed by adjusting the callisites' distribution factors.
      
      Reviewed By: wmi
      
      Differential Revision: https://reviews.llvm.org/D93264
      
      (cherry picked from commit 3d89b3cb)
      a9157c56
    • Wenlei He's avatar
      [CSSPGO] Factor out common part for CSSPGO inline and AFDO inline · c2f3f45b
      Wenlei He authored
      Refactoring SampleProfileLoader::inlineHotFunctions to use helpers from CSSPGO inlining and reduce similar code in the inlining loop, plus minor cleanup for AFDO path.
      
      This is resubmit of D95024, with build break and overtighten assertion fixed.
      
      Test Plan:
      
      (cherry picked from commit 1645f465)
      c2f3f45b
    • Wenlei He's avatar
      [CSSPGO] Call site prioritized inlining for sample PGO · 27ff658e
      Wenlei He authored
      This change implemented call site prioritized BFS profile guided inlining for sample profile loader. The new inlining strategy maximize the benefit of context-sensitive profile as mentioned in the follow up discussion of CSSPGO RFC. The change will not affect today's AutoFDO as it's opt-in. CSSPGO now defaults to the new FDO inliner, but can fall back to today's replay inliner using a switch (`-sample-profile-prioritized-inline=0`).
      
      Motivation
      
      With baseline AutoFDO, the inliner in sample profile loader only replays previous inlining, and the use of profile is only for pruning previous inlining that turned out to be cold. Due to the nature of replay, the FDO inliner is simple with hotness being the only decision factor. It has the following limitations that we're improving now for CSSPGO.
       - It doesn't take inline candidate size into account. Since it's doing replay, the size growth is bounded by previous CGSCC inlining. With context-sensitive profile, FDO inliner is no longer limited by previous inlining, so we need to take size into account to avoid significant size bloat.
       - The way it looks at hotness is not accurate. It uses total samples in an inlinee as proxy for hotness, while what really matters for an inline decision is the call site count. This is an unfortunate fall back because call site count and callee entry count are not reliable due to dwarf based correlation, especially for inlinees. Now paired with pseudo-probe, we have accurate call site count and callee's entry count, so we can use that to gauge hotness more accurately.
       - It treats all call sites from a block as hot as long as there's one call site considered hot. This is normally true, but since total samples is used as hotness proxy, this transitiveness within block magnifies the inacurate hotness heuristic. With pseduo-probe and the change above, this is no longer an issue for CSSPGO.
      
      New FDO Inliner
      
      Putting all the requirement for CSSPGO together, we need a top-down call site prioritized BFS inliner. Here're reasons why each component is needed.
       - Top-down: We need a top-down inliner to better leverage context-sensitive profile, so inlining is driven by accurate context profile, and post-inline is also accurate. This is already implemented in https://reviews.llvm.org/D70655.
       - Size Cap: For top-down inliner, taking function size into account for inline decision alone isn't sufficient to control size growth. We also need to explicitly cap size growth because with top-down inlining, we can grow inliner size significantly with large number of smaller inlinees even if each individually passes the cost/size check.
       - Prioritize call sites: With size cap, inlining order also becomes important, because if we stop inlining due to size budget limit, we'd want to use budget towards the most beneficial call sites.
       - BFS inline: Same as call site prioritization, if we stop inlining due to size budget limit, we want a balanced inline tree, rather than going deep on one call path.
      
      Note that the new inliner avoids repeatedly evaluating same set of call site, so it should help with compile time too. For this reason, we could transition today's FDO inliner to use a queue with equal priority to avoid wasted reevaluation of same call site (TODO).
      
      Speculative indirect call promotion and inlining is also supported now with CSSPGO just like baseline AutoFDO.
      
      Tunings and knobs
      
      I created tuning knobs for size growth/cap control, and for hot threshold separate from CGSCC inliner. The default values are selected based on initial tuning with CSSPGO.
      
      Results
      
      Evaluated with an internal LLVM fork couple months ago, plus another change to adjust hot-threshold cutoff for context profile (will send up after this one), the new inliner show ~1% geomean perf win on spec2006 with CSSPGO, while reducing code size too. The measurement was done using train-train setup, MonoLTO w/ new pass manager and pseudo-probe. Note that this is just a starting point - we hope that the new inliner will open up more opportunity with CSSPGO, but it will certainly take more time and effort to make it fully calibrated and ready for bigger workloads (we're working on it).
      
      Differential Revision: https://reviews.llvm.org/D94001
      
      (cherry picked from commit 6bae5973)
      27ff658e
    • Hongtao Yu's avatar
      [CSSPGO] Passing the clang driver switch -fpseudo-probe-for-profiling to the linker. · b9fa16f2
      Hongtao Yu authored
      As titled.
      
      Reviewed By: wmi, wenlei
      
      Differential Revision: https://reviews.llvm.org/D95271
      
      (cherry picked from commit d3e2e374)
      b9fa16f2
    • Hongtao Yu's avatar
      [CSSPGO] Tweaking inlining with pseudo probes. · f2cabaac
      Hongtao Yu authored
      Fixing up a couple places where `getCallSiteIdentifier` is needed to support pseudo-probe-based callsites.
      
      Also fixing an issue in the extbinary profile reader where the metadata section is not fully scanned based on the number of profiles loaded only for the current module.
      
      Reviewed By: wmi, wenlei
      
      Differential Revision: https://reviews.llvm.org/D95791
      
      (cherry picked from commit 224fee82)
      f2cabaac
    • Hongtao Yu's avatar
      [CSSPGO] Support of CS profiles in extended binary format. · 7d096f9b
      Hongtao Yu authored
      This change brings up support of context-sensitive profiles in the format of extended binary. Existing sample profile reader/writer/merger code is being tweaked to reflect the fact of bracketed input contexts, like (`[...]`). The paired brackets are also needed in extbinary profiles because we don't yet have an otherwise good way to tell calling contexts apart from regular function names since the context delimiter `@` can somehow serve as a part of the C++ mangled names.
      
      Reviewed By: wmi, wenlei
      
      Differential Revision: https://reviews.llvm.org/D95547
      
      (cherry picked from commit 7e99bddf)
      7d096f9b
    • Shilei Tian's avatar
      [OpenMP] Disabled profiling in `libomp` by default to unblock link errors · f5602e0b
      Shilei Tian authored
      Link error occurred when time profiling in libomp is enabled by default
      because `libomp` is assumed to be a C library but the dependence on
      `libLLVMSupport` for profiling is a C++ library. Currently the issue blocks all
      OpenMP tests in Phabricator.
      
      This patch set a new CMake option `OPENMP_ENABLE_LIBOMP_PROFILING` to
      enable/disable the feature. By default it is disabled. Note that once time
      profiling is enabled for `libomp`, it becomes a C++ library.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95585
      
      (cherry picked from commit c571b168)
      f5602e0b
    • Stephen Kelly's avatar
      2a917b70
    • Richard Smith's avatar
      PR44325 (and duplicates): don't issue -Wzero-as-null-pointer-constant · 678c259d
      Richard Smith authored
      when rewriting 'a < b' as '(a <=> b) < 0'.
      
      It's pretty common for comparison category types to use a pointer or
      pointer-to-member type as their '0' parameter.
      
      (cherry picked from commit 1f06f419)
      678c259d
    • Joseph Huber's avatar
      [OpenMP] Fix seg fault in libomptarget when using Info with multiple threads · 922e4149
      Joseph Huber authored
      Summary:
      One option for the LIBOMPTARGET_INFO environment variable is to print the current status of the device's data mappings. These are a shared resource among threads so this needs to be protected when using multiple streams.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95786
      
      (cherry picked from commit fda48539)
      922e4149
    • Shilei Tian's avatar
      [OpenMP][NFC] Added release note for new `deviceRTLs` and hidden helper task · 255f7398
      Shilei Tian authored
      Added release note for new `deviceRTLs` and hidden helper task for LLVM
      12.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95584
      
      (cherry picked from commit 7bc31018)
      255f7398
    • Shilei Tian's avatar
      [OpenMP][deviceRTLs] Added `[[clang::loader_uninitialized]]` explicitly · 5d926bb3
      Shilei Tian authored
      `[[clang::loader_uninitialized]]` is in macro `SHARED` but it doesn't
      work for array like `parallelLevel`, so the variable will be zero initialized.
      There is also a similar issue for `omptarget_nvptx_device_State` which is in
      global address space. Its c'tor is also generated, which was not in the past when
      building the `deviceRTLs` with CUDA. In this patch, we added the attribute to
      the two variables explicitly.
      
      Reviewed By: jdoerfert
      
      Differential Revision: https://reviews.llvm.org/D95550
      
      (cherry picked from commit 19248d30)
      5d926bb3
    • Shilei Tian's avatar
      [OpenMP][NVPTX] Added the missing -O1 when building NVPTX bitcode libraries · 4d0874c7
      Shilei Tian authored
      In the past `-O1` was used when building NVPTX bitcode libraries. After
      we switched to OpenMP, `-O1` was missing by mistake, leading to a huge performance
      regression.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D95545
      
      (cherry picked from commit 5a64794b)
      4d0874c7
    • Atmn Patel's avatar
      [OpenMP][Libomptarget] Fix conditional in CMake for remote plugin · 12b6579b
      Atmn Patel authored
      The remote offloading plugin's CMakeLists was trying to build if its
      flag was enabled even if it didn't find gRPC/protobuf. The conditional
      was wrong, it's fixed by this.
      
      Differential Revision: https://reviews.llvm.org/D95574
      
      (cherry picked from commit 8a770562)
      12b6579b
    • Haowei Wu's avatar
      [elfabi] Fix tests which failed on different timezones · e2d822c3
      Haowei Wu authored
      This patch fixes elfabi tests on machines using a GMT+X timezone
      settings.
      
      Differential Revision: https://reviews.llvm.org/D95641
      
      (cherry picked from commit 771b3596)
      e2d822c3
    • Andrew Ng's avatar
      [X86] Fix disassembly of x86-64 GDTLS code sequence · b15f3fc5
      Andrew Ng authored
      For x86-64 the REX.w prefix takes precedence over any other size
      override (i.e. 0x66). Therefore, for x86-64 when REX.w is present set
      'hasOpSize' to false to ensure that any size override is ignored.
      
      Fixes PR48901.
      
      Differential Revision: https://reviews.llvm.org/D95682
      
      (cherry picked from commit 94fedd26)
      b15f3fc5
    • Cullen Rhodes's avatar
      [LV] Fix crash when computing max VF too early · c5904f5c
      Cullen Rhodes authored
      D90687 introduced a crash:
      
        llvm::LoopVectorizationCostModel::computeMaxVF(llvm::ElementCount, unsigned int):
          Assertion `WideningDecisions.empty() && Uniforms.empty() && Scalars.empty() &&
          "No decisions should have been taken at this point"' failed.
      
      when compiling the following C code:
      
        typedef struct {
        char a;
        } b;
      
        b *c;
        int d, e;
      
        int f() {
          int g = 0;
          for (; d; d++) {
            e = 0;
            for (; e < c[d].a; e++)
              g++;
          }
          return g;
        }
      
      with:
      
        clang -Os -target hexagon -mhvx -fvectorize -mv67 testcase.c -S -o -
      
      This occurred since prior to D90687 computeFeasibleMaxVF would only be
      called in computeMaxVF when a scalar epilogue was allowed, but now it's
      always called. This causes the assert above since computeFeasibleMaxVF
      collects all viable VFs larger than the default MaxVF, and for each VF
      calculates the register usage which results in analysis being done the
      assert above guards against. This can occur in computeFeasibleMaxVF if
      TTI.shouldMaximizeVectorBandwidth and this target hook is implemented in
      the hexagon backend to always return true.
      
      Reported by @iajbar.
      
      Reviewed By: fhahn
      
      Differential Revision: https://reviews.llvm.org/D94869
      
      (cherry picked from commit 8cda2274)
      c5904f5c