1. Sep 17, 2019
    • Dan Albert's avatar
      Revert "Implement std::condition_variable via pthread_cond_clockwait() where available" · c1c519d2
      Dan Albert authored
      This reverts commit 5e37d7f9.
      
      llvm-svn: 372034
      c1c519d2
    • Bardia Mahjour's avatar
      [NFC] Test commit access · 474c713f
      Bardia Mahjour authored
      llvm-svn: 372033
      474c713f
    • DeForest Richards's avatar
      [Docs] Bug fix for docs homepage · 3b27f4c0
      DeForest Richards authored
      Removes reference to non-existent Reference Documentation page.
      
      llvm-svn: 372032
      3b27f4c0
    • DeForest Richards's avatar
      [Docs] Adds Getting Started/Tutorials, Reference to LLVM docs homepage · e151cb7c
      DeForest Richards authored
      Adds a section for Getting Started/Tutorials and Reference topics to the LLVM docs homepage.
      
      llvm-svn: 372031
      e151cb7c
    • Lei Huang's avatar
      [PowerPC] Cust lower fpext v2f32 to v2f64 from extract_subvector v4f32 · bfb197d7
      Lei Huang authored
      This is a follow up patch from https://reviews.llvm.org/D57857 to handle
      extract_subvector v4f32.  For cases where we fpext of v2f32 to v2f64 from
      extract_subvector we currently generate on P9 the following:
      
        lxv 0, 0(3)
        xxsldwi 1, 0, 0, 1
        xscvspdpn 2, 0
        xxsldwi 3, 0, 0, 3
        xxswapd 0, 0
        xscvspdpn 1, 1
        xscvspdpn 3, 3
        xscvspdpn 0, 0
        xxmrghd 0, 0, 3
        xxmrghd 1, 2, 1
        stxv 0, 0(4)
        stxv 1, 0(5)
      
      This patch custom lower it to the following sequence:
      
        lxv 0, 0(3)       # load the v4f32 <w0, w1, w2, w3>
        xxmrghw 2, 0, 0   # Produce the following vector <w0, w0, w1, w1>
        xxmrglw 3, 0, 0   # Produce the following vector <w2, w2, w3, w3>
        xvcvspdp 2, 2     # FP-extend to <d0, d1>
        xvcvspdp 3, 3     # FP-extend to <d2, d3>
        stxv 2, 0(5)      # Store <d0, d1> (%vecinit11)
        stxv 3, 0(4)      # Store <d2, d3> (%vecinit4)
      
      Differential Revision: https://reviews.llvm.org/D61961
      
      llvm-svn: 372029
      bfb197d7
    • Jonas Devlieghere's avatar
      [NFC] Move dumping into GDBRemotePacket · 4e053ff1
      Jonas Devlieghere authored
      This moves the dumping logic from the GDBRemoteCommunicationHistory
      class into the GDBRemotePacket so that it can be reused from the
      reproducer command object.
      
      llvm-svn: 372028
      4e053ff1
    • Dan Albert's avatar
      Open fstream files in O_CLOEXEC mode when possible. · a7e90599
      Dan Albert authored
      Reviewers: EricWF, mclow.lists, ldionne
      
      Reviewed By: ldionne
      
      Subscribers: smeenai, dexonsmith, christof, ldionne, libcxx-commits
      
      Tags: #libc
      
      Differential Revision: https://reviews.llvm.org/D59839
      
      llvm-svn: 372027
      a7e90599
    • Lubos Lunak's avatar
      do not emit -Wunused-macros warnings in -frewrite-includes mode (PR15614) · a507a5ec
      Lubos Lunak authored
      -frewrite-includes calls PP.SetMacroExpansionOnlyInDirectives() to avoid
      macro expansions that are useless in that mode, but this can lead
      to -Wunused-macros false positives. As -frewrite-includes does not emit
      normal warnings, block -Wunused-macros too.
      
      Differential Revision: https://reviews.llvm.org/D65371
      
      llvm-svn: 372026
      a507a5ec
    • Vedant Kumar's avatar
      [Coverage] Speed up file-based queries for coverage info, NFC · 413647d7
      Vedant Kumar authored
      Speed up queries for coverage info in a file by reducing the amount of
      time spent determining whether a function record corresponds to a file.
      
      This gives a 36% speedup when generating a coverage report for `llc`.
      The reduction is entirely in user time.
      
      rdar://54758110
      
      Differential Revision: https://reviews.llvm.org/D67575
      
      llvm-svn: 372025
      413647d7
    • Vedant Kumar's avatar
      [Coverage] Assert that filenames in a TU are unique, NFC · 95de2497
      Vedant Kumar authored
      llvm-svn: 372024
      95de2497
    • Steven Wu's avatar
      [lld] Update lld driver to use new LTO APIs to handle libcall symbols · dd63b9f5
      Steven Wu authored
      NFC. Remove duplicated code in ELF/COFF driver and libLTO legacy
      interfaces.
      
      llvm-svn: 372022
      dd63b9f5
    • Steven Wu's avatar
      [LTO][Legacy] Add new C inferface to query libcall functions · 34d80461
      Steven Wu authored
      Summary:
      This is needed to implemented the same approach as lld (implemented in r338434)
      for how to handling symbols that can be generated by LTO code generator
      but not present in the symbol table for linker that uses legacy C APIs.
      
      libLTO is in charge of providing the list of symbols. Linker is in
      charge of implementing the eager loading from static libraries using
      the list of symbols.
      
      rdar://problem/52853974
      
      Reviewers: tejohnson, bd1976llvm, deadalnix, espindola
      
      Reviewed By: tejohnson
      
      Subscribers: emaste, arichardson, hiraditya, MaskRay, dang, kledzik, mehdi_amini, inglorion, jkorous, dexonsmith, ributzka, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D67568
      
      llvm-svn: 372021
      34d80461
    • Reid Kleckner's avatar
      [PGO] Use linkonce_odr linkage for __profd_ variables in comdat groups · 32837a0c
      Reid Kleckner authored
      This fixes relocations against __profd_ symbols in discarded sections,
      which is PR41380.
      
      In general, instrumentation happens very early, and optimization and
      inlining happens afterwards. The counters for a function are calculated
      early, and after inlining, counters for an inlined function may be
      widely referenced by other functions.
      
      For C++ inline functions of all kinds (linkonce_odr &
      available_externally mainly), instr profiling wants to deduplicate these
      __profc_ and __profd_ globals. Otherwise the binary would be quite
      large.
      
      I made __profd_ and __profc_ comdat in r355044, but I chose to make
      __profd_ internal. At the time, I was only dealing with coverage, and in
      that case, none of the instrumentation needs to reference __profd_.
      However, if you use PGO, then instrumentation passes add calls to
      __llvm_profile_instrument_range which reference __profd_ globals. The
      solution is to make these globals externally visible by using
      linkonce_odr linkage for data as was done for counters.
      
      This is safe because PGO adds a CFG hash to the names of the data and
      counter globals, so if different TUs have different globals, they will
      get different data and counter arrays.
      
      Reviewers: xur, hans
      
      Differential Revision: https://reviews.llvm.org/D67579
      
      llvm-svn: 372020
      32837a0c
    • Roman Lebedev's avatar
      [ARM][Codegen] Autogenerate arm-cgp-casts.ll test. · 69911b8d
      Roman Lebedev authored
      Apparently it got broken by r372009 while i thought it was r372012.
      
      llvm-svn: 372019
      69911b8d
    • Raphael Isemann's avatar
      [lldb] Remove SetCount/ClearCount from Flags · 0d8a0086
      Raphael Isemann authored
      Summary:
      These functions are only used in tests where we should test the actual flag values instead of counting all bits for an approximate check.
      Also these popcount implementation aren't very efficient and doesn't seem to be optimised to anything fast.
      
      Reviewers: davide, JDevlieghere
      
      Reviewed By: davide, JDevlieghere
      
      Subscribers: abidh, JDevlieghere, lldb-commits
      
      Tags: #lldb
      
      Differential Revision: https://reviews.llvm.org/D67540
      
      llvm-svn: 372018
      0d8a0086
    • Raphael Isemann's avatar
      [lldb][NFC] Make ApplyObjcCastHack less scary · 21641a2f
      Raphael Isemann authored
      llvm-svn: 372017
      21641a2f
    • Dan Albert's avatar
      Implement std::condition_variable via pthread_cond_clockwait() where available · 5e37d7f9
      Dan Albert authored
      std::condition_variable is currently implemented via
      pthread_cond_timedwait() on systems that use pthread. This is
      problematic, since that function waits by default on CLOCK_REALTIME
      and libc++ does not provide any mechanism to change from this
      default.
      
      Due to this, regardless of if condition_variable::wait_until() is
      called with a chrono::system_clock or chrono::steady_clock parameter,
      condition_variable::wait_until() will wait using CLOCK_REALTIME. This
      is not accurate to the C++ standard as calling
      condition_variable::wait_until() with a chrono::steady_clock parameter
      should use CLOCK_MONOTONIC.
      
      This is particularly problematic because CLOCK_REALTIME is a bad
      choice as it is subject to discontinuous time adjustments, that may
      cause condition_variable::wait_until() to immediately timeout or wait
      indefinitely.
      
      This change fixes this issue with a new POSIX function,
      pthread_cond_clockwait() proposed on
      http://austingroupbugs.net/view.php?id=1216. The new function is
      similar to pthread_cond_timedwait() with the addition of a clock
      parameter that allows it to wait using either CLOCK_REALTIME or
      CLOCK_MONOTONIC, thus allowing condition_variable::wait_until() to
      wait using CLOCK_REALTIME for chrono::system_clock and CLOCK_MONOTONIC
      for chrono::steady_clock.
      
      pthread_cond_clockwait() is implemented in glibc (2.30 and later) and
      Android's bionic (Android API version 30 and later).
      
      This change additionally makes wait_for() and wait_until() with clocks
      other than chrono::system_clock use CLOCK_MONOTONIC.<Paste>
      
      llvm-svn: 372016
      5e37d7f9
    • Roman Lebedev's avatar
      [Clang][Codegen] Disable arm_acle.c test. · 6fcd4e08
      Roman Lebedev authored
      This test is broken by design. Clang codegen tests should not depend
      on llvm middle-end behaviour, they should *only* test clang codegen.
      Yet this test runs whole optimization pipeline.
      I've really tried to fix it, but there isn't just a few things
      that depend on passes, but everything there does.
      
      llvm-svn: 372015
      6fcd4e08
    • Roman Lebedev's avatar
      [Clang][Codegen] Relax available-externally-suppress.c test · b9909ffe
      Roman Lebedev authored
      That test is broken by design.
      It depends on llvm middle-end behavior.
      No clang codegen test should be doing that.
      This one is salvageable by relaxing check lines.
      
      llvm-svn: 372014
      b9909ffe
    • Simon Pilgrim's avatar
      [X86][AVX] matchShuffleWithSHUFPD - add support for zeroable operands · 3df0dadd
      Simon Pilgrim authored
      Determine if all of the uses of LHS/RHS operands can be replaced with a zero vector.
      
      llvm-svn: 372013
      3df0dadd
    • David Green's avatar
      [ARM] A predicate cast of a predicate cast is a predicate cast · 8d21460d
      David Green authored
      The adds some very basic folding of PREDICATE_CASTS, removing cases when they
      are chained together. These would already be removed eventually, as these are
      lowered to copies. This just allows it to happen earlier, which can help other
      simplifications.
      
      Differential Revision: https://reviews.llvm.org/D67591
      
      llvm-svn: 372012
      8d21460d
    • Alexey Bataev's avatar
      [OPENMP]Fix parsing/sema for function templates with declare simd. · a0063078
      Alexey Bataev authored
      Need to return original declaration group with FunctionTemplateDecl, not
      the inner FunctionDecl, to correctly handle parsing of directives with
      the templates parameters.
      
      llvm-svn: 372011
      a0063078
    • Roman Lebedev's avatar
      [SimplifyCFG] FoldTwoEntryPHINode(): consider *total* speculation cost, not per-BB cost · 10151f66
      Roman Lebedev authored
      Summary:
      Previously, if the threshold was 2, we were willing to speculatively
      execute 2 cheap instructions in both basic blocks (thus we were willing
      to speculatively execute cost = 4), but weren't willing to speculate
      when one BB had 3 instructions and other one had no instructions,
      even thought that would have total cost of 3.
      
      This looks inconsistent to me.
      I don't think `cmov`-like instructions will start executing
      until both of it's inputs are available: https://godbolt.org/z/zgHePf
      So i don't see why the existing behavior is the correct one.
      
      Also, let's add it's own `cl::opt` for this threshold,
      with default=4, so it is not stricter than the previous threshold:
      will allow to fold when there are 2 BB's each with cost=2.
      And since the logic has changed, it will also allow to fold when
      one BB has cost=3 and other cost=1, or there is only one BB with cost=4.
      
      This is an alternative solution to D65148:
      This fix is mainly motivated by `signbit-like-value-extension.ll` test.
      That pattern comes up in JPEG decoding, see e.g.
      `Figure F.12 – Extending the sign bit of a decoded value in V`
      of `ITU T.81` (JPEG specification).
      That branch is not predictable, and it is within the innermost loop,
      so the fact that that pattern ends up being stuck with a branch
      instead of `select` (i.e. `CMOV` for x86) is unlikely to be beneficial.
      
      This has great results on the final assembly (vanilla test-suite + RawSpeed): (metric pass - D67240)
      | metric                                 |     old |     new | delta |      % |
      | x86-mi-counting.NumMachineFunctions    |   37720 |   37721 |     1 |  0.00% |
      | x86-mi-counting.NumMachineBasicBlocks  |  773545 |  771181 | -2364 | -0.31% |
      | x86-mi-counting.NumMachineInstructions | 7488843 | 7486442 | -2401 | -0.03% |
      | x86-mi-counting.NumUncondBR            |  135770 |  135543 |  -227 | -0.17% |
      | x86-mi-counting.NumCondBR              |  423753 |  422187 | -1566 | -0.37% |
      | x86-mi-counting.NumCMOV                |   24815 |   25731 |   916 |  3.69% |
      | x86-mi-counting.NumVecBlend            |      17 |      17 |     0 |  0.00% |
      
      We significantly decrease basic block count, notably decrease instruction count,
      significantly decrease branch count and very significantly increase `cmov` count.
      
      Performance-wise, unsurprisingly, this has great effect on
      target RawSpeed benchmark. I'm seeing 5 **major** improvements:
      ```
      Benchmark                                                                                             Time             CPU      Time Old      Time New       CPU Old       CPU New
      ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
      Samsung/NX3000/_3184416.SRW/threads:8/process_time/real_time_pvalue                                 0.0000          0.0000      U Test, Repetitions: 49 vs 49
      Samsung/NX3000/_3184416.SRW/threads:8/process_time/real_time_mean                                  -0.3064         -0.3064      226.9913      157.4452      226.9800      157.4384
      Samsung/NX3000/_3184416.SRW/threads:8/process_time/real_time_median                                -0.3057         -0.3057      226.8407      157.4926      226.8282      157.4828
      Samsung/NX3000/_3184416.SRW/threads:8/process_time/real_time_stddev                                -0.4985         -0.4954        0.3051        0.1530        0.3040        0.1534
      Kodak/DCS760C/86L57188.DCR/threads:8/process_time/real_time_pvalue                                  0.0000          0.0000      U Test, Repetitions: 49 vs 49
      Kodak/DCS760C/86L57188.DCR/threads:8/process_time/real_time_mean                                   -0.1747         -0.1747       80.4787       66.4227       80.4771       66.4146
      Kodak/DCS760C/86L57188.DCR/threads:8/process_time/real_time_median                                 -0.1742         -0.1743       80.4686       66.4542       80.4690       66.4436
      Kodak/DCS760C/86L57188.DCR/threads:8/process_time/real_time_stddev                                 +0.6089         +0.5797        0.0670        0.1078        0.0673        0.1062
      Sony/DSLR-A230/DSC08026.ARW/threads:8/process_time/real_time_pvalue                                 0.0000          0.0000      U Test, Repetitions: 49 vs 49
      Sony/DSLR-A230/DSC08026.ARW/threads:8/process_time/real_time_mean                                  -0.1598         -0.1598      171.6996      144.2575      171.6915      144.2538
      Sony/DSLR-A230/DSC08026.ARW/threads:8/process_time/real_time_median                                -0.1598         -0.1597      171.7109      144.2755      171.7018      144.2766
      Sony/DSLR-A230/DSC08026.ARW/threads:8/process_time/real_time_stddev                                +0.4024         +0.3850        0.0847        0.1187        0.0848        0.1175
      Canon/EOS 77D/IMG_4049.CR2/threads:8/process_time/real_time_pvalue                                  0.0000          0.0000      U Test, Repetitions: 49 vs 49
      Canon/EOS 77D/IMG_4049.CR2/threads:8/process_time/real_time_mean                                   -0.0550         -0.0551      280.3046      264.8800      280.3017      264.8559
      Canon/EOS 77D/IMG_4049.CR2/threads:8/process_time/real_time_median                                 -0.0554         -0.0554      280.2628      264.7360      280.2574      264.7297
      Canon/EOS 77D/IMG_4049.CR2/threads:8/process_time/real_time_stddev                                 +0.7005         +0.7041        0.2779        0.4725        0.2775        0.4729
      Canon/EOS 5DS/2K4A9929.CR2/threads:8/process_time/real_time_pvalue                                  0.0000          0.0000      U Test, Repetitions: 49 vs 49
      Canon/EOS 5DS/2K4A9929.CR2/threads:8/process_time/real_time_mean                                   -0.0354         -0.0355      316.7396      305.5208      316.7342      305.4890
      Canon/EOS 5DS/2K4A9929.CR2/threads:8/process_time/real_time_median                                 -0.0354         -0.0356      316.6969      305.4798      316.6917      305.4324
      Canon/EOS 5DS/2K4A9929.CR2/threads:8/process_time/real_time_stddev                                 +0.0493         +0.0330        0.3562        0.3737        0.3563        0.3681
      ```
      
      That being said, it's always best-effort, so there will likely
      be cases where this worsens things.
      
      Reviewers: efriedma, craig.topper, dmgreen, jmolloy, fhahn, Carrot, hfinkel, chandlerc
      
      Reviewed By: jmolloy
      
      Subscribers: xbolva00, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D67318
      
      llvm-svn: 372009
      10151f66
    • Ilya Biryukov's avatar
      [clangd] Simplify semantic highlighting visitor · 685d8a95
      Ilya Biryukov authored
      Summary:
      - Functions to compute highlighting kinds for things are separated from
        the ones that add highlighting tokens.
        This keeps each of them more focused on what they're doing: getting
        locations and figuring out the kind of the entity, correspondingly.
      
      - Less special cases in visitor for various nodes.
      
      This change is an NFC.
      
      Reviewers: hokein
      
      Reviewed By: hokein
      
      Subscribers: MaskRay, jkorous, arphaman, kadircet, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D67341
      
      llvm-svn: 372008
      685d8a95
    • Sanjay Patel's avatar
      [InstCombine] remove unneeded one-use checks for icmp fold · 3961a143
      Sanjay Patel authored
      Related folds were added in:
      rL125734
      ...the code comment about register pressure is discussed in
      more detail in:
      https://bugs.llvm.org/show_bug.cgi?id=2698
      
      But 10 years later, perf testing bzip2 with this change now
      shows a slight (0.2% average) improvement on Haswell although
      that's probably within test noise.
      
      Given that this is IR canonicalization, we shouldn't be worried
      about register pressure though; the backend should be able to
      adjust for that as needed.
      
      This is part of solving PR43310 the theoretically right way:
      https://bugs.llvm.org/show_bug.cgi?id=43310
      ...ie, if we don't cripple basic transforms, then we won't
      need to add special-case code to detect larger patterns.
      
      rL371940 and rL371981 are related patches in this series.
      
      llvm-svn: 372007
      3961a143
  2. Sep 16, 2019