1. Jan 19, 2017
    • Peter Collingbourne's avatar
      LowerTypeTests: Implement exporting of type identifiers. · 22d9d3cd
      Peter Collingbourne authored
      Type identifiers are exported by:
      - Adding coarse-grained information about how to test the type
        identifier to the summary.
      - Creating symbols in the object file (aliases and absolute symbols)
        containing fine-grained information about the type identifier.
      
      Differential Revision: https://reviews.llvm.org/D28424
      
      llvm-svn: 292462
      22d9d3cd
    • Justin Bogner's avatar
      GlobalISel: Implement narrowing for G_LOAD · d09c3ce6
      Justin Bogner authored
      llvm-svn: 292461
      d09c3ce6
    • Justin Bogner's avatar
      GlobalISel: Fix text wrapping in a comment. NFC · 1f5c5054
      Justin Bogner authored
      llvm-svn: 292460
      1f5c5054
    • Matthias Braun's avatar
      Use an actual valid register in test · 58f99615
      Matthias Braun authored
      llvm-svn: 292459
      58f99615
    • Dehao Chen's avatar
      Add -fdebug-info-for-profiling to emit more debug info for sample pgo profile collection · b3a70de7
      Dehao Chen authored
      Summary:
      SamplePGO uses profile with debug info to collect profile. Unlike the traditional debugging purpose, sample pgo needs more accurate debug info to represent the profile. We add -femit-accurate-debug-info for this purpose. It can be combined with all debugging modes (-g, -gmlt, etc). It makes sure that the following pieces of info is always emitted:
      
      * start line of all subprograms
      * linkage name of all subprograms
      * standalone subprograms (functions that has neither inlined nor been inlined)
      
      The impact on speccpu2006 binary size (size increase comparing with -g0 binary, also includes data for -g binary, which does not change with this patch):
      
                     -gmlt(orig) -gmlt(patched) -g
      433.milc       4.68%       5.40%          19.73%
      444.namd       8.45%       8.93%          45.99%
      447.dealII     97.43%      115.21%        374.89%
      450.soplex     27.75%      31.88%         126.04%
      453.povray     21.81%      26.16%         92.03%
      470.lbm        0.60%       0.67%          1.96%
      482.sphinx3    5.77%       6.47%          26.17%
      400.perlbench  17.81%      19.43%         73.08%
      401.bzip2      3.73%       3.92%          12.18%
      403.gcc        31.75%      34.48%         122.75%
      429.mcf        0.78%       0.88%          3.89%
      445.gobmk      6.08%       7.92%          42.27%
      456.hmmer      10.36%      11.25%         35.23%
      458.sjeng      5.08%       5.42%          14.36%
      462.libquantum 1.71%       1.96%          6.36%
      464.h264ref    15.61%      16.56%         43.92%
      471.omnetpp    11.93%      15.84%         60.09%
      473.astar      3.11%       3.69%          14.18%
      483.xalancbmk  56.29%      81.63%         353.22%
      geomean        15.60%      18.30%         57.81%
      
      Debug info size change for -gmlt binary with this patch:
      
      433.milc       13.46%
      444.namd       5.35%
      447.dealII     18.21%
      450.soplex     14.68%
      453.povray     19.65%
      470.lbm        6.03%
      482.sphinx3    11.21%
      400.perlbench  8.91%
      401.bzip2      4.41%
      403.gcc        8.56%
      429.mcf        8.24%
      445.gobmk      29.47%
      456.hmmer      8.19%
      458.sjeng      6.05%
      462.libquantum 11.23%
      464.h264ref    5.93%
      471.omnetpp    31.89%
      473.astar      16.20%
      483.xalancbmk  44.62%
      geomean        16.83%
      
      Reviewers: davidxl, andreadb, rob.lougher, dblaikie, echristo
      
      Reviewed By: dblaikie, echristo
      
      Subscribers: hfinkel, rob.lougher, andreadb, gbedwell, cfe-commits, probinson, llvm-commits, mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D25435
      
      llvm-svn: 292458
      b3a70de7
    • Dehao Chen's avatar
      Add -debug-info-for-profiling to emit more debug info for sample pgo profile collection · 1ce8d6ca
      Dehao Chen authored
      Summary:
      SamplePGO binaries built with -gmlt to collect profile. The current -gmlt debug info is limited, and we need some additional info:
      
      * start line of all subprograms
      * linkage name of all subprograms
      * standalone subprograms (functions that has neither inlined nor been inlined)
      
      This patch adds these information to the -gmlt binary. The impact on speccpu2006 binary size (size increase comparing with -g0 binary, also includes data for -g binary, which does not change with this patch):
      
                     -gmlt(orig) -gmlt(patched) -g
      433.milc       4.68%       5.40%          19.73%
      444.namd       8.45%       8.93%          45.99%
      447.dealII     97.43%      115.21%        374.89%
      450.soplex     27.75%      31.88%         126.04%
      453.povray     21.81%      26.16%         92.03%
      470.lbm        0.60%       0.67%          1.96%
      482.sphinx3    5.77%       6.47%          26.17%
      400.perlbench  17.81%      19.43%         73.08%
      401.bzip2      3.73%       3.92%          12.18%
      403.gcc        31.75%      34.48%         122.75%
      429.mcf        0.78%       0.88%          3.89%
      445.gobmk      6.08%       7.92%          42.27%
      456.hmmer      10.36%      11.25%         35.23%
      458.sjeng      5.08%       5.42%          14.36%
      462.libquantum 1.71%       1.96%          6.36%
      464.h264ref    15.61%      16.56%         43.92%
      471.omnetpp    11.93%      15.84%         60.09%
      473.astar      3.11%       3.69%          14.18%
      483.xalancbmk  56.29%      81.63%         353.22%
      geomean        15.60%      18.30%         57.81%
      
      Debug info size change for -gmlt binary with this patch:
      
      433.milc       13.46%
      444.namd       5.35%
      447.dealII     18.21%
      450.soplex     14.68%
      453.povray     19.65%
      470.lbm        6.03%
      482.sphinx3    11.21%
      400.perlbench  8.91%
      401.bzip2      4.41%
      403.gcc        8.56%
      429.mcf        8.24%
      445.gobmk      29.47%
      456.hmmer      8.19%
      458.sjeng      6.05%
      462.libquantum 11.23%
      464.h264ref    5.93%
      471.omnetpp    31.89%
      473.astar      16.20%
      483.xalancbmk  44.62%
      geomean        16.83%
      
      Reviewers: davidxl, echristo, dblaikie
      
      Reviewed By: echristo, dblaikie
      
      Subscribers: aprantl, probinson, llvm-commits, mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D25434
      
      llvm-svn: 292457
      1ce8d6ca
    • Michael Kuperstein's avatar
      [LV] Run loop-simplify and LCSSA explicitly instead of "requiring" them · 230867e5
      Michael Kuperstein authored
      This changes the vectorizer to explicitly use the loopsimplify and lcssa utils,
      instead of "requiring" the transformations as if they were analyses.
      
      This is not NFC, since it changes the LCSSA behavior - we no longer run LCSSA
      for all loops, but rather only for the loops we expect to modify.
      
      Differential Revision: https://reviews.llvm.org/D28868
      
      llvm-svn: 292456
      230867e5
    • Matthias Braun's avatar
      LiveIntervalAnalysis: Cleanup; NFC · 9f21a8d7
      Matthias Braun authored
      - Fix doxygen comments: Do not repeat name, remove duplicated doxygen
        comment (on declaration + implementation), etc.
      - Use more range based for
      
      llvm-svn: 292455
      9f21a8d7
    • Jason Molenda's avatar
      Fix a problem with the new dyld interface code -- when a new process · 848c7be0
      Jason Molenda authored
      starts up, we need to clear the target's image list and only add
      the binaries into the target that are actually present in this
      process run.
      
      <rdar://problem/29857613> 
      
      llvm-svn: 292454
      848c7be0
    • Artem Belevich's avatar
      [NVPTX] Fix lowering of fp16 ISD::FNEG. · 3d3f6190
      Artem Belevich authored
      There's no neg.f16 instruction, so negation has to
      be done via subtraction from zero.
      
      Differential Revision: https://reviews.llvm.org/D28876
      
      llvm-svn: 292452
      3d3f6190
    • Peter Collingbourne's avatar
      Add llvm-dis dependency to check-clang. · 87cdfa76
      Peter Collingbourne authored
      llvm-svn: 292450
      87cdfa76
    • Eli Friedman's avatar
      [SCEV] Make getUDivExactExpr handle non-nuw multiplies correctly. · f1f49c82
      Eli Friedman authored
      To avoid regressions, make ScalarEvolution::createSCEV a bit more
      clever.
      
      Also get rid of some useless code in ScalarEvolution::howFarToZero
      which was hiding this bug.
      
      No new testcase because it's impossible to actually expose this bug:
      we don't have any in-tree users of getUDivExactExpr besides the two
      functions I just mentioned, and they both dodged the problem. I'll
      try to add some interesting users in a followup.
      
      Differential Revision: https://reviews.llvm.org/D28587
      
      llvm-svn: 292449
      f1f49c82
    • Peter Collingbourne's avatar
      Move vtable type metadata emission behind a cc1-level flag. · 1e1475ac
      Peter Collingbourne authored
      In ThinLTO mode, type metadata will require the module to be written as a
      multi-module bitcode file, which is currently incompatible with the Darwin
      linker. It is also useful to be able to enable or disable multi-module bitcode
      for testing purposes. This introduces a cc1-level flag, -f{,no-}lto-unit,
      which is used by the driver to enable multi-module bitcode on all but
      Darwin+ThinLTO, and can also be used to enable/disable the feature manually.
      
      Differential Revision: https://reviews.llvm.org/D28877
      
      llvm-svn: 292448
      1e1475ac
    • Eli Friedman's avatar
      Preserve domtree and loop-simplify for runtime unrolling. · 0a217453
      Eli Friedman authored
      Mostly straightforward changes; we just didn't do the computation before.
      One sort of interesting change in LoopUnroll.cpp: we weren't handling
      dominance for children of the loop latch correctly, but
      foldBlockIntoPredecessor hid the problem for complete unrolling.
      
      Currently punting on loop peeling; made some minor changes to isolate
      that problem to LoopUnrollPeel.cpp.
      
      Adds a flag -unroll-verify-domtree; it verifies the domtree immediately
      after we finish updating it. This is on by default for +Asserts builds.
      
      Differential Revision: https://reviews.llvm.org/D28073
      
      llvm-svn: 292447
      0a217453
    • Krzysztof Parzyszek's avatar
      de44c9d8
    • Krzysztof Parzyszek's avatar
      954dd8d9
    • Michael Kuperstein's avatar
      Revert r291670 because it introduces a crash. · d3d29259
      Michael Kuperstein authored
      r291670 doesn't crash on the original testcase from PR31589,
      but it crashes on a slightly more complex one.
      
      PR31589 has the new reproducer.
      
      llvm-svn: 292444
      d3d29259
    • Stephan T. Lavavej's avatar
      [libcxx] [test] Add msvc_stdlib_force_include.hpp. · d6c0b35c
      Stephan T. Lavavej authored
      No functional change; nothing includes this, instead our test harness
      injects it via the /FI compiler option.
      
      No code review; blessed in advance by EricWF.
      
      llvm-svn: 292443
      d6c0b35c
    • Mehdi Amini's avatar
      Improve the `-filter-print-funcs` option to skip the banner for CGSCC pass... · 062b3fed
      Mehdi Amini authored
      Improve the `-filter-print-funcs` option to skip the banner for CGSCC pass when nothing is to be printed
      
      Before, it would print a sequence of:
      
        *** IR Dump After Function Integration/Inlining ******
        *** IR Dump After Function Integration/Inlining ******
        *** IR Dump After Function Integration/Inlining ******
        ...
      
      for every single function in the module.
      
      llvm-svn: 292442
      062b3fed
    • Sanjay Patel's avatar
      [InstCombine] add tests for shl nsw with icmp eq/ne; NFCI · cfb8a459
      Sanjay Patel authored
      These should be fixed with D28406.
      
      llvm-svn: 292441
      cfb8a459
    • Sanjay Patel's avatar
      ae23d65a
    • David Blaikie's avatar
      Remove now redundant code that ensured debug info for class definitions was... · 75ed8ad6
      David Blaikie authored
      Remove now redundant code that ensured debug info for class definitions was emitted under certain circumstances
      
      Introduced in r181561 - it may've been subsumed by work done to allow
      emission of declarations for vtable types while still emitting some of
      their member functions correctly for those declarations. Whatever the
      reason, the tests pass without this code now.
      
      llvm-svn: 292439
      75ed8ad6
    • Haicheng Wu's avatar
      [CodeGenPrepare] Fix a typo in the comment. NFC. · 8ce2d143
      Haicheng Wu authored
      encode => endcode.
      
      Differential Revision: https://reviews.llvm.org/D28866
      
      llvm-svn: 292438
      8ce2d143
    • Arpith Chacko Jacob's avatar
      [OpenMP] Support for the if-clause on the combined directive 'target parallel'. · fe4890a6
      Arpith Chacko Jacob authored
      The if-clause on the combined directive potentially applies to both the
      'target' and the 'parallel' regions.  Codegen'ing the if-clause on the
      combined directive requires additional support because the expression in
      the clause must be captured by the 'target' capture statement but not
      the 'parallel' capture statement.  Note that this situation arises for
      other clauses such as num_threads.
      
      The OMPIfClause class inherits OMPClauseWithPreInit to support capturing
      of expressions in the clause.  A member CaptureRegion is added to
      OMPClauseWithPreInit to indicate which captured statement (in this case
      'target' but not 'parallel') captures these expressions.
      
      To ensure correct codegen of captured expressions in the presence of
      combined 'target' directives, OMPParallelScope was added to 'parallel'
      codegen.
      
      Reviewers: ABataev
      Differential Revision: https://reviews.llvm.org/D28781
      
      llvm-svn: 292437
      fe4890a6
    • Graydon Hoare's avatar
      [ASTReader] Add a DeserializationListener callback for IMPORTED_MODULES · 9c982440
      Graydon Hoare authored
      Summary:
      Add a callback from ASTReader to DeserializationListener when the former
      reads an IMPORTED_MODULES block. This supports Swift in using PCH for
      bridging headers.
      
      Reviewers: doug.gregor, manmanren, bruno
      
      Reviewed By: manmanren
      
      Subscribers: cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D28779
      
      llvm-svn: 292436
      9c982440
    • Graydon Hoare's avatar
      [Modules] Correct test comment from obsolete earlier version of code. NFC · dc0405f7
      Graydon Hoare authored
      Summary:
      Code committed in rL290219 went through a few iterations; test wound up with
      stale comment.
      
      Reviewers: doug.gregor, manmanren
      
      Reviewed By: manmanren
      
      Subscribers: cfe-commits
      
      Differential Revision: https://reviews.llvm.org/D28790
      
      llvm-svn: 292435
      dc0405f7
    • Stephan T. Lavavej's avatar
      [libcxx] [test] Fix comment typos, strip trailing whitespace. · a730ed31
      Stephan T. Lavavej authored
      No functional change, no code review.
      
      llvm-svn: 292434
      a730ed31
    • Sanjay Patel's avatar
      [InstCombine] remove a redundant check; NFCI · 589de5ea
      Sanjay Patel authored
      I missed deleting this check when I refactored this chunk in:
      https://reviews.llvm.org/rL292260
      
      llvm-svn: 292433
      589de5ea
    • Stephan T. Lavavej's avatar
      [libcxx] [test] Fix MSVC warnings C4127 and C6326 about constants. · 3d26ee29
      Stephan T. Lavavej authored
      MSVC has compiler warnings C4127 "conditional expression is constant" (enabled
      by /W4) and C6326 "Potential comparison of a constant with another constant"
      (enabled by /analyze). They're potentially useful, although they're slightly
      annoying to library devs who know what they're doing. In the latest version of
      the compiler, C4127 is suppressed when the compiler sees simple tests like
      "if (name_of_thing)", so extracting comparison expressions into named
      constants is a workaround. At the same time, using std::integral_constant
      avoids C6326, which doesn't look at template arguments.
      
      test/std/containers/sequences/vector.bool/emplace.pass.cpp
      Replace 1 == 1 with true, which is the same as far as the library is concerned.
      
      Fixes D28837.
      
      llvm-svn: 292432
      3d26ee29
    • Peter Collingbourne's avatar
      ThinLTOBitcodeWriter: Clear comdats on filtered globals. · 20a00933
      Peter Collingbourne authored
      Differential Revision: https://reviews.llvm.org/D28839
      
      llvm-svn: 292431
      20a00933
    • Peter Collingbourne's avatar
      Cloning: Copy comdats when cloning globals. · 10e3b12c
      Peter Collingbourne authored
      Differential Revision: https://reviews.llvm.org/D28838
      
      llvm-svn: 292430
      10e3b12c
    • Arpith Chacko Jacob's avatar
      [OpenMP] Codegen for the 'target parallel' directive on the NVPTX device. · 44a87c9f
      Arpith Chacko Jacob authored
      This patch adds codegen for the 'target parallel' directive on the NVPTX
      device.  We term offload OpenMP directives such as 'target parallel' and
      'target teams distribute parallel for' as SPMD constructs.  SPMD constructs,
      in contrast to Generic ones like the plain 'target', can never contain
      a serial region.
      
      SPMD constructs can be handled more efficiently on the GPU and do not
      require the Warp Loop of the Generic codegen scheme. This patch adds
      SPMD codegen support for 'target parallel' on the NVPTX device and can
      be reused for other SPMD constructs.
      
      Reviewers: ABataev
      Differential Revision: https://reviews.llvm.org/D28755
      
      llvm-svn: 292428
      44a87c9f
    • Richard Smith's avatar
      PR9551: Implement DR1004 (http://wg21.link/cwg1004). · 11255ec7
      Richard Smith authored
      This rule permits the injected-class-name of a class template to be used as
      both a template type argument and a template template argument, with no extra
      syntax required to disambiguate.
      
      llvm-svn: 292426
      11255ec7
    • Michael Kuperstein's avatar
      Fix up a comment. NFC. · 0de990da
      Michael Kuperstein authored
      llvm-svn: 292425
      0de990da
    • Michael Kuperstein's avatar
      [LV] Allow reductions that have several uses outside the loop · 7cefb409
      Michael Kuperstein authored
      We currently check whether a reduction has a single outside user. We don't
      really need to require that - we just need to make sure a single value is
      used externally. The number of external users of that value shouldn't actually
      matter.
      
      Differential Revision: https://reviews.llvm.org/D28830
      
      llvm-svn: 292424
      7cefb409
    • Justin Bogner's avatar
      cmake: Only sanitize use-after-scope if the host compiler supports it · 2ceeb30e
      Justin Bogner authored
      In r292256, we started adding -fsanitize-use-after-scope when using
      the address sanitizer, but that flag wasn't always available. This
      fixes the config to only add the flag if the host compiler supports
      it.
      
      llvm-svn: 292423
      2ceeb30e
    • Evandro Menezes's avatar
      [AArch64] Generate literals by the little end · 7960b2e1
      Evandro Menezes authored
      ARM seems to prefer that long literals be formed from their little end in
      order to promote the fusion of the instrs pairs MOV/MOVK and MOVK/MOVK on
      Cortex A57 and others (v.  "Cortex A57 Software Optimisation Guide", section
      4.14).
      
      Differential revision: https://reviews.llvm.org/D28697
      
      llvm-svn: 292422
      7960b2e1
    • Davide Italiano's avatar
      [NewGVN] We don't use postdom info anymore. Update. · bca9d733
      Davide Italiano authored
      Differential Revision:  https://reviews.llvm.org/D28842
      
      llvm-svn: 292421
      bca9d733
    • Mehdi Amini's avatar
      [ThinLTO] Add a recursive step in Metadata lazy-loading · 67d2cc1f
      Mehdi Amini authored
      Summary:
      Without this, we're stressing the RAUW of unique nodes,
      which is a costly operation. This is intended to limit
      the number of RAUW, and is very effective on the total
      link-time of opt with ThinLTO, before:
      
        real 4m4.587s  user 15m3.401s  sys 0m23.616s
      
      after:
      
        real 3m25.261s user 12m22.132s sys 0m24.152s
      
      Reviewers: tejohnson, pcc
      
      Subscribers: llvm-commits
      
      Differential Revision: https://reviews.llvm.org/D28751
      
      llvm-svn: 292420
      67d2cc1f
    • Arpith Chacko Jacob's avatar
      [OpenMP] Codegen support for 'target parallel' on the host. · 19b911cb
      Arpith Chacko Jacob authored
      This patch adds support for codegen of 'target parallel' on the host.
      It is also the first combined directive that requires two or more
      captured statements.  Support for this functionality is included in
      the patch.
      
      A combined directive such as 'target parallel' has two captured
      statements, one for the 'target' and the other for the 'parallel'
      region.  Two captured statements are required because each has
      different implicit parameters (see SemaOpenMP.cpp).  For example,
      the 'parallel' has 'global_tid' and 'bound_tid' while the 'target'
      does not.  The patch adds support for handling multiple captured
      statements based on the combined directive.
      
      When codegen'ing the 'target parallel' directive, the 'target'
      outlined function is created using the outer captured statement
      and the 'parallel' outlined function is created using the inner
      captured statement.
      
      Reviewers: ABataev
      Differential Revision: https://reviews.llvm.org/D28753
      
      llvm-svn: 292419
      19b911cb