1. Feb 04, 2020
    • Peter Collingbourne's avatar
      scudo: Use more size classes in the malloc_free_loop benchmarks. · 47cda0cb
      Peter Collingbourne authored
      As a result of recent changes to the Android size classes, the malloc_free_loop
      benchmark started exhausting the 8192 size class at 32768 iterations. To avoid
      this problem (and to make the test more realistic), change the benchmark to
      use a variety of size classes.
      
      Differential Revision: https://reviews.llvm.org/D73918
      47cda0cb
    • Artem Dergachev's avatar
      [analyzer] Suppress linker invocation in scan-build tests. · 4b05fc24
      Artem Dergachev authored
      This should fix PS4 buildbots.
      4b05fc24
    • Artem Dergachev's avatar
      [analyzer] Re-land 0aba69eb "Add test directory for scan-build." · 5521236a
      Artem Dergachev authored
      The tool is now looked for in the source directory rather than in the
      install directory, which should exclude the problems with not being able
      to find it.
      
      The tests still aren't being run on Windows, but they hopefully will run
      on other platforms that have shell, which hopefully also means Perl.
      
      Differential Revision: https://reviews.llvm.org/D69781
      5521236a
    • Matt Arsenault's avatar
      AMDGPU: Analyze divergence of inline asm · cb7b661d
      Matt Arsenault authored
      cb7b661d
    • Mitch Phillips's avatar
      [GWP-ASan] Allow late initialisation if single-threaded. · 0d6fccb4
      Mitch Phillips authored
      Summary:
      This patch allows for late initialisation of the GWP-ASan allocator. Previously, if late initialisation occurred, the sample counter was never updated, meaning we would end up having to wait for 2^32 allocations before getting a sampled allocation.
      
      Now, we initialise the sampling mechanism in init() as well. We require init() to be called single-threaded, so this isn't a problem.
      
      Reviewers: eugenis
      
      Reviewed By: eugenis
      
      Subscribers: merge_guards_bot, mgorny, #sanitizers, llvm-commits, cferris
      
      Tags: #sanitizers, #llvm
      
      Differential Revision: https://reviews.llvm.org/D73896
      0d6fccb4
    • Matt Arsenault's avatar
      2758ae41
    • Matt Arsenault's avatar
      AMDGPU: Fix splitting wide f32 s.buffer.load intrinsics · 726446a0
      Matt Arsenault authored
      This would witch f32 to i32, and produce an invald concat_vectors from
      i32 pieces to an f32 vector.
      726446a0
    • Petr Hosek's avatar
      Revert "[clang-doc] Improving Markdown Output" · 80e63c17
      Petr Hosek authored
      This reverts commit 0fbaf3a7 as tests
      are failing on some bots.
      80e63c17
    • David Tenty's avatar
      [AIX] Don't use a zero fill with a second parameter · 77e71c52
      David Tenty authored
      Summary:
      The AIX assembler .space directive can't take a second non-zero argument to fill
      with. But LLVM emitFill currently assumes it can. We add a flag to the AsmInfo
      to check if non-zero fill is supported, and if we can't zerofill non-zero values
      we just splat the .byte directives.
      
      Reviewers: stevewan, sfertile, DiggerLin, jasonliu, Xiangling_L
      
      Reviewed By: jasonliu
      
      Subscribers: Xiangling_L, wuzish, nemanjai, hiraditya, kbarton, jsji, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73554
      77e71c52
    • Petr Hosek's avatar
      [clang-doc] Improving Markdown Output · 0fbaf3a7
      Petr Hosek authored
      This change has two components. The moves the generated file
      for a namespace to the directory named after the namespace in
      a file named 'index.<format>'. This greatly improves the browsing
      experience since the index page is shown by default for a directory.
      
      The second improves the markdown output by adding the links to the
      referenced pages for children objects and the link back to the source
      code.
      
      Patch By: Clayton
      
      Differential Revision: https://reviews.llvm.org/D72954
      0fbaf3a7
    • Alexander Belyaev's avatar
      [MLIR][Linalg] Use GenericLoopNestRangeBuilder in tiling code. · 0da755df
      Alexander Belyaev authored
      Preparation for adding support for tiling to parallel loops.
      
      Differential Revision: https://reviews.llvm.org/D73872
      0da755df
    • Alexander Belyaev's avatar
      [MLIR][Linalg] Allow fusion of more than 2 linalg ops. · eda6b2e2
      Alexander Belyaev authored
      LinalgDependenceGraph was not updated after successful producer-consumer
      fusion for linalg ops. In this patch it is fixed by reconstructing
      LinalgDependenceGraph on every iteration. This is very ineffective and
      should be improved by updating LDGraph only when it is necessary.
      eda6b2e2
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Walk through G_AND in TB(N)Z bit calculation · 2bd46444
      Jessica Paquette authored
      Given
      
      ```
      tb(n)z (and x, m), b
      ```
      
      Where the `b`-th bit of `m` is 1,
      
      ```
      tb(n)z (and x, m), b == tb(n)z x, b
      ```
      
      So, we can walk past a `G_AND` in this case.
      
      Also add test/CodeGen/AArch64/GlobalISel/opt-fold-and-tbz-tbnz.mir to test this.
      
      Differential Revision: https://reviews.llvm.org/D73790
      2bd46444
    • Amara Emerson's avatar
      [AArch64][GlobalISel] Don't reconvert to p0 in convertPtrAddToAdd(). · b911b990
      Amara Emerson authored
      convertPtrAddToAdd improved overall code size and quality by a significant amount,
      but on -O0 we generate some cross-class copies due to the fact that we emitted
      G_PTRTOINT and G_INTTOPTR around the G_ADD. Unfortunately at -O0 we don't run any
      register coalescing, so these cross class copies end up escaping as moves, and
      we ended up regressing 3 benchmarks on CTMark (though still a winner overall).
      
      This patch changes the lowering to instead directly emit the G_ADD into the
      destination register, and then force changes the dest LLT to s64 from p0. This
      should be ok, as all uses of the register should now be selected and therefore
      the LLT doesn't matter for the users. It does however matter for the importer
      patterns, which will fail to select a G_ADD if there's a p0 LLT.
      
      I'm not able to get rid of the G_PTRTOINT on the source yet however. We can't
      use the same trick of breaking the type system since that could break the
      selection of the defining instruction. Thus with -O0 we still end up with a
      cross class copy on source.
      
      Code size improvements on -O0:
      Program                                         baseline      new         diff
       test-suite :: CTMark/Bullet/bullet.test        965520       949164      -1.7%
       test-suite...TMark/7zip/7zip-benchmark.test    1069456      1052600     -1.6%
       test-suite...ark/tramp3d-v4/tramp3d-v4.test    1213692      1199804     -1.1%
       test-suite...:: CTMark/sqlite3/sqlite3.test    421680       419736      -0.5%
       test-suite...-typeset/consumer-typeset.test    837076       833380      -0.4%
       test-suite :: CTMark/lencod/lencod.test        799712       796976      -0.3%
       test-suite...:: CTMark/ClamAV/clamscan.test    688264       686132      -0.3%
       test-suite :: CTMark/kimwitu++/kc.test         1002344      999648      -0.3%
       test-suite...Mark/mafft/pairlocalalign.test    422296       421768      -0.1%
       test-suite :: CTMark/SPASS/SPASS.test          656792       656532      -0.0%
       Geomean difference                                                      -0.6%
      
      Differential Revision: https://reviews.llvm.org/D73910
      b911b990
    • Matt Arsenault's avatar
      GlobalISel: Implement fewerElementsVector for G_SEXT_INREG · cd7650c1
      Matt Arsenault authored
      Start using a new strategy with a combination of merge and unmerges.
      
      This allows scalarizing before lowering, which in cases like
      <2 x s128> avoids producing giant illegal shifts.
      cd7650c1
    • Quentin Colombet's avatar
      [TargetRegisterInfo] Make the heuristic to skip region split overridable by the target · f26ff8c9
      Quentin Colombet authored
      RegAllocGreedy uses a fairly compile time intensive splitting heuristic
      called region splitting. This heuristic was disabled via another heuristic
      when it is likely that it won't be worth the compile time. The only way
      to control this other heuristic was via a command line option (huge-size-for-split).
      
      This commit gives more control on this heuristic by making it overridable
      by the target using a target hook in TargetRegisterInfo called
      shouldRegionSplitForVirtReg.
      
      The default implementation of this hook keeps the heuristic as it was
      before this patch.
      f26ff8c9
    • Nico Weber's avatar
      Fix a -Wbitwise-conditional-parentheses warning in _LIBUNWIND_ARM_EHABI libunwind builds · 221c5af4
      Nico Weber authored
      ```
      src/UnwindCursor.hpp:1344:51: error: operator '?:' has lower precedence than '|';
          '|' will be evaluated first [-Werror,-Wbitwise-conditional-parentheses]
        _info.flags = isSingleWordEHT ? 1 : 0 | scope32 ? 0x2 : 0;  // Use enum?
                                            ~~~~~~~~~~~ ^
      src/UnwindCursor.hpp:1344:51: note: place parentheses around the '|' expression
          to silence this warning
        _info.flags = isSingleWordEHT ? 1 : 0 | scope32 ? 0x2 : 0;  // Use enum?
                                                        ^
                                            (          )
      src/UnwindCursor.hpp:1344:51: note: place parentheses around the '?:' expression
          to evaluate it first
        _info.flags = isSingleWordEHT ? 1 : 0 | scope32 ? 0x2 : 0;  // Use enum?
                                                        ^
                                                (                )
      ```
      
      But `0 |` is a no-op for either of those two interpretati...
      221c5af4
    • Reid Kleckner's avatar
      Add PassManagerImpl.h to hide implementation details · 105642af
      Reid Kleckner authored
      ClangBuildAnalyzer results show that a lot of time is spent
      instantiating AnalysisManager::getResultImpl across the code base:
      
      **** Templates that took longest to instantiate:
       50445 ms: llvm::AnalysisManager<llvm::Function>::getResultImpl (412 times, avg 122 ms)
       47797 ms: llvm::AnalysisManager<llvm::Function>::getResult<llvm::TargetLibraryAnalysis> (389 times, avg 122 ms)
       46894 ms: std::tie<const unsigned long long, const bool> (2452 times, avg 19 ms)
       43851 ms: llvm::BumpPtrAllocatorImpl<llvm::MallocAllocator, 4096, 4096>::Allocate (3228 times, avg 13 ms)
       33911 ms: std::tie<const unsigned int, const unsigned int, const unsigned int, const unsigned int> (897 times, avg 37 ms)
       33854 ms: std::tie<const unsigned long long, const unsigned long long> (1897 times, avg 17 ms)
       27886 ms: std::basic_string<char, std::char_traits<char>, std::allocator<char> >::basic_string (11156 times, avg 2 ms)
      
      I mentioned this result to @chandlerc, and he suggested this direction.
      
      AnalysisManager is already explicitly instantiated, and getResultImpl
      doesn't need to be inlined. Move the definition to an Impl header, and
      include that header in files that explicitly instantiate
      AnalysisManager. There are only four (real) IR units:
      - function
      - module
      - loop
      - cgscc
      
      Looking at a specific transform (ArgumentPromotion.cpp), here are three
      compilations before & after this change:
      
      BEFORE:
      $ for i in $(seq 3) ; do ./ccit.bat ; done
      peak memory: 258.15MB
      real: 0m6.297s
      peak memory: 257.54MB
      real: 0m5.906s
      peak memory: 257.47MB
      real: 0m6.219s
      
      AFTER:
      $ for i in $(seq 3) ; do ./ccit.bat ; done
      peak memory: 235.35MB
      real: 0m5.454s
      peak memory: 234.72MB
      real: 0m5.235s
      peak memory: 234.39MB
      real: 0m5.469s
      
      The 20MB of memory saved seems real, and the time improvement seems like
      it is there.
      
      Reviewed By: MaskRay
      
      Differential Revision: https://reviews.llvm.org/D73817
      105642af
    • Reid Kleckner's avatar
      Revert "[SVE] Fix bug in simplification of scalable vector instructions" · a0544103
      Reid Kleckner authored
      This reverts commit 31574d38.
      
      The newly added shufflevector test does not pass locally on either of my
      workstations.
      a0544103
    • Michael Trent's avatar
      [llvm-objdump] Suppress spurious warnings when parsing Mach-O binaries. · 0ad18bf3
      Michael Trent authored
      Summary:
      llvm-objdump started warning when asked to disassemble a section that
      isn't present in the input files, in Yuanfang Chen's change:
      d16c162c. The problem is that the
      logic was restricted only to the generic llvm-objdump parser, not to the
      Mach-O-specific parser used for Apple toolchain compatibility. The
      solution is to log section names from the Mach-O parser.
      
      The macho-cstring-dump.test has been updated to fail if it encounters
      this new warning in the future.
      
      Reviewers: pete, ab, lhames, jhenderson, grimar, MaskRay, ychen
      
      Reviewed By: jhenderson, grimar
      
      Subscribers: rupprecht, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73586
      0ad18bf3
    • Alex Langford's avatar
      [lldb] Remove unused parameter from ValueObject::GetExpressionPath · 3014efe0
      Alex Langford authored
      I previously removed the code in ValueObject::GetExpressionPath that
      took advantage of the parameter `qualify_cxx_base_classes`. As a result,
      this is now unused and can be removed.
      3014efe0
    • Alex Langford's avatar
      [lldb] Delete ClangForward.h · 5b0c8dd3
      Alex Langford authored
      Summary:
      I think that there are very few things from clang that actually need forward
      declaration, so not having a ClangForward header makes sense.
      
      Differential Revision: https://reviews.llvm.org/D73827
      5b0c8dd3
    • Luboš Luňák's avatar
      [clang] detect switch fallthrough marked by a comment (PR43465) · 398b4ed8
      Luboš Luňák authored
      The regex can be extended if needed, but this should probably handle
      most of the cases.
      
      Differential Revision: https://reviews.llvm.org/D73852
      398b4ed8
    • Alina Sbirlea's avatar
      [LoopUtils] Make duplicate method a utility. [NFCI] · 388de9df
      Alina Sbirlea authored
      Summary:
      Method appendLoopsToWorklist is duplicate in LoopUnroll and in the
      LoopPassManager as an internal method. Make it an utility.
      
      Reviewers: dmgreen, chandlerc, fedor.sergeev, yamauchi
      
      Subscribers: mehdi_amini, hiraditya, zzheng, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73569
      388de9df
    • Christopher Tetreault's avatar
      [SVE] Fix bug in simplification of scalable vector instructions · 31574d38
      Christopher Tetreault authored
      Summary:
      * Most of the simplifications in SimplifyShuffleVectorInst depend on the
      concrete value of, or the length of the mask vector. For scalable
      vectors, this cannot be known at compile time.
      ** for these tests, detect if the vector is scalable before attempting
      the transformation
      * The functions ShuffleVectorInst::getMaskValue and
      ShuffleVectorInst::getShuffleMask access the value of the constant mask.
      However, since the length of the mask is unknown at compile time, these
      function do not work for scalable vectors. Add asserts to ensure that
      the input mask is not scalable
      
      Reviewers: efriedma, sdesmalen, apazos, chrisj, huihuiz
      
      Reviewed By: efriedma
      
      Subscribers: tschuett, hiraditya, rkruppe, psnobl, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73555
      31574d38
    • Nikita Popov's avatar
      [SimplifyLibCalls] Remove unused IRBuilder argument; NFC · 575a975a
      Nikita Popov authored
      isLocallyOpenedFile() does not use IRBuilder.
      575a975a
    • Nikita Popov's avatar
      [IRBuilder] Add missing NoFolder::CreatePointerBitCastOrAddrSpaceCast(); NFC · 23e3c3df
      Nikita Popov authored
      Split out from D73835. This method was added to ConstantFolder and
      TargetFolder, but not NoFolder.
      23e3c3df
    • Fangrui Song's avatar
      Revert "[CodeGenModule] Assume dso_local for -fpic -fno-semantic-interposition" · dbc96b51
      Fangrui Song authored
      This reverts commit 789a46f2.
      
      Accidentally committed.
      dbc96b51
    • Nikita Popov's avatar
      [IRBuilder] Remove unnecessary NoFolder methods; NFCI · 7c3becf4
      Nikita Popov authored
      Split out from D73835: These methods are not part of the
      ConstantFolder API and as such don't serve a purpose.
      7c3becf4
    • Simon Pilgrim's avatar
    • Nikita Popov's avatar
      [InstCombine] Add replaceOperand() helper · 878cb38a
      Nikita Popov authored
      Adds a replaceOperand() helper, which is like Instruction.setOperand()
      but adds the old operand to the worklist. This reduces the amount of
      missing or incorrect worklist management.
      
      This only applies the helper to a relatively small subset of
      setOperand() calls in InstCombine, namely those of the pattern
      `I.setOperand(); return &I;`, where it is most obviously applicable.
      
      Differential Revision: https://reviews.llvm.org/D73803
      878cb38a
    • Nikita Popov's avatar
      [InstCombine] Rename worklist methods; NFC · e6c9ab4f
      Nikita Popov authored
      This renames Worklist.AddDeferred() to Worklist.add() and
      Worklist.Add() to Worklist.push(). The intention here is that
      Worklist.add() should be the go-to method for explicit worklist
      management, while the raw Worklist.push() is mostly for
      InstCombine internals. I will then migrate uses of Worklist.push()
      to Worklist.add() in followup changes.
      
      As suggested by spatel on D73411 I'm also changing the remaining
      method names to lowercase first character, in line with current
      coding standards.
      
      Differential Revision: https://reviews.llvm.org/D73745
      e6c9ab4f
    • Fangrui Song's avatar
      [CodeGenModule] Assume dso_local for -fpic -fno-semantic-interposition · 789a46f2
      Fangrui Song authored
      Summary:
      Clang -fpic defaults to -fno-semantic-interposition (GCC -fpic defaults
      to -fsemantic-interposition).
      Users need to specify -fsemantic-interposition to get semantic
      interposition behavior.
      
      Semantic interposition is currently a best-effort feature. There may
      still be some cases where it is not handled well.
      
      Reviewers: peter.smith, rnk, serge-sans-paille, sfertile, jfb, jdoerfert
      
      Subscribers: dschuff, jyknight, dylanmckay, nemanjai, jvesely, kbarton, fedor.sergeev, asb, rbar, johnrusso, simoncook, sabuasal, niosHD, jrtc27, zzheng, edward-jones, atanasyan, rogfer01, MartinMosbeck, brucehoult, the_o, arphaman, PkmX, jocewei, jsji, Jim, lenary, s.egerton, pzheng, sameer.abuasal, apazos, luismarques, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D73865
      789a46f2
    • Nikita Popov's avatar
      [ARM] Expand vector reduction intrinsics on soft float · 1cc4f8d1
      Nikita Popov authored
      Followup to D73135. If the target doesn't have hard float (default
      for ARM), then we assert when trying to soften the result of vector
      reduction intrinsics. This patch marks these for expansion as well.
      (A bit odd to use vectors on a target without hard float ... but
      that's where you end up if you expose target-independent vector types.)
      
      Differential Revision: https://reviews.llvm.org/D73854
      1cc4f8d1
    • Nikita Popov's avatar
      [Examples] Link BitReader in ThinLtoJIT example · 9eb74f60
      Nikita Popov authored
      D72486 broke the shared library build.
      9eb74f60
    • Nikita Popov's avatar
      a5995405
    • Alexey Bataev's avatar
    • Alexey Bataev's avatar
      [OPENMP50]Codegen support for order(concurrent) clause. · a7815218
      Alexey Bataev authored
      Emit llvm parallel access metadata for the loops if they are marked as
      order(concurrent).
      a7815218
    • Teresa Johnson's avatar
      [ThinLTO] More efficient export computation (NFC) · bed4d9c8
      Teresa Johnson authored
      Summary:
      A recent change to enable more importing of global variables with
      references exposed some efficiency issues with export computation.
      See D73724 for more information and detailed analysis.
      
      The first was specific to variable importing. The code was marking every
      copy of a referenced value (from possibly thousands of files in the case
      of linkonce_odr) as exported, and we only need to mark the copy in the
      module containing the variable def being imported as exported. The
      reason is that this is tracking what values are newly exported as a
      result of importing. Anything that was defined in another module and
      simply used in the exporting module is already exported, and would have
      been identified by the caller (e.g. the LTO API implementations).
      
      The second issue is that the code was re-adding previously exported
      values (along with all references). It is easy to identify when a
      variable was already imported into the same module (via the
      import list insert call return value), and we already did this for
      function importing. However, what we weren't doing for either function
      or variable importing was avoiding a re-insertion when it was previously
      exported into a different importing module. The reason we couldn't do
      this is there was no way of telling from the export list whether it was
      previously inserted there because its definition was exported (in which
      case we already marked all its references as exported) from when it was
      inserted there because it was referenced by another exported value (in
      which case we haven't yet inserted its own references).
      
      To address this we can restructure the way the export list is
      constructed. This patch only adds the actual imported definitions
      (variable or function) to the export list for its module during the
      import computation. After import computation is complete, where we were
      already post-processing the export list we go ahead and add all
      references made by those exported values to the export list.
      
      These changes speed up the thin link not only with constant variable
      importing enabled, but also without (due to the efficiency improvement
      in function importing).
      
      Some thin link user time measurements for one large application, average
      of 5 runs:
      
      With constant variable importing enabled:
      - without this patch: 479.5s
      - with this patch: 74.6s
      
      Without constant variable importing enabled:
      - without this patch: 80.6s
      - with this patch: 70.3s
      
      Note I have not re-enabled constant variable importing here, as I would
      like to do additional compile time measurements with these fixes first.
      
      Reviewers: evgeny777
      
      Subscribers: mehdi_amini, inglorion, hiraditya, dexonsmith, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73851
      bed4d9c8
    • Jay Foad's avatar
      [AMDGPU] getMemOperandsWithOffset: add resource operand for BUF instructions · 05297b7c
      Jay Foad authored
      Summary:
      This prevents unwanted clustering of BUF instructions with the same
      vaddr but different resource descriptors.
      
      Reviewers: rampitec, arsenm, nhaehnle
      
      Subscribers: kzhuravl, jvesely, wdng, yaxunl, dstuttard, tpr, t-tye, hiraditya, kerbowa, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D73867
      05297b7c