1. Jun 02, 2020
    • Craig Topper's avatar
      [X86] Fix a few recursivelyDeleteUnusedNodes calls that were trying to delete... · e51d5bc7
      Craig Topper authored
      [X86] Fix a few recursivelyDeleteUnusedNodes calls that were trying to delete nodes before their user was really gone.
      
      We looked through a truncate to get to the load. So we should be
      deleting the truncate first.
      
      There is a check that the node is really unused before deleting
      so this didn't cause a functional issue.
      e51d5bc7
    • Yevgeny Rouban's avatar
      3bb0d95f
    • Kostya Serebryany's avatar
      [asan] fix a comment typo · 801d823b
      Kostya Serebryany authored
      801d823b
    • Kostya Serebryany's avatar
      add debug code to chase down a rare crash in asan/lsan... · 2e6c3e3e
      Kostya Serebryany authored
      add debug code to chase down a rare crash in asan/lsan https://github.com/google/sanitizers/issues/1193
      
      Summary: add debug code to chase down a rare crash in asan/lsan https://github.com/google/sanitizers/issues/1193
      
      Reviewers: vitalybuka
      
      Subscribers: #sanitizers, llvm-commits
      
      Tags: #sanitizers
      
      Differential Revision: https://reviews.llvm.org/D80967
      2e6c3e3e
    • John McCall's avatar
      Fix how cc1 command line options are mapped into FP options. · 8a8d703b
      John McCall authored
      Canonicalize on storing FP options in LangOptions instead of
      redundantly in CodeGenOptions.  Incorporate -ffast-math directly
      into the values of those LangOptions rather than considering it
      separately when building FPOptions.  Build IR attributes from
      those options rather than a mix of sources.
      
      We should really simplify the driver/cc1 interaction here and have
      the driver pass down options that cc1 directly honors.  That can
      happen in a follow-up, though.
      
      Patch by Michele Scandale!
      https://reviews.llvm.org/D80315
      8a8d703b
    • Reid Kleckner's avatar
      [COFF] Free some memory used for chunks · 11d1aa0b
      Reid Kleckner authored
      First, do not reserve numSections in the Chunks array. In cases where
      there are many non-prevailing sections, this will overallocate memory
      which will not be used.
      
      Second, free the memory for sparseChunks after initializeSymbols. After
      that, it is never used.
      
      This saves 50MB of 627MB for my use case without affecting performance.
      11d1aa0b
    • Adrian Prantl's avatar
      Fix UB in EmulateInstructionARM64.cpp · a0b674fd
      Adrian Prantl authored
      This fixes an unhandled signed integer overflow in AddWithCarry() by
      using the llvm::checkedAdd() function. Thats to Vedant Kumar for the
      suggestion!
      
      <rdar://problem/60926115>
      
      Differential Revision: https://reviews.llvm.org/D80955
      a0b674fd
    • Vedant Kumar's avatar
      [os_log][test] Remove -O1 from a test, NFC · a66e1d2a
      Vedant Kumar authored
      a66e1d2a
    • Vedant Kumar's avatar
      [docs] Sketch outline for HowToUpdateDebugInfo.rst · b429a0fe
      Vedant Kumar authored
      Summary:
      Sketch the outline for a new document that explains how to update debug
      info in various kinds of code transformations.
      
      Some of the guidelines that belong in HowToUpdateDebugInfo.rst were in
      SourceLevelDebugging.rst already under the debugify section. It seems
      like the distinction between the two docs ought to be that the former is
      more prescriptive, while the latter is more descriptive.
      
      To that end I've consolidated the "how to update debug info" guidelines
      which were in SourceLevelDebugging.rst into the new doc, along with the
      information about using "debugify" to test transformations. Since we've
      added a mir-debugify pass, I've described that as well.
      
      Reviewers: aprantl, jmorse, chrisjackson, dsanders
      
      Subscribers: llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D80052
      b429a0fe
    • Amara Emerson's avatar
      [AArch64][GlobalISel] Split G_GLOBAL_VALUE into ADRP + G_ADD_LOW and optimize. · f573d489
      Amara Emerson authored
      The concept of G_GLOBAL_VALUE is nice and simple, but always using it as the
      representation for global var addressing until selection time creates some
      problems in optimizing accesses in certain code/relocation models.
      
      The problem comes from trying to optimize adrp -> add -> load/store sequences
      in the most common "small" code model. These accesses can be optimized into an
      adrp -> load with the add offset being folded into the load's immediate field.
      If we try to keep all global var references as a single generic instruction
      then by the time we get to the complex operand trying to match these, we end up
      generating an adrp at the point of use. The real issue here is that we don't
      have any form of CSE during selection, so the code size will bloat from many
      redundant adrp's.
      
      This patch custom legalizes small code mode non-GOT G_GLOBALs into target ADRP
      and a new "target specific generic opcode" G_ADD_LOW. We also teach the
      localizer to localize these instructions via the custom hook that was added
      recently. Finally, the complex pattern for indexed loads/stores is extended to
      try to fold these G_ADD_LOW instructions into the load immediate.
      
      On -O0 CTMark, we see a 0.8% geomean code size improvement. We should also see
      some minor performance improvements too.
      
      Differential Revision: https://reviews.llvm.org/D78465
      f573d489
    • Amara Emerson's avatar
      [AArch64] Fix CollectLOH creating an AdrpAdd LOH when there's a live used reg · 19ff00da
      Amara Emerson authored
      between the two instructions.
      
      If there's a pattern like:
      $xA = ADRP foo @PAGE
      [some killing use of reg Xb]
      $Xb = ADDXri $Xa, 0, @PAGEOFF
      
      CollectLOH would create an AdrpAdd LOH that resulted in the linker optimizing
      this sequence into:
      $xB = ADR foo
      [some killing use of reg $Xb]
      ... and therefore clobbers the live $Xb register that was used by the
      instruction in between.
      
      This was discovered by a GlobalISel patch D78465 which broke up global variable
      accesses into two pseudos, which in some cases could be moved apart.
      
      Differential Revision: https://reviews.llvm.org/D80834
      19ff00da
    • Vedant Kumar's avatar
      [LiveDebugValues] Remove early-exit when testing regmasks, NFC · 776708b0
      Vedant Kumar authored
      In transferRegisterDef, if the instruction has a regmask attached, we'll
      check if any currently used register is clobbered by the regmask.
      
      The early exit in this scan isn't necessary, costs a set lookup, and is
      almost never taken [1]. Delete it.
      
      [1]
      http://lab.llvm.org:8080/coverage/coverage-reports/coverage/Users/buildslave/jenkins/workspace/coverage/llvm-project/llvm/lib/CodeGen/LiveDebugValues.cpp.html#L1136
      776708b0
    • Matt Arsenault's avatar
      AMDGPU: Change internal tracking of wave size · a8f72092
      Matt Arsenault authored
      Store the log2 wave size instead of forcing division and log2
      operations when querying either.
      a8f72092
    • Olivier Giroux's avatar
    • Akira Hatanaka's avatar
      Clean up clang/test/CodeGenObjC/os_log.m · 959517ac
      Akira Hatanaka authored
      Don't run optimization passes at -O2 and remove unneeded #ifdef and test
      cases.
      959517ac
    • Kirstóf Umann's avatar
      [analyzer][MallocChecker] Fix the incorrect retrieval of the from argument in realloc() · 6bedfaf5
      Kirstóf Umann authored
      In the added testfile, the from argument was recognized as
      &Element{SymRegion{reg_$0<long * global_a>},-1 S64b,long}
      instead of
      reg_$0<long * global_a>.
      6bedfaf5
    • Louis Dionne's avatar
      [libc++] Add assertions on OOB accesses in std::array when the debug mode is enabled · 23776a17
      Louis Dionne authored
      Like we do for empty std::array, make sure we have assertions in place
      for obvious out-of-bounds issues in std::array when the debug mode is
      enabled (which isn't by default).
      23776a17
    • Lei Huang's avatar
      [PowerPC] Add clang option -m[no-]pcrel · 7cfded35
      Lei Huang authored
      Summary:
      Add user-facing front end option to turn off pc-relative memops.
      This will be compatible with gcc.
      
      Reviewers: stefanp, nemanjai, hfinkel, power-llvm-team, #powerpc, NeHuang, saghir
      
      Reviewed By: stefanp, NeHuang, saghir
      
      Subscribers: saghir, wuzish, shchenz, cfe-commits, kbarton, echristo
      
      Tags: #clang, #powerpc
      
      Differential Revision: https://reviews.llvm.org/D80757
      7cfded35
    • Louis Dionne's avatar
      66a14d15
    • Joseph Huber's avatar
      [OpenMP] Replace Clang's OpenMP RTL Definitions with OMPKinds.def · 1a4fb2ed
      Joseph Huber authored
      Summary: This changes Clang's generation of OpenMP runtime functions to use the types and functions defined in OpenMPKinds and OpenMPConstants. New OpenMP runtime function information should now be added to OMPKinds.def. This patch also changed the definitions of __kmpc_push_num_teams and __kmpc_copyprivate to match those found in the runtime.
      
      Reviewers: jdoerfert
      
      Reviewed By: jdoerfert
      
      Subscribers: jfb, AndreyChurbanov, openmp-commits, fghanim, hiraditya, sstefan1, cfe-commits, llvm-commits
      
      Tags: #openmp, #clang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D80222
      1a4fb2ed
    • Reid Kleckner's avatar
      [PDB] Share code to relocate .debug$[SF] sections, NFC · 45fd3e46
      Reid Kleckner authored
      Sink relocateDebugChunk near the only call site.
      45fd3e46
    • Sterling Augustine's avatar
      For --relativenames, ignore directory 0, which is the comp_dir. · f027cfa3
      Sterling Augustine authored
      Update for upstream comments. Improve test by writing all the debug
      info by hand.
      
      Reviewers: dblaikie, jhenderson
      
      Subscribers: hiraditya, MaskRay, rupprecht, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D80168
      f027cfa3
    • Jonas Devlieghere's avatar
    • Mircea Trofin's avatar
      [llvm][NFC] Cache FAM in InlineAdvisor · 999ea25a
      Mircea Trofin authored
      Summary:
      This simplifies the interface by storing the function analysis manager
      with the InlineAdvisor, and, thus, not requiring it be passed each time
      we inquire for an advice.
      
      Reviewers: davidxl, asbirlea
      
      Subscribers: eraman, hiraditya, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D80405
      999ea25a
    • Daniel Grumberg's avatar
    • Paula Toth's avatar
      [libc] Expose APIGenerator. · 1ab092b7
      Paula Toth authored
      Summary: This is split off from D79192 and exposes APIGenerator (renames to APIIndexer) for use in generating the integrations tests.
      
      Reviewers: sivachandra
      
      Reviewed By: sivachandra
      
      Subscribers: tschuett, ecnelises, libc-commits
      
      Tags: #libc-project
      
      Differential Revision: https://reviews.llvm.org/D80832
      1ab092b7
    • Reid Kleckner's avatar
      [PDB] Use inlinee file checksum offsets directly · 8f0a6600
      Reid Kleckner authored
      The inlinees section contains references to the file checksum table. The
      file checksum table in the PDB must have the same layout as the file
      checksum table in the object file, so all the existing file id
      references should stay valid.
      
      Previously, we would do this:
        for all inlined functions:
          - lookup filename from checksum and string table
          - make that filename absolute
          - look up the new file id for that filename up in the new checksum
            table
      
      This lead to pdbMakeAbsolute and remove_dots ending up in the hot path.
      We should only need to absolutify the source path once, not once every
      time we process an inline function from that source file.
      
      This speeds up linking chrome PGO stage 1 net_unittests.exe from 9.203s
      to 8.500s (-7.6%). Looking just at time to process symbol records, it
      goes from ~2000ms to ~1300ms, which is consistent with the overall
      speedup of about 700ms. This will be less noticeable in debug builds,
      which have fewer inlined functions records.
      8f0a6600
    • Florian Hahn's avatar
      [Matrix] Implement matrix index expressions ([][]). · 8f3f88d2
      Florian Hahn authored
      This patch implements matrix index expressions
      (matrix[RowIdx][ColumnIdx]).
      
      It does so by introducing a new MatrixSubscriptExpr(Base, RowIdx, ColumnIdx).
      MatrixSubscriptExprs are built in 2 steps in ActOnMatrixSubscriptExpr. First,
      if the base of a subscript is of matrix type, we create a incomplete
      MatrixSubscriptExpr(base, idx, nullptr). Second, if the base is an incomplete
      MatrixSubscriptExpr, we create a complete
      MatrixSubscriptExpr(base->getBase(), base->getRowIdx(), idx)
      
      Similar to vector elements, it is not possible to take the address of
      a MatrixSubscriptExpr.
      For CodeGen, a new MatrixElt type is added to LValue, which is very
      similar to VectorElt. The only difference is that we may need to cast
      the type of the base from an array to a vector type when accessing it.
      
      Reviewers: rjmccall, anemet, Bigcheese, rsmith, martong
      
      Reviewed By: rjmccall
      
      Differential Revision: https://reviews.llvm.org/D76791
      8f3f88d2
    • Martin Liska's avatar
      Move internal_uname to #if SANITIZER_LINUX scope. · b638b63b
      Martin Liska authored
      Remove it from target-specific scope which corresponds
      to sanitizer_linux.cpp where it lives in the same macro
      scope.
      
      Differential Revision: https://reviews.llvm.org/D80864
      b638b63b
    • Fangrui Song's avatar
      [ELF] Refine --export-dynamic-symbol semantics to be compatible GNU ld 2.35 · 751f18e7
      Fangrui Song authored
      GNU ld from binutils 2.35 onwards will likely support
      --export-dynamic-symbol but with different semantics.
      https://sourceware.org/pipermail/binutils/2020-May/111302.html
      
      Differences:
      
      1. -export-dynamic-symbol is not supported
      2. --export-dynamic-symbol takes a glob argument
      3. --export-dynamic-symbol can suppress binding the references to the definition within the shared object if (-Bsymbolic or -Bsymbolic-functions)
      4. --export-dynamic-symbol does not imply -u
      
      I don't think the first three points can affect any user.
      For the fourth point, Not implying -u can lead to some archive members unfetched.
      Add -u foo to restore the previous behavior.
      
      Exact semantics:
      
      * -no-pie or -pie: matched non-local defined symbols will be added to the dynamic symbol table.
      * -shared: matched non-local STV_DEFAULT symbols will not be bound to definitions within the shared object
        even if they would otherwise be due to -Bsymbolic, -Bsymbolic-fun...
      751f18e7
    • Sanjay Patel's avatar
      [InstCombine] fix use of base VectorType; NFC · 26ebe936
      Sanjay Patel authored
      SimplifyDemandedVectorElts() bails out on ScalableVectorType
      anyway, but we can exit faster with the external check.
      
      Move this to a helper function because there are likely other
      vector folds that we can try here.
      26ebe936
    • Matt Arsenault's avatar
      AMDGPU: Fix not emitting nofpexcept on fdiv expansion · 89d48cca
      Matt Arsenault authored
      In this awkward case, we have to emit custom pseudo-constrained FP
      wrappers. InstrEmitter concludes that since a mayRaiseFPException
      instruction had a chain, it can't add nofpexcept.
      
      Test deferred until mayRaiseFPException is really set on everything.
      89d48cca
    • Vedant Kumar's avatar
      11c617c4
    • Vedant Kumar's avatar
      [LiveDebugValues] Speed up removeEntryValue, NFC · 2ecaf935
      Vedant Kumar authored
      Summary:
      Instead of iterating over all VarLoc IDs in removeEntryValue(), just
      iterate over the interval reserved for entry value VarLocs. This changes
      the iteration order, hence the test update -- otherwise this is NFC.
      
      This appears to give an ~8.5x wall time speed-up for LiveDebugValues when
      compiling sqlite3.c 3.30.1 with a Release clang (on my machine):
      
      ```
                ---User Time---   --System Time--   --User+System--   ---Wall Time--- --- Name ---
        Before: 2.5402 ( 18.8%)   0.0050 (  0.4%)   2.5452 ( 17.3%)   2.5452 ( 17.3%) Live DEBUG_VALUE analysis
         After: 0.2364 (  2.1%)   0.0034 (  0.3%)   0.2399 (  2.0%)   0.2398 (  2.0%) Live DEBUG_VALUE analysis
      ```
      
      The change in removeEntryValue() is the only one that appears to affect
      wall time, but for consistency (and to resolve a pending TODO), I made
      the analogous changes for iterating over SpillLocKind VarLocs.
      
      Reviewers: nikic, aprantl, jmorse, djtodoro
      
      Subscribers: hiraditya, dexonsmith, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D80684
      2ecaf935
    • Matt Arsenault's avatar
      DAG: Fix getNode dropping flags if there's a glue output · 836c7dcf
      Matt Arsenault authored
      The AMDGPU non-strict fdiv lowering needs to introduce an FP mode
      switch in some cases, and has custom nodes to provide chain/glue for
      the intermediate FP operations. We need to propagate nofpexcept here,
      but getNode was dropping the flags.
      
      Adding nofpexcept in the AMDGPU custom lowering is left to a future
      patch.
      
      Also fix a second case where flags were dropped, but in this case it
      seems it just didn't handle this number of operands.
      
      Test will be included in future AMDGPU patch.
      836c7dcf
    • Julian Lettner's avatar
      [Darwin] Add and adopt a way to query the Darwin kernel version · f97a609b
      Julian Lettner authored
      This applies the learnings from [1].  What I intended as a simple
      cleanup made me realize that the compiler-rt version checks have two
      separate issues:
      
      1) In some places (e.g., mmap flag setting) what matters is the kernel
         version, not the OS version.
      2) OS version checks are implemented by querying the kernel version.
         This is not necessarily correct inside the simulators if the
         simulator runtime isn't aligned with the host macOS.
      
      This commit tackles 1) by adopting a separate query function for the
      Darwin kernel version.  2) (and cleanups) will be dealt with in
      follow-ups.
      
      [1] https://reviews.llvm.org/D78942
      
      rdar://63031937
      
      Reviewed By: delcypher
      
      Differential Revision: https://reviews.llvm.org/D79965
      f97a609b
    • Hiroshi Yamauchi's avatar
      [PGO] Improve the working set size heuristics under the partial sample PGO. · 6c27c61d
      Hiroshi Yamauchi authored
      Summary:
      The working set size heuristics (ProfileSummaryInfo::hasHugeWorkingSetSize)
      under the partial sample PGO may not be accurate because the profile is partial
      and the number of hot profile counters in the ProfileSummary may not reflect the
      actual working set size of the program being compiled.
      
      To improve this, the (approximated) ratio of the the number of profile counters
      of the program being compiled to the number of profile counters in the partial
      sample profile is computed (which is called the partial profile ratio) and the
      working set size of the profile is scaled by this ratio to reflect the working
      set size of the program being compiled and used for the working set size
      heuristics.
      
      The partial profile ratio is approximated based on the number of the basic
      blocks in the program and the NumCounts field in the ProfileSummary and computed
      through the thin LTO indexing. This means that there is the limitation that the
      sca...
      6c27c61d
    • Matt Arsenault's avatar
      AMDGPU: Fix test in code directory · 20793b2a
      Matt Arsenault authored
      20793b2a
    • Matt Arsenault's avatar
      AMDGPU: Remove dead file · ed08c4fb
      Matt Arsenault authored
      ed08c4fb
    • hsmahesha's avatar
      [AMDGPU/MemOpsCluster] Let mem ops clustering logic also consider number of clustered bytes · 0ed2c046
      hsmahesha authored
      Summary:
      While clustering mem ops, AMDGPU target needs to consider number of clustered bytes
      to decide on max number of mem ops that can be clustered. This patch adds support to pass
      number of clustered bytes to target mem ops clustering logic.
      
      Reviewers: foad, rampitec, arsenm, vpykhtin, javedabsar
      
      Reviewed By: foad
      
      Subscribers: MatzeB, kzhuravl, jvesely, wdng, nhaehnle, yaxunl, dstuttard, tpr, t-tye, hiraditya, javed.absar, kerbowa, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D80545
      0ed2c046