1. Aug 04, 2020
    • Fangrui Song's avatar
      [MC] Set sh_link to 0 if the associated symbol is undefined · 11bb7c22
      Fangrui Song authored
      Part of https://bugs.llvm.org/show_bug.cgi?id=41734
      
      LTO can drop externally available definitions. Such AssociatedSymbol is
      not associated with a symbol. ELFWriter::writeSection() will assert.
      
      Allow a SHF_LINK_ORDER section to have sh_link=0.
      
      We need to give sh_link a syntax, a literal zero in the linked-to symbol
      position, e.g. `.section name,"ao",@progbits,0`
      
      Reviewed By: pcc
      
      Differential Revision: https://reviews.llvm.org/D72899
      11bb7c22
    • Akira Hatanaka's avatar
      [CodeGen][ObjC] Mark calls to objc_unsafeClaimAutoreleasedReturnValue as · 41b1e97b
      Akira Hatanaka authored
      notail on x86-64
      
      This is needed because the epilogue code inserted before tail calls on
      x86-64 breaks the handshake between the caller and callee.
      
      Calls to objc_retainAutoreleasedReturnValue used to have the same
      problem, which was fixed in https://reviews.llvm.org/D59656.
      
      rdar://problem/66029552
      
      Differential Revision: https://reviews.llvm.org/D84540
      41b1e97b
    • Daniel Sanders's avatar
      Allow .dSYM's to be directly placed in an alternate directory · 7209f831
      Daniel Sanders authored
      Once available in the relevant toolchains this will allow us to implement
      LLVM_EXTERNALIZE_DEBUGINFO_OUTPUT_DIR after D84127 by directly placing the dSYM
      in the desired location instead of emitting next to the output file and moving
      it.
      
      Reviewed By: JDevlieghere
      
      Differential Revision: https://reviews.llvm.org/D84572
      7209f831
    • Jon Roelofs's avatar
      Fix typo: s/epomymous/eponymous/ NFC · 7f1556f2
      Jon Roelofs authored
      7f1556f2
    • Eli Friedman's avatar
      [AArch64] Add missing isel patterns for fcvtzs/u intrinsic on v1f64. · dca23ed8
      Eli Friedman authored
      Fixes test-suite compile failure caused by 8dfb5d76.
      
      While I'm in the area, add some more test coverage to related
      operations, to make sure we aren't missing any other patterns.
      dca23ed8
    • Lang Hames's avatar
      [llvm-jitlink] Add support for static archives and MachO universal archives. · 777824b4
      Lang Hames authored
      Archives can now be specified as input files the same way that object
      files are. Archives will always be linked after all objects (regardless
      of the relative order of the inputs) but before any dynamic libraries or
      process symbols.
      
      This patch also relaxes matching for slice triples in
      StaticLibraryDefinitionGenerator in order to support this feature:
      Vendors need not match if the source vendor is unknown.
      777824b4
    • Hiroshi Yamauchi's avatar
      [PGO] Enable the extended value profile buckets for mem op sizes. · 3e89cbf3
      Hiroshi Yamauchi authored
      Following up D81682 and enable the new, extended value profile buckets for mem
      op sizes.
      
      Differential Revision: https://reviews.llvm.org/D83903
      3e89cbf3
    • Tim Keith's avatar
      [flang] Fix bug detecting intrinsic function · 0d454e8e
      Tim Keith authored
      Don't set the INTRINSIC attribute on a dummy procedure.
      
      Differential Revision: https://reviews.llvm.org/D85136
      0d454e8e
    • Sanjay Patel's avatar
    • Sanjay Patel's avatar
      7efd9ceb
    • Arthur Eubanks's avatar
      Fix layering violation Transforms/Utils -> Scalar · 456f38a9
      Arthur Eubanks authored
      Introduced in D85063.
      456f38a9
    • Jian Cai's avatar
      [X86] support .nops directive · c6334db5
      Jian Cai authored
      Add support of .nops on X86. This addresses llvm.org/PR45788.
      
      Reviewed By: craig.topper
      
      Differential Revision: https://reviews.llvm.org/D82826
      c6334db5
    • Florian Hahn's avatar
      [ArgPromotion] Replace all md uses of promoted values with undef. · 1e392fc4
      Florian Hahn authored
      Currently, ArgPromotion may leave metadata uses of promoted values,
      which will end up in the wrong function, creating invalid IR.
      
      PR33641 fixed this for dead arguments, but it can be also be triggered
      arguments with users that are promoted (see the updated test case).
      
      We also have to drop uses to them after promoting them. We need to do
      this after dealing with the non-metadata uses, so I also moved the empty
      use case to the loop that deals with updating the arguments of the new
      function.
      
      Reviewed By: aprantl
      
      Differential Revision: https://reviews.llvm.org/D85127
      1e392fc4
    • LLVM GN Syncbot's avatar
      [gn build] Port f78f509c · c12bd8da
      LLVM GN Syncbot authored
      c12bd8da
    • Hiroshi Yamauchi's avatar
      [PGO] Extend the value profile buckets for mem op sizes. · f78f509c
      Hiroshi Yamauchi authored
      Extend the memop value profile buckets to be more flexible (could accommodate a
      mix of individual values and ranges) and to cover more value ranges (from 11 to
      22 buckets).
      
      Disabled behind a flag (to be enabled separately) and the existing code to be
      removed later.
      
      Differential Revision: https://reviews.llvm.org/D81682
      f78f509c
    • Rainer Orth's avatar
      [compiler-rt][profile] Fix various InstrProf tests on Solaris · 39494d9c
      Rainer Orth authored
      Currently, several InstrProf tests `FAIL` on Solaris (both sparc and x86):
      
        Profile-i386 :: Posix/instrprof-visibility.cpp
        Profile-i386 :: instrprof-merging.cpp
        Profile-i386 :: instrprof-set-file-object-merging.c
        Profile-i386 :: instrprof-set-file-object.c
      
      On sparc there's also
      
        Profile-sparc :: coverage_comments.cpp
      
      The failure mode is always the same:
      
        error: /var/llvm/local-amd64/projects/compiler-rt/test/profile/Profile-i386/Posix/Output/instrprof-visibility.cpp.tmp: Failed to load coverage: Malformed coverage data
      
      The error is from `llvm/lib/ProfileData/Coverage/CoverageMappingReader.cpp`
      (`loadBinaryFormat`), l.926:
      
        InstrProfSymtab ProfileNames;
        std::vector<SectionRef> NamesSectionRefs = *NamesSection;
        if (NamesSectionRefs.size() != 1)
          return make_error<CoverageMapError>(coveragemap_error::malformed);
      
      where .size() is 2 instead.
      
      Looking at the executable, I find (with `elfdump -c -N __llvm_prf_names`):
      
        Section Header[15]:  sh_name: __llvm_prf_names
            sh_addr:      0x8053ca5       sh_flags:   [ SHF_ALLOC ]
            sh_size:      0x86            sh_type:    [ SHT_PROGBITS ]
            sh_offset:    0x3ca5          sh_entsize: 0
            sh_link:      0               sh_info:    0
            sh_addralign: 0x1
      
        Section Header[31]:  sh_name: __llvm_prf_names
            sh_addr:      0x8069998       sh_flags:   [ SHF_WRITE SHF_ALLOC ]
            sh_size:      0               sh_type:    [ SHT_PROGBITS ]
            sh_offset:    0x9998          sh_entsize: 0
            sh_link:      0               sh_info:    0
            sh_addralign: 0x1
      
      Unlike GNU `ld` (which primarily operates on section names) the Solaris
      linker, following the ELF spirit, only merges input sections into an output
      section if both section name and section flags match, so two separate
      sections are maintained.
      
      The read-write one comes from `lib/clang/12.0.0/lib/sunos/libclang_rt.profile-i386.a(InstrProfilingPlatformLinux.c.o)`
      while the read-only one is generated by
      `llvm/lib/Transforms/Instrumentation/InstrProfiling.cpp` (`InstrProfiling::emitNameData`)
      at l.1004 where `isConstant = true`.
      
      The easiest way to avoid the mismatch is to change the definition in
      `compiler-rt/lib/profile/InstrProfilingPlatformLinux.c` to `const`.
      
      This fixes all failures observed.
      
      Tested on `amd64-pc-solaris2.11`, `sparcv9-sun-solaris2.11`, and
      `x86_64-pc-linux-gnu`.
      
      Differential Revision: https://reviews.llvm.org/D85116
      39494d9c
    • Joao Moreira's avatar
      [X86] Make ENDBR instruction a scheduling boundary · f208c659
      Joao Moreira authored
      Instructions should not be scheduled across ENDBR instructions, as this would result in the ENDBR being displaced, breaking the parity needed for the Indirect Branch Tracking feature of CET.
      
      Currently, the X86IndirectBranchTracking pass is later than the instruction scheduling in the pipeline, what causes the bug to be unnoticeable and very hard (if not unfeasible) to be triggered while compiling C files with the standard LLVM setup. Yet, for correctness and to prevent issues in future changes, the compiler should prevent the such scheduling.
      
      Differential Revision: https://reviews.llvm.org/D84862
      f208c659
    • Simon Pilgrim's avatar
      [X86][SSE] Shuffle combine blends to OR(X,Y) if the relevant elements are known zero. · 219f32f4
      Simon Pilgrim authored
      This allows us to remove the (depth violating) code in getFauxShuffleMask where we were combining the OR(SHUFFLE,SHUFFLE) shuffle inputs as well, and not just the OR().
      
      This is a minor step toward being able to shuffle combine from/to SELECT/BLENDV as a faux shuffle.
      219f32f4
    • Arthur Eubanks's avatar
      [NewPM][LoopVersioning] Port LoopVersioning to NPM · 7c19c89d
      Arthur Eubanks authored
      Reviewed By: ychen, fhahn
      
      Differential Revision: https://reviews.llvm.org/D85063
      7c19c89d
    • Kevin P. Neal's avatar
      [FPEnv] IRBuilder fails to add strictfp attribute · d535a91d
      Kevin P. Neal authored
      The strictfp attribute is required on all function calls in a function
      that is itself marked with the strictfp attribute. The IRBuilder knows
      this and has a method for adding the attribute to function call instructions.
      
      If a function being called has the strictfp attribute itself then the
      IRBuilder will refuse to add the attribute to the calling instruction
      despite being asked to add it. Eliminate this error.
      
      Differential Revision: https://reviews.llvm.org/D84878
      d535a91d
    • Fangrui Song's avatar
      [PGO] Change a `NumVSites == 0` workaround to assert · 317e00dc
      Fangrui Song authored
      The root cause was fixed by 3d6f5301.
      The workaround added in 99ad956f can be changed
      to an assert now. (In case the fix regresses, there will be a heap-use-after-free.)
      317e00dc
    • Craig Topper's avatar
      [X86] Use h-register for final XOR of __builtin_parity on 64-bit targets. · ac82b918
      Craig Topper authored
      This adds an isel pattern and special XOR8rr_NOREX instruction
      to enable the use of h-registers for __builtin_parity. This avoids
      a copy and a shift instruction. The NOREX instruction is in case
      register allocation doesn't use the matching l-register for some
      reason. If a R8-R15 register gets picked instead, we won't be
      able to encode the instruction since an h-register can't be used
      with a REX prefix.
      
      Fixes PR46954
      ac82b918
    • MaheshRavishankar's avatar
      [mlir][DialectConversion] Remove usage of std::distance to track position. · 32f3a9a9
      MaheshRavishankar authored
      Remove use of iterator::difference_type to know where to insert a
      moved or erased block during undo actions.
      
      Differential Revision: https://reviews.llvm.org/D85066
      32f3a9a9
    • MaheshRavishankar's avatar
    • Nicolas Vasilache's avatar
      [mlir][Vector] Add transformation + pattern to split vector.transfer_read into... · d313e9c1
      Nicolas Vasilache authored
      [mlir][Vector] Add transformation + pattern to split vector.transfer_read into full and partial copies.
      
      This revision adds a transformation and a pattern that rewrites a "maybe masked" `vector.transfer_read %view[...], %pad `into a pattern resembling:
      
      ```
         %1:3 = scf.if (%inBounds) {
            scf.yield %view : memref<A...>, index, index
          } else {
            %2 = vector.transfer_read %view[...], %pad : memref<A...>, vector<...>
            %3 = vector.type_cast %extra_alloc : memref<...> to
            memref<vector<...>> store %2, %3[] : memref<vector<...>> %4 =
            memref_cast %extra_alloc: memref<B...> to memref<A...> scf.yield %4 :
            memref<A...>, index, index
         }
         %res= vector.transfer_read %1#0[%1#1, %1#2] {masked = [false ... false]}
      ```
      where `extra_alloc` is a top of the function alloca'ed buffer of one vector.
      
      This rewrite makes it possible to realize the "always full tile" abstraction where vector.transfer_read operations are guaranteed to read from a padded full buffer.
      The extra work only occurs on the boundary tiles.
      
      Differential Revision: https://reviews.llvm.org/D84631
      d313e9c1
    • Mircea Trofin's avatar
      [llvm] Add a parser from JSON to TensorSpec · 4b1b109c
      Mircea Trofin authored
      A JSON->TensorSpec utility we will use subsequently to specify
      additional outputs needed for certain training scenarios.
      
      Differential Revision: https://reviews.llvm.org/D84976
      4b1b109c
    • Gui Andrade's avatar
    • Gui Andrade's avatar
      [MSAN] Instrument freeze instruction by clearing shadow · 3ebd1ba6
      Gui Andrade authored
      Freeze always returns a defined value. This also prevents msan from
      checking the input shadow, which happened because freeze wasn't
      explicitly visited.
      
      Differential Revision: https://reviews.llvm.org/D85040
      3ebd1ba6
    • Florian Hahn's avatar
      [SCEV] If Start>=RHS, simplify (Start smin RHS) = RHS for trip counts. · ee1c1270
      Florian Hahn authored
      In some cases, it seems like we can get rid of unnecessary s/umins by
      using information from the loop guards (unless I am missing something).
      
      One place where this seems to be helpful in practice is when computing
      loop trip counts. This patch just changes howManyGreaterThans for now.
      Note that this requires a loop for which we can check 'is guarded'.
      
      On SPEC2000/SPEC2006/MultiSource, there are some notable changes for
      some programs in the number of loops unrolled and trip counts computed.
      
      ```
      Same hash: 179 (filtered out)
      Remaining: 58
      Metric: scalar-evolution.NumTripCountsComputed
      
      Program                                        base    patch   diff
       test-suite...langs-C/compiler/compiler.test    25.00   31.00  24.0%
       test-suite.../Applications/SPASS/SPASS.test   2020.00 2323.00 15.0%
       test-suite...langs-C/allroots/allroots.test    29.00   32.00  10.3%
       test-suite.../Prolangs-C/loader/loader.test    17.00   18.00   5.9%
       test-suite...fice-ispell/office-ispell.test   253.00  265.00   4.7%
       test-suite...006/450.soplex/450.soplex.test   3552.00 3692.00  3.9%
       test-suite...chmarks/MallocBench/gs/gs.test   453.00  470.00   3.8%
       test-suite...ngs-C/assembler/assembler.test    29.00   30.00   3.4%
       test-suite.../Benchmarks/Ptrdist/bc/bc.test   263.00  270.00   2.7%
       test-suite...rks/FreeBench/pifft/pifft.test   722.00  741.00   2.6%
       test-suite...count/automotive-bitcount.test    41.00   42.00   2.4%
       test-suite...0/253.perlbmk/253.perlbmk.test   1417.00 1451.00  2.4%
       test-suite...000/197.parser/197.parser.test   387.00  396.00   2.3%
       test-suite...lications/sqlite3/sqlite3.test   1168.00 1189.00  1.8%
       test-suite...000/255.vortex/255.vortex.test   173.00  176.00   1.7%
      
      Metric: loop-unroll.NumUnrolled
      
      Program                                        base   patch  diff
       test-suite...langs-C/compiler/compiler.test     1.00   3.00 200.0%
       test-suite.../Applications/SPASS/SPASS.test   134.00 234.00 74.6%
       test-suite...count/automotive-bitcount.test     3.00   4.00 33.3%
       test-suite.../Prolangs-C/loader/loader.test     3.00   4.00 33.3%
       test-suite...langs-C/allroots/allroots.test     3.00   4.00 33.3%
       test-suite...Source/Benchmarks/sim/sim.test    10.00  12.00 20.0%
       test-suite...fice-ispell/office-ispell.test    21.00  25.00 19.0%
       test-suite.../Benchmarks/Ptrdist/bc/bc.test    32.00  38.00 18.8%
       test-suite...006/450.soplex/450.soplex.test   300.00 352.00 17.3%
       test-suite...rks/FreeBench/pifft/pifft.test    60.00  69.00 15.0%
       test-suite...chmarks/MallocBench/gs/gs.test    57.00  63.00 10.5%
       test-suite...ngs-C/assembler/assembler.test    10.00  11.00 10.0%
       test-suite...0/253.perlbmk/253.perlbmk.test   145.00 157.00  8.3%
       test-suite...000/197.parser/197.parser.test    43.00  46.00  7.0%
       test-suite...TimberWolfMC/timberwolfmc.test   205.00 214.00  4.4%
       Geomean difference                                           7.6%
      ```
      
      Fixes https://bugs.llvm.org/show_bug.cgi?id=46939
      Fixes https://bugs.llvm.org/show_bug.cgi?id=46924 on X86.
      
      Reviewed By: mkazantsev
      
      Differential Revision: https://reviews.llvm.org/D85046
      ee1c1270
    • Mehdi Amini's avatar
      Revert "[mlir][Vector] Add transformation + pattern to split... · 7ba82a73
      Mehdi Amini authored
      Revert "[mlir][Vector] Add transformation + pattern to split vector.transfer_read into full and partial copies."
      
      This reverts commit 35b65be0.
      
      Build is broken with -DBUILD_SHARED_LIBS=ON with some undefined
      references like:
      
      VectorTransforms.cpp:(.text._ZN4llvm12function_refIFvllEE11callback_fnIZL24createScopedInBoundsCondN4mlir25VectorTransferOpInterfaceEE3$_8EEvlll+0xa5): undefined reference to `mlir::edsc::op::operator+(mlir::Value, mlir::Value)'
      7ba82a73
  2. Aug 03, 2020