1. Feb 27, 2021
    • Heejin Ahn's avatar
      [WebAssembly] Fix reverse mapping in WasmEHFuncInfo · aa097ef8
      Heejin Ahn authored
      D97247 added the reverse mapping from unwind destination to their
      source, but it had a critical bug; sources can be multiple, because
      multiple BBs can have a single BB as their unwind destination.
      
      This changes `WasmEHFuncInfo::getUnwindSrc` to `getUnwindSrcs` and makes
      it return a vector rather than a single BB. It does not return the const
      reference to the existing vector but creates a new vector because
      `WasmEHFuncInfo` stores not `BasicBlock*` or `MachineBasicBlock*` but
      `PointerUnion` of them. Also I hoped to unify those methods for
      `BasicBlock` and `MachineBasicBlock` into one using templates to reduce
      duplication, but failed because various usages require `BasicBlock*` to
      be `const` but it's hard to make it `const` for `MachineBasicBlock`
      usages.
      
      Fixes https://github.com/emscripten-core/emscripten/issues/13514.
      (More precisely, fixes
      https://github.com/emscripten-core/emscripten/issues/13514#issuecomment-784708744)
      
      Reviewed By: dschuff, tlively
      
      Differential Revision: https://reviews.llvm.org/D97583
      aa097ef8
    • Sam Clegg's avatar
      [lld][WebAssembly] Rename methods/members to match ELF backend. NFC. · 14ffbb84
      Sam Clegg authored
      Specifically:
      
      - InputChunk::outputOffset -> outSecOffset
      - Symbol::get/setVirtualAddress -> get/setVA
      - add InputChunk::getOffset helper that takes an offset
      
      These are mostly in preparation for adding support for
      SHF_MERGE/SHF_STRINGS but its also good to align with ELF where
      possible.
      
      Differential Revision: https://reviews.llvm.org/D97595
      14ffbb84
    • Kevin Zhou's avatar
      [Polly] Refactoring IsInnermostParallel() in ISL to take the C++ wrapper object. NFC · 1ab2753d
      Kevin Zhou authored
      Currently, the IslAst library is a C library that would be incompatible with the rest of the LLVM because LLVM is written in C++.
      I took one function, IsInnermostParallel(), and refactored it so that it would take the C++ wrapper object instead of using reference counters with the C ISL library. As well, all the references that use IsInnermostParallel() will use manage_copy() since they are still expecting the C object.
      
      Reviewed By: Meinersbur
      
      Differential Revision: https://reviews.llvm.org/D97425
      1ab2753d
    • Fangrui Song's avatar
      ELF: Create unique SHF_GNU_RETAIN sections for llvm.used global objects · 47c5576d
      Fangrui Song authored
      If a global object is listed in `@llvm.used`, place it in a unique section with
      the `SHF_GNU_RETAIN` flag. The section is a GC root under `ld --gc-sections`
      with LLD>=13 or GNU ld>=2.36.
      
      For front ends which do not expect to see multiple sections of the same name,
      consider emitting `@llvm.compiler.used` instead of `@llvm.used`.
      
      SHF_GNU_RETAIN is restricted to ELFOSABI_GNU and ELFOSABI_FREEBSD in
      binutils. We don't do the restriction - see the rationale in D95749.
      
      The integrated assembler has supported SHF_GNU_RETAIN since D95730.
      GNU as>=2.36 supports section flag 'R'.
      We don't need to worry about GNU ld support because older GNU ld just ignores
      the unknown SHF_GNU_RETAIN.
      
      With this change, `__attribute__((retain))` functions/variables emitted
      by clang will get the SHF_GNU_RETAIN flag.
      
      Differential Revision: https://reviews.llvm.org/D97448
      47c5576d
    • Fangrui Song's avatar
      Add GNU attribute 'retain' · 8afdacba
      Fangrui Song authored
      For ELF targets, GCC 11 will set SHF_GNU_RETAIN on the section of a
      `__attribute__((retain))` function/variable to prevent linker garbage
      collection. (See AttrDocs.td for the linker support).
      
      This patch adds `retain` functions/variables to the `llvm.used` list, which has
      the desired linker GC semantics. Note: `retain` does not imply `used`,
      so an unused function/variable can be dropped by Sema.
      
      Before 'retain' was introduced, previous ELF solutions require inline asm or
      linker tricks, e.g.  `asm volatile(".reloc 0, R_X86_64_NONE, target");`
      (architecture dependent) or define a non-local symbol in the section and use
      `ld -u`. There was no elegant source-level solution.
      
      With D97448, `__attribute__((retain))` will set `SHF_GNU_RETAIN` on ELF targets.
      
      Differential Revision: https://reviews.llvm.org/D97447
      8afdacba
    • Kazu Hirata's avatar
      233ba270
    • Jianzhou Zhao's avatar
      [msan] Use non-transparent-huge-page at SetShadow · c0dc885d
      Jianzhou Zhao authored
      This prevents from getting THP ranges more and more.
      
      Did not see any issues in practice, just found this by code review.
      
      Reviewed By: eugenis, vitalybuka
      
      Differential Revision: https://reviews.llvm.org/D97593
      c0dc885d
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Import FMOV patterns rather than manually selecting it · f5d5a7d7
      Jessica Paquette authored
      There are existing patterns for FMOVHi, FMOVSi, and FMOVDi in
      AArch64InstrFormats.td.
      
      Importing these allows us to remove the manual selection code for FMOV.
      
      It also allows us to select FMOVHi for non-zero constants when we have full
      fp-16 support.
      
      Refactor some of the code in AArch64InstrFormats.td so that we can create
      equivalent custom renderers in GlobalISel.
      
      Differential Revision: https://reviews.llvm.org/D97511
      f5d5a7d7
    • Fangrui Song's avatar
      [test] Fix PGOProfile/comdat_internal.ll · 1d7f8c75
      Fangrui Song authored
      1d7f8c75
    • Jacques Pienaar's avatar
      [mlir] Add regions to OpAdaptor · 91ab48ea
      Jacques Pienaar authored
      Allows querying regions too via OpAdaptor's generated. This does not yet move region verification to adaptor nor require regions for ops where needed.
      
      Differential Revision: https://reviews.llvm.org/D97519
      91ab48ea
    • Ryan Prichard's avatar
      Reland "[builtins] Define fmax and scalbn inline" · d2022014
      Ryan Prichard authored
      This reverts commit 680f836c.
      
      Disable the non-default-rounding-mode scalbn[f] tests when we're using
      the MSVC libraries.
      
      Differential Revision: https://reviews.llvm.org/D91841
      d2022014
    • Vladimir Vereschaka's avatar
      [Driver] Print process statistics report on CC_PRINT_PROC_STAT env variable. · 155c49e0
      Vladimir Vereschaka authored
      Added supporting CC_PRINT_PROC_STAT and CC_PRINT_PROC_STAT_FILE
      environment variables to trigger clang driver reporting the process
      statistics into specified file (alternate for -fproc-stat-report
      option).
      
      Differential Revision: https://reviews.llvm.org/D97094
      155c49e0
    • Fangrui Song's avatar
      [InstrProfiling] Use llvm.compiler.used instead of llvm.used for ELF · bf176c49
      Fangrui Song authored
      Many optimizers (e.g.  GlobalOpt/ConstantMerge) do not respect linker semantics
      for comdat and may not discard the sections as a unit.
      
      The interconnected `__llvm_prf_{cnts,data}` sections (in comdat for ELF)
      are similar to D97432: `__profd_` is not directly referenced, so
      `__profd_` may be discarded while `__profc_` is retained, breaking the
      interconnection.  We currently conservatively add all such sections to
      `llvm.used` and let the linker do GC for ELF.
      
      In D97448, we will change GlobalObject's in the llvm.used list to use SHF_GNU_RETAIN,
      causing the metadata sections to be unnecessarily retained (some `check-profile` tests check for GC).
      Use `llvm.compiler.used` to retain the current GC behavior.
      
      Differential Revision: https://reviews.llvm.org/D97585
      bf176c49
    • Eric Schweitz's avatar
      5077d42c
    • Matheus Izvekov's avatar
      [clang] implicitly delete space ship operator with function pointers · 4a8530fc
      Matheus Izvekov authored
      See bug #48856
      
      Definitions of classes with member function pointers and default
      spaceship operator were getting accepted with no diagnostic on
      release build, and triggering assert on builds with runtime checks
      enabled. Diagnostics were only produced when actually comparing
      instances of such classes.
      
      This patch makes it so Spaceship and Less operators are not considered
      as builtin operator candidates for function pointers, producing
      equivalent diagnostics for the cases where pointers to member function
      and pointers to data members are used instead.
      
      Reviewed By: rsmith
      
      Differential Revision: https://reviews.llvm.org/D95409
      4a8530fc
    • Rob Suderman's avatar
      [MLIR][TOSA] Lower tosa.identity and tosa.identitiyn to linalg · f685c9ac
      Rob Suderman authored
      Both identity ops can be loweried by replacing their results with their
      inputs. We keep this as a linalg lowering as other backends may choose to
      create copies.
      
      Differential Revision: https://reviews.llvm.org/D97517
      f685c9ac
    • Arthur Eubanks's avatar
      [docs] Add documentation on using the new pass manager · 016f0ee6
      Arthur Eubanks authored
      And clarify in the "writing a pass" docs that both the legacy and new
      PMs are being used for the codegen/optimization pipelines.
      
      Reviewed By: ychen, asbirlea
      
      Differential Revision: https://reviews.llvm.org/D97515
      016f0ee6
    • Matt Arsenault's avatar
      AMDGPU: Use kill instruction to hint soft clause live ranges · 81b2c23b
      Matt Arsenault authored
      Previously we would use a bundle to hint the register allocator to not
      overwrite the pointers in a sequence of loads to avoid breaking soft
      clauses. This bundling was based on a fuzzy register pressure
      heuristic, so we could not guarantee using more registers than are
      really available. This would result in register allocator failing on
      unsatisfiable bundles. Use a kill to artificially extend the live
      ranges, so we can always succeed at register allocation even if it
      means extra spills in the worst case.
      
      This seems to capture most of the benefit of the bundle while avoiding
      most of the risk presented by the bundle. However the lit tests do
      show a handful of regressions. In some cases with sequences of
      volatile loads, unused load components end up getting reallocated to
      the next load which forces a wait between. There are also a few small
      scheduling regressions where a hazard used to be avoided, and one
      spill torture test which for some reason nearly doubles the stack
      usage. There is also a bit of noise from leftover kills (it may make
      sense for post-RA pseudos to strip all of these out).
      81b2c23b
    • Craig Topper's avatar
      [DAGCombiner] Optimize SMULO/UMULO if we can prove that overflow is impossible. · eea53b14
      Craig Topper authored
      Using ComputeNumSignBits or computeKnownBits we might be able
      to determine that overflow is impossible.
      
      This especially helps after type legalization if the type was
      promoted from a type with half the bits or more. Type legalization
      conservatively creates a promoted smulo/umulo and an overflow
      check for the promoted bits. The overflow from the promoted
      smulo/umulo is ORed with the result of the promoted bits
      overflow check. Proving that the promoted smulo/umulo can never
      overflow will leave us with just the promoted bits overflow check.
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D97160
      eea53b14
    • Peter Steinfeld's avatar
      [flang] Detect circularly defined interfaces of procedures · 07de0846
      Peter Steinfeld authored
      It's possible to define a procedure whose interface depends on a procedure
      which has an interface that depends on the original procedure.  Such a circular
      definition was causing the compiler to fall into an infinite loop when
      resolving the name of the second procedure.  It's also possible to create
      circular dependency chains of more than two procedures.
      
      I fixed this by adding the function HasCycle() to the class DeclarationVisitor
      and calling it from DeclareProcEntity() to detect procedures with such
      circularly defined interfaces.  I marked the associated symbols of such
      procedures by calling SetError() on them.  When processing subsequent
      procedures, I called HasError() before attempting to analyze their interfaces.
      Unfortunately, this did not work.
      
      With help from Tim, we determined that the SymbolSet used to track the
      erroneous symbols was instantiated using a "<" operator which was
      defined using the name of the procedure.  But the procedure name was
      being changed by a call to ReplaceName() between the times that the
      calls to SetError() and HasError() were made.  This caused HasError() to
      incorrectly report that a symbol was not in the set of erroneous
      symbols.  I fixed this by making SymbolSet be an ordered set, which does
      not use the "<" operator.
      
      I also added tests that will crash the compiler without this change.
      And I fixed the formatting on an error message from a previous update.
      
      Differential Revision: https://reviews.llvm.org/D97201
      07de0846
    • George Balatsouras's avatar
      [dfsan] Record dfsan metadata in globals · c9075a1c
      George Balatsouras authored
      This will allow identifying exactly how many shadow bytes were used
      during compilation, for when fast8 mode is introduced.
      
      Also, it will provide a consistent matching point for instrumentation
      tests so that the exact llvm type used (i8 or i16) for the shadow can
      be replaced by a pattern substitution. This is handy for tests with
      multiple prefixes.
      
      Reviewed by: stephan.yichao.zhao, morehouse
      
      Differential Revision: https://reviews.llvm.org/D97409
      c9075a1c
    • Vitaly Buka's avatar
      [sanitizers][NFC] Change typesto avoid warnings · 812a9061
      Vitaly Buka authored
      Warning was enabled by D94640
      812a9061
    • Vitaly Buka's avatar
      [NFC][libc++] Suppress "warning: ignoring return value" · 3744ba24
      Vitaly Buka authored
      According to the comment on the next line
      it's expected behaviour.
      3744ba24
    • Vitaly Buka's avatar
      e29063b1
    • Aart Bik's avatar
      [mlir][vector] add higher dimensional support to gather/scatter · df5ccf5a
      Aart Bik authored
      Similar to mask-load/store and compress/expand, the gather and
      scatter operation now allow for higher dimension uses. Note that
      to support the mixed-type index, the new syntax is:
         vector.gather %base [%i,%j] [%kvector] ....
      The first client of this generalization is the sparse compiler,
      which needs to define scatter and gathers on dense operands
      of higher dimensions too.
      
      Reviewed By: bixia
      
      Differential Revision: https://reviews.llvm.org/D97422
      df5ccf5a
    • Dan Gohman's avatar
      [WebAssembly] Avoid `bit_cast` when printing f32 and f64 immediates · c62dabc3
      Dan Gohman authored
      Use `APInt` to convert a 32-bit or 64-bit immediate to an `APFloat` rather than
      `bit_cast` to a `float` or `double` to avoid going through host floating-point and
      potentially changing the bit pattern of NaNs.
      
      Differential Revision: https://reviews.llvm.org/D97490
      c62dabc3
    • Nico Weber's avatar
      [lld/mac] Add some support for dynamic lookup symbols, and implement -U · cafb6cd1
      Nico Weber authored
      Dynamic lookup symbols are symbols that work like dynamic symbols
      in ELF: They're not bound to a dylib like normal Mach-O twolevel lookup
      symbols, but they live in a global pool and dyld resolves them against
      exported symbols from all loaded dylibs.
      
      This adds support for dynamical lookup symbols to lld/mac. They are
      represented as DylibSymbols with file set to nullptr.
      
      This also uses this support to implement the -U flag, which makes
      a specific symbol that's undefined at the end of the link a
      dynamic lookup symbol.
      
      For -U, it'd be sufficient to just to a pass over remaining undefined symbols
      at the end of the link and to replace them with dynamic lookup symbols then.
      But I'd like to use this code to implement flat_namespace too, and that will
      require real support for resolving dynamic lookup symbols in SymbolTable. So
      this patch adds this now already.
      
      While writing tests for this, I noticed that we didn't set N_WEAK_DEF in the
      symbol table for DylibSymbols, so this fixes that too.
      
      Differential Revision: https://reviews.llvm.org/D97521
      cafb6cd1
    • Casey Carter's avatar
      [libcxx][test] Don't require Container<cv T> extension on non-libc++ · 30cd3dd0
      Casey Carter authored
      ... when testing `default_initializable`. Also, include `<memory>` for `unique_ptr`.
      30cd3dd0
    • Heejin Ahn's avatar
      [WebAssembly] Fix remapping branch dests in fixCatchUnwindMismatches · d8b3dc5a
      Heejin Ahn authored
      This is a case D97178 tried to solve but missed. D97178 could not handle
      the case when
      multiple consecutive delegates are generated:
      - Before:
      ```
      block
        br (a)
        try
        catch
        end_try
      end_block
                <- (a)
      ```
      
      - After
      ```
      block
        br (a)
        try
          ...
          try
            try
            catch
            end_try
                  <- (a)
          delegate
        delegate
      end_block
                <- (b)
      ```
      (The `br` should point to (b) now)
      
      D97178 assumed `end_block` exists two BBs later than `end_try`, because
      it assumed the order as `end_try` BB -> `delegate` BB -> `end_block` BB.
      But it turned out there can be multiple `delegate`s in between. This
      patch changes the logic so we just search from `end_try` BB until we
      find `end_block`.
      
      Fixes https://github.com/emscripten-core/emscripten/issues/13515.
      (More precisely, fixes
      https://github.com/emscripten-core/emscripten/issues/13515#issuecomment-784711318.)
      
      Reviewed By: dschuff, tlively
      
      Differential Revision: https://reviews.llvm.org/D97569
      d8b3dc5a
    • Philip Reames's avatar
      [tests] Precommit for upcoming patch · 83bc7815
      Philip Reames authored
      83bc7815
    • Rob Suderman's avatar
      [MLIR][TOSA] Lower tosa.reshape to linalg.reshape · caccddc5
      Rob Suderman authored
      Lowering from the tosa.reshape op to linalg.reshape. For same-rank or
      non-collapsed/expanded cases two linalg.reshapes are inserted.
      
      Differential Revision: https://reviews.llvm.org/D97439
      caccddc5
    • Stanislav Mekhanoshin's avatar
      [AMDGPU] Avoid second rescheduling for some regions · 799c50fe
      Stanislav Mekhanoshin authored
      If a region was not constrained by a high register pressure
      and was not rescheduled without clustering we can skip
      rescheduling it ClusteredLowOccupancyReschedule stage.
      
      This improves scheduling speed by 25% on some kernels.
      
      Differential Revision: https://reviews.llvm.org/D97506
      799c50fe
    • Stanislav Mekhanoshin's avatar
      [AMDGPU] Skip unclusterd rescheduling w/o ld/st · 635993f0
      Stanislav Mekhanoshin authored
      We are attempting rescheduling without load store clustering
      if occupancy limits were not met with clustering. Skip this
      for regions which do not have any loads or stores at all.
      
      In a set of kernels I am experimenting with this improves
      scheduling time by ~30%.
      
      Differential Revision: https://reviews.llvm.org/D97342
      635993f0
    • Anirudh Prasad's avatar
      [SystemZ] Introducing assembler dialects for the Z backend · bcc1aba6
      Anirudh Prasad authored
      - This patch introduces a different assembler dialect ("hlasm") for z/OS.
        The default dialect has now been given the "att" dialect name. For this
        appropriate changes have been added to SystemZ.td.
      - This patch also makes a few changes to SystemZInstrFormats.td which
        restrict a few condition code mnemonics to just the "att" dialect
        variant (he, le, lh, nhe, nle, nlh). These extended condition code
        mnemonics are not available in HLASM.
      - A new private function has been introduced in SystemZAsmParser.cpp to
        return the assembler dialect set in SystemZMCAsmInfo.cpp. The reason we
        couldn't/haven't explicitly queried the overriden getAssemblerDialect
        function from AsmParser is outlined in this thread here. This returned
        dialect is directly passed onto the relevant matcher functions which taken
        in a variantID, so that the matcher functions can appropriately choose an
        instruction based on the variant.
      
      Reviewed By: uweigand
      
      Differential Revision: https://reviews.llvm.org/D94250
      bcc1aba6
    • James Y Knight's avatar
      Use getAlign() on atomicrmw/cmpxchg instructions, now that it's available. · 6de64557
      James Y Knight authored
      These locations were missed as part of adding alignment to the
      instructions, and were still making their own alignment assumptions.
      6de64557
    • Philip Reames's avatar
    • Jianzhou Zhao's avatar
      [dfsan] Do not test origin-tracking in atomic.cpp · c5c316f6
      Jianzhou Zhao authored
      This would cause linking errors after https://reviews.llvm.org/D97483
      that introduced new prefixes for ABI wrappers with origin tracking mode.
      We will renable this after the full origin tracking is checked in.
      c5c316f6
    • Craig Topper's avatar
      [RISCV] Call SelectBaseAddr on the base pointer in the custom isel for vector loads and stores. · b183cbfa
      Craig Topper authored
      This will allow FrameIndex as the base address instead of
      emitting a separate ADDI from isel. eliminateFrameIndex will likely turn
      it back into an ADDI, but this makes things consistent with the
      SDPatterns and VLPatterns.
      
      I only tested one case for simplicity. I can test more if reviewers
      want.
      
      Reviewed By: frasercrmck
      
      Differential Revision: https://reviews.llvm.org/D97221
      b183cbfa
    • Philip Reames's avatar
      Be more mathematicly precise about definition of recurrence [NFC] · f2cfef35
      Philip Reames authored
      This clarifies the interface of the matchSimpleRecurrence helper introduced in 8020be0b for non-commutative operators.  After ebd3aeba, I realized the original way I framed the routine was inconsistent.  For shifts, we only matched the the LHS form, but for sub we matched both and the caller wanted that information.  So, instead, we now consistently match both forms for non-commutative operators and the caller becomes responsible for filtering if needed.  I tried to put a clear warning in the header because I suspect the RHS form of e.g. a sub recurrence is non-obvious for most folks.  (It was for me.)
      f2cfef35
    • Leonard Chan's avatar
      [scudo][test] Disable -Wfree-nonheap-object · bed88824
      Leonard Chan authored
      As of 4f395db8 which contains updates to
      -Wfree-nonheap-object, a line in this test will trigger the warning. This
      particular line is ok though since it's meant to test a free on a bad pointer.
      
      Differential Revision: https://reviews.llvm.org/D97516
      bed88824