1. Feb 02, 2024
    • ManuelvOK's avatar
      [Coverage] Map regions from system headers (#76950) · c07fcd45
      ManuelvOK authored
      In 21551951, the
      "system-headers-coverage" option has been added but not used in all
      necessary places.
      
      This is the recommit since it has been reverted in
      faef68bc
      
      
      
      Potential reviewers: @gulfemsavrun @petrhosek
      
      Co-authored-by: default avatarManuel Kalettka <manuel.kalettka@kernkonzept.com>
      c07fcd45
    • Matthias Springer's avatar
      [mlir][IR] Do not trigger `notifyOperationInserted` for unlinked ops (#80278) · a792cb6e
      Matthias Springer authored
      This commit changes `OpBuilder::create` and `OpBuilder::createOrFold`
      such that `notifyOperationInserted` is no longer triggered if no
      insertion point is set. In such a case, an unlinked operation is created
      but not inserted, so `notifyOperationInserted` should not be triggered.
      
      Note: Inserting another op into a block that belongs to an unlinked op
      (e.g., by the builder of the unlinked op) will trigger a notification.
      a792cb6e
    • Matthias Springer's avatar
      [mlir][IR] Notify about block insertion when cloning an op (#80262) · 237a799e
      Matthias Springer authored
      `OpBuilder::clone(Operation &)` should trigger not only
      `notifyOperationInserted` but also `notifyBlockInserted` (for all block
      contained in `op`).
      237a799e
    • Maciej Gabka's avatar
      [TLI][AArch64] Adjust TLI mappings to vector functions taking linear pointers (#80296) · 0f26441c
      Maciej Gabka authored
      The masked symbols in SLEEF are incorrectly implemented as calls to non
      masked variants, what only works fine for functions which do not modify
      memory.
      For vector variants which modify memory we can only use a non masked
      symbols for now.
      The SVE ArmPL mappings need to be removed for now as well.
      0f26441c
    • Timm Bäder's avatar
      [clang][Interp][NFC] Implement dumping Invalid/Valid results · 75c4339e
      Timm Bäder authored
      This was just an omission from an earlier commit, clearly
      we can print them.
      75c4339e
    • Vlad Serebrennikov's avatar
      [clang] Update documentation for `#pragma diagnostic` (#78095) · 3be79790
      Vlad Serebrennikov authored
      
      
      GCC has changed over the past decade, and we're not implementing
      everything they do.
      Fixes #51472
      
      ---------
      
      Co-authored-by: default avatarAaron Ballman <aaron@aaronballman.com>
      3be79790
    • Timm Bäder's avatar
      0be39155
    • Matthew Devereau's avatar
      [AArch64][SME] Implement inline-asm clobbers for za/zt0 (#79276) · d9c20e43
      Matthew Devereau authored
      This enables specifing "za" or "zt0" to the clobber list for inline asm.
      This complies with the acle SME addition to the asm extension here:
      https://github.com/ARM-software/acle/pull/276
      d9c20e43
    • Timm Bäder's avatar
      [clang][Interp][NFC] Add a broken test case · a8b5994b
      Timm Bäder authored
      The LHS of the subtraction returns 16 right now, but should
      return 0.
      a8b5994b
    • Kai Sasaki's avatar
      [mlir] Skip invalid test on big endian platform (s390x) (#80246) · 65ac8c16
      Kai Sasaki authored
      The buildbot test running on s390x platform keeps failing since [this
      time](https://lab.llvm.org/buildbot/#/builders/199/builds/31136). This
      is because of the dependency on the endianness of the platform. It
      expects the format invalid in the big endian platform (s390x). We can
      simply skip it.
      
      See: https://discourse.llvm.org/t/mlir-s390x-linux-failure/76695
      65ac8c16
    • Timm Bäder's avatar
      [clang][Interp] Ignore LValueToRValue casts before doing the load · a2da7d06
      Timm Bäder authored
      If the SubExpr results in an invalid pointer, we will otherwise
      reject the constant expression.
      a2da7d06
    • Timm Bäder's avatar
      [clang][Interp] Support ChooseExprs · 58ceefe0
      Timm Bäder authored
      58ceefe0
    • Timm Bäder's avatar
      [clang][Interp] Not all TypeTraitExprs are of bool type · 2147a2a4
      Timm Bäder authored
      In C, they return an integer, so emit their value as such.
      2147a2a4
    • Yuta Mukai's avatar
      [MachinePipeliner] Fix missing requirements for tests (#80386) · 374a600d
      Yuta Mukai authored
      Add asserts requirements for tests that verify debug output.
      374a600d
    • Fangrui Song's avatar
      [ELF] Fix compareSections assertion failure when OutputDescs in sectionCommands are non-contiguous · dee8786f
      Fangrui Song authored
      In a `--defsym y0=0 -T a.lds` link where a.lds contains only INSERT
      commands, the `script->sectionCommands` layout may be:
      ```
      orphan sections
      SymbolAssignment due to --defsym
      sections created by INSERT commands
      ```
      
      The `OutputDesc` objects are not contiguous in sortInputSections, and
      `compareSections` will be called with a SymbolAssignment argument,
      leading to an assertion failure.
      dee8786f
    • Jason Molenda's avatar
      [lldb] NFC fixes addressing David's feedback · 7dd790db
      Jason Molenda authored
      David Spickett had several suggestions for
      https://github.com/llvm/llvm-project/pull/79962 after I'd
      already merged it.  Address those.
      7dd790db
    • Shengchen Kan's avatar
      [X86] X86InstrInfo.cpp - Remove dead code for memory folding, NFCI · e270ec47
      Shengchen Kan authored
      `commuteInstruction(MI, false, OpNum, CommuteOpIdx2)` should never create
      any new instruction, so we don't need to check and erase it.
      e270ec47
    • Maksim Panchenko's avatar
      [BOLT] Remove duplicate expression (#80380) · 082fe9a5
      Maksim Panchenko authored
      Reported by cpp check static analyzer in #80111.
      
      Fixes #80111.
      082fe9a5
    • Craig Topper's avatar
      [RISCV] Add -march support for many of the S extensions mentioned in the... · 58c494f4
      Craig Topper authored
      [RISCV] Add -march support for many of the S extensions mentioned in the profile specification. (#79399)
      
      This is a good portion of the extensions mentioned in the RVA23 profile
      here
      https://github.com/riscv/riscv-profiles/blob/main/rva23-profile.adoc
      
      I don't believe these add any new CSRs. Sstc does add new CSRs, but we
      already added them without the extension name a while back.
      
      I tried to keep the descriptions in RISCVFeatures.td fairly short since
      the strings show up in `-print-supported-extensions`.
      58c494f4
    • Philip Reames's avatar
    • Rahman Lavaee's avatar
      [SHT_LLVM_BB_ADDR_MAP] Allow basic-block-sections and labels be used together... · acec6419
      Rahman Lavaee authored
      [SHT_LLVM_BB_ADDR_MAP] Allow basic-block-sections and labels be used together by decoupling the handling of the two features. (#74128)
      
      Today `-split-machine-functions` and `-fbasic-block-sections={all,list}`
      cannot be combined with `-basic-block-sections=labels` (the labels
      option will be ignored).
      The inconsistency comes from the way basic block address map -- the
      underlying mechanism for basic block labels -- encodes basic block
      addresses
      (https://lists.llvm.org/pipermail/llvm-dev/2020-July/143512.html).
      Specifically, basic block offsets are computed relative to the function
      begin symbol. This relies on functions being contiguous which is not the
      case for MFS and basic block section binaries. This means Propeller
      cannot use binary profiles collected from these binaries, which limits
      the applicability of Propeller for iterative optimization.
          
      To make the `SHT_LLVM_BB_ADDR_MAP` feature work with basic block section
      binaries, we propose modifying the encoding of this section as follows.
      
      First let us review the current encoding which emits the address of each
      function and its number of basic blocks, followed by basic block entries
      for each basic block.
      
      | | |
      |--|--|
      | Address of the function | Function Address |
      |  Number of basic blocks in this function | NumBlocks |
      |  BB entry 1
      |  BB entry 2
      |   ...
      |  BB entry #NumBlocks
          
      To make this work for basic block sections, we treat each basic block
      section similar to a function, except that basic block sections of the
      same function must be encapsulated in the same structure so we can map
      all of them to their single function.
          
      We modify the encoding to first emit the number of basic block sections
      (BB ranges) in the function. Then we emit the address map of each basic
      block section section as before: the base address of the section, its
      number of blocks, and BB entries for its basic block. The first section
      in the BB address map is always the function entry section.
      | | |
      |--|--|
      |  Number of sections for this function   | NumBBRanges |
      | Section 1 begin address                     | BaseAddress[1]  |
      | Number of basic blocks in section 1 | NumBlocks[1]    |
      | BB entries for Section 1
      |..................|
      | Section #NumBBRanges begin address | BaseAddress[NumBBRanges] |
      | Number of basic blocks in section #NumBBRanges |
      NumBlocks[NumBBRanges] |
      | BB entries for Section #NumBBRanges
          
      The encoding of basic block entries remains as before with the minor
      change that each basic block offset is now computed relative to the
      begin symbol of its containing BB section.
          
      This patch adds a new boolean codegen option `-basic-block-address-map`.
      Correspondingly, the front-end flag `-fbasic-block-address-map` and LLD
      flag `--lto-basic-block-address-map` are introduced.
      Analogously, we add a new TargetOption field `BBAddrMap`. This means BB
      address maps are either generated for all functions in the compiling
      unit, or for none (depending on `TargetOptions::BBAddrMap`).
          
      This patch keeps the functionality of the old
      `-fbasic-block-sections=labels` option but does not remove it. A
      subsequent patch will remove the obsolete option.
      
      We refactor the `BasicBlockSections` pass by separating the BB address
      map and BB sections handing to their own functions (named
      `handleBBAddrMap` and `handleBBSections`). `handleBBSections` renumbers
      basic blocks and places them in their assigned sections.
      `handleBBAddrMap` is invoked after `handleBBSections` (if requested) and
      only renumbers the blocks.
        - New tests added:
      - Two tests basic-block-address-map-with-basic-block-sections.ll and
      basic-block-address-map-with-mfs.ll to exercise the combination of
      `-basic-block-address-map` with `-basic-block-sections=list` and
      '-split-machine-functions`.
      - A driver sanity test for the `-fbasic-block-address-map` option
      (basic-block-address-map.c).
      - An LLD test for testing the `--lto-basic-block-address-map` option.
      This reuses the LLVM IR from `lld/test/ELF/lto/basic-block-sections.ll`.
      - Renamed and modified the two existing codegen tests for basic block
      address map (`basic-block-sections-labels-functions-sections.ll` and
      `basic-block-sections-labels.ll`)
      - Removed `SHT_LLVM_BB_ADDR_MAP_V0` tests. Full deprecation of
      `SHT_LLVM_BB_ADDR_MAP_V0` and `SHT_LLVM_BB_ADDR_MAP` version less than 2
      will happen in a separate PR in a few months.
      acec6419
    • Yuta Mukai's avatar
      [AArch64][MachinePipeliner] Add pipeliner support for AArch64 (#79589) · 70eab122
      Yuta Mukai authored
      Add AArch64 implementations for the interfaces of MachinePipeliner pass.
      The pass is disabled by default for AArch64. It is enabled by specifying
      --aarch64-enable-pipeliner.
      
      5 tests in llvm-test-suites show performance improvement by more than 5%
      on a Neoverse V1 processor.
      
      | test | improvement |
      | ---------------------------------------------------------------- |
      -----------:|
      | MultiSource/Benchmarks/TSVC/Recurrences-dbl/Recurrences-dbl.test | 16%
      |
      | MultiSource/Benchmarks/TSVC/Recurrences-dbl/Recurrences-flt.test | 16%
      |
      | SingleSource/Benchmarks/Adobe-C++/loop_unroll.test | 14% |
      | SingleSource/Benchmarks/Misc/flops-5.test | 13% |
      | SingleSource/Benchmarks/BenchmarkGame/spectral-norm.test | 6% |
      
      (base flags: -mcpu=neoverse-v1 -O3 -mrecip, flags for pipelining: -mllvm
      -aarch64-enable-pipeliner -mllvm
      -pipeliner-max-stages=100 -mllvm -pipeliner-max-mii=100 -mllvm
      -pipeliner-enable-copytophi=0)
      
      On the other hand, there are cases of significant performance
      degradation. Algorithm improvements and adding the option/pragma will be
      needed in the future.
      70eab122
    • Aiden Grossman's avatar
      cc0d752f
    • Jacques Pienaar's avatar
      59eadcd2
    • Nico Weber's avatar
      [gn] port ecb5a1b0 · ff319403
      Nico Weber authored
      ff319403
    • michaelrj-google's avatar
      [libc][bazel] disable epoll_pwait2 (#80362) · 4d89356f
      michaelrj-google authored
      Similar to #80051. The epoll_pwait2 syscall isn't available on all
      target platforms, and this is causing downstream test failures. This
      patch disables it until it can be detected whether or not it is
      available.
      4d89356f
    • Jakub Kuderski's avatar
      [mlir][spirv][memref] Calculate alignment for `PhysicalStorageBuffer`s (#80243) · 8fd0bce4
      Jakub Kuderski authored
      The SPIR-V spec requires that memory accesses to
      `PhysicalStorageBuffer`s are annotated with appropriate alignment
      attributes [1]. Calculate these based on memref alignment attributes or
      scalar type sizes.
      
      [1] Otherwise spirv-val complains:
      ```
      [VULKAN] ! Validation Error: [ VUID-VkShaderModuleCreateInfo-pCode-01379 ] | MessageID = 0x2a1bf17f | SPIR-V module not valid: [VUID-StandaloneSpirv-PhysicalStorageBuffer64-04708] Memory accesses with PhysicalStorageBuffer must use Aligned.
        %48 = OpLoad %float %47
      ```
      8fd0bce4
    • Peiming Liu's avatar
    • Greg Clayton's avatar
    • Peiming Liu's avatar
    • Philip Reames's avatar
    • Hana Dusíková's avatar
    • michaelrj-google's avatar
      [libc] Support epoll_wait using epoll_pwait (#80224) · ecdbffe5
      michaelrj-google authored
      The epoll_wait syscall is equivalent to calling epoll_pwait with a null
      sigset. This is useful to support systems that have epoll_pwait but not
      epoll_wait.
      ecdbffe5
    • Kyungwoo Lee's avatar
      [lld-macho] icf objc stubs (#79730) · 39139317
      Kyungwoo Lee authored
      This supports icf for objc stubs.
      39139317
    • Philip Reames's avatar
    • Greg Clayton's avatar
      [lldb] Fix a crash when using .dwp files and make type lookup reliable with... · 9258f3e6
      Greg Clayton authored
      [lldb] Fix a crash when using .dwp files and make type lookup reliable with the index cache (#79544)
      
      When using split DWARF with .dwp files we had an issue where sometimes
      the DWO file within the .dwp file would be parsed _before_ the skeleton
      compile unit. The DWO file expects to be able to always be able to get a
      link back to the skeleton compile unit. Prior to this fix, the only time
      the skeleton compile unit backlink would get set, was if the unit
      headers for the main executable have been parsed _and_ if the unit DIE
      was parsed in that DWARFUnit. This patch ensures that we can always get
      the skeleton compile unit for a DWO file by adding a function:
      
      ```
      DWARFCompileUnit *DWARFUnit::GetSkeletonUnit();
      ```
      
      Prior to this fix DWARFUnit had some unsafe accessors that were used to
      store two different things:
      
      ```
        void *DWARFUnit::GetUserData() const;
        void DWARFUnit::SetUserData(void *d);
      ```
      
      This was used by SymbolFileDWARF to cache the `lldb_private::CompileUnit
      *` for a SymbolFileDWARF and was also used to store the `DWARFUnit *`
      for SymbolFileDWARFDwo. This patch clears up this unsafe usage by adding
      two separate accessors and ivars for this:
      ```
      lldb_private::CompileUnit *DWARFUnit::GetLLDBCompUnit() const { return m_lldb_cu; }
      void DWARFUnit::SetLLDBCompUnit(lldb_private::CompileUnit *cu) { m_lldb_cu = cu; }
      DWARFCompileUnit *DWARFUnit::GetSkeletonUnit();
      void DWARFUnit::SetSkeletonUnit(DWARFUnit *skeleton_unit);
      ```
      This will stop anyone from calling `void *DWARFUnit::GetUserData()
      const;` and casting the value to an incorrect value.
      
      A crash could occur in `SymbolFileDWARF::GetCompUnitForDWARFCompUnit()`
      when the `non_dwo_cu`, which is a backlink to the skeleton compile unit,
      was not set and was NULL. There is an assert() in the code, and then the
      code just will kill the program if the assert isn't enabled because the
      code looked like:
      ```
        if (dwarf_cu.IsDWOUnit()) {
          DWARFCompileUnit *non_dwo_cu =
              static_cast<DWARFCompileUnit *>(dwarf_cu.GetUserData());
          assert(non_dwo_cu);
          return non_dwo_cu->GetSymbolFileDWARF().GetCompUnitForDWARFCompUnit(
              *non_dwo_cu);
        }
      ```
      This is now fixed by calling the `DWARFUnit::GetSkeletonUnit()` which
      will correctly always get the skeleton compile uint for a DWO file
      regardless of if the skeleton unit headers have been parse or if the
      skeleton unit DIE wasn't parsed yet.
      
      To implement the ability to get the skeleton compile units, I added code
      the DWARFDebugInfo.cpp/.h that make a map of DWO ID -> skeleton
      DWARFUnit * that gets filled in for DWARF5 when the unit headers are
      parsed. The `DWARFUnit::GetSkeletonUnit()` will end up parsing the unit
      headers of the main executable to fill in this map if it already hasn't
      been done. For DWARF4 and earlier we maintain a separate map that gets
      filled in only for any DWARF4 compile units that have a DW_AT_dwo_id or
      DW_AT_gnu_dwo_id attributes. This is more expensive, so this is done
      lazily and in a thread safe manor. This allows us to be as efficient as
      possible when using DWARF5 and also be backward compatible with DWARF4 +
      split DWARF.
      
      There was also an issue that stopped type lookups from succeeding in
      `DWARFDIE SymbolFileDWARF::GetDIE(const DIERef &die_ref)` where it
      directly was accessing the `m_dwp_symfile` ivar without calling the
      accessor function that could end up needing to locate and load the .dwp
      file. This was fixed by calling the
      `SymbolFileDWARF::GetDwpSymbolFile()` accessor to ensure we always get a
      valid value back if we can find the .dwp file. Prior to this fix it was
      down which APIs were called and if any APIs were called that loaded the
      .dwp file, it worked fine, but it might not if no APIs were called that
      did cause it to get loaded.
      
      When we have valid debug info indexes and when the lldb index cache was
      enabled, this would cause this issue to show up more often.
      
      I modified an existing test case to test that all of this works
      correctly and doesn't crash.
      9258f3e6
    • Natalie Chouinard's avatar
      [docs] Add beginner-focused office hours (#80308) · 5d228eaf
      Natalie Chouinard authored
      These are initially being hosted by a rotating cast of: @danakj
      @gburgessiv @nickdesaulniers @sudonatalie
      5d228eaf
    • Aart Bik's avatar
      [mlir][sparse] external entry method wrapper for sparse tensors (#80326) · 33b463ad
      Aart Bik authored
      Similar to the emit_c_interface, this pull request adds a pass that
      converts public entry methods that use sparse tensors as input
      parameters and/or output return values into wrapper functions that
      [dis]assemble the individual tensors that constitute the actual storage
      used externally into MLIR sparse tensors. This pass can be used to
      prepare the public entry methods of a program that is compiled by the
      MLIR sparsifier to interface with an external runtime, e.g., when
      passing sparse tensors as numpy arrays from and to Python. Note that
      eventual bufferization decisions (e.g. who [de]allocates the underlying
      memory) should be resolved in agreement with the external runtime
      (Python, PyTorch, JAX, etc.)
      33b463ad
    • Craig Topper's avatar
      [StackSlotColoring] Ignore non-spill objects in RemoveDeadStores. (#80242) · 5cf0fb43
      Craig Topper authored
      The stack slot coloring pass is concerned with optimizing spill
      slots. If any change is a pass is made over the function to remove
      stack stores that use the same register and stack slot as an
      immediately preceding load.
          
      The register check is too simple for constant registers like AArch64
      and RISC-V's zero register. This register can be used as the result
      of a load if we want to discard the result, but still have the memory
      access performed. Like for a volatile or atomic load.
          
      If the code sees a load from the zero register followed by a store
      of the zero register at the same stack slot, the pass mistakenly
      believes the store isn't needed.
          
      Since the main stack coloring optimization is only concerned with
      spill slots, it seems reasonable that RemoveDeadStores should
      only be concerned with spills. Since we never generate a reload of
      x0, this avoids the issue seen by RISC-V.
          
      Test case concept is adapted from pr30821.mir from X86. That test
      had to be updated to mark the stack slot as a spill slot.
          
      Fixes #80052.
      5cf0fb43
    • Nick Desaulniers's avatar
      [libc][stdbit] fix return types (#80337) · edbd93d3
      Nick Desaulniers authored
      All of the functions I've previously implemented return an unsigned int; not
      the parameter type.
      edbd93d3