1. Feb 02, 2024
    • Timm Bäder's avatar
      [clang][Interp] Support ChooseExprs · 58ceefe0
      Timm Bäder authored
      58ceefe0
    • Timm Bäder's avatar
      [clang][Interp] Not all TypeTraitExprs are of bool type · 2147a2a4
      Timm Bäder authored
      In C, they return an integer, so emit their value as such.
      2147a2a4
    • Yuta Mukai's avatar
      [MachinePipeliner] Fix missing requirements for tests (#80386) · 374a600d
      Yuta Mukai authored
      Add asserts requirements for tests that verify debug output.
      374a600d
    • Fangrui Song's avatar
      [ELF] Fix compareSections assertion failure when OutputDescs in sectionCommands are non-contiguous · dee8786f
      Fangrui Song authored
      In a `--defsym y0=0 -T a.lds` link where a.lds contains only INSERT
      commands, the `script->sectionCommands` layout may be:
      ```
      orphan sections
      SymbolAssignment due to --defsym
      sections created by INSERT commands
      ```
      
      The `OutputDesc` objects are not contiguous in sortInputSections, and
      `compareSections` will be called with a SymbolAssignment argument,
      leading to an assertion failure.
      dee8786f
    • Jason Molenda's avatar
      [lldb] NFC fixes addressing David's feedback · 7dd790db
      Jason Molenda authored
      David Spickett had several suggestions for
      https://github.com/llvm/llvm-project/pull/79962 after I'd
      already merged it.  Address those.
      7dd790db
    • Shengchen Kan's avatar
      [X86] X86InstrInfo.cpp - Remove dead code for memory folding, NFCI · e270ec47
      Shengchen Kan authored
      `commuteInstruction(MI, false, OpNum, CommuteOpIdx2)` should never create
      any new instruction, so we don't need to check and erase it.
      e270ec47
    • Maksim Panchenko's avatar
      [BOLT] Remove duplicate expression (#80380) · 082fe9a5
      Maksim Panchenko authored
      Reported by cpp check static analyzer in #80111.
      
      Fixes #80111.
      082fe9a5
    • Craig Topper's avatar
      [RISCV] Add -march support for many of the S extensions mentioned in the... · 58c494f4
      Craig Topper authored
      [RISCV] Add -march support for many of the S extensions mentioned in the profile specification. (#79399)
      
      This is a good portion of the extensions mentioned in the RVA23 profile
      here
      https://github.com/riscv/riscv-profiles/blob/main/rva23-profile.adoc
      
      I don't believe these add any new CSRs. Sstc does add new CSRs, but we
      already added them without the extension name a while back.
      
      I tried to keep the descriptions in RISCVFeatures.td fairly short since
      the strings show up in `-print-supported-extensions`.
      58c494f4
    • Philip Reames's avatar
    • Rahman Lavaee's avatar
      [SHT_LLVM_BB_ADDR_MAP] Allow basic-block-sections and labels be used together... · acec6419
      Rahman Lavaee authored
      [SHT_LLVM_BB_ADDR_MAP] Allow basic-block-sections and labels be used together by decoupling the handling of the two features. (#74128)
      
      Today `-split-machine-functions` and `-fbasic-block-sections={all,list}`
      cannot be combined with `-basic-block-sections=labels` (the labels
      option will be ignored).
      The inconsistency comes from the way basic block address map -- the
      underlying mechanism for basic block labels -- encodes basic block
      addresses
      (https://lists.llvm.org/pipermail/llvm-dev/2020-July/143512.html).
      Specifically, basic block offsets are computed relative to the function
      begin symbol. This relies on functions being contiguous which is not the
      case for MFS and basic block section binaries. This means Propeller
      cannot use binary profiles collected from these binaries, which limits
      the applicability of Propeller for iterative optimization.
          
      To make the `SHT_LLVM_BB_ADDR_MAP` feature work with basic block section
      binaries, we propose modifying the encoding of this section as follows.
      ...
      acec6419
    • Yuta Mukai's avatar
      [AArch64][MachinePipeliner] Add pipeliner support for AArch64 (#79589) · 70eab122
      Yuta Mukai authored
      Add AArch64 implementations for the interfaces of MachinePipeliner pass.
      The pass is disabled by default for AArch64. It is enabled by specifying
      --aarch64-enable-pipeliner.
      
      5 tests in llvm-test-suites show performance improvement by more than 5%
      on a Neoverse V1 processor.
      
      | test | improvement |
      | ---------------------------------------------------------------- |
      -----------:|
      | MultiSource/Benchmarks/TSVC/Recurrences-dbl/Recurrences-dbl.test | 16%
      |
      | MultiSource/Benchmarks/TSVC/Recurrences-dbl/Recurrences-flt.test | 16%
      |
      | SingleSource/Benchmarks/Adobe-C++/loop_unroll.test | 14% |
      | SingleSource/Benchmarks/Misc/flops-5.test | 13% |
      | SingleSource/Benchmarks/BenchmarkGame/spectral-norm.test | 6% |
      
      (base flags: -mcpu=neoverse-v1 -O3 -mrecip, flags for pipelining: -mllvm
      -aarch64-enable-pipeliner -mllvm
      -pipeliner-max-stages=100 -mllvm -pipeliner-max-mii=100 -mllvm
      -pipeliner-enable-copytophi=0)
      
      On the other hand, there are cases of significant performance
      degradation. Algorithm improvements and adding the option/pragma will be
      needed in the future.
      70eab122
    • Aiden Grossman's avatar
      cc0d752f
    • Jacques Pienaar's avatar
      59eadcd2
    • Nico Weber's avatar
      [gn] port ecb5a1b0 · ff319403
      Nico Weber authored
      ff319403
    • michaelrj-google's avatar
      [libc][bazel] disable epoll_pwait2 (#80362) · 4d89356f
      michaelrj-google authored
      Similar to #80051. The epoll_pwait2 syscall isn't available on all
      target platforms, and this is causing downstream test failures. This
      patch disables it until it can be detected whether or not it is
      available.
      4d89356f
    • Jakub Kuderski's avatar
      [mlir][spirv][memref] Calculate alignment for `PhysicalStorageBuffer`s (#80243) · 8fd0bce4
      Jakub Kuderski authored
      The SPIR-V spec requires that memory accesses to
      `PhysicalStorageBuffer`s are annotated with appropriate alignment
      attributes [1]. Calculate these based on memref alignment attributes or
      scalar type sizes.
      
      [1] Otherwise spirv-val complains:
      ```
      [VULKAN] ! Validation Error: [ VUID-VkShaderModuleCreateInfo-pCode-01379 ] | MessageID = 0x2a1bf17f | SPIR-V module not valid: [VUID-StandaloneSpirv-PhysicalStorageBuffer64-04708] Memory accesses with PhysicalStorageBuffer must use Aligned.
        %48 = OpLoad %float %47
      ```
      8fd0bce4
    • Peiming Liu's avatar
    • Greg Clayton's avatar
    • Peiming Liu's avatar
    • Philip Reames's avatar
    • Hana Dusíková's avatar
    • michaelrj-google's avatar
      [libc] Support epoll_wait using epoll_pwait (#80224) · ecdbffe5
      michaelrj-google authored
      The epoll_wait syscall is equivalent to calling epoll_pwait with a null
      sigset. This is useful to support systems that have epoll_pwait but not
      epoll_wait.
      ecdbffe5
    • Kyungwoo Lee's avatar
      [lld-macho] icf objc stubs (#79730) · 39139317
      Kyungwoo Lee authored
      This supports icf for objc stubs.
      39139317
    • Philip Reames's avatar
    • Greg Clayton's avatar
      [lldb] Fix a crash when using .dwp files and make type lookup reliable with... · 9258f3e6
      Greg Clayton authored
      [lldb] Fix a crash when using .dwp files and make type lookup reliable with the index cache (#79544)
      
      When using split DWARF with .dwp files we had an issue where sometimes
      the DWO file within the .dwp file would be parsed _before_ the skeleton
      compile unit. The DWO file expects to be able to always be able to get a
      link back to the skeleton compile unit. Prior to this fix, the only time
      the skeleton compile unit backlink would get set, was if the unit
      headers for the main executable have been parsed _and_ if the unit DIE
      was parsed in that DWARFUnit. This patch ensures that we can always get
      the skeleton compile unit for a DWO file by adding a function:
      
      ```
      DWARFCompileUnit *DWARFUnit::GetSkeletonUnit();
      ```
      
      Prior to this fix DWARFUnit had some unsafe accessors that were used to
      store two different things:
      
      ```
        void *DWARFUnit::GetUserData() const;
        void DWARFUnit::SetUserData(void *d);
      ```
      
      This was used by SymbolFileDWARF to cache the `lldb_private::CompileUnit
      *` for a SymbolFileDWARF and was also used to store the `DWARFUnit *`
      for SymbolFileDWARFDwo. This patch clears up this unsafe usage by adding
      two separate accessors and ivars for this:
      ```
      lldb_private::CompileUnit *DWARFUnit::GetLLDBCompUnit() const { return m_lldb_cu; }
      void DWARFUnit::SetLLDBCompUnit(lldb_private::CompileUnit *cu) { m_lldb_cu = cu; }
      DWARFCompileUnit *DWARFUnit::GetSkeletonUnit();
      void DWARFUnit::SetSkeletonUnit(DWARFUnit *skeleton_unit);
      ```
      This will stop anyone from calling `void *DWARFUnit::GetUserData()
      const;` and casting the value to an incorrect value.
      
      A crash could occur in `SymbolFileDWARF::GetCompUnitForDWARFCompUnit()`
      when the `non_dwo_cu`, which is a backlink to the skeleton compile unit,
      was not set and was NULL. There is an assert() in the code, and then the
      code just will kill the program if the assert isn't enabled because the
      code looked like:
      ```
        if (dwarf_cu.IsDWOUnit()) {
          DWARFCompileUnit *non_dwo_cu =
              static_cast<DWARFCompileUnit *>(dwarf_cu.GetUserData());
          assert(non_dwo_cu);
          return non_dwo_cu->GetSymbolFileDWARF().GetCompUnitForDWARFCompUnit(
              *non_dwo_cu);
        }
      ```
      This is now fixed by calling the `DWARFUnit::GetSkeletonUnit()` which
      will correctly always get the skeleton compile uint for a DWO file
      regardless of if the skeleton unit headers have been parse or if the
      skeleton unit DIE wasn't parsed yet.
      
      To implement the ability to get the skeleton compile units, I added code
      the DWARFDebugInfo.cpp/.h that make a map of DWO ID -> skeleton
      DWARFUnit * that gets filled in for DWARF5 when the unit headers are
      parsed. The `DWARFUnit::GetSkeletonUnit()` will end up parsing the unit
      headers of the main executable to fill in this map if it already hasn't
      been done. For DWARF4 and earlier we maintain a separate map that gets
      filled in only for any DWARF4 compile units that have a DW_AT_dwo_id or
      DW_AT_gnu_dwo_id attributes. This is more expensive, so this is done
      lazily and in a thread safe manor. This allows us to be as efficient as
      possible when using DWARF5 and also be backward compatible with DWARF4 +
      split DWARF.
      
      There was also an issue that stopped type lookups from succeeding in
      `DWARFDIE SymbolFileDWARF::GetDIE(const DIERef &die_ref)` where it
      directly was accessing the `m_dwp_symfile` ivar without calling the
      accessor function that could end up needing to locate and load the .dwp
      file. This was fixed by calling the
      `SymbolFileDWARF::GetDwpSymbolFile()` accessor to ensure we always get a
      valid value back if we can find the .dwp file. Prior to this fix it was
      down which APIs were called and if any APIs were called that loaded the
      .dwp file, it worked fine, but it might not if no APIs were called that
      did cause it to get loaded.
      
      When we have valid debug info indexes and when the lldb index cache was
      enabled, this would cause this issue to show up more often.
      
      I modified an existing test case to test that all of this works
      correctly and doesn't crash.
      9258f3e6
    • Natalie Chouinard's avatar
      [docs] Add beginner-focused office hours (#80308) · 5d228eaf
      Natalie Chouinard authored
      These are initially being hosted by a rotating cast of: @danakj
      @gburgessiv @nickdesaulniers @sudonatalie
      5d228eaf
    • Aart Bik's avatar
      [mlir][sparse] external entry method wrapper for sparse tensors (#80326) · 33b463ad
      Aart Bik authored
      Similar to the emit_c_interface, this pull request adds a pass that
      converts public entry methods that use sparse tensors as input
      parameters and/or output return values into wrapper functions that
      [dis]assemble the individual tensors that constitute the actual storage
      used externally into MLIR sparse tensors. This pass can be used to
      prepare the public entry methods of a program that is compiled by the
      MLIR sparsifier to interface with an external runtime, e.g., when
      passing sparse tensors as numpy arrays from and to Python. Note that
      eventual bufferization decisions (e.g. who [de]allocates the underlying
      memory) should be resolved in agreement with the external runtime
      (Python, PyTorch, JAX, etc.)
      33b463ad
    • Craig Topper's avatar
      [StackSlotColoring] Ignore non-spill objects in RemoveDeadStores. (#80242) · 5cf0fb43
      Craig Topper authored
      The stack slot coloring pass is concerned with optimizing spill
      slots. If any change is a pass is made over the function to remove
      stack stores that use the same register and stack slot as an
      immediately preceding load.
          
      The register check is too simple for constant registers like AArch64
      and RISC-V's zero register. This register can be used as the result
      of a load if we want to discard the result, but still have the memory
      access performed. Like for a volatile or atomic load.
          
      If the code sees a load from the zero register followed by a store
      of the zero register at the same stack slot, the pass mistakenly
      believes the store isn't needed.
          
      Since the main stack coloring optimization is only concerned with
      spill slots, it seems reasonable that RemoveDeadStores should
      only be concerned with spills. Since we never generate a reload of
      x0, this avoids the issue seen by RISC-V.
          
      Test case concept is adapted from pr30821.mir from X86. That test
      had to be updated to mark the stack slot as a spill slot.
          
      Fixes #80052.
      5cf0fb43
    • Nick Desaulniers's avatar
      [libc][stdbit] fix return types (#80337) · edbd93d3
      Nick Desaulniers authored
      All of the functions I've previously implemented return an unsigned int; not
      the parameter type.
      edbd93d3
    • Philip Reames's avatar
      Revert "[RISCV] Refine cost on Min/Max reduction" (#80340) · 59e55906
      Philip Reames authored
      Reverts llvm/llvm-project#79402. Crash reported. On closer inspection,
      this patch does not handle Intrinsic::maximum and Intrinsic::minimum.
      59e55906
    • Alexey Bataev's avatar
      [TTI]Add support for strided loads/stores. · 8ad14b6d
      Alexey Bataev authored
      Added basic legality check and cost estimation functions for strided loads and stores.
      
      These interfaces will be built upon in https://github.com/llvm/llvm-project/pull/80310.
      
      Reviewers: preames
      
      Reviewed By: preames
      
      Pull Request: https://github.com/llvm/llvm-project/pull/80329
      8ad14b6d
    • Artem Dergachev's avatar
      [analyzer][HTMLRewriter] Cache partial rewrite results. (#80220) · 243bfed6
      Artem Dergachev authored
      This is a follow-up for 721dd3bc [analyzer] NFC: Don't regenerate
      duplicate HTML reports.
      
      Because HTMLRewriter re-runs the Lexer for syntax highlighting and macro
      expansion purposes, it may get fairly expensive when the rewriter is
      invoked multiple times on the same file. In the static analyzer (which
      uses HTMLRewriter for HTML output mode) we only get away with this
      because there are usually very few reports emitted per file. But if loud
      checkers are enabled, such as `webkit.*`, this may explode in complexity
      and even cause the compiler to run over the 32-bit SourceLocation
      addressing limit.
      
      This patch caches intermediate results so that re-lexing only needed to
      happen once.
      
      As the clever __COUNTER__ test demonstrates, "once" is still too many.
      Ideally we shouldn't re-lex anything at all, which remains a TODO.
      243bfed6
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][openacc][openmp] Use #0 from hlfir.declare value when generating bound ops (#80317) · fe408eb5
      `getDataOperandBaseAddr` retrieve the address of a value when we need to
      generate bound operations. When switching to HLFIR, we did not really
      handle the fact that this value was then pointing to the result of a
      hlfir.declare. Because of that the `#1` value was being used. `#0` value
      is carrying the correct information about lowerbounds and should be
      used. This patch updates the `getDataOperandBaseAddr` function to use
      the correct result value from hlfir.declare.
      fe408eb5
    • Anatoly Trosinenko's avatar
      [AArch64][PAC] Expand blend(reg, imm) operation in aarch64-pauth pass (#74729) · 08fccf80
      Anatoly Trosinenko authored
      In preparation for implementing code generation for more @llvm.ptrauth.* intrinsics, move the expansion of blend(register, small integer) variant of @llvm.ptrauth.blend to the AArch64PointerAuth pass, where most other PAuth-related code generation takes place.
      08fccf80
    • Micah Weston's avatar
      [SHT_LLVM_BB_ADDR_MAP][llvm-readobj] Implements llvm-readobj handling for PGOAnalysisMap. (#79520) · aaaff74f
      Micah Weston authored
      Adds raw printing of PGOAnalysisMap in llvm-readobj.
      
      I'm leaving the fixme's for a later patch that will provide a 'pretty'
      printing for BBFreq and BrProb (i.e. relative frequencies and
      probabilities) that will apply to both llvm-readobj and llvm-objdump.
      aaaff74f
    • michaelrj-google's avatar
      [libc] add bazel support for most of unistd (#80078) · 7a7d5481
      michaelrj-google authored
      Much of unistd involves modifying files. The tests for these functions
      need to use libc_make_test_file_path which didn't exist when they were
      first implemented. This patch adds most of unistd to the bazel along
      with the corresponding tests. Tests that modify directories had to be
      disabled since bazel doesn't seem to handle them properly.
      7a7d5481
    • Carlos Galvez's avatar
      [clang-tidy] Remove enforcement of rule C.48 from cppcoreguidelines-prefer-member-init (#80330) · 6f32d6a4
      Carlos Galvez authored
      This functionality already exists in
      cppcoreguidelines-use-default-member-init. It was deprecated from this
      check in clang-tidy 17.
      
      This allows us to fully decouple this check from the corresponding
      modernize check, which has an unhealthy dependency.
      
      Fixes https://github.com/llvm/llvm-project/issues/62169
      
      
      
      ---------
      
      Co-authored-by: default avatarCarlos Gálvez <carlos.galvez@zenseact.com>
      6f32d6a4
    • Kelvin Li's avatar
      [OpenMP] Fix typo (NFC) (#80332) · a063df20
      Kelvin Li authored
      a063df20
    • Maksim Panchenko's avatar
      [BOLT] Enable re-writing of Linux kernel binary (#80228) · a693ae53
      Maksim Panchenko authored
      Write modified Linux kernel binary to disk. The output is not supposed
      to be functional at the moment, but it will allow for future patches to
      test the output binary.
      a693ae53
    • Maksim Panchenko's avatar
      [BOLT] Adjust section sizes based on file offsets (#80226) · 116e801a
      Maksim Panchenko authored
      When we adjust section sizes while rewriting a binary, we should be
      using section offsets and not addresses to determine if section overlap.
      NFC for existing binaries.
      116e801a