1. Feb 25, 2024
    • Jacek Caban's avatar
      [llvm-ar] Use COFF archive format for COFF targets. (#82642) · cf9201cf
      Jacek Caban authored
      Detect COFF files by default and allow specifying it with --format
      argument.
      
      This is important for ARM64EC, which uses a separated symbol map for EC
      symbols. Since K_COFF is mostly compatible with K_GNU, this shouldn't
      really make a difference for other targets.
      cf9201cf
  2. Feb 24, 2024
    • Simon Pilgrim's avatar
      MachineInstr - update TargetRegisterInfo arguments comments. NFC. · d877ab1b
      Simon Pilgrim authored
      "TargetRegisterInfo is passed" -> "TargetRegisterInfo is non-null" - matches the term in the rest of the header.
      d877ab1b
    • Artem Tyurin's avatar
      [InstCombine] Handle more even/odd math functions (#81324) · 1901f442
      Artem Tyurin authored
      At the moment this PR adds support only for `erf` function.
      
      Fixes #77220.
      1901f442
    • Serge Pavlov's avatar
      [AArch64] Intrinsics aarch64_{get,set}_fpsr (#81867) · 00c0638b
      Serge Pavlov authored
      Two new intrinsics are introduced to read/write FPSR. They are similar
      to the existing intrinsics aarch64_{get,set}_fpcr.
      00c0638b
    • Aiden Grossman's avatar
      [llvm-exegesis] Fix typos in README · 60a904b2
      Aiden Grossman authored
      60a904b2
    • Matthias Springer's avatar
      [mlir] Use `OpBuilder::createBlock` in op builders and patterns (#82770) · 91d5653e
      Matthias Springer authored
      When creating a new block in (conversion) rewrite patterns,
      `OpBuilder::createBlock` must be used. Otherwise, no
      `notifyBlockInserted` notification is sent to the listener.
      
      Note: The dialect conversion relies on listener notifications to keep
      track of IR modifications. Creating blocks without the builder API can
      lead to memory leaks during rollback.
      91d5653e
    • yingopq's avatar
      [Mips] Fix unable to handle inline assembly ends with compat-branch o… (#77291) · 96abee5e
      yingopq authored
      …n MIPS
      
      Modify:
      Add a global variable 'CurForbiddenSlotAttr' to save current
      instruction's forbidden slot and whether set reorder. This is the
      judgment condition for whether to add nop. We would add a couple of
      '.set noreorder' and '.set reorder' to wrap the current instruction and
      the next instruction.
      Then we can get previous instruction`s forbidden slot attribute and
      whether set reorder by 'CurForbiddenSlotAttr'.
      If previous instruction has forbidden slot and .set reorder is active
      and current instruction is CTI. Then emit a NOP after it.
      
      Fix https://github.com/llvm/llvm-project/issues/61045.
      
      Because https://reviews.llvm.org/D158589 was 'Needs Review' state, not
      ending, so we commit pull request again.
      96abee5e
    • MalaySanghiIntel's avatar
      Convert argument to reference. (#82741) · 330af6ed
      MalaySanghiIntel authored
      Avoid copy of large object
      330af6ed
    • Owen Pan's avatar
      [clang-format][NFC] Enable RemoveSemicolon for clang-format style (#82735) · b0d2a52c
      Owen Pan authored
      Also insert separators for decimal integers longer than 4 digits.
      b0d2a52c
    • Peter Klausler's avatar
      [flang] Ensure USE-associated objects can be in NAMELIST (#82846) · 31ab2c4f
      Peter Klausler authored
      The name resolution for NAMELIST objects didn't allow for symbols that
      are not ObjectEntityDetails symbols.
      
      Fixes https://github.com/llvm/llvm-project/issues/82574.
      31ab2c4f
    • Shilei Tian's avatar
    • Michael Spencer's avatar
      [clang][ScanDeps] Allow PCHs to have different VFS overlays (#82294) · de3b2c29
      Michael Spencer authored
      It turns out it's not that uncommon for real code to pass a different
      set of VFSs while building a PCH than while using the PCH. This can
      cause problems as seen in `test/ClangScanDeps/optimize-vfs-pch.m`. If
      you scan `compile-commands-tu-no-vfs-error.json` without -Werror and run
      the resulting commands, Clang will emit a fatal error while trying to
      emit a note saying that it can't find a remapped header.
      
      This also adds textual tracking of VFSs for prebuilt modules that are
      part of an included PCH, as the same issue can occur in a module we are
      building if we drop VFSs. This has to be textual because we have no
      guarantee the PCH had the same list of VFSs as the current TU.
      
      This uses the `PrebuiltModuleListener` to collect `VFSOverlayFiles`
      instead of trying to extract it out of a `serialization::ModuleFile`
      each time it's needed. There's not a great way to just store a pointer
      to the list of strings in the serialized AST.
      de3b2c29
    • Michael Spencer's avatar
      reland: [clang][ScanDeps] Canonicalize -D and -U flags (#82568) · d42de86e
      Michael Spencer authored
      Canonicalize `-D` and `-U` flags by sorting them and only keeping the
      last instance of a given name.
      
      This optimization will only fire if all `-D` and `-U` flags start with a
      simple identifier that we can guarantee a simple analysis of can
      determine if two flags refer to the same identifier or not. See the
      comment on `getSimpleMacroName()` for details of what the issues are.
      
      Previous version of this had issues with sed differences between macOS,
      Linux, and Windows. This test doesn't check paths, so just don't run
      sed.
      Other tests should use `sed -E 's:\\\\?:/:g'` to get portable behavior.
      
      Windows has different command line parsing behavior than Linux for
      compilation databases, so the test has been adjusted to ignore that
      difference.
      d42de86e
    • Jeffrey Byrnes's avatar
      [AMDGPU] Introduce iglp_opt(2): Generalized exp/mfma interleaving for select kernels (#81342) · 8f2bd8ae
      Jeffrey Byrnes authored
      This implements the basic pipelining structure of exp/mfma interleaving
      for better extensibility. While it does have improved extensibility,
      there are controls which only enable it for DAGs with certain
      characteristics (matching the DAGs it has been designed against).
      8f2bd8ae
    • Jorge Gorbe Moya's avatar
      Revert "[RemoveDIs] Enable DPLabels conversion [3b/3] (#82639)" · c862e612
      Jorge Gorbe Moya authored
      This reverts commit 71d47a0b because
      it causes clang to crash in some cases. See repro posted at
      https://github.com/llvm/llvm-project/commit/71d47a0b00e9f48dc740556d7f452ffadf308731
      c862e612
    • Tom Stellard's avatar
      [llvm-shlib] Change libLLVM-$MAJOR.so symlink to point to versioned SO (#82660) · 10c48a77
      Tom Stellard authored
      This symlink was added in 91a38462 to
      maintain backwards compatibility, but it needs to point to
      libLLVM.so.$MAJOR.$MINOR rather than libLLVM.so. This works better for
      distros that ship libLLVM.so and libLLVM.so.$MAJOR.$MINOR in separate
      packages and also prevents mistakes like
      libLLVM-19.so -> libLLVM.so -> libLLVM.so.18.1
      
      Fixes #82647
      10c48a77
    • Visoiu Mistrih Francis's avatar
      [RISCV] Add scheduling info for Zcmp (#82719) · 775bd603
      Visoiu Mistrih Francis authored
      The order of the entries in the list is:
      
      outs, ins, Defs, Uses, implicit-defs, implicit uses, where the last two
      are added programatically during codegen depending on the registers
      saved/restored and are not described in the TD files.
      775bd603
    • Adrian Prantl's avatar
      Revert "Replace ArchSpec::PiecewiseCompare() with Triple::operator==()" · 3f91bdfd
      Adrian Prantl authored
      This reverts commit 5e6bed8c0ea2f7fe380127763c8f753adae0fc1b while investigating the bots.
      3f91bdfd
    • Jason Molenda's avatar
      [lldb] Correctly annotate threads at a bp site as hitting it (#82709) · 87fadb39
      Jason Molenda authored
      This is next in my series of "fix the racey tests that fail on
      greendragon" addressing the failure of TestConcurrentManyBreakpoints.py
      where we set a breakpoint in a function that 100 threads execute, and we
      check that we hit the breakpoint 100 times. But sometimes it is only hit
      99 times, and the test fails.
      
      When we hit a software breakpoint, the pc value for the thread is the
      address of the breakpoint instruction - as if it had not been hit yet.
      And because a user might ADD a breakpoint for the current pc from the
      commandline, when we go to resume execution, any thread that is sitting
      at a breakpoint site will be silently advanced past the breakpoint
      instruction (disable bp, instruction step that thread, re-enable bp)
      before resuming -- whether that thread has hit its breakpoint or not.
      
      What this test is exposing is that there is another corner case, a
      thread that is sitting at a breakpoint site but has not yet executed the
      breakpoint instruction. The thread will have no stop reason, no mach
      exception, so it will not be recorded as having hit the breakpoint
      (because it hasn't yet). But when we resume execution, because it is
      sitting at a breakpoint site, we advance past it and miss the breakpoint
      hit.
      
      In 2016 Abhishek Aggarwal handled a similar issue with a patch in
      `ProcessGDBRemote::SetThreadStopInfo()`, adding a breakpoint StopInfo
      for a thread sitting at a breakpoint site that has no stop reason.
      debugserver's `jThreadsInfo` would not correctly execute Abhishek's code
      though because it would respond with `"reason":"none"` for a thread with
      no stop reason, and `SetThreadStopInfo()` expected an empty reason here.
      The first part of my patch is to clear the `reason` if it is `"none"` so
      we flow through the code correctly.
      
      On Darwin, though, our stop reply packet (Txx...) includes the
      `threads`, `thread-pcs`, and `jstopinfo` keys, which give us the tids
      for all current threads, the pc values for those threads, and
      `jstopinfo` has a JSON dictionary with the mach exceptions for all
      threads that have a mach exception. In
      `ProcessGDBRemote::CalculateThreadStopInfo()` we set the StopInfo for
      each thread for a private stop and if we have `jstopinfo` it is the
      source of all the StopInfos. I have to add the same logic here, to give
      the thread a breakpoint StopInfo even though it hasn't executed the
      breakpoint yet. In this case we are very early in thread construction
      and I only have the information in the Txx stop reply packet -- tids,
      pcs, and jstopinfo, so I can't use the normal general mechanisms of
      going through the RegisterContext to get the pc, it's a bit different.
      
      If I hack debugserver to not issue `jstopinfo`,
      `CalculateThreadStopInfo` will fall back to sending `qThreadStopInfo`
      for each thread and going through
      `ProcessGDBRemote::SetThreadStopInfo()` to set the stop infos (and with
      the `reason:none` fix, use Abhishek's code).
      
      rdar://110549165
      87fadb39
    • Joseph Huber's avatar
      [libc][NFC] Remove all trailing spaces from libc (#82831) · 69c0b2fe
      Joseph Huber authored
      Summary:
      There are a lot of random training spaces on various lines. This patch
      just got rid of all of them with `sed 's/\ \+$//g'.
      69c0b2fe
    • Adrian Prantl's avatar
      Replace ArchSpec::PiecewiseCompare() with Triple::operator==() (#82804) · 25940956
      Adrian Prantl authored
      Looking ast the definition of both functions this is *almost* an NFC
      change, except that Triple also looks at the SubArch (important) and
      ObjectFormat (less so).
      
      This fixes a bug that only manifests with how Xcode uses the SBAPI to
      attach to a process by name: it guesses the architecture based on the
      system. If the system is arm64 and the Process is arm64e Target fails to
      update the triple because it deemed the two to be equivalent.
      
      rdar://123338218
      25940956
    • agozillon's avatar
      [OpenMP][MLIR][OMPIRBuilder] Add a small optional constant alloca raise... · dcf4ca55
      agozillon authored
      [OpenMP][MLIR][OMPIRBuilder] Add a small optional constant alloca raise function pass to finalize, utilised in convertTarget (#78818)
      
      This patch seeks to add a mechanism to raise constant (not ConstantExpr
      or runtime/dynamic) sized allocations into the entry block for select
      functions that have been inserted into a list for processing. This
      processing occurs during the finalize call, after OutlinedInfo regions
      have completed. This currently has only been utilised for
      createOutlinedFunction, which is triggered for TargetOp generation in
      the OpenMP MLIR dialect lowering to LLVM-IR.
      
      This currently is required for Target kernels generated by
      createOutlinedFunction to avoid subsequent optimization passes doing
      some unintentional malformed optimizations for AMD kernels (unsure if it
      occurs for other vendors). If the allocas are generated inside of the
      kernel and are not in the entry block and are subsequently passed to a
      function this can lead to required instructions being erased or
      manipulated in a way that causes the kernel to run into a HSA access
      error.
      
      This fix is related to a series of problems found in:
      https://github.com/llvm/llvm-project/issues/74603
      
      This problem primarily presents itself for Flang's HLFIR AssignOp
      currently, when utilised with a scalar temporary constant on the RHS and
      a descriptor type on the LHS. It will generate a call to a runtime
      function, wrap the RHS temporary in a newly allocated descriptor (an
      llvm struct), and pass both the LHS and RHS descriptor into the runtime
      function call. This will currently be
      embedded into the middle of the target region in the user entry block,
      which means the allocas are also embedded in the middle, which seems to
      pose
      issues when later passes are executed. This issue may present itself in
      other HLFIR operations or unrelated operations that generate allocas as
      a by product, but for the moment, this one test case is the only
      scenario I've found this problem.
      
      Perhaps this is not the appropriate fix, I am very open to other
      suggestions, I've tried a few others (at varying levels of the
      flang/mlir compiler flow), but this one is the smallest and least
      intrusive change set. The other two, that come to mind (but I've not
      fully looked into, the former I tried a little with blocks but it had a
      few issues I'd need to think through):
      
      - Having a proper alloca only block (or region) generated for TargetOps
      that we could merge into the entry block that's generated by
      convertTarget's createOutlinedFunction.
      - Or diverging a little from Clang's current target generation and using
      the CodeExtractor to generate the user code as an outlined function
      region invoked from the kernel we make, with our kernel arguments passed
      into it. Similar to the current parallel generation. I am not sure how
      well this would intermingle with the existing parallel generation though
      that's layered in.
      
      Both of these methods seem like quite a divergence from the current
      status quo, which I am not entirely sure is merited for the small test
      this change aims to fix.
      dcf4ca55
    • Krzysztof Parzyszek's avatar
      [flang][OpenMP] Set OpenMP attributes in MLIR module in bbc before lo… (#82774) · 47aee8b5
      Krzysztof Parzyszek authored
      …wering
      
      Right now attributes like OpenMP version or target attributes for
      offload are set after lowering in bbc. The flang frontend sets them
      before lowering, making them available in the lowering process.
      
      This change sets them before lowering in bbc as well.
      47aee8b5
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][cuda] Allow object with SHARED attribute as definable (#82822) · 5c90527b
      A semantic error was raised in device subprogram like: 
      
      ```
      attributes(global) subroutine devsubr2()
         real, shared :: rs
         rs = 1
      end subroutine
      ```
      
      Object with the SHARED attribute can be can be read or written by all
      threads in the block.
      
      
      https://docs.nvidia.com/hpc-sdk/archive/24.1/compilers/cuda-fortran-prog-guide/index.html#cfpg-var-qual-attr-shared
      5c90527b
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][cuda] Fix semantic for the CONSTANT attribute (#82821) · 99f31bab
      Object with the CONSTANT attribute cannot be declared in the host
      subprogram.
      
      It can be declared in a module or a device subprogram.
      
      Adapt the semantic check to trigger the error in host subprogram.
      99f31bab
    • Timothy Herchen's avatar
      [X86][MC] Reject out-of-range control and debug registers encoded with APX (#82584) · ae91a427
      Timothy Herchen authored
      Fixes #82557. APX specification states that the high bits found in REX2
      used to encode GPRs can also be used to encode control and debug
      registers, although all of them will #UD. Therefore, when disassembling
      we reject attempts to create control or debug registers with a value of
      16 or more.
      
      See page 22 of the
      [specification](https://www.intel.com/content/www/us/en/developer/articles/technical/advanced-performance-extensions-apx.html):
      
      > Note that the R, X and B register identifiers can also address non-GPR
      register types, such as vector registers, control registers and debug
      registers. When any of them does, the highest-order bits REX2.R4,
      REX2.X4 or REX2.B4 are generally ignored, except when the register being
      addressed is a control or debug register. [...] The exception is that
      REX2.R4 and REX2.R3 [*sic*] are not ignored when the R register
      identifier addresses a control or debug register. Furthermore, if any
      attempt is made to access a non-existent control register (CR*) or debug
      register (DR*) using the REX2 prefix and one of the following
      instructions:
      “MOV CR*, r64”, “MOV r64, CR*”, “MOV DR*, r64”, “MOV r64, DR*”. #UD is
      raised.
      
      The invalid encodings are 64-bit only because `0xd5` is a valid
      instruction in 32-bit mode.
      ae91a427
    • Aart Bik's avatar
    • Joseph Huber's avatar
      [Clang] Append target search paths for direct offloading compilation (#82699) · 99660082
      Joseph Huber authored
      Summary:
      Recent changes to the `libc` project caused the headers to be installed
      to `include/<triple>` for the GPU and the libraries to be in
      `lib/<triple>`. This means we should automatically append these search
      paths so they can be found by default. This allows the following to work
      targeting AMDGPU.
      
      ```shell
      $ clang foo.c -flto -mcpu=native --target=amdgcn-amd-amdhsa -lc <install>/lib/amdgcn-amd-amdhsa/crt1.o
      $ amdhsa-loader a.out
      ```
      99660082
    • Joseph Huber's avatar
      [libc] Install a single LLVM-IR version of the GPU library (#82791) · b43dd08a
      Joseph Huber authored
      Summary:
      Recent patches have allowed us to treat these libraries as direct
      builds. This makes it easier to simply build them to a single LLVM-IR
      file. This matches the way these files are presented by the ROCm and
      CUDA toolchains and makes it easier to work with.
      b43dd08a
    • Joseph Huber's avatar
      [libc] Remove 'llvm-gpu-none' directory from build (#82816) · 1a2ecbb3
      Joseph Huber authored
      Summary:
      This directory is leftover from when we handled both AMDGPU and NVPTX in
      the same build and merged them into a pseudo triple. Now the only thing
      it contains is the RPC server header. This gets rid of it, but now that
      it's in the base install directory we should make it clear that it's an
      LLVM libc header.
      1a2ecbb3
    • Joseph Huber's avatar
      [libc] Remove use of BlockStore for GPU atexit (#82823) · a3a316e2
      Joseph Huber authored
      Summary:
      The GPU backends have restrictions on the kinds of initializers they can
      produce. The use of BlockStore here currently breaks the backends
      through the use of recursive initializers. This prevents it from
      actually being included in any builds. This patchs changes it to just
      use a fixed size of 64 slots .The chances of someone exceeding the 64
      slots in practice is very, very low.
      
      However, this is primarily a bandaid solution as a real solution will
      need to use a lock free data structure to push work in parallel.
      Currently the mutexes on the GPU build do nothing, so they only work if
      the user guards the use themselves.
      a3a316e2
    • Florian Mayer's avatar
      [NFC] Make RingBuffer an atomic pointer (#82547) · 6dd6d487
      Florian Mayer authored
      This will allow us to atomically swap out RingBuffer and StackDepot.
      
      Patched into AOSP and ran debuggerd_tests.
      6dd6d487
    • Michael Halkenhäuser's avatar
      [llvm-link] Improve missing file error message (#82514) · a64ff963
      Michael Halkenhäuser authored
      Add error messages showing the missing filenames.
      
      Currently, we only get 'No such file or directory' without any(!)
      further info. This patch will (only upon ENOENT error) iterate over all
      requested files and print which ones are actually missing.
      a64ff963
    • David Goldman's avatar
      [clangd] Fix renaming single argument ObjC methods (#82396) · 59e5519c
      David Goldman authored
      Use the legacy non-ObjC rename logic when dealing with selectors that
      have zero or one arguments. In addition, make sure we don't add an extra
      `:` during the rename.
      
      Add a few more tests to verify this works (thanks to @ahoppen for the
      tests and finding this bug).
      59e5519c
    • LLVM GN Syncbot's avatar
      [gn build] Port 5874874c · 07fd5ca3
      LLVM GN Syncbot authored
      07fd5ca3
    • Min-Yih Hsu's avatar
      [SelectionDAG] Introducing the SelectionDAG pattern matching framework (#78654) · 5874874c
      Min-Yih Hsu authored
      Akin to `llvm::PatternMatch` and `llvm::MIPatternMatch`, the
      `llvm::SDPatternMatch` introduced in this patch provides a DSL-alike
      framework to match SDValue / SDNode with a more succinct syntax.
      5874874c
    • Aart Bik's avatar
      [mlir][sparse] cleanup sparse runtime library (#82807) · f8ce460e
      Aart Bik authored
      remove some obsoleted APIs from the library that have been fully
      replaced with actual direct IR codegen
      f8ce460e
    • Krzysztof Parzyszek's avatar
      [flang][bbc] Fix dangling reference to `envDefaults` (#82800) · a24421fe
      Krzysztof Parzyszek authored
      The lowering bridge stores the evvironment defaults (passed to the
      constructor) as a reference. In the call to the constructor in bbc, the
      defaults were passed as `{}`, which creates a temporary whose lifetime
      ends immediately after the call.
      
      The flang driver passes a member of the compilation instance to the
      constructor, which presumably remains alive long enough, so storing the
      reference in the bridge is justified. To avoid the dangling reference,
      create an actual object `envDefaults` in bbc.
      a24421fe
    • Jay Foad's avatar
      [AMDGPU] Simplify AMDGPUDisassembler::getInstruction by removing Res. (#82775) · 42f6f95e
      Jay Foad authored
      Remove all the code that set and tested Res. Change all convert*
      functions to return void since none of them can fail. getInstruction
      only has one main point of failure, after all calls to tryDecodeInst
      have failed.
      42f6f95e
    • Craig Topper's avatar
      [SelectionDAG] Remove unused VP strided load/store creation functions that build an MMO. (#82676) · 962a6970
      Craig Topper authored
      The base case of these call InferPtrInfo. This is dangerous due to
      #82657, but it turns out none of these are used.
      
      It seemed best to reduce the surface area until these are needed.
      962a6970