1. Feb 14, 2024
    • lntue's avatar
      [libc] Allow BigInt class to use base word types other than uint64_t. (#81634) · 4e005515
      lntue authored
      This will allow DyadicFloat class to replace NormalFloat class.
      4e005515
    • Craig Topper's avatar
      [TypePromotion] Remove an unreachable 'return false'. NFC · d0a1bf8b
      Craig Topper authored
      The if and the else above this both return so this is unreachable.
      Delete it and remove the else after return.
      d0a1bf8b
    • James Y Knight's avatar
      [Sparc] limit MaxAtomicSizeInBitsSupported to 32 for 32-bit Sparc. (#81655) · c1a99b2c
      James Y Knight authored
      When in 32-bit mode, the backend doesn't currently implement 64-bit
      atomics, even though the hardware is capable if you have specified a V9
      CPU. Thus, limit the width to 32-bit, for now, leaving behind a TODO.
      
      This fixes a regression triggered by PR #73176.
      c1a99b2c
    • Xing Xue's avatar
      [OpenMP][AIX]Define struct kmp_base_tas_lock with the order of two members... · ac97562c
      Xing Xue authored
      [OpenMP][AIX]Define struct kmp_base_tas_lock with the order of two members swapped for big-endian (#79188)
      
      The direct lock data structure has bit `0` (the least significant bit)
      of the first 32-bit word set to `1` to indicate it is a direct lock. On
      the other hand, the first word (in 32-bit mode) or first two words (in
      64-bit mode) of an indirect lock are the address of the entry allocated
      from the indirect lock table. The runtime checks bit `0` of the first
      32-bit word to tell if this is a direct or an indirect lock. This works
      fine for 32-bit and 64-bit little-endian because its memory layout of a
      64-bit address is (`low word`, `high word`). However, this causes
      problems for big-endian where the memory layout of a 64-bit address is
      (`high word`, `low word`). If an address of the indirect lock table
      entry is something like `0x110035300`, i.e., (`0x1`, `0x10035300`), it
      is treated as a direct lock. This patch defines `struct
      kmp_base_tas_lock` with the ordering of the two 32-bit members flippe...
      ac97562c
    • Jeffrey Byrnes's avatar
      [SeparateConstOffsetFromGEP] Fix test after 1b65742f · ec0aa164
      Jeffrey Byrnes authored
      Change-Id: I7ced7774c80997d21969ab7886fc30c0c1e1cc81
      ec0aa164
    • Kazu Hirata's avatar
      [mlir] Fix a warning · f5cc9612
      Kazu Hirata authored
      This patch fixes:
      
        mlir/lib/Target/LLVMIR/AttrKindDetail.h:65:1: error: unused function
        'getAttrNameToKindMapping' [-Werror,-Wunused-function]
      f5cc9612
    • Kiran Chandramohan's avatar
      27726920
    • Krystian Stasiowski's avatar
      [Clang][Sema] Diagnose friend declarations with enum elaborated-type-specifier... · 3a48630a
      Krystian Stasiowski authored
      [Clang][Sema] Diagnose friend declarations with enum elaborated-type-specifier in all language modes (#80171)
      
      According to [dcl.type.elab] p4:
      > If an _elaborated-type-specifier_ appears with the `friend` specifier
      as an entire _member-declaration_, the _member-declaration_ shall have
      one of the following forms:
      >     `friend` _class-key_ _nested-name-specifier_(opt) _identifier_ `;`
      >     `friend` _class-key_ _simple-template-id_ `;`
      > `friend` _class-key_ _nested-name-specifier_ `template`(opt)
      _simple-template-id_ `;`
      
      Notably absent from this list is the `enum` form of an
      _elaborated-type-specifier_ "`enum` _nested-name-specifier_(opt)
      _identifier_", which appears to be intentional per the resolution of
      CWG2363.
      
      Most major implementations accept these declarations, so the diagnostic
      is a pedantic warning across all C++ versions.
      
      In addition to the trivial cases previously diagnosed in C++98, we now
      diagnose cases where the _elaborated-type-specifier_ has a dependent
      _nested-name-specifier_:
      ```
      template<typename T>
      struct A
      {
          enum class E;
      };
      
      struct B
      {
          template<typename T>
          friend enum A<T>::E; // pedantic warning: elaborated enumeration type cannot be a friend
      };
      
      template<typename T>
      struct C
      {
          friend enum T::E;  // pedantic warning: elaborated enumeration type cannot be a friend
      };
      ```
      3a48630a
    • Jeffrey Byrnes's avatar
      [SeparateConstOffsetFromGEP] Reorder trivial GEP chains to separate constants (#73056) · 1b65742f
      Jeffrey Byrnes authored
      In this case, a trivial GEP chain has the form:
      
      ```
      %ptr = getelementptr sameType, %base, constant
      %val = getelementptr sameType, %ptr, %variable
      ```
      
      That is, a one-index GEP consumes another (of the same basis and result
      type) one-index GEP, where the inner GEP uses a constant index and the
      outer GEP uses a variable index. For chains of this type, it is trivial
      to reorder them (by simply swapping the indexes). The result of doing so
      is better AddrMode matching for users of the ultimate ptr produced by
      GEP chain.
      
      Future patches can extend this to support non-trivial GEP chains (e.g.
      those with different basis types and/or multiple indices).
      1b65742f
    • Richard Dzenis's avatar
    • David Truby's avatar
      [mlir][flang][openmp] Rework wsloop reduction operations (#80019) · be9f8ffd
      David Truby authored
      
      
      This patch reworks the way that wsloop reduction operations function to
      better match the expected semantics from the OpenMP specification,
      following the rework of parallel reductions.
      
      The new semantics create a private reduction variable as a block
      argument which should be used normally for all operations on that
      variable in the region; this private variable is then combined with the
      others into the shared variable. This way no special omp.reduction
      operations are needed inside the region. These block arguments follow
      the loop control block arguments.
      
      ---------
      
      Co-authored-by: default avatarKiran Chandramohan <kiran.chandramohan@arm.com>
      be9f8ffd
    • jimingham's avatar
      Add the ability to define a Python based command that uses CommandObjectParsed (#70734) · a69ecb24
      jimingham authored
      This allows you to specify options and arguments and their definitions
      and then have lldb handle the completions, help, etc. in the same way
      that lldb does for its parsed commands internally.
      
      This feature has some design considerations as well as the code, so I've
      also set up an RFC, but I did this one first and will put the RFC
      address in here once I've pushed it...
      
      Note, the lldb "ParsedCommand interface" doesn't actually do all the
      work that it should. For instance, saying the type of an option that has
      a completer doesn't automatically hook up the completer, and ditto for
      argument values. We also do almost no work to verify that the arguments
      match their definition, or do auto-completion for them. This patch
      allows you to make a command that's bug-for-bug compatible with built-in
      ones, but I didn't want to stall it on getting the auto-command checking
      to work all the way correctly.
      
      As an overall...
      a69ecb24
    • jimingham's avatar
      Don't count all the frames just to skip the current inlined ones. (#80918) · a04c6366
      jimingham authored
      The algorithm to find the DW_OP_entry_value requires you to find the
      nearest non-inlined frame. It did that by counting the number of stack
      frames so that it could use that as a loop stopper.
      
      That is unnecessary and inefficient. Unnecessary because GetFrameAtIndex
      will return a null frame when you step past the oldest frame, so you
      already have the "got to the end" signal without counting all the stack
      frames.
      And counting all the stack frames can be expensive.
      a04c6366
    • Mark de Wever's avatar
      [libc++][modules] Re-add build dir CMakeLists.txt. (#81370) · fc0e9c83
      Mark de Wever authored
      This CMakeLists.txt is used to build modules without build system
      support. This was removed in d06ae33e.
      This is used in the documentation how to use modules.
      
      Made some minor changes to make it work with the std.compat module using
      the std module.
      
      Note the CMakeLists.txt in the build dir should be removed once build
      system support is generally available.
      fc0e9c83
    • Matt Arsenault's avatar
      InstCombine: Enable SimplifyDemandedUseFPClass and remove flag (#81108) · 9dd2c593
      Matt Arsenault authored
      This completes the unrevert of ef388334.
      9dd2c593
    • Danila Malyutin's avatar
      [StatepointLowering] Use Constant instead of TargetConstant for undef value (#81635) · e20462a0
      Danila Malyutin authored
      
      
      Prevents isel errors when trying to lower gc relocate of undef value
      (which turns into CopyToReg of TargetConstant). Such relocates may occur
      after DCE (e.g. after GVN removes some dead blocks) if there are not
      passes like instcombine scheduled after to clean them up.
      
      Fixes #80294
      
      ---------
      
      Co-authored-by: default avatarMatt Arsenault <arsenm2@gmail.com>
      e20462a0
    • Prabhuk's avatar
      Revert "[clang] Remove #undef alloca workaround" (#81649) · f79f58d5
      Prabhuk authored
      Reverts llvm/llvm-project#81534
      
      llvm/llvm-project#81534 breaks building (Fuchsia) Clang toolchain on
      Windows.
      
      Log:
      https://logs.chromium.org/logs/fuchsia/buildbucket/cr-buildbucket/8756186536543250705/+/u/clang/install/stdout
      Builder:
      https://ci.chromium.org/ui/p/fuchsia/builders/toolchain.ci/clang-windows-x64/b8756186536543250705/overview
      
      ```
      FAILED: tools/clang/tools/extra/clang-include-fixer/tool/CMakeFiles/clang-include-fixer.dir/ClangIncludeFixer.cpp.obj 
      C:\b\s\w\ir\x\w\cipd\bin\clang-cl.exe  /nologo -TP -DCLANG_REPOSITORY_STRING=\"https://llvm.googlesource.com/llvm-project\" -DGTEST_HAS_RTTI=0 -DUNICODE -D_CRT_NONSTDC_NO_DEPRECATE -D_CRT_NONSTDC_NO_WARNINGS -D_CRT_SECURE_NO_DEPRECATE -D_CRT_SECURE_NO_WARNINGS -D_GLIBCXX_ASSERTIONS -D_HAS_EXCEPTIONS=0 -D_SCL_SECURE_NO_DEPRECATE -D_SCL_SECURE_NO_WARNINGS -D_UNICODE -D__STDC_CONSTANT_MACROS -D__STDC_FORMAT_MACROS -D__STDC_LIMIT_MACROS -IC:\b\s\w\ir\x\w\llvm_build\tools\clang\tools\extra\clang-include-fixer\tool -IC:\b\s\w\ir\x\w\llvm-llvm-project\clang-tools-extra\clang-include-fixer\tool -IC:\b\s\w\ir\x\w\llvm-llvm-project\clang\include -IC:\b\s\w\ir\x\w\llvm_build\tools\clang\include -IC:\b\s\w\ir\x\w\recipe_cleanup\tensorflow-venv\store\python_venv-q9i5kpsp0iun0ktmqgab125ti8\contents\Lib\site-packages\tensorflow\include -IC:\b\s\w\ir\x\w\llvm_build\include -IC:\b\s\w\ir\x\w\llvm-llvm-project\llvm\include -IC:\b\s\w\ir\x\w\llvm-llvm-project\clang-tools-extra\clang-include-fixer\tool\.. -imsvcC:\b\s\w\ir\x\w\zlib_install_target\include -imsvcC:\b\s\w\ir\x\w\zstd_install\include /DWIN32 /D_WINDOWS   /Zc:inline /Zc:__cplusplus /Oi /Brepro /bigobj /permissive- /W4  -Wextra -Wno-unused-parameter -Wwrite-strings -Wcast-qual -Wmissing-field-initializers -Wimplicit-fallthrough -Wcovered-switch-default -Wno-noexcept-type -Wnon-virtual-dtor -Wdelete-non-virtual-dtor -Wsuggest-override -Wstring-conversion -Wmisleading-indentation -Wctad-maybe-unsupported /Gw -no-canonical-prefixes /O2 /Ob2  -std:c++17 -MT  /EHs-c- /GR- -UNDEBUG /showIncludes /Fotools\clang\tools\extra\clang-include-fixer\tool\CMakeFiles\clang-include-fixer.dir\ClangIncludeFixer.cpp.obj /Fdtools\clang\tools\extra\clang-include-fixer\tool\CMakeFiles\clang-include-fixer.dir\ -c -- C:\b\s\w\ir\x\w\llvm-llvm-project\clang-tools-extra\clang-include-fixer\tool\ClangIncludeFixer.cpp
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang-tools-extra\clang-include-fixer\tool\ClangIncludeFixer.cpp:11:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang-tools-extra\clang-include-fixer\tool\..\IncludeFixer.h:15:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/Sema/ExternalSemaSource.h:15:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/AST/ExternalASTSource.h:18:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/AST/DeclBase.h:18:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/AST/DeclarationName.h:18:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/Basic/IdentifierTable.h:18:
      In file included from C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/Basic/Builtins.h:63:
      C:\b\s\w\ir\x\w\llvm_build\tools\clang\include\clang/Basic/Builtins.inc(151,1): error: redefinition of enumerator 'BI_alloca'
        151 | LANGBUILTIN(_alloca, "v*z", "n", ALL_MS_LANGUAGES)
            | ^
      C:\b\s\w\ir\x\w\llvm_build\tools\clang\include\clang/Basic/Builtins.inc(15,54): note: expanded from macro 'LANGBUILTIN'
         15 | #  define LANGBUILTIN(ID, TYPE, ATTRS, BUILTIN_LANG) BUILTIN(ID, TYPE, ATTRS)
            |                                                      ^
      C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/Basic/Builtins.h(62,34): note: expanded from macro 'BUILTIN'
         62 | #define BUILTIN(ID, TYPE, ATTRS) BI##ID,
            |                                  ^
      <scratch space>(72,1): note: expanded from here
         72 | BI_alloca
            | ^
      C:\b\s\w\ir\x\w\llvm_build\tools\clang\include\clang/Basic/Builtins.inc(150,1): note: previous definition is here
        150 | LIBBUILTIN(alloca, "v*z", "fn", STDLIB_H, ALL_GNU_LANGUAGES)
            | ^
      C:\b\s\w\ir\x\w\llvm_build\tools\clang\include\clang/Basic/Builtins.inc(11,61): note: expanded from macro 'LIBBUILTIN'
         11 | #  define LIBBUILTIN(ID, TYPE, ATTRS, HEADER, BUILTIN_LANG) BUILTIN(ID, TYPE, ATTRS)
            |                                                             ^
      C:\b\s\w\ir\x\w\llvm-llvm-project\clang\include\clang/Basic/Builtins.h(62,34): note: expanded from macro 'BUILTIN'
         62 | #define BUILTIN(ID, TYPE, ATTRS) BI##ID,
            |                                  ^
      <scratch space>(71,1): note: expanded from here
         71 | BI_alloca
            | ^
      ```
      f79f58d5
    • Noah Goldstein's avatar
      [InstCombine] Extend `(lshr/shl (shl/lshr -1, x), x)` -> `(lshr/shl -1, x)` for multi-use · 79ce9331
      Noah Goldstein authored
      We previously did this iff the inner `(shl/lshr -1, x)` was
      one-use. No instructions are added even if the inner `(shl/lshr -1,
      x)` is multi-use and this canonicalization both makes the resulting
      instruction easier to analyze and shrinks its dependency chain.
      
      Closes #81576
      79ce9331
    • Mingming Liu's avatar
      [NFC][InstrProf]Factor out getCanonicalName to compute the canonical name... · 2422e969
      Mingming Liu authored
      [NFC][InstrProf]Factor out getCanonicalName to compute the canonical name given a pgo name. (#81547)
      
      - Also update the `InstrProf::addFuncWithName` to call the newly added
      `getCanonicalName`.
      2422e969
    • Joseph Huber's avatar
      [libc] Remove leftover target dependent intrinsic · c830c120
      Joseph Huber authored
      Summary:
      I forgot to remove these because I thought I did it already. This caused
      the build to fail when actually linked.
      c830c120
    • Giuseppe Rossini's avatar
      [mlir][ROCDL] Add synchronization primitives (#80888) · 16140ff2
      Giuseppe Rossini authored
      This PR adds two LLVM intrinsics to MLIR:
      - llvm.amdgcn.s.setprio which sets the priority of a wave for the GPU
      scheduler
      - llvm.amdgcn.sched.barrier which sets a software barrier so that the
      scheduler cannot move instructions around
      16140ff2
    • Joseph Huber's avatar
      [libc] Remove remaining GPU architecture dependent instructions (#81612) · 63198e06
      Joseph Huber authored
      Summary:
      Recent patches have added solutions to the remaining sources of
      divergence. This patch simply removes the last occures of things like
      `has_builtin`, `ifdef` or builtins with feature requirements. The one
      exception here is `nanosleep`, but I made changes in the
      `__nvvm_reflect` pass to make usage like this actually work at O0.
      
      Depends on https://github.com/llvm/llvm-project/pull/81331
      63198e06
    • Alex Langford's avatar
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][cuda] Lower cluster_dims values (#81636) · 5e3c7e3a
      This PR adds a new attribute to carry over the information from
      `cluster_dims`. The new attribute `CUDAClusterDimsAttr` holds 3 integer
      attributes and is added to `func.func` operation.
      5e3c7e3a
    • Craig Topper's avatar
      [RISCV] Enable the TypePromotion pass from AArch64/ARM. · 7d40ea85
      Craig Topper authored
      This pass looks for unsigned icmps that have illegal types and tries
      to widen the use/def graph to improve the placement of the zero
      extends that type legalization would need to insert.
      
      I've explicitly disabled it for i32 by adding a check for
      isSExtCheaperThanZExt to the pass.
      
      The generated code isn't perfect, but my data shows a net
      dynamic instruction count improvement on spec2017 for both base and
      Zba+Zbb+Zbs.
      7d40ea85
    • Craig Topper's avatar
      9838c851
    • Arthur Eubanks's avatar
      [clang] Remove #undef alloca workaround (#81534) · 742a06f5
      Arthur Eubanks authored
      Added in 26670dcb to workaround #4885.
      
      Windows CI and a local Windows build are happy with this change, so it
      seems like this has been properly fixed at some point. If this does
      break somebody, this can be easily reverted. (Also, Linux does the same
      `#define alloca` in system headers, so I'm not sure why it'd be
      different on Windows)
      
      This is tech debt that caused breakages, see comments on #71709.
      742a06f5
    • Craig Topper's avatar
      [IRGen][AArch64][RISCV] Generalize bitcast between i1 predicate vector and i8... · 9be7b0a5
      Craig Topper authored
      [IRGen][AArch64][RISCV] Generalize bitcast between i1 predicate vector and i8 fixed vector. (#76548)
      
      Instead of only handling vscale x 16 x i1 predicate vectors, handle any
      scalable i1 vector where the known minimum is divisible by 8.
      
      This is used on RISC-V where we have multiple sizes of predicate
      types.
      9be7b0a5
    • Jay Foad's avatar
      a7cebadc
    • Jay Foad's avatar
      e847abc5
    • Joseph Huber's avatar
      [libc] Round up time for GPU nanosleep implementation (#81630) · 1dacfd11
      Joseph Huber authored
      Summary:
      The GPU `nanosleep` tests would occasionally fail. This was due to the
      fact that we used integer division to determine how many ticks we had to
      sleep for. This would then truncate, leaving us with a value just
      slightly below the requested value. This would then occasionally leave
      us with a return value of `-1`. This patch just changes the code to
      round up by 1 so we always sleep for at least the requested value.
      1dacfd11
    • Valentin Clement (バレンタイン クレメン)'s avatar
      [flang][cuda] Lower launch_bounds values (#81537) · d79c3c50
      This PR adds a new attribute to carry over the information from
      `launch_bounds`. The new attribute `CUDALaunchBoundsAttr` holds 2 to 3
      integer attrinbutes and is added to `func.func` operation.
      d79c3c50
    • Joseph Huber's avatar
      [libc] Rework the RPC interface to accept runtime wave sizes (#80914) · f879ac03
      Joseph Huber authored
      Summary:
      The RPC interface needs to handle an entire warp or wavefront at once.
      This is currently done by using a compile time constant indicating the
      size of the buffer, which right now defaults to some value on the client
      (GPU) side. However, there are currently attempts to move the `libc`
      library to a single IR build. This is problematic as the size of the
      wave fronts changes between ISAs on AMDGPU. The builitin
      `__builtin_amdgcn_wavefrontsize()` will return the appropriate value,
      but it is only known at runtime now.
      
      In order to support this, this patch restructures the packet. Now
      instead of having an array of arrays, we simply have a large array of
      buffers and slice it according to the runtime value if we don't know it
      ahead of time. This also somewhat has the advantage of making the buffer
      contiguous within a page now that the header has been moved out of it.
      f879ac03
    • Andrzej Warzyński's avatar
      [mlir][nfc] Add tests for linalg.mmt4d (#81422) · 7a471133
      Andrzej Warzyński authored
      linalg.mmt4d was added a while back (https://reviews.llvm.org/D105244),
      but there are virtually no tests in-tree. In the spirit of documenting
      through test, this PR adds a few basic examples.
      7a471133
    • David Spickett's avatar
      [clang][docs] Fix warning in LanguageExtensions · 7a5c1a4a
      David Spickett authored
      build-llvm/tools/clang/docs/LanguageExtensions.rst:2768: WARNING: Title underline too short.
      7a5c1a4a
    • Zequan Wu's avatar
      [lldb-dap][NFC] Add Breakpoint struct to share common logic. (#80753) · d58c128b
      Zequan Wu authored
      This adds a layer between `SounceBreakpoint`/`FunctionBreakpoint` and
      `BreakpointBase` to have better separation and encapsulation so we are
      not directly operating on `SBBreakpoint`.
      
      I basically moved the `SBBreakpoint` and the methods that requires it
      from `BreakpointBase` to `Breakpoint`. This allows adding support for
      data watchpoint easier by sharing the logic inside `BreakpointBase`.
      d58c128b
    • David Spickett's avatar
      1d847922
    • S. Bharadwaj Yadavalli's avatar
      [DirectX][NFC] Change specification of overload types and attribute in DXIL.td (#81184) · 8ba4ff39
      S. Bharadwaj Yadavalli authored
      - Specify overload types of DXIL Operation as list of types instead of a
      string.
      - Add supported DXIL type record definitions to `DXIL.td` leveraging
      `LLVMType` to avoid duplicate definitions.
       - Spell out DXIL Operation Attribute specification string.
       - Make corresponding changes to process the records in DXILEmitter.cpp
      8ba4ff39
    • Jay Foad's avatar
      [TableGen] Do not speculatively grow RegUnitSets. NFC. · 1f90af18
      Jay Foad authored
      This seems to be a trick to avoid copying a RegUnitSet, but it can be
      done more simply using std::move.
      1f90af18
    • Joseph Huber's avatar
      [LLVM] Add `__builtin_readsteadycounter` intrinsic and builtin for realtime clocks (#81331) · 11fcae69
      Joseph Huber authored
      Summary:
      This patch adds a new intrinsic and builtin function mirroring the
      existing `__builtin_readcyclecounter`. The difference is that this
      implementation targets a separate counter that some targets have which
      returns a fixed frequency clock that can be used to determine elapsed
      time, this is different compared to the cycle counter which often has
      variable frequency.
      
      This patch only adds support for the NVPTX and AMDGPU targets.
      
      This is done as a new and separate builtin rather than an argument to
      `readcyclecounter` to avoid needing to change existing code and to make
      the separation more explicit.
      11fcae69