1. Nov 30, 2022
    • Matt Arsenault's avatar
    • Matt Arsenault's avatar
      AMDGPU: Fix creating illegal f16 fp_class · c08d5562
      Matt Arsenault authored
      We were missing legality checks. The device library build was broken
      for targets without f16 support. Technically the first pattern isn't
      tested by this patch; it only triggers with the isBeforeLegalize check
      in performAndCombine removed. I'm not sure how to trick this into
      appearing post-legalization.
      c08d5562
    • Matt Arsenault's avatar
    • Matt Arsenault's avatar
      AMDGPU: Bulk update some generic intrinsic tests to opaque pointers · fb1d166e
      Matt Arsenault authored
      Done purely with the script.
      fb1d166e
    • Matt Arsenault's avatar
      AMDGPU: Convert amdgpu-alias-analysis.ll to opaque pointers · a1ac8902
      Matt Arsenault authored
      This one was slightly tricky. The AA debug printing usually, but not
      always, uses the old pointer syntax. Also, we need to stop folding out
      0 index GEPs in a few of these cases.
      a1ac8902
    • Matt Arsenault's avatar
      AMDGPU: Convert some fp op tests to opaque issues · 177ff42d
      Matt Arsenault authored
      fmax_legacy.ll had one test that produced "ptraddrspace(1)", since
      somehow "i1addrspace(1)*" used to parse.
      177ff42d
    • Benjamin Kramer's avatar
    • Brett Wilson's avatar
      [clang-doc] Move file layout to the generators. · 7b8c7e02
      Brett Wilson authored
      Previously file naming and directory layout was handled on a per Info
      object basis by ClangDocMain and the generators blindly wrote to the
      files given. This means all generators must use the same file layout and
      caused problems where multiple objects mapped to the same file. The
      object collision problem happens most easily with template
      specializations because the template parameters are not part of the
      "name".
      
      This patch moves the responsibility for output file organization to the
      generators. Currently HTML and MD use the same structure as before. But
      they now collect all objects that map to a given file and combine them,
      avoiding the corruption problems.
      
      Converts the YAML generator to naming files based on USR in one
      directory. This is easier for downstream tools to manage and avoids the
      naming problems with template specializations. Since this change
      requires backward-incompatible output changes to referenced files anyway
      (since each one is now an array), this is a good time to introduce this
      change.
      
      Differential Revision: https://reviews.llvm.org/D138073
      7b8c7e02
    • varconst's avatar
      [libc++][ranges][NFC] Revamp the Ranges status page · 2c5a548b
      varconst authored
      Focus on the not-yet-implemented features: remove most details about the
      already-implemented C++20 stuff, list out the major C++23 additions.
      
      Differential Revision: https://reviews.llvm.org/D136657
      2c5a548b
    • Alex Lorenz's avatar
    • Krzysztof Parzyszek's avatar
      [Hexagon] Further improve code generation for shuffles · 073d5e59
      Krzysztof Parzyszek authored
      * Concatenate partial shuffles into longer ones whenever possible:
      In selection DAG, shuffle's operands and return type must all agree. This
      is not the case in LLVM IR, and non-conforming IR-level shuffles will be
      rewritten to match DAG's requirements. This can also make a shuffle that
      can be matched to a single HVX instruction become shuffles that require
      more complex handling. Example: anything that takes two single vectors
      and returns a pair (e.g. V6_vshuffvdd).
      This is avoided by concatenating such shuffles into ones that take a vector
      pair, and an undef pair, and produce a vector pair.
      
      * Recognize perfect shuffles when masks contain `undef` values.
      
      * Use funnel shifts for contracting shuffles.
      
      * Recognize rotations as a separate step.
      
      These changes go into a single commit, because each one on their own
      introduced some regressions.
      073d5e59
    • Wael Yehia's avatar
      [AIX][LTO] Properly respect LDR_CNTRL and set MAXDATA32 to 0xA0000000@DSA. · d4cb3928
      Wael Yehia authored
      Reviewed By: rzurob
      
      Differential Revision: https://reviews.llvm.org/D138944
      d4cb3928
    • Nicolai Hähnle's avatar
      update_test_checks: fix typos · cba252de
      Nicolai Hähnle authored
      Found by our downstream CI.
      cba252de
    • Konstantin Varlamov's avatar
      [libc++] Add a missing include to `swap_allocator.h`. · e67903eb
      Konstantin Varlamov authored
      Also add tests for the file.
      
      Reviewed By: #libc, ldionne
      
      Differential Revision: https://reviews.llvm.org/D135635
      e67903eb
    • William Huang's avatar
      [InstCombine] Revert D125845 · be4b1dd3
      William Huang authored
      Reverting D125845 `[InstCombine] Canonicalize GEP of GEP by swapping constant-indexed GEP to the back` because multiple users reported performance regression
      
      Reviewed By: davidxl
      
      Differential Revision: https://reviews.llvm.org/D138950
      be4b1dd3
    • Martin Storsjö's avatar
      [clang] [test] Fix recently pushed mingw tests in some environments · 3a37c112
      Martin Storsjö authored
      Account for backslashes in paths in mingw.cpp.
      
      Testing clang with the <triple>-clang form seems to require the
      x86 target to be enabled, when the triple is an x86 triple. Just
      skip that aspect of the test, since the "clang --target=<triple>"
      form should give enough test coverage here.
      3a37c112
    • Benjamin Kramer's avatar
      AMDGPU: Remove unused variables. NFC · 3a86931f
      Benjamin Kramer authored
      3a86931f
    • Maryam Moghadas's avatar
      [PowerPC] Fix vperm codegen · 7614ba0a
      Maryam Moghadas authored
      Commit rG934d5fa2 changed the vperm codegen
      for cases that vperm is not replaced by xxperm, this patch is to revert that.
      
      Reviewed By: stefanp
      
      Differential Revision: https://reviews.llvm.org/D138736
      7614ba0a
    • Ron Lieberman's avatar
      Revert "enable code-object-version=5" · ca856fff
      Ron Lieberman authored
      very sorry wrong repo.
      
      This reverts commit d882ba7a.
      ca856fff
    • Ron Lieberman's avatar
      Revert "Add mean_anyway to hpc config" · b09a5e5c
      Ron Lieberman authored
      my bad, wrong repo ,so sorry.
      
      This reverts commit 0b9350f3.
      b09a5e5c
    • Alex Lorenz's avatar
      [clang][driver][darwin] Enforce consistent major version limit for any Darwin OS · 2a670144
      Alex Lorenz authored
      Limit can also be bumped up to 999 to allow OS versions over 100
      2a670144
    • Martin Storsjö's avatar
    • Martin Storsjö's avatar
      Reapply [openmp] [test] XFAIL many-microtask-args.c on ARM · 2bd2734f
      Martin Storsjö authored
      On ARM, a C fallback version of __kmp_invoke_microtask is used,
      which only handles up to a fixed number of arguments - while
      many-microtask-args.c tests that the function can handle an
      arbitrarily large number of arguments (the testcase produces 17
      arguments).
      
      On the CMake level, we can't add ${LIBOMP_ARCH} directly to
      OPENMP_TEST_COMPILER_FEATURES in OpenMPTesting.cmake, since
      that file is parsed before LIBOMP_ARCH is set. Instead
      convert the feature list into a proper CMake list, and append
      ${LIBOMP_ARCH} into it before serializing it to an Python array.
      
      Reapply: Make sure OPENMP_TEST_COMPILER_FEATURES is defined
      properly in all other test subdirectories other than
      runtime/test too.
      
      Differential Revision: https://reviews.llvm.org/D138738
      2bd2734f
    • Martin Storsjö's avatar
      [clang] [MinGW] Improve detection of libstdc++ headers on Fedora · 98454e38
      Martin Storsjö authored
      There's some variation in where different toolchain distributions
      (and linux distributions) package the mingw sysroots - this is
      so far handled by adding specific known subdirectory paths
      to the include and lib directory lists.
      
      There are multiple degrees of combinatorics involved here though;
      the distros may use different locations such as
      /usr/x86_64-w64-mingw32/include or
      /usr/x86_64-w64-mingw32/sys-root/mingw/include.
      
      So far, this setup has been treated as base=/usr, subdir=x86_64-w64-mingw32,
      and the driver tries to add further subdirectories such as
      <base>/<subdir>/include, <base>/<subdir>/sys-root/mingw/include.
      
      When it comes to libstdc++ (and libc++), each of these come with
      a large number of potential subdirectories. Instead of further
      exploding the combinatorics another step by adding all combinations
      of all paths, check whether <base>/<subdir>/sys-root/mingw/include
      exists, and if it does, append that subpath into the subdir variable.
      
      This allows finding libstdc++ headers in e.g.
      /usr/x86_64-w64-mingw32/sys-root/mingw/include/c++/x86_64-w64-mingw32
      on Fedora.
      
      The same logic (where everything belonging to this target fits
      under one expanded <subdir> path, with just /include and /lib
      under it) doesn't seem to apply on Gentoo, where the includes
      are found in <base>/<subdir>/usr/include while the libraries
      are in <base>/<subdir>/mingw/lib (see
      8e218026). But apparently
      the libstdc++ headers aren't installed under
      <base>/<subdir>/usr/include, so that path hierarchy quirk doesn't
      need to be taken into account in AddClangCXXStdlibIncludeArgs.
      
      Differential Revision: https://reviews.llvm.org/D138693
      98454e38
    • Martin Storsjö's avatar
      [clang] [MinGW] Improve/extend the gcc/sysroot detection logic · 02b25bd9
      Martin Storsjö authored
      There are three functions that try to detect the right implicit
      sysroot and libgcc directory setup to use
      - One which looks for mingw sysroots located in
        <clangbin>/../<sysrootname>
      - One which looks for a mingw-targeting gcc executables in the PATH
      - One which looks in the <gccroot>/lib/gcc directory to find the
        right one to use, and the right specific triple used for arch
        specific directories in the gcc/libstdc++ install
      
      These have mostly tried to look for executables named
      "<arch>-w64-mingw32-gcc" or "mingw32-gcc" or subdirectories
      named "<arch>-w64-mingw32" or "mingw32".
      
      In the case of findClangRelativeSysroot, it also has looked
      for directories with the name of the actual triple. This
      was added in deff7536,
      with the intent of looking for a directory matching exactly
      the user provided literal triple - however the triple here
      is the normalized one, not the one provided by the user on
      the command line.
      
      Improve and unify this logic somewhat:
      - Always first look for things based on the literal triple
        provided by the user.
      - Secondly look for things based on the normalized triple
        (which usually ends up as e.g. x86_64-w64-windows-gnu),
        accessed via the Triple which is passed to the constructor
      - Then look for the common triple form <arch>-w64-mingw32
      
      The literal triple provided by the user is available via
      Driver::getTargetTriple(), but computeTargetTriple() may
      change e.g. the architecture of it, so we need to
      reapply the effective architecture on the literal triple
      spelling from Driver::getTargetTriple().
      
      Do this consistently for all of findGcc, findClangRelativeSysroot
      and findGccLibDir (while keeping the existing plain "mingw32"
      cases in findGcc and findGccLibDir too).
      
      Fedora 37 started shipping mingw sysroots targeting UCRT,
      in addition to the traditional msvcrt.dll, and these use
      triples in the form <arch>-w64-mingw32ucrt - see
      https://fedoraproject.org/wiki/Changes/F37MingwUCRT.
      
      Thus, in addition to the existing default tested triples,
      try looking for triples in the form <arch>-w64-mingw32ucrt,
      to automatically find the UCRT sysroots on Fedora 37.
      By explicitly setting a specific target on the Clang command
      line, the user can be more explicit with which flavour is
      to be preferred.
      
      This should fix the main issue in
      https://github.com/llvm/llvm-project/issues/59001.
      
      Differential Revision: https://reviews.llvm.org/D138692
      02b25bd9
    • Nicolai Hähnle's avatar
      AMDGPU: Remove BufferPseudoSourceValue · 43b86bf9
      Nicolai Hähnle authored
      The use of a PSV for buffer intrinsics is misleading because it may be
      misinterpreted as all buffer intrinsics accessing the same address in
      memory, which is clearly not true.
      
      Instead, build MachineMemOperands without a pointer value but with an
      address space, so that address space-based alias analysis can still
      work.
      
      There is a lot of test churn because previously address space 4
      (constant address space) was used as an address space for buffer
      intrinsics. This doesn't make much sense and seems to have been an
      accident -- see the change in
      AMDGPUTargetMachine::getAddressSpaceForPseudoSourceKind.
      
      Differential Revision: https://reviews.llvm.org/D138711
      43b86bf9
    • Ron Lieberman's avatar
      Add mean_anyway to hpc config · 0b9350f3
      Ron Lieberman authored
      0b9350f3
    • Ron Lieberman's avatar
      enable code-object-version=5 · d882ba7a
      Ron Lieberman authored
      d882ba7a
    • Peter Rong's avatar
      [FuzzMutate] Fix a bug in `connectToSink` which might invalidate the whole module. · 50921a21
      Peter Rong authored
      `connectToSink` uses a value by putting it in a future instruction.
      It will replace the operand of a future instruction with the current value.
      
      However, if current value is an `Instruction` and put into a switch case, the module is invalid.
      We fix that by only connecting to Br/Switch's condition, and don't touch other operands.
      
      Will have other strategies to mutate other Br/Switch operands to be patched once this patch is passed
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D138890
      50921a21
    • Joseph Huber's avatar
      [libc] Fix test not including 'free' · 21d9f725
      Joseph Huber authored
      Summary:
      A previous change removed a transient inclusion of `stdlib.h` from the
      `string_utils.h` file which this test depended on. Include it directly
      here.
      21d9f725
    • Thurston Dang's avatar
      [msan] Increase size of app/shadow/origin mappings on aarch64 · b726df1b
      Thurston Dang authored
      msan's app memory mappings for aarch64 are constrained by
      the MEM_TO_SHADOW constant to 64GB or less, and some app
      memory mappings (in kMemoryLayout) are even smaller in
      practice. This will lead to a crash with the error message
      "MemorySanitizer can not mmap the shadow memory" if the
      executable's memory mappings (e.g., libraries) extend
      beyond msan's app memory mappings.
      
      This patch makes the app/shadow/origin memory mappings
      considerably larger, along with corresponding changes to
      the MEM_TO_SHADOW and SHADOW_TO_ORIGIN constants.
      
      Note that this deprecates compatibility with 39- and 42-bit
      VMAs.
      
      Differential Revision: https://reviews.llvm.org/D137666
      b726df1b
    • Joseph Huber's avatar
      [libc][docs] Add documentation for the new GPU mode · 194788b2
      Joseph Huber authored
      This patch introduces documentation for the new GPU mode added in
      D138608. The documentation includes instructions for building and using
      the library, along with a description of the supported functions and
      headers.
      
      Reviewed By: sivachandra, lntue, michaelrj
      
      Differential Revision: https://reviews.llvm.org/D138856
      194788b2
    • Joseph Huber's avatar
      [libc] Add initial support for a libc implementation for the GPU · 55151e13
      Joseph Huber authored
      This patch contains the initial support for building LLVM's libc as a
      target for the GPU. Currently this only supports a handful of very basic
      functions that can be implemented without an operating system. The GPU
      code is build using the existing OpenMP toolchain. This allows us to
      minimally change the existing codebase and get a functioning static
      library. This patch allows users to create a static library called
      `libcgpu.a` that contains fat binaries containing device IR.
      
      Current limitations are the lack of test support and the fact that only
      one target OS can be built at a time. That is, the user cannot get a
      `libc` for Linux and one for the GPU simultaneously.
      
      This introduces two new CMake variables to control the behavior
      `LLVM_LIBC_TARET_OS` is exported so the user can now specify it to equal
      `"gpu"`. `LLVM_LIBC_GPU_ARCHITECTURES` is also used to configure how
      many targets to build for at once.
      
      Depends on D138607
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D138608
      55151e13
    • Joseph Huber's avatar
      [libc] Move strdup implementation to a new header · d85699eb
      Joseph Huber authored
      The `strdup` family of functions rely on `malloc` to be implemented.
      Its presence in the `string_utils.h` header meant that compiling many of
      the string functions relied on `malloc` being implementated as well.
      This patch simply moves the implementation into a new file to avoid
      including `stdlib.h` from the other string functions. This was a barrier
      for compiling string functions for the GPU where there is no malloc
      currently.
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D138607
      d85699eb
    • LLVM GN Syncbot's avatar
      [gn build] Port 185d4964 · 1114d1a9
      LLVM GN Syncbot authored
      1114d1a9
    • Dave Lee's avatar
      [lldb] Introduce dwim-print command · 185d4964
      Dave Lee authored
      Implements `dwim-print`, a printing command that chooses the most direct,
      efficient, and resilient means of printing a given expression.
      
      DWIM is an acronym for Do What I Mean. From Wikipedia, DWIM is described as:
      
        > attempt to anticipate what users intend to do, correcting trivial errors
        > automatically rather than blindly executing users' explicit but
        > potentially incorrect input
      
      The `dwim-print` command serves as a single print command for users who don't
      yet know, or prefer not to know, the various lldb commands that can be used to
      print, and when to use them.
      
      This initial implementation is the base foundation for `dwim-print`. It accepts
      no flags, only an expression. If the expression is the name of a variable in
      the frame, then effectively `frame variable` is used to get, and print, its
      value. Otherwise, printing falls back to using `expression` evaluation. In this
      initial version, frame variable paths will be handled with `expression`.
      
      Following this, there are a number of improvements that can be made. Some
      improvements include supporting `frame variable` expressions or registers.
      
      To provide transparency, especially as the `dwim-print` command evolves, a new
      setting is also introduced: `dwim-print-verbosity`. This setting instructs
      `dwim-print` to optionally print a message showing the effective command being
      run. For example `dwim-print var.meth()` can print a message such as: "note:
      ran `expression var.meth()`".
      
      See https://discourse.llvm.org/t/dwim-print-command/66078 for the proposal and
      discussion.
      
      Differential Revision: https://reviews.llvm.org/D138315
      185d4964
    • bixia1's avatar
      101a0c84
    • Paul Robinson's avatar
      [Windows] Convert llvm/test/ExecutionEngine/MCJIT/remote tests to check 'target=<triple>' · 284b77f1
      Paul Robinson authored
      Part of the project to eliminate special handling for triples in lit
      expressions.
      284b77f1
    • Krzysztof Parzyszek's avatar
      [Hexagon] Simplify logic for generating code for contracting shuffles · 359174c3
      Krzysztof Parzyszek authored
      Add functions that generate masks for the HVX instructions we were
      targeting. This is both simpler than analyzing the masks, and these
      functions may also be used in other places.
      359174c3
    • Paul Robinson's avatar
      [Windows] Convert clang/test/Modules tests to check 'target=<triple>' · f2c0c729
      Paul Robinson authored
      Part of the project to eliminate special handling for triples in lit
      expressions.
      f2c0c729