- Nov 30, 2022
-
-
Matt Arsenault authored
-
Matt Arsenault authored
We were missing legality checks. The device library build was broken for targets without f16 support. Technically the first pattern isn't tested by this patch; it only triggers with the isBeforeLegalize check in performAndCombine removed. I'm not sure how to trick this into appearing post-legalization.
-
Matt Arsenault authored
-
Matt Arsenault authored
Done purely with the script.
-
Matt Arsenault authored
This one was slightly tricky. The AA debug printing usually, but not always, uses the old pointer syntax. Also, we need to stop folding out 0 index GEPs in a few of these cases.
-
Matt Arsenault authored
fmax_legacy.ll had one test that produced "ptraddrspace(1)", since somehow "i1addrspace(1)*" used to parse.
-
Benjamin Kramer authored
-
Brett Wilson authored
Previously file naming and directory layout was handled on a per Info object basis by ClangDocMain and the generators blindly wrote to the files given. This means all generators must use the same file layout and caused problems where multiple objects mapped to the same file. The object collision problem happens most easily with template specializations because the template parameters are not part of the "name". This patch moves the responsibility for output file organization to the generators. Currently HTML and MD use the same structure as before. But they now collect all objects that map to a given file and combine them, avoiding the corruption problems. Converts the YAML generator to naming files based on USR in one directory. This is easier for downstream tools to manage and avoids the naming problems with template specializations. Since this change requires backward-incompatible output changes to referenced files anyway (since each one is now an array), this is a good time to introduce this change. Differential Revision: https://reviews.llvm.org/D138073
-
varconst authored
Focus on the not-yet-implemented features: remove most details about the already-implemented C++20 stuff, list out the major C++23 additions. Differential Revision: https://reviews.llvm.org/D136657
-
Alex Lorenz authored
-
Krzysztof Parzyszek authored
* Concatenate partial shuffles into longer ones whenever possible: In selection DAG, shuffle's operands and return type must all agree. This is not the case in LLVM IR, and non-conforming IR-level shuffles will be rewritten to match DAG's requirements. This can also make a shuffle that can be matched to a single HVX instruction become shuffles that require more complex handling. Example: anything that takes two single vectors and returns a pair (e.g. V6_vshuffvdd). This is avoided by concatenating such shuffles into ones that take a vector pair, and an undef pair, and produce a vector pair. * Recognize perfect shuffles when masks contain `undef` values. * Use funnel shifts for contracting shuffles. * Recognize rotations as a separate step. These changes go into a single commit, because each one on their own introduced some regressions.
-
Wael Yehia authored
Reviewed By: rzurob Differential Revision: https://reviews.llvm.org/D138944
-
Nicolai Hähnle authored
Found by our downstream CI.
-
Konstantin Varlamov authored
Also add tests for the file. Reviewed By: #libc, ldionne Differential Revision: https://reviews.llvm.org/D135635
-
William Huang authored
Reverting D125845 `[InstCombine] Canonicalize GEP of GEP by swapping constant-indexed GEP to the back` because multiple users reported performance regression Reviewed By: davidxl Differential Revision: https://reviews.llvm.org/D138950
-
Martin Storsjö authored
Account for backslashes in paths in mingw.cpp. Testing clang with the <triple>-clang form seems to require the x86 target to be enabled, when the triple is an x86 triple. Just skip that aspect of the test, since the "clang --target=<triple>" form should give enough test coverage here.
-
Benjamin Kramer authored
-
Maryam Moghadas authored
Commit rG934d5fa2 changed the vperm codegen for cases that vperm is not replaced by xxperm, this patch is to revert that. Reviewed By: stefanp Differential Revision: https://reviews.llvm.org/D138736
-
Ron Lieberman authored
very sorry wrong repo. This reverts commit d882ba7a.
-
Ron Lieberman authored
my bad, wrong repo ,so sorry. This reverts commit 0b9350f3.
-
Alex Lorenz authored
Limit can also be bumped up to 999 to allow OS versions over 100
-
Martin Storsjö authored
Differential Revision: https://reviews.llvm.org/D138818
-
Martin Storsjö authored
On ARM, a C fallback version of __kmp_invoke_microtask is used, which only handles up to a fixed number of arguments - while many-microtask-args.c tests that the function can handle an arbitrarily large number of arguments (the testcase produces 17 arguments). On the CMake level, we can't add ${LIBOMP_ARCH} directly to OPENMP_TEST_COMPILER_FEATURES in OpenMPTesting.cmake, since that file is parsed before LIBOMP_ARCH is set. Instead convert the feature list into a proper CMake list, and append ${LIBOMP_ARCH} into it before serializing it to an Python array. Reapply: Make sure OPENMP_TEST_COMPILER_FEATURES is defined properly in all other test subdirectories other than runtime/test too. Differential Revision: https://reviews.llvm.org/D138738 -
Martin Storsjö authored
There's some variation in where different toolchain distributions (and linux distributions) package the mingw sysroots - this is so far handled by adding specific known subdirectory paths to the include and lib directory lists. There are multiple degrees of combinatorics involved here though; the distros may use different locations such as /usr/x86_64-w64-mingw32/include or /usr/x86_64-w64-mingw32/sys-root/mingw/include. So far, this setup has been treated as base=/usr, subdir=x86_64-w64-mingw32, and the driver tries to add further subdirectories such as <base>/<subdir>/include, <base>/<subdir>/sys-root/mingw/include. When it comes to libstdc++ (and libc++), each of these come with a large number of potential subdirectories. Instead of further exploding the combinatorics another step by adding all combinations of all paths, check whether <base>/<subdir>/sys-root/mingw/include exists, and if it does, append that subpath into the subdir variable. This allows finding libstdc++ headers in e.g. /usr/x86_64-w64-mingw32/sys-root/mingw/include/c++/x86_64-w64-mingw32 on Fedora. The same logic (where everything belonging to this target fits under one expanded <subdir> path, with just /include and /lib under it) doesn't seem to apply on Gentoo, where the includes are found in <base>/<subdir>/usr/include while the libraries are in <base>/<subdir>/mingw/lib (see 8e218026). But apparently the libstdc++ headers aren't installed under <base>/<subdir>/usr/include, so that path hierarchy quirk doesn't need to be taken into account in AddClangCXXStdlibIncludeArgs. Differential Revision: https://reviews.llvm.org/D138693
-
Martin Storsjö authored
There are three functions that try to detect the right implicit sysroot and libgcc directory setup to use - One which looks for mingw sysroots located in <clangbin>/../<sysrootname> - One which looks for a mingw-targeting gcc executables in the PATH - One which looks in the <gccroot>/lib/gcc directory to find the right one to use, and the right specific triple used for arch specific directories in the gcc/libstdc++ install These have mostly tried to look for executables named "<arch>-w64-mingw32-gcc" or "mingw32-gcc" or subdirectories named "<arch>-w64-mingw32" or "mingw32". In the case of findClangRelativeSysroot, it also has looked for directories with the name of the actual triple. This was added in deff7536, with the intent of looking for a directory matching exactly the user provided literal triple - however the triple here is the normalized one, not the one provided by the user on the command line. Improve and unify this logic somewhat: - Always first look for things based on the literal triple provided by the user. - Secondly look for things based on the normalized triple (which usually ends up as e.g. x86_64-w64-windows-gnu), accessed via the Triple which is passed to the constructor - Then look for the common triple form <arch>-w64-mingw32 The literal triple provided by the user is available via Driver::getTargetTriple(), but computeTargetTriple() may change e.g. the architecture of it, so we need to reapply the effective architecture on the literal triple spelling from Driver::getTargetTriple(). Do this consistently for all of findGcc, findClangRelativeSysroot and findGccLibDir (while keeping the existing plain "mingw32" cases in findGcc and findGccLibDir too). Fedora 37 started shipping mingw sysroots targeting UCRT, in addition to the traditional msvcrt.dll, and these use triples in the form <arch>-w64-mingw32ucrt - see https://fedoraproject.org/wiki/Changes/F37MingwUCRT. Thus, in addition to the existing default tested triples, try looking for triples in the form <arch>-w64-mingw32ucrt, to automatically find the UCRT sysroots on Fedora 37. By explicitly setting a specific target on the Clang command line, the user can be more explicit with which flavour is to be preferred. This should fix the main issue in https://github.com/llvm/llvm-project/issues/59001. Differential Revision: https://reviews.llvm.org/D138692
-
Nicolai Hähnle authored
The use of a PSV for buffer intrinsics is misleading because it may be misinterpreted as all buffer intrinsics accessing the same address in memory, which is clearly not true. Instead, build MachineMemOperands without a pointer value but with an address space, so that address space-based alias analysis can still work. There is a lot of test churn because previously address space 4 (constant address space) was used as an address space for buffer intrinsics. This doesn't make much sense and seems to have been an accident -- see the change in AMDGPUTargetMachine::getAddressSpaceForPseudoSourceKind. Differential Revision: https://reviews.llvm.org/D138711
-
Ron Lieberman authored
-
Ron Lieberman authored
-
Peter Rong authored
`connectToSink` uses a value by putting it in a future instruction. It will replace the operand of a future instruction with the current value. However, if current value is an `Instruction` and put into a switch case, the module is invalid. We fix that by only connecting to Br/Switch's condition, and don't touch other operands. Will have other strategies to mutate other Br/Switch operands to be patched once this patch is passed Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D138890
-
Joseph Huber authored
Summary: A previous change removed a transient inclusion of `stdlib.h` from the `string_utils.h` file which this test depended on. Include it directly here.
-
Thurston Dang authored
msan's app memory mappings for aarch64 are constrained by the MEM_TO_SHADOW constant to 64GB or less, and some app memory mappings (in kMemoryLayout) are even smaller in practice. This will lead to a crash with the error message "MemorySanitizer can not mmap the shadow memory" if the executable's memory mappings (e.g., libraries) extend beyond msan's app memory mappings. This patch makes the app/shadow/origin memory mappings considerably larger, along with corresponding changes to the MEM_TO_SHADOW and SHADOW_TO_ORIGIN constants. Note that this deprecates compatibility with 39- and 42-bit VMAs. Differential Revision: https://reviews.llvm.org/D137666
-
Joseph Huber authored
This patch introduces documentation for the new GPU mode added in D138608. The documentation includes instructions for building and using the library, along with a description of the supported functions and headers. Reviewed By: sivachandra, lntue, michaelrj Differential Revision: https://reviews.llvm.org/D138856
-
Joseph Huber authored
This patch contains the initial support for building LLVM's libc as a target for the GPU. Currently this only supports a handful of very basic functions that can be implemented without an operating system. The GPU code is build using the existing OpenMP toolchain. This allows us to minimally change the existing codebase and get a functioning static library. This patch allows users to create a static library called `libcgpu.a` that contains fat binaries containing device IR. Current limitations are the lack of test support and the fact that only one target OS can be built at a time. That is, the user cannot get a `libc` for Linux and one for the GPU simultaneously. This introduces two new CMake variables to control the behavior `LLVM_LIBC_TARET_OS` is exported so the user can now specify it to equal `"gpu"`. `LLVM_LIBC_GPU_ARCHITECTURES` is also used to configure how many targets to build for at once. Depends on D138607 Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D138608
-
Joseph Huber authored
The `strdup` family of functions rely on `malloc` to be implemented. Its presence in the `string_utils.h` header meant that compiling many of the string functions relied on `malloc` being implementated as well. This patch simply moves the implementation into a new file to avoid including `stdlib.h` from the other string functions. This was a barrier for compiling string functions for the GPU where there is no malloc currently. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D138607
-
LLVM GN Syncbot authored
-
Dave Lee authored
Implements `dwim-print`, a printing command that chooses the most direct, efficient, and resilient means of printing a given expression. DWIM is an acronym for Do What I Mean. From Wikipedia, DWIM is described as: > attempt to anticipate what users intend to do, correcting trivial errors > automatically rather than blindly executing users' explicit but > potentially incorrect input The `dwim-print` command serves as a single print command for users who don't yet know, or prefer not to know, the various lldb commands that can be used to print, and when to use them. This initial implementation is the base foundation for `dwim-print`. It accepts no flags, only an expression. If the expression is the name of a variable in the frame, then effectively `frame variable` is used to get, and print, its value. Otherwise, printing falls back to using `expression` evaluation. In this initial version, frame variable paths will be handled with `expression`. Following this, there are a number of improvements that can be made. Some improvements include supporting `frame variable` expressions or registers. To provide transparency, especially as the `dwim-print` command evolves, a new setting is also introduced: `dwim-print-verbosity`. This setting instructs `dwim-print` to optionally print a message showing the effective command being run. For example `dwim-print var.meth()` can print a message such as: "note: ran `expression var.meth()`". See https://discourse.llvm.org/t/dwim-print-command/66078 for the proposal and discussion. Differential Revision: https://reviews.llvm.org/D138315
-
bixia1 authored
Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D138823
-
Paul Robinson authored
Part of the project to eliminate special handling for triples in lit expressions.
-
Krzysztof Parzyszek authored
Add functions that generate masks for the HVX instructions we were targeting. This is both simpler than analyzing the masks, and these functions may also be used in other places.
-
Paul Robinson authored
Part of the project to eliminate special handling for triples in lit expressions.
-