- Jul 12, 2022
-
-
Xiang1 Zhang authored
-
David Blaikie authored
-
Joseph Huber authored
Summary: We use the `--host-triple=` argument to manually set the target triple. This was changed to include the `=` previously but was not included in these additional test cases, causing it for fail on some unsupported systems.
-
David Blaikie authored
-
Rafael Auler authored
Turn off execution of tests that use UNIX-specific features. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126933
-
Argyrios Kyrtzidis authored
[DependencyScanningTool.cpp] Use `using namespace` instead of wrapping the `.cpp` file contents in namespaces, NFC This makes the file consistent with the coding style of the rest of LLVM.
-
Rafael Auler authored
Add -experimental-shrink-wrapping flag to control when we want to move callee-saved registers even when addresses of the stack frame are captured and used in pointer arithmetic, making it more challenging to do alias analysis to prove that we do not access optimized stack positions. This alias analysis is not yet implemented, hence, it is experimental. In practice, though, no compiler would emit code to do pointer arithmetic to access a saved callee-saved register unless there is a memory bug or we are failing to identify a callee-saved reg, so I'm not sure how useful it would be to formally prove that. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126115
-
Rafael Auler authored
Change shrink-wrapping to try a priority list of save positions, instead of trying the best one and giving up if it doesn't work. This also increases coverage. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126114
-
Rafael Auler authored
Add the option to run -equalize-bb-counts before shrink wrapping to avoid unnecessarily optimizing some CFGs where profile is inaccurate but we can prove two blocks have the same frequency. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126113
-
Rafael Auler authored
Refactor isStackAccess() to reflect updates by D126116. Now we only handle simple stack accesses and delegate the rest of the cases to getMemDataSize. Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126112
-
Rafael Auler authored
Change how function score is calculated and provide more detailed statistics when reporting back frame optimizer and shrink wrapping results. In this new statistics, we provide dynamic coverage numbers. The main metric for shrink wrapping is the number of executed stores that were saved because of shrink wrapping (push instructions that were either entirely moved away from the hot block or converted to a stack adjustment instruction). There is still a number of reduced load instructions (pop) that we are not counting at the moment. Also update alloc combiner to report dynamic numbers, as well as frame optimizer. For debugging purposes, we also include a list of top 10 functions optimized by shrink wrapping. These changes are aimed at better understanding the impact of shrink wrapping in a given binary. We also remove an assertion in dataflow analysis to do not choke on empty functions (which makes no sense). Reviewed By: Amir Differential Revision: https://reviews.llvm.org/D126111
-
Nico Weber authored
-
Michael Jones authored
Move the constants for printf's return values into core_structs, and update the converters to match. Reviewed By: lntue Differential Revision: https://reviews.llvm.org/D128767
-
Raphael Isemann authored
LLDB supports having globbing regexes in the process launch arguments that will be resolved using the user's shell. This requires that we pass the launch args to the shell and then read back the expanded arguments using LLDB's argdumper utility. As the shell will not just expand the globbing regexes but all special characters, we need to escape all non-globbing charcters such as $, &, <, >, etc. as those otherwise are interpreted and removed in the step where we expand the globbing characters. Also because the special characters are shell-specific, LLDB needs to maintain a list of all the characters that need to be escaped for each specific shell. This patch adds the list of special characters that need to be escaped for fish. Without this patch on systems where fish is the user's shell having any of these special characters in your arguments or path to the binary will cause the process launch to fail. E.g., `lldb -- ./calc 1<2` is failing without this patch. The same happens if the absolute path to calc is in a directory that contains for example parentheses or other special characters. Differential revision: https://reviews.llvm.org/D104635
-
Jonas Devlieghere authored
Add a test that ensures we always prioritize exact triple matches when creating platforms. This is a regression test for a (now resolved) bug that that resulted in the remote tvOS platform being selected for a tvOS simulator binary because the ArchSpecs are compatible.
-
Florian Hahn authored
Add test that hits the limit introduced in 4796b4ae.
-
Florian Hahn authored
-
Alex Brachet authored
It seems like the `sed` on Windows is not particularly smart. It's not actually needed in this place, so I've removed it's usage and just created an invalid yaml another way.
-
Craig Topper authored
Only one caller didn't already have an MVT and that was easy to fix. Since the return type is MVT and it uses MVT::getVectorVT, taking an MVT as input makes the most sense.
-
Christopher Bate authored
Differential Revision: https://reviews.llvm.org/D129333
-
Jonas Devlieghere authored
Make sure we use the libc++ from the build dir. Currently, by passing -stdlib=libc++, we might pick up the system libc++. This change ensures that if LLVM_LIBS_DIR is set, we try to use the libc++ from there. Differential revision: https://reviews.llvm.org/D129166
-
Aart Bik authored
A previous revision implemented expand/collapse reshaping between dense and sparse tensors for sparse2dense and dense2sparse since those could use the "cheap" view reshape on the already materialized dense tensor (at either the input or output side), and do some reshuffling from or to sparse. The dense2dense case, as always, is handled with a "cheap" view change. This revision implements the sparse2sparse cases. Lacking any "view" support on sparse tensors this operation necessarily has to perform data reshuffling on both ends. Tracker for improving this: https://github.com/llvm/llvm-project/issues/56477 Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D129416
-
Petr Hosek authored
This matches the standard behavior on other platforms. Differential Revision: https://reviews.llvm.org/D129512
-
Alex Brachet authored
Error message is not capitalized on Windows
-
Alex Brachet authored
This patch adds a new flag vfsoverlay similar to clang’s ivfsoverlay flag. This is helpful when compiling on case sensitive file systems when cross compiling to Windows. Particularly when compiling third party code containing \#pragma comment(“linker”, “/defaultlib:...”) which can’t be easily changed. Differential Revision: https://reviews.llvm.org/D125800
-
Alex Brachet authored
Differential Revision: https://reviews.llvm.org/D129517
-
Jonas Devlieghere authored
This reverts commit ef0fa9f0 as a follow up to b19d3ee7 which reverted commit ac507102. See https://reviews.llvm.org/D126189 for more details.
-
Craig Topper authored
As far as I can tell what was happening in the original code is that the getNode call receives the same operands as the original node with different SDNodeFlags. The logic inside getNode detects that the node already exists and intersects the flags into the existing node and returns it. This results in Op and NewOp for the TLO.CombineTo call always being the same node. We may have already called CombineTo as part of the recursive handling. A second call to CombineTo as we unwind the recursion overwrites the previous CombineTo. I think this means any time we updated the poison flags that was the only change that ends up getting made and we relied on DAGCombiner to revisit and call SimplifyDemandedBits again. The second time the poison flags wouldn't need to be dropped and we would keep the CombineTo call from further down the recursion. We can instead call setFlags to drop the poison flags and remove the call to TLO.CombineTo. This way we keep the CombineTo from deeper in the recursion which should be more efficient. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D129511
-
Martin Storsjö authored
This fixes https://github.com/llvm/llvm-project/issues/56095. Differential Revision: https://reviews.llvm.org/D129455
-
Piotr Sobczak authored
Fix a regression introduced in D128865. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D129375
-
Aiden Grossman authored
Currently the autogenerated regalloc model will sometimes output an incorrect LR index to evict instead of the first LR with with the mask set to 1. This trips an assertion within the MLRegallocAdvisor that the evicted LR has a mask of 1. This patch, made possible by https://reviews.llvm.org/D124565, simplifies the autogenerated model by taking away all unnecessary features and getting rid of the functions that were previously to mix in all the necessary inputs so they wouldn't get pruned by the Tensorflow XLA AOT compiler. This is no longer necessary after the previously mentioned patch. This also fixes the nondeterministic behavior that is sometimes observed where the autogenerated model will simply output 0 instead of the correct index. Reviewed By: yundiqian Differential Revision: https://reviews.llvm.org/D129254
-
George Petterson authored
Reviewed By: silvas Differential Revision: https://reviews.llvm.org/D128880
-
George Petterson authored
-
Hui Xie authored
For some reason the pre-commit CI of https://reviews.llvm.org/D129233 was all green so I didn't spot this https://reviews.llvm.org/B174525 Reviewed By: #libc, philnik, Mordante Differential Revision: https://reviews.llvm.org/D129503
-
Kai Sasaki authored
Lower complex.log to corresponding function call with libm. Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D129417
-
Fangrui Song authored
[sanitizer] Remove #include <linux/fs.h> to resolve fsconfig_command/mount_attr conflict with glibc 2.36 It is generally not a good idea to mix usage of glibc headers and Linux UAPI headers (https://sourceware.org/glibc/wiki/Synchronizing_Headers). In glibc since 7eae6a91e9b1670330c9f15730082c91c0b1d570 (milestone: 2.36), sys/mount.h defines `fsconfig_command` which conflicts with linux/mount.h: .../usr/include/linux/mount.h:95:6: error: redeclaration of ‘enum fsconfig_command’ Remove #include <linux/fs.h> which pulls in linux/mount.h. Expand its 4 macros manually. Android sys/mount.h doesn't define BLKBSZGET and it still needs linux/fs.h. In the long term we should move Linux specific definitions to sanitizer_platform_limits_linux.cpp but this commit is easy to cherry pick into older compiler-rt releases. Fix https://github.com/llvm/llvm-project/issues/56421 Reviewed By: #sanitizers, vitalybuka, zatrazz Differential Revision: https://reviews.llvm.org/D129471
-
Fangrui Song authored
Revert "[sanitizer] Remove #include <linux/fs.h> to resolve fsconfig_command/mount_attr conflict with glibc 2.36" This reverts commit b379129c. Breaks Android build. Android sys/mount.h doesn't define macros like BLKBSZGET.
-
Joseph Huber authored
This patch adds the necessary changes required to bundle and wrap HIP files. The bundling is done using `clang-offload-bundler` currently to mimic `fatbinary` and the wrapping is done using very similar runtime calls to CUDA. This still does not support managed / surface / texture variables, that would require some additional information in the entry. One difference in the codegeneration with AMD is that I don't check if the handle is null before destructing it, I'm not sure if that's required. With this we should be able to support HIP with the new driver. Depends on D128850 Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D128914
-
Joseph Huber authored
This patch adds the small change required to output offloading entried for HIP instead of CUDA. These should be placed in different sections so because they need to be distinct to the offloading toolchain, otherwise we'd have HIP trying to register CUDA kernels or vice-versa. This patch will precede support for HIP in the linker wrapper. Reviewed By: yaxunl, tra Differential Revision: https://reviews.llvm.org/D128850
-