- Sep 13, 2023
-
-
Matt Arsenault authored
This reverts commit d9333e36.
-
Matt Arsenault authored
Ported from old amdgcn intrinsic which will soon be deleted. https://reviews.llvm.org/D149587
-
Kai Luo authored
In Orc runtime, we use `dlopen(nullptr, ...)` to open current executable and use `dlsym` to find addresses of symbols, this requires `-rdynamic` flag. As `llvm/CMakeLists.txt` suggests ``` # Make sure we don't get -rdynamic in every binary. For those that need it, # use export_executable_symbols(target). ``` This patch exports symbols in `ClangReplInterpreterExceptionTests`. This also fixes `ClangReplInterpreterExceptionTests` is skipped on ppc64 when jitlink is used. Reviewed By: v.g.vassilev Differential Revision: https://reviews.llvm.org/D159167
-
Pravin Jagtap authored
[D156301](https://reviews.llvm.org/D156301 ) introduced atomic optimizations for FAdd/FSub. For FSub, reduction/scan needs to be performed using add operation (`not sub`) and memory location will be updated by reduced value using atomic sub later by only one lane. --------- Authored-by:
Pravin Jagtap <Pravin.Jagtap@amd.com>
-
Hao Jin authored
While creating a temporary alloca for a box in OpenMp region, the insertion point should be the OpenMP region block instead of the function entry block.
-
Farzon Lotfi authored
Updates GetInstructionSize to account for arm64 instruction sizes. ARM64 instruction are always 4 bytes long but GetInstructionSize in interception_win.cpp assumes x86_64 which has mixed sizes. Fix is for: https://github.com/llvm/llvm-project/issues/64319 Before the changeclang_rt.asan_dynamic-aarch64.dll would crash at: OverrideFunction -> OverrideFunctionWithHotPatch -> GetInstructionSize:825 After the change: dllthunkintercept -> dllthunkgetrealaddressordie -> InternalGetProcAddress
-
LLVM GN Syncbot authored
-
Peiming Liu authored
-
Nico Weber authored
-
Jeffrey Byrnes authored
This reverts commit 7fda1b74.
-
Ellis Hoag authored
This fixes a build error introduced by https://reviews.llvm.org/D153587 when using an old version of GCC. See https://reviews.llvm.org/D153587#4644735 for details.
-
Louis Dionne authored
Since we use C++20 to build the dylib, we can use a lambda to do the first-time initialization instead of emulating std::bind. This should not change the behavior of the code at all, it merely simplifies it. This removes a symbol from the dylib, however that symbol was only ever used inside the dylib so it shouldn't break the ABI for anyone. I confirmed that by searching for that symbol on the ABI boundary of a large number of programs and couldn't find any references to that function.
-
Aaron Jarmusch authored
This reverts commit e831a32c.
-
Tobias Stadler authored
Actually pass along the depth parameter of getKnownBits to computeKnownBitsImpl. Reviewed By: aemerson Differential Revision: https://reviews.llvm.org/D159321
-
Alex authored
Instead of creating a copy of the vector, we should just pass a reference along. The only method that calls this Ctor also holds onto a non-mutable reference to the vector of strings so a copy should be unnecessary.
-
Mohammed Keyvanzadeh authored
- Remove usages of the non-existent `ignore-forks` field, conditions in jobs already exist to prevent the jobs from running in forks. - Don't use variables in the `printf` format string. Use `printf "..%s.." "$foo"`. ([SC2059](https://www.shellcheck.net/wiki/SC2059)) - Double quote variable expansion to prevent globbing and word splitting. ([SC2086](https://www.shellcheck.net/wiki/SC2086)) - Prefer `[ p ] || [ q ]` as `[ p -o q ]` is not well defined. ([SC2166](https://www.shellcheck.net/wiki/SC2166)) - Consider `{ cmd1; cmd2; } >> file` instead of individual redirects. ([SC2129](https://www.shellcheck.net/wiki/SC2129)) - Use `$(...)` notation instead of legacy notation `...`. ([SC2006](https://www.shellcheck.net/wiki/SC2006)) - Use `./*glob*` or `-- *glob*` so names with dashes won't become options. ([SC2035](https://www.shellcheck.net/wiki/SC2035)) - Refactor JavaScript code in certain workflows. - Change workflow variable substitution style of some wor...
-
Armando Martín authored
This "bug" was probably not noticed because it doesn't affect any integer type we currently support. It requires integers with more than 2x the size of `unsigned long long`. However, with such types, the algorithm used to break down the large integer into groups of size `unsigned long long` didn't work because we rotated in the wrong direction. For example, the 256 bit number (1 << 255) would yield the wrong answer when used with the algorithm before this patch. In particular, note that the current rotation happens to work for 128 bit integers because it just swaps the halves in this case. Differential Revision: https://reviews.llvm.org/D134625 Co-authored-by:
Louis Dionne <ldionne.2@gmail.com>
-
Jonas Devlieghere authored
Adopt the new markup overload, introduced in 77d10325, in the ARM backend. This commit completes the migration and removes the old overload.
-
Joseph Huber authored
Summary: This patch improves the implementation of the standard `rand()` function by implementing it in terms of the xorshift64star pRNG as described in https://en.wikipedia.org/wiki/Xorshift#xorshift*. This is a good, general purpose random number generator that is sufficient for most applications that do not require an extremely long period. This patch also correctly initializes the seed to be `1` as described by the standard. We also increase the `RAND_MAX` value to be `INT_MAX` as the standard only specifies that it can be larger than 32768.
-
Louis Dionne authored
-
Rodrigo Ceccato de Freitas authored
This commit removes an optimization that skips the initialization of the reduction struct if the number of threads in a team is 1. This optimization caused a bug with Hidden Helper Threads. When the task group is initially initialized by the master thread but a Hidden Helper Thread executes a target nowait region, it requires the reduction struct initialization to properly accumulate the data. This commit also adds a LIT test for issue #57522 to ensure that the issue is properly addressed and that the optimization removal does not introduce any regressions. Fixes: #57522
-
Jakub Kuderski authored
- Check `MakePointer*` load/store attribute values. - Support coop matrix types in `MatrixTimesScalar` verification. - Add test cases for all the remaining ops that accept coop matrix types. - Split NV and KHR tests.
-
Kazu Hirata authored
This patch fixes: llvm/lib/Target/AMDGPU/AMDGPU.h:297:18: error: private field 'TM' is not used [-Werror,-Wunused-private-field]
-
Aaron Ballman authored
This addresses issues introduced by efe4a548
-
Amir Ayupov authored
AutoFDO profile has no leading 0x in hex dumps. Reviewed By: #bolt, rafauler Differential Revision: https://reviews.llvm.org/D159507
-
Xiang Li authored
Add const qualifier to avoid warning "_cast from 'const char *' to 'char*' drops const qualifier [-Werror,-Wcast-qual]_" This will allow enable LLVM_ENABLE_WERROR when build with clang-cl on Windows. This is imported from https://github.com/openbsd/src/commit/b81763002452802e4f7304ea60f121253bd94 Also removed the gcc change for the cast-qual warning which is not needed with const qualifier added.
-
Aaron Jarmusch authored
-
kkwli authored
This patch is to add the support of declaring a Cray pointer in a module.
-
Paul T Robinson authored
-
Matthias Braun authored
Propagate "branch_weights" metadata whe turning a select into a conditional branch in tryToUnfoldSelectInCurrBB
-
Artem Belevich authored
Fixes https://github.com/llvm/llvm-project/issues/57544
-
jwanggit86 authored
This patch ports the AMDGPURewriteUndefForPHI pass to the new pass manager. With this, the pass is supported under both the legacy and the new pass managers. --------- Co-authored-by:Jun Wang <jun.wang7@amd.com>
-
Justin Bogner authored
It's weird that we're specifying this in both the dxil and the LLVM IR way, but if we are we should at least be consistent about it.
-
Aiden Grossman authored
These tests are still relatively flaky on certain buildbots and on some developer machines. This patch hopefully reduces the noise produced by these tests in those environments by setting the retry count to 2 so that they will hopefully flaky pass rather than just fail while I perform more investigation.
-
-
-
Matt Arsenault authored
We want the !fpmath metadata to be attached to the sqrt intrinsic to make it to the backend lowering. Emit an available_externally definition which uses the builtin, which emits the !fpmath. Fixes #64264 https://reviews.llvm.org/D156743
-
Matt Arsenault authored
Make codegen emit correctly rounded sqrt by default. Emit the fast but only kind of fast expansion in AMDGPUCodeGenPrepare based on !fpmath, like the fdiv case. Hack around visitation ordering problems from AMDGPUCodeGenPrepare using forward iteration instead of a well behaved combiner. https://reviews.llvm.org/D158129
-
Tom Stellard authored
This will reduce the number of notifications created when a pull request label is added. Each team will only get a notification when their team's label is added and not when other teams' labels are added.
-
Vitaly Buka authored
This loop is wrong, most of targets are not defined yet. Also if we build with LLVM_ENABLE_RUNTIMES, these deps are irrelevant.
-