- Sep 28, 2021
-
-
Roman Lebedev authored
The only sched models that for cpu's that support avx2 but not avx512 are: haswell, broadwell, skylake, zen1-3 For load we have: https://godbolt.org/z/5EYc6r9nh - for intels `Block RThroughput: =6.0`; for ryzens, `Block RThroughput: <=3.0` So pick cost of `6`. For store we have: https://godbolt.org/z/z61e5d6GE - for intels `Block RThroughput: =2.0`; for ryzens, `Block RThroughput: <=1.0` So pick cost of `2`. I'm directly using the shuffling asm the llc produced, without any manual fixups that may be needed to ensure sequential execution. Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D110536
-
Chris Bieneman authored
I always forget that new line...
-
Chris Bieneman authored
This patch adds a new preprocessor extension ``#pragma clang final`` which enables warning on undefinition and re-definition of macros. The intent of this warning is to extend beyond ``-Wmacro-redefined`` to warn against any and all alterations to macros that are marked `final`. This warning is part of the ``-Wpedantic-macros`` diagnostics group. Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D108567
-
Sanjay Patel authored
-
Sanjay Patel authored
-
Aart Bik authored
This integration tests runs a fused and non-fused version of sampled matrix multiplication. Both should eventually have the same performance! NOTE: relies on pending tensor.init fix! Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D110444
-
Nico Weber authored
This reverts commit 0b1eff1b. Breaks check-clangd on Windows, see comments on https://reviews.llvm.org/D110386
-
Jon Chesterfield authored
This reverts commit 1a761e5b. Failed CI, albeit with a different failure mode to BZ51982
-
Jon Chesterfield authored
Fixes 51982. Minor refactor to remove `return x = y` construct. Test case derived from https://github.com/ROCm-Developer-Tools/aomp/\ blob/aomp-dev/test/smoke/nest_call_par2/nest_call_par2.c by deleting parts while checking the assertion failure still occurred. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D110556
-
Aart Bik authored
This revision makes sure that when the output buffer materializes locally (in contrast with the passing in of output tensors either in-place or not in-place), the zero initialization assumption is preserved. This also adds a bit more documentation on our sparse kernel assumption (viz. TACO assumptions). Reviewed By: bixia Differential Revision: https://reviews.llvm.org/D110442
-
Aaron Ballman authored
-
Sanjay Patel authored
This bug was introduced with the refactoring in: 9075edc8 ...but there were no tests to detect it.
-
Sanjay Patel authored
-
Jameson Nash authored
We see that it might otherwise do: %10 = getelementptr {}**, <2 x {}***> %9, <2 x i32> <i32 10, i32 4> %11 = bitcast <2 x {}***> %10 to <2 x i64*> ... %27 = extractelement <2 x i64*> %11, i32 0 %28 = bitcast i64* %27 to <2 x i64>* store <2 x i64> %22, <2 x i64>* %28, align 4, !tbaa !2 Which is an out-of-bounds store (the extractelement got offset 10 instead of offset 4 as intended). With the fix, we correctly generate extractelement for i32 1 and generate correct code. Differential Revision: https://reviews.llvm.org/D106613 -
Carlos Galvez authored
Fixes https://bugs.llvm.org/show_bug.cgi?id=51790. The check triggers incorrectly with non-type template parameters. A bisect determined that the bug was introduced here: https://github.com/llvm/llvm-project/commit/ea2225a10be986d226e041d20d36dff17e78daed Unfortunately that patch can no longer be reverted on top of the main branch, so add a fix instead. Add a unit test to avoid regression in the future.
-
peter klausler authored
Enforce constraints C1034 & C1038, which disallow the use of otherwise valid statements as branch targets when they appear in FORALL &/or WHERE constructs. (And make the diagnostic message somewhat more user-friendly.) Differential Revision: https://reviews.llvm.org/D109936
-
Praveen Velliengiri authored
HSA runtime fails to find the symbols for Init and Fini kernels as they mark with internal linkage, changing the linkage to external to fix those errors. Differential Revision: https://reviews.llvm.org/D110054
-
Sumesh Udayakumaran authored
New mode option that allows for either running the default fusion kind that happens today or doing either of producer-consumer or sibling fusion. This will also be helpful to minimize the compile-time of the fusion tests. Reviewed By: bondhugula, dcaballe Differential Revision: https://reviews.llvm.org/D110102
-
Quinn Pham authored
This patch fixes the pattern for the P10 instructions Vector Shift Left Double by Bit Immediate VN-form and Vector Shift Right Double by Bit Immediate VN-form. The third argument should be a target constant (`timm`) instead of an `i32` because an immediate is expected. Reviewed By: lei Differential Revision: https://reviews.llvm.org/D109920
-
Yaxun (Sam) Liu authored
HIP currently uses -mlink-builtin-bitcode to link all bitcode libraries, which changes the linkage of functions to be internal once they are linked in. This works for common bitcode libraries since these functions are not intended to be exposed for external callers. However, the functions in the sanitizer bitcode library is intended to be called by instructions generated by the sanitizer pass. If their linkage is changed to internal, their parameters may be altered by optimizations before the sanitizer pass, which renders them unusable by the sanitizer pass. To fix this issue, HIP toolchain links the sanitizer bitcode library with -mlink-bitcode-file, which does not change the linkage. A struct BitCodeLibraryInfo is introduced in ToolChain as a generic approach to pass the bitcode library information between ToolChain and Tool. Reviewed by: Artem Belevich Differential Revision: https://reviews.llvm.org/D110304
-
William S. Moses authored
Address post-commit review in https://reviews.llvm.org/D108524 to add appropriate diagnostics. Reviewed By: mehdi_amini Differential Revision: https://reviews.llvm.org/D110566
-
peter klausler authored
A defined assignment subroutine invoked in the context of a WHERE statement or construct must necessarily be elemental (C1032). Differential Revision: https://reviews.llvm.org/D109932
-
Craig Topper authored
This can avoid a loss of decoupling with the scalar unit on cores with decoupled scalar and vector units. We should support FP too, but those use extract_element and not a custom ISD node so it is a little different. I also left a FIXME in the test for i64 extract and store on RV32. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D109482
-
Fangrui Song authored
Fix PR51961 Reviewed By: jhenderson Differential Revision: https://reviews.llvm.org/D110490
-
Kazu Hirata authored
-
Bixia Zheng authored
The sparse constant provides a constant tensor in coordinate format. We first split the sparse constant into a constant tensor for indices and a constant tensor for values. We then generate a loop to fill a sparse tensor in coordinate format using the tensors for the indices and the values. Finally, we convert the sparse tensor in coordinate format to the destination sparse tensor format. Add tests. Reviewed By: aartbik Differential Revision: https://reviews.llvm.org/D110373
-
@vladaindjic authored
The minor code refactorization introduces the TASK_TIED constant inside kmp_gsupprot.cpp as a replacement for the literal value 1. The mentioned constant is now used in both kmp_tasking.cpp and kmp_gsupport.cpp files. Differential Revision: https://reviews.llvm.org/D110441
-
Craig Topper authored
If one input of a fixed vector multiply is a sign/zero extend and the other operand is a splat of a scalar, we can use a widening multiply if the scalar value has sufficient sign/zero bits. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D110028
-
Daniil Fukalov authored
1. Convert to generated tests. 2. Added code-size case in few places.
-
Sanjay Patel authored
This is no-functional-change-intended, but it hopefully makes things slightly clearer and more efficient to have transforms that require 'shl' be called only from visitShl(). Further cleanup is possible.
-
Pavel Labath authored
we need to drop nuls from the end of the string.
-
- Sep 27, 2021
-
-
Kazu Hirata authored
Note that getTheLanaiTarget is declared in TargetInfo/LanaiTargetInfo.h, which LanaiDisassembler.cpp includes. Identified with readability-redundant-declaration.
-
Kirill Bobyrev authored
Preparation for D108194. Reviewed By: sammccall Differential Revision: https://reviews.llvm.org/D110386
-
Joseph Huber authored
This path defines the newly added `__kmpc_disitrute_static_init` functions in the device runtime library. These functions are currently exact copies of the current worksharing method but can be tuned later. Depends on D110429 Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D110430
-
Joseph Huber authored
This patch adds a new RTL function for worksharing. Currently we use `__kmpc_for_static_init` for both the `distribute` and `parallel` portion of the loop clause. This patch replaces the `distribute` portion with a new runtime call `__kmpc_distribute_static_init`. Currently this will be used exactly the same way, but will make it easier in the future to fine-tune the distribute and parallel portion of the loop. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D110429
-
Raphael Isemann authored
Apparently macOS is padding the name result with several padding zeroes at the end. Just strip them all to pretend it's a C-string. Thanks to Pavel for suggesting this fix.
-
Nico Weber authored
-
Jake Egan authored
The tests only specify -march, so when the tests are run on AIX the target OS defaults to AIX, which causes the tests to misbehave. This patch constrains the tests by specifying -mtriple instead of -march. Reviewed By: daltenty, jsji, MaskRay Differential Revision: https://reviews.llvm.org/D110186
-
Nico Weber authored
-
Nico Weber authored
-