- Jan 24, 2023
-
-
Stanislav Mekhanoshin authored
Unlike older ASICs GFX10+ have a lot of VGPRs. Therefore, it is possible to achieve high occupancy even with all or almost all addressable VGPRs used. Our scheduler was never tuned for this scenario. The VGPR Critical Limit threshold always comes very high, even if maximum occupancy is targeted. For example on gfx1100 it is set to 192 registers even with the requested occupancy 16. As a result scheduler starts prioritizing register pressure reduction very late and we easily end up spilling. This patch makes VGPR critical limit similar to what we would have on pre-gfx10 targets with much more limited VGPR budget while still trying to maintain occupancy as it does now. Pre-gfx10 ASICs shall not be affected as the limit shall be the same as before, and on gfx10+ it shall only affect regions where we have to spill. Fixes: SWDEV-377300 Differential Revision: https://reviews.llvm.org/D141876
-
David Green authored
The Armv8.6-a and later architecture definitions included AES, SHA2, SHA3 and SM4, but this did not have an effect when specifying -march=armv8.6-a. The did not set preprocessor features (https://godbolt.org/z/1YKad6M8e) or enable the relevant instructions (like eor3 from sha3: https://godbolt.org/z/vY9v4MqvG). Similarly architectures armv8 to armv8.5 defined +crypto, but this did not effect the -march's, only the -mcpu with those architectures. I believe this was working as intended. After D141411 we now add the default features for architectures except for +crypto, which has had the effect of enabling aes/sha2/sha3/sm4 when -march=armv8.6-a is used. This patch removed those crypto features again, going back to how things were before. It also removes the AEK_CRYPTO feature from lower architecture levels, moving it to the cpus that use it. This shouldn't make any changes, but a few extra tests have been added for preprocessor features that have improved since llvm 15. The -mcpu=ampere1 cpu is the only armv8.6+ cpu at present. For that, the AES, SHA2 and SHA3 features have been re-added to the CPU definition to keep it in-line with the gcc definition from https://github.com/gcc-mirror/gcc/commit/db2f5d661239737157cf131de7d4df1c17d8d88d. Differential Revision: https://reviews.llvm.org/D141606
-
Pavel Iliin authored
Differential Revision: https://reviews.llvm.org/D142265
-
Han Zhu authored
Differential Revision: https://reviews.llvm.org/D142039
-
Pavel Iliin authored
This reverts commit 5474d7d9. Wrong differential revision link was used.
-
Shivam Gupta authored
When compiling clang/Lex/DirectoryLookup.h with option -Wbitfield-enum-conversion, we get the following warning: DirectoryLookup.h:77:17: warning: bit-field 'DirCharacteristic' is not wide enough to store all enumerators of 'CharacteristicKind' [-Wbitfield-enum-conversion] : u(Map), DirCharacteristic(DT), LookupType(LT_HeaderMap), DirCharacteristic is a bitfield with 2 bits (4 values) /// DirCharacteristic - The type of directory this is: this is an instance of /// SrcMgr::CharacteristicKind. unsigned DirCharacteristic : 2; Whereas SrcMgr::CharacterKind is an enum with 5 values: enum CharacteristicKind { C_User, C_System, C_ExternCSystem, C_User_ModuleMap, C_System_ModuleMap }; Solution is to increase DirCharacteristic bitfield from 2 to 3. Patch by Dimitri van Heesch Reviewed By: aaron.ballman Differential Revision: https://reviews.llvm.org/D142304 -
Craig Topper authored
This reduces RISCV.td to mainly being a top level include file. Reviewed By: asb, luismarques Differential Revision: https://reviews.llvm.org/D142239
-
Adrian Prantl authored
-
Pavel Iliin authored
Differential Revision: https://reviews.llvm.org/D141606
-
Caroline Concatto authored
Add the following intrinsic: SQCVT SQCVTU UQCVT NOTE: These intrinsics are still in development and are subject to future changes. Reviewed By: kmclaughlin Differential Revision: https://reviews.llvm.org/D142035
-
Florian Hahn authored
This patch moves a couple of helper functions from the global llvm:: namespace into the SCCPSolver class. This reduces the need for separate SCCPSolver arguments and also limits the scope of those functions that have quite generic names. (The remaining isConstant and isOverdefined should ideally be removed) Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D142370
-
Florian Hahn authored
-
Mark de Wever authored
-
Caroline Concatto authored
Add the following intrinsic: FCVT BFCVT FCVTZS FCVTZU SCVTF UCVTF This patch also adds SelectCVTIntrinsic to handle the cases when the intrinsic returns multiple (two or four) outputs NOTE: These intrinsics are still in development and are subject to future changes. Reviewed By: kmclaughlin Differential Revision: https://reviews.llvm.org/D142032
-
Zahira Ammarguellat authored
This option is useful for clang and clang-cl. Differential Revision: https://reviews.llvm.org/D142367
-
Xiang Li authored
Fixes #60184 https://github.com/llvm/llvm-project/issues/60184 Differential Revision: https://reviews.llvm.org/D142295
-
Guillaume Chatelet authored
This patch reduces CMake configuration time drastically by removing a non-linear behavior. Time to execute CMake configure step goes from 45s to 15s. Differential Revision: https://reviews.llvm.org/D142374
-
Lucas Prates authored
Update the clang driver to include the following features as default for the v8.9-A/v9.4-A architecture versions: * FEAT_SPECRES2 * FEAT_CSSC * FEAT_RASv2 Patch by Sam Elliott. Reviewed By: lenary, tmatheson Differential Revision: https://reviews.llvm.org/D141404
-
Lucas Prates authored
This introduces command line support (`+ite`) for the v9.4-A's Instrumentation Extension (FEAT_ITE). Patch by Son Tuan Vu. Reviewed By: lenary, tmatheson Differential Revision: https://reviews.llvm.org/D141403
-
Florian Hahn authored
-
Ties Stuij authored
Reviewed By: peter.smith Differential Revision: https://reviews.llvm.org/D142229
-
- Jan 23, 2023
-
-
Quentin Colombet authored
collapse/expand_shape are supposed to be expanded before we hit the lowering code. The expansion is done with the pass called expand-strided-metadata. This patch is NFC in spirit but not in practice because expand-strided-metadata won't try to accomodate for "invalid" strides for dynamic sizes that are 1 at runtime. The previous code was broken in that respect too, but differently: it handled only the case of row-major layouts. That whole part is being reworked separately. Differential Revision: https://reviews.llvm.org/D136483
-
Krasimir Georgiev authored
bazel: adapt for https://github.com/llvm/llvm-project/commit/a4699a43e42615281c96599d20977cabf10bfb9c
-
Archibald Elliott authored
This patch contains several related changes: 1. We move to using TARGET_BUILTIN for the 128-bit system register builtins to give better error messages when d128 has not been enabled, or has been enabled in a per-function manner. 2. We now validate the inputs to the 128-bit system register builtins, like we validate the other system register builtins. 3. We update the list of named PSTATE accessors for MSR (immediate), and now correctly enforce the expected ranges of the immediates. There is a long comment about how we chose to do this to comply with the ACLE when most of the PSTATE accessors for MSR (immediate) have aliased system registers for MRS/MSR which expect different values. In short, the MSR (immediate) names are prioritised, rather than falling-back to the register form when the value is out of range. Differential Revision: https://reviews.llvm.org/D140222
-
Alex Zinenko authored
The verification of affine value classification for symbols was expecting, incorrectly, that the dimension operand of `memref.dim` was being produced by a constant-like operation. This is legacy of the dimension being an attribute originally, and was never updated after it was switched to be an operation. Treat such cases conservatively and classify the value as non-symbol. A more advanced version could attempt to check that the value would be a valid symbol for all possible values the dimension attribute could take, but this does not seem immediately useful. Fixes #59993. Reviewed By: bondhugula Differential Revision: https://reviews.llvm.org/D142204
-
Alex Zinenko authored
It should have an "Allocate" effect on entry block arguments of all regions in addition to consuming the operand. Also relax the assertion in transform-dialect-check-uses until we can properly support region-based control flow. Fixes #60075. Reviewed By: springerm Differential Revision: https://reviews.llvm.org/D142200
-
Xiang Li authored
Fixes #60058 https://github.com/llvm/llvm-project/issues/60058 It hit assert when legalizePatternResult on success of ExpandIfCondition which did nothing just return success when if condition is constant. Added RemoveConstantIfCondition to remove the if cond by getCanonicalizationPatterns. Also remove the check for constant if cond in ExpandIfCondition and change check ifCond to assert because only op with ifCond will need legalize in ConvertOpenACCToSCFPass Differential Revision: https://reviews.llvm.org/D142286
-
Lucas Prates authored
This adds support for the v8.9-A/v9.4-A architectural extensions to be used in .arch_extension assembly directives. Patch by Sam Elliott. Reviewed By: lenary, tmatheson Differential Revision: https://reviews.llvm.org/D141402
-
Lucas Prates authored
This adds support for the missing `PIRE0_EL12` system register, part of v8.9-A/v9.4-A's Permission Indirection Extension. Patch by Son Tuan Vu. Reviewed By: tmatheson Differential Revision: https://reviews.llvm.org/D141400
-
Joseph Huber authored
Summary: Fix a few minor warnings that show up in `libomptarget`.
-
Joseph Huber authored
Summary: Recently AMD moved the "hsa.h" include to "hsa/hsa.h". This causes several warning. This patch checks to see if we can include that one instead. This should hopefully keep things backwards compatible while silencing the warnings.
-
Joseph Huber authored
Summary: These warnings are very loud considering they get repeated at least 30 times each build. This patch just silences them.
-
David Spickett authored
Make the first line a title and relative link to the Markdown of the demo notebook.
-
Shivam Gupta authored
Found by PVS-Studio - https://pvs-studio.com/en/blog/posts/cpp/1003/, N37. The code you is using the bit mask NullabilityKindMask which is 0x3 (00000011 in binary) to clear the bits in the NullabilityPayload variable. Since NullabilityPayload is a 64-bit variable and NullabilityKindMask is only a 8-bit variable(0x3), it will only affect the last 8 bits of the variable. The higher 56 bits will remain unchanged. Differential Revision: https://reviews.llvm.org/D142334
-
Jay Foad authored
The new methods return a range for easier iteration. Use them everywhere instead of getImplicitUses, getNumImplicitUses, getImplicitDefs and getNumImplicitDefs. A future patch will remove the old methods. In some use cases the new methods are less efficient because they always have to scan the whole uses/defs array to count its length, but that will be fixed in a future patch by storing the number of implicit uses/defs explicitly in MCInstrDesc. At that point there will be no need to 0-terminate the arrays. Differential Revision: https://reviews.llvm.org/D142215
-
Valentin Clement authored
When a `!fir.box<>` is passed as an actual argument to an optional `!fir.class<>` dummy it needs a `fir.rebox` in order to propagate the dynamic type information. The `fir.rebox` needs to happen only on present argument. Reviewed By: jeanPerier Differential Revision: https://reviews.llvm.org/D142340
-
Florian Hahn authored
-
Simon Pilgrim authored
Also, move the XformToShuffleWithZero and combineCarryDiamond folds later after some of the more basic canonicalizations/combines (such as this) have had a chance to occur Fixes the v8i1-masks.ll regression from D127115
-
Kadir Cetinkaya authored
Introduce signals to rank providers of a symbol. Differential Revision: https://reviews.llvm.org/D139921
-