- Jan 30, 2024
-
-
Peter Klausler authored
The runtime type information table generator couldn't handle a null pointer returned correctly for a original (not instantiated) derived type with kind parameters. Fixes https://github.com/llvm/llvm-project/issues/79590.
-
Nilanjana Basu authored
[LV] Update interleaving count computation when scalar epilogue loop needs to run at least once (#79651) Update loop interleaving count computation to address loops that require at least one scalar iteration in the epilogue loop. For this case, the available trip count for interleaving the loop is one less.
-
Fangrui Song authored
The interaction between --warn-backrefs was not tested, but if --defsym-created reference causes archive member extraction, it seems reasonable to suppress the diagnostic, which was the behavior before #78944.
-
Joseph Huber authored
Summary: This patch adds support for `globaltimer` to match `clock` and `clock64`. See the PTX ISA reference for details. This patch does not implement the `hi` or `lo` variants for brevity as they can be obtained from this with the cost of an additional register. https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#special-registers-globaltimer-globaltimer-lo-globaltimer-hi
-
Joseph Huber authored
Summary: The PTX ISA has always supported the 'exit' instruction to terminate individual threads. This patch adds a builtin to handle it. See the PTX documentation for further details. https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#control-flow-instructions-exit
-
Joseph Huber authored
Summary: This patch adds a builtin for the `nanosleep` PTX function. It takes either an immediate or a register and sleeps for [0, 2t] nanoseconds given t. More information at the documentation: https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#miscellaneous-instructions-nanosleep
-
Joseph Huber authored
Summary: This patch adds support for getting the 'activemask' instruction's value without needing to use inline assembly. See the relevant PTX reference for details. https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#parallel-synchronization-and-communication-instructions-activemask
-
Fangrui Song authored
Update code from https://reviews.llvm.org/D138847 `buildTestVector` is a standard DFS (walking a reduced ordered binary decision diagram). Avoid shouldCopyOffTestVectorFor{True,False}Path complexity and redundant `Map[ID]` lookups. `findIndependencePairs` unnecessarily uses four nested loops (n<=6) to find independence pairs. Instead, enumerate the two execution vectors and find the number of mismatches. This algorithm can be optimized using the marking function technique described in _Efficient Test Coverage Measurement for MC/DC, 2013_, but this may be overkill.
-
Aiden Grossman authored
This patch removes the llvm:: prefix within llvm-exegesis where it is not necessary. This is most occurrences of the prefix within exegesis as exegesis is within the llvm namespace. This patch makes things more consistent as the vast majority of the code did not use the llvm:: prefix for anything.
-
David Majnemer authored
-
Louis Dionne authored
-
Louis Dionne authored
The upload-artifact@v3 action is using Node 16, which is reaching EOL. As a result, we are getting warnings prompting us to move our jobs over to the latest version of upload-artifact.
-
Louis Dionne authored
Fixes #79783
-
Craig Topper authored
Using dyn_cast allows us to use CastInst::getOperand instead of Instruction::getOperand. This is more efficient since CastInst::getOperand doesn't need to check how the operands are stored. Instruction::getOperand has to consider HungOffUses.
-
Jason Molenda authored
Also revert this patch until Ismail can re-land. This reverts commit febb4c42.
-
Adrian Prantl authored
-
Hristo Hristov authored
Deleted the offending test case. `libcxx/test/std/utilities/format/format.arguments/format.arg/visit.return_type.pass.cpp` lines: 134-135: > test<Context, bool, long>(true, 192812079084L); test<Context, bool, long>(false, 192812079084L); Relands: https://github.com/llvm/llvm-project/pull/76449 Reverted in: https://github.com/llvm/llvm-project/commit/02f95b77515fe18ed1076b94cbb850ea0cf3c77e --------- Co-authored-by:Zingam <zingam@outlook.com>
-
Alexey Bataev authored
demote the tree entry. Need to check if all user nodes are marked for demotion before demoting the node. Otherwise, some data info might be lost after vectorization.
-
Alexander Shaposhnikov authored
Add tests for llvm.abs >= 0. This is a preparation for https://github.com/llvm/llvm-project/pull/79070 .
-
Jason Molenda authored
Temporarily revert to unblock the CI bots, this is breaking the -DLLVM_ENABLE_MODULES=On modules style build. I've notified Ismail. This reverts commit 888501bc.
-
Alexey Bataev authored
-
Adrian Prantl authored
-
Nilanjana Basu authored
[Tests][LV][AArch64] Pre-commit tests for changing loop interleaving count computation for loops that need to run scalar iterations (#79640) This patch contains a set of pre-commit tests for changing the loop interleaving count computation in a subsequent patch in order to address loops that need to execute at least a single scalar iteration in the epilogue.
-
Arthur Eubanks authored
The very beginning already talks about how to git clone the repo. The section about checking out specific versions doesn't really belong in GettingStarted and seems unnecessary.
-
Shafik Yaghmour authored
In Sema in `BuildReturnStmt(...)` when we try to determine is the type is move eligible or copy elidable we don't currently check of the init of the `VarDecl` contain errors or not. This can lead to a crash since we may send a type that is not complete into `getTypeInfo(...)` which does not allow this. This fixes: https://github.com/llvm/llvm-project/issues/63244 https://github.com/llvm/llvm-project/issues/79745
-
Antonio Frighetto authored
Simplify multi-use `and`/`or`/`xor` when these last do not affect the demanded bits being considered. Fixes: https://github.com/llvm/llvm-project/issues/78596. Proofs: https://alive2.llvm.org/ce/z/EjuWHa.
-
Antonio Frighetto authored
-
Simon Pilgrim authored
We try to only use X32 for gnux32 triple tests.
-
Simon Pilgrim authored
32/64-bit triples and check prefixes were inverted, and missing unwind attribute to strip cfi noise
-
Justin Bogner authored
Pull Request: https://github.com/llvm/llvm-project/pull/78225
-
Will Hawkins authored
According to internally agreed upon best practices, type template parameter names representing iterator types should be named `Iter`. For type template parameters representing sentinel types, they should be named `Sent`. Signed-off-by:Will Hawkins <hawkinsw@obs.cr>
-
jeanPerier authored
Compiler was rewriting SIZE(PACK(x, MASK)) to COUNT(MASK). It was wrapping the COUNT call without a KIND argument (leading to INTEGER(4) result in the characteristics) in an Expr<ExtentType> (implying INTEGER(8) result), this lead to inconsistencies that later hit verifier errors in lowering. Set the KIND argument to the KIND of ExtentType to ensure the built expression is consistent. This requires giving access to some safe place where the "kind" name can be saved and turned into a CharBlock (count has a DIM argument that require using the KIND keyword here). For the FoldingContext that belong to SemanticsContext, this is the same string set as the one used by SemanticsContext for similar purposes.
-
Justin Bogner authored
This is a good place to put all of the ABI-sensitive DXIL values that we'll need in both reading and writing contexts. Pull Request: https://github.com/llvm/llvm-project/pull/78224
-
Slava Zakharin authored
The existing type size computation in LoopVersioning does not work for REAL*10, because the compute element size is 10 bytes, which violates the power-of-two assertion. We'd better use the DataLayout for computing the storage size of each element of an array of the given type.
-
Joseph Huber authored
This reverts commit c9a6e993. This breaks HIP code that incorrectly depended on GPU-specific macros to be set. The code is totally wrong as using `__WAVEFRTONSIZE__` on the host is absolutely meaningless, but it seems this entire corner of the toolchain is fundmentally broken. Reverting for now to avoid breakages.
-
Shengchen Kan authored
-
Simon Pilgrim authored
We try to only use X32 for gnux32 triple tests.
-
David Green authored
The condition for allowing integer complex number support could also allow neon fixed length complex numbers if +sve2 was specified. This tightens the condition to only allow integer complex number support for scalable vectors. We could generalize this in the future to generate SVE intrinsics for fixed-length vectors, but for the moment this opts for the simpler fix.
-
Gheorghe-Teodor Bercea authored
This patch outlines the SPMD code path into a separate function that can be called directly.
-
Alexandros Lamprineas authored
With a690e867 we added -mcpu/mtune=native support to handle the Microsoft Azure Cobalt 100 CPU as a Neoverse N2. This patch adds a CPU alias in TargetParser to maintain compatibility with GCC.
-