- May 23, 2024
-
-
Vlad Serebrennikov authored
This patch remove 36 checks for compiler flags that are done via invoking the compiler across LLVM, Clang, and LLDB. It's was made possible by raising the bar for supported compilers that has been happening over the years since the checks were added. This is going to improve CMake configuration times. This topic was highlighted in https://discourse.llvm.org/t/cmake-compiler-flag-checks-are-really-slow-ideas-to-speed-them-up/78882.
-
David Spickett authored
| is a reserved character in regex.
-
Matheus Izvekov authored
-
2LoS authored
Removed redundant template in '__delete_node()' member function of '__forward_list_base' and '__list_imp' classes. (#84323)
-
Abid Qadeer authored
This PR adds supports for conversion of complex type to corresponding DITypeAttr. Both fir and mlir types are supported. Apart from lit testing, I have also tested the types in debugger and they work correctly. An exception is 128 bit complex which somehow requires that its name be different from `complex`. I am going to open a separate PR to add (kind=n) in the type names similar to what gfortran does.
-
Timm Bäder authored
No inline descriptor means we can't do that.
-
Timm Bäder authored
-
Hari Limaye authored
This patch extends support for more efficient lowering of the experimental.cttz.elts intrinsic to fixed-width vector types, by first creating an SVE predicate register mask from the fixed-width vector.
-
Jay Foad authored
-
Jan Patrick Lehr authored
While we investigate the issue, we disable the test on host-offloading so the buildbots are back to more useful state. Issue is tracked: https://github.com/llvm/llvm-project/issues/93173
-
Simon Pilgrim authored
Improve test checks for better codegen review of #92576
-
Tom Eccles authored
This operation did not model the behaviour of reductions in the openmp standard. It has since been replaced by block arguments on the outer operation. See https://github.com/llvm/llvm-project/pull/79308 and https://github.com/llvm/llvm-project/pull/80019
-
Timm Bäder authored
Since we now have type info for dummy pointers, we don't need this check anymore and can also have the same output for the test case in records.cpp.
-
Ramkumar Ramachandra authored
Implement NFC improvements spotted during a cursory reading of LoopAccessAnalysis.
-
LLVM GN Syncbot authored
-
Balázs Kéri authored
The "cert" package looks not useful and the checker has not a meaningful name with the old naming scheme. Additionally tests and documentation is updated.
-
LLVM GN Syncbot authored
-
Chuanqi Xu authored
Revert "[Coroutines] Always set the calling convention of generated resuming call from 'llvm.coro.await.suspend.handle' as fast" This reverts commit 31f1590e. It looks like some bots are not happy about the FileChecks
-
Pierre van Houtryve authored
This enables the --lto-partitions option to work more consistently. This module splitting logic is fully aware of AMDGPU modules and their specificities and takes advantage of them to split modules in a way that avoids compilation issue (such as resource usage being incorrectly represented). This also includes a logging system that's more elaborate than just LLVM_DEBUG which allows printing logs to uniquely named files, and optionally with all value names hidden so they can be safely shared without leaking informatiton about the source. Logs can also be enabled through an environment variable, which avoids the sometimes complicated process of passing a -mllvm option all the way from clang driver to the offload linker that handles full LTO codegen.
-
Chuanqi Xu authored
[Coroutines] Always set the calling convention of generated resuming call from 'llvm.coro.await.suspend.handle' as fast See the post commit message in https://github.com/llvm/llvm-project/pull/89751 We met a regression due to a change of calling convention of this patch. Previously, the calling convention of indirect resume calls is always fast. And in this patch, although we tried to take care of it in the cloner, we forget the case that we have to update the resuming calls in the ramp functions. So this is the root cause of the downstream failure. This patch tries to mark the generated resuming calls as fast immediately after they got created to make sure the calling convention is correct.
-
Guray Ozen authored
The test used the check generated ptx with `CHECK-PTX`, but does not check that anymore. The PR removes these lines.
-
Simon Pilgrim authored
If the comparison results are allbits masks, we can expand as `abd(lhs, rhs) -> sub(cmpgt(lhs, rhs), xor(sub(lhs, rhs), cmpgt(lhs, rhs)))`, replacing a sub+sub+select pattern with the simpler sub+xor+sub pattern. This allows us to remove a lot of X86 specific legalization code, and will be useful in future generic expansion for the legalization work in #92576 Alive2: https://alive2.llvm.org/ce/z/sj863C
-
Dmitry Vasilyev authored
Sometimes this test failed on the assert `The thread exited` in case of a remote target. Increase the timeout to 1 second to avoid a racing condition.
-
Dmitry Vasilyev authored
The TestBreakpointCommand test is incorrectly disabled for Windows target. We can disable it for Windows host instead or just fix the issue. This patch fixes the path separator in BreakpointResolverFileLine::DeduceSourceMapping() and the Windows specific absolute path in the test in case of the Windows host.
-
Timm Bäder authored
-
Nikita Popov authored
-
Timm Bäder authored
We still need to handle Inits.size() == 0, but we can do that earlier.
-
Sander de Smalen authored
The intrinsics are currently defined as: ``` __aio __attribute__((target("sve"))) svint8_t svreinterpret_s8(svuint8_t op) __arm_streaming_compatible { return __builtin_sve_reinterpret_s8_u8(op); } ``` which doesn't work when calling it from an __arm_streaming function when only +sme is available. By defining it in the same way as we've defined all the other intrinsics, we can leave it to the code in SemaChecking to verify that either +sve or +sme is available. This PR also fixes the target guards for the svreinterpret_c and svreinterpret_b intrinsics, that convert between svcount_t and svbool_t, as these are available both in SME2 and SVE2p1. -
Tom Eccles authored
The pass constructor can be generated automatically. This pass is module-level and then runs on all relevant intrinsic operations inside of the module, no matter what top level operation they are inside of.
-
csstormq authored
Patch co-authored by AtariDreams (gfunni234@gmail.com). Fixes #38037. [AMDGPU] Update test results to fix build (#92982)
-
Pavel Labath authored
We currently cannot represent abbreviation codes with more than 16 bits, and we were lldb-asserting if we ever ran into one. While I haven't seen any real DWARF with these kinds of abbreviations, it is possible to hit this with handcrafted evil dwarf, due some sort of corruptions, or just bugs (the addition of PeekDIEName makes these bugs more likely, as the function blindly dereferences offsets within the debug info section) . Missing abbreviations were already reporting an error. This patch turns sure that large abbreviations into an error as well, and adds a test for both cases.
-
Alexandros Lamprineas authored
Fixes the following bug: namespace Name { int __attribute((target_version("default"))) foo() { return 0; } } namespace Name { int __attribute((target_version("sve"))) foo() { return 1; } } int bar() { return Name::foo(); } error: redefinition of 'foo' int __attribute((target_version("sve"))) foo() { return 1; } note: previous definition is here int __attribute((target_version("default"))) foo() { return 0; } While fixing this I also found that in the absence of default version declaration, the one we implicitly create has incorrect mangling if we are in a namespace: namespace OtherName { int __attribute((target_version("sve"))) foo() { return 2; } } int baz() { return OtherName::foo(); } In this example instead of creating a declaration for the symbol @_ZN9OtherName3fooEv.default we are creating one for the symbol @_Z3foov.default (the namespace mangling prefix is omitted). This has now been fixed. -
Mubashar Ahmad authored
The deinterleave operation constructs two vectors from a single input vector. The first result vector contains the elements from even indexes of the input, and the second contains elements from odd indexes. This is the inverse of a `vector.interleave` operation. Each output's trailing dimension is half of the size of the input vector's trailing dimension. This operation requires the input vector to have a rank > 0 and an even number of elements in its trailing dimension. The operation supports scalable vectors. Example: ```mlir %0, %1 = vector.deinterleave %a : vector<8xi8> -> vector<4xi8> %2, %3 = vector.deinterleave %b : vector<2x8xi8> -> vector<2x4xi8> %4, %5 = vector.deinterleave %c : vector<2x8x4xi8> -> vector<2x8x2xi8> %6, %7 = vector.deinterleave %d : vector<[8]xf32> -> vector<[4]xf32> %8, %9 = vector.deinterleave %e : vector<2x[6]xf64> -> vector<2x[3]xf64> %10, %11 = vector.deinterleave %f : vector<2x4x[6]xf64> -> vector<2x4x[3]xf64> ``` -
Med Ismail Bennani authored
Revert "[lldb] Make use of Scripted{Python,}Interface for ScriptedThreadPlan (Reland #70392)" (#93153) Reverts llvm/llvm-project#93149 since it breaks https://lab.llvm.org/buildbot/#/builders/68/builds/74799 -
LLVM GN Syncbot authored
-
Med Ismail Bennani authored
This patch makes ScriptedThreadPlan conforming to the ScriptedInterface & ScriptedPythonInterface facilities by introducing 2 ScriptedThreadPlanInterface & ScriptedThreadPlanPythonInterface classes. This allows us to get rid of every ScriptedThreadPlan-specific SWIG method and re-use the same affordances as other scripting offordances, like Scripted{Process,Thread,Platform} & OperatingSystem. To do so, this adds new transformer methods for `ThreadPlan`, `Stream` & `Event`, to allow the bijection between C++ objects and their python counterparts. This just re-lands #70392 after fixing test failures. Signed-off-by:Med Ismail Bennani <ismail@bennani.ma>
-
Timm Bäder authored
-
Timm Bäder authored
Now that we call this more often, try to keep pointer chasing to a minimum.
-
Anchu Rajendran S authored
Named location attribute added to `tgt_offload_entry` shall be used by runtime calls like `ompx_dump_mapping_tables` to print the information of variables that are mapped to the device. `ompx_dump_mapping_tables` was printing the wrong location information and this change fixes it. A sample execution of example before the change: ``` omptarget device 0 info: OpenMP Host-Device pointer mappings after block at libomptarget:0:0: omptarget device 0 info: Host Ptr Target Ptr Size (B) DynRefCount HoldRefCount Declaration omptarget device 0 info: 0x0000000000206df0 0x00007f02cdc00000 20000000 1 0 <program-file-loc> at unknown:18:35 ``` The change replaces unknown to the mapped symbol and location to the declaration location.
-
Dhruv Chawla authored
This combine matches the existing fold in InstCombine, i.e. InstCombinerImpl::pushFreezeToPreventPoisonFromPropagating. It tries to push freeze through an operand if the operand has only one maybe-poison operand and all other operands are guaranteed non-poison, and if the operation itself cannot generate poison (eg. add with nsw can generate poison, even with non-poison operands). This is beneficial because it can potentially enable other optimizations to occur that would otherwise be blocked because of the freeze.
-