- Nov 29, 2023
-
-
Dinar Temirbulatov authored
Thing change add builtins for SME2: sclamp.single.x2 uclamp.single.x2 fclamp.single.x2 sclamp.single.x4 uclamp.single.x4 fclamp.single.x4 Patch by: Hassnaa Hamdi <hassnaa.hamdi@arm.com> -
Louis Dionne authored
Locale objects use atomic reference counting, which may be very expensive in parallel applications. The classic locale is used by default by all streams and can be very contended. But it's never destroyed, so the reference counting is also completely pointless on the classic locale. Currently ~70% of time in the parallel stringstream benchmarks is spent in locale ctor/dtor. And the execution radically slows down with more threads. Avoid reference counting on the classic locale. With this change parallel benchmarks start to scale with threads. This is a re-application of f8afc53d (aka PR #72112) which was reverted in 4e0c48b9 because it broke the sanitizer builds due to an initialization order fiasco. This issue has now been fixed by ensuring that the locale is constinit'ed. Co-authored-by:
Dmitry Vyukov <dvyukov@google.com>
-
Louis Dionne authored
We were detecting which sanitizer flags to use when building libc++.dylib but we were never actually adding those flags to the targets, which means that our sanitized builds would basically build the dylib without any sanitizers enabled.
-
Guray Ozen authored
This PR introduce `setmaxregister.sync.aligned` Op to increase or decrease the register size. https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#miscellaneous-instructions-setmaxnreg
-
Stephan T. Lavavej authored
VSCode's Pylance extension informed me, and text searching confirmed, that these imports are unused. I believe we should be able to remove them harmlessly.
-
Andrzej Warzyński authored
This patch refactors tests for: vector.contract -> vector.outerproduct for matvec operations (b += Ax). Summary of changes: * names of LIT variables are unified, * "plain" tests (i.e. without masking and with fixed-width vectors) are moved to the top of their respective sections, * missing "plain" cases are added. This is a part of a larger effort to add cases with scalable vectors to tests for the Vector dialect. I am refactoring these tests so that it's easier to identify what cases are tested and where to add tests for scalable vectors. Implements #72834. -
Simon Pilgrim authored
Split out of getTargetConstantFromBasePtr
-
Dominik Adamski authored
Add new implementation of workshare loop functions. These functions will be used by OpenMPIRBuilder to support handling of OpenMP workshare loops for the target region. --------- Co-authored-by:Johannes Doerfert <johannes@jdoerfert.de>
-
Jan Patrick Lehr authored
Limit the use to two SDMA engines which are optimized for such transfers.
-
Jeremy Morse authored
Debugify is extremely useful as a testing and debugging tool, and a good number of LLVM-IR transform tests use it. We need it to support "new" non-instruction debug-info to get test coverage, but it's not important enough to completely convert right now (and it'd be a large undertaking). Thus: convert to/from dbg.value/DPValue mode on entry and exit of the pass, which gives us the functionality without any further work. The cost is compile-time, but again this is only happening during tests. Tested by: the large set of debugify tests enabled here. Note the InstCombine test (cast-mul-select.ll) that hasn't been fully enabled: this is because there's a debug-info sinking piece of code there that hasn't been instrumented.
-
Aaron Ballman authored
The internals manual seems like a more obvious home for the details instead of hiding them away in a header file and relying on doxygen output to document it publicly.
-
jeanPerier authored
LOC is a GNU extension, so gfortran is the reference for it, and it accepts absent OPTIONAL and returns zero. Support this use case in flang too. Update the LOC test to use HLFIR while touching it. Fixes https://github.com/llvm/llvm-project/issues/72823.
-
Pete Steinfeld authored
After merge request #73124, the flang test Driver/ctofortran started failing because both the C and the Fortran code had main programs. This update fixes that by eliminating the C main program in the test.
-
Francesco Petrogalli authored
According to the code in `SelectionDAG::getMaskedStore`, the Mask operator is in position 4, not 3: SDValue Ops[] = {Chain, Val, Base, Offset, Mask}; -
Ben Shi authored
-
Paschalis Mpeis authored
Auto-generate test `armpl-intrinsics.ll` and simplify tests: - Eliminate scalar tail with no tail-folding flag. - Use active lane mask for shorter check lines (no long `shufflevectors`). - Eliminate scalar loops by providing `noalias` to relevant arguments and run `simplifycfg` to drop them. - Update script now use `@llvm.compiler.used` instead of a longer regex.
-
Simon Pilgrim authored
Just leave the (zext (trunc (and x, c))) pattern which is still being used to create some zext_inreg patterns.
-
Simon Pilgrim authored
Thanks to @yubingex007-a11y for the original test case
-
Nikita Popov authored
-
Nikita Popov authored
-
Nikita Popov authored
It looks like this function is actually unused.
-
Nikita Popov authored
-
Aiden Grossman authored
This reverts commit 9eb80ab3. This is causing failures on multiple builders. Pulling it out until I have time to fix it.
-
Alex Bradbury authored
This adds minimal support for load clustering, but disables it by default. The intent is to iterate on the precise heuristic and the question of turning this on by default in a separate PR. Although previous discussion indicates hope that the MachineScheduler would replace most uses of the SelectionDAG scheduler, it does seem most targets aren't using MachineScheduler load clustering right now: PPC+AArch64 seem to just use it to help with paired load/store formation and although AMDGPU uses it for general clustering it also implements ShouldScheduleLoadsNear for the SelectionDAG scheduler's clustering.
-
Nikita Popov authored
-
Tim Northover authored
These are still v8.6a and have no real changes as far as LLVM cares, so it's mostly just a copy/paste job.
-
Alex Bradbury authored
As the same hook is called for both load and store clustering, NumLoads is a misleading name. Use ClusterSize instead.
-
Sandeep Kosuri authored
- Removed an unnecessary check that was preventing `nothing` to work properly inside `metadirective`.
-
Aiden Grossman authored
Before this patch, in subprocess mode, llvm-exegesis setup the stack pointer register with the rest of the registers when it was requested by the user. This would cause a segfault when the instructions to start the perf counter ran as they use the stack to preserve the three registers needed to make the syscall. This patch moves the setup of the stack register to after the configuration of the perf counter to fix this issue so that we have a valid stack pointer for all the preceeding operations. Regression test added. This fixes #72193.
-
Guillaume Chatelet authored
-
jeanPerier authored
This operation allows computing the address of descriptor fields. It is needed to help attaching descriptors in OpenMP/OpenACC target region. The pointers inside the descriptor structure must be mapped too, but the fir.box is abstract, so these fields cannot be computed with fir.coordinate_of. To preserve the abstraction of the descriptor layout in FIR, introduce an operation specifically to !fir.ref<fir.box<>> address fields based on field names (base_addr or derived_type).
-
paperchalice authored
- Sort all passes in alphabetical order. - Format `llvm/lib/Passes/PassRegistry.def`. - Remove redundant semicolon.
-
Tim Northover authored
We pretended they were v8.5a in the past because LLVM's modelling used to fold SM4 crypto support into v8.6a (which the CPUs don't actually have). That's changed in the last year so we can use the real value. This is mostly a tidy-up commit before one that'll bring in A17 and M3.
-
Dominik Adamski authored
Information about code object version can be configured by the user for AMD GPU target and it needs to be placed in LLVM IR generated by Flang. Information about code object version in MLIR generated by the parser can be reused by other tools. There is no need to specify extra flags if we want to invoke MLIR tools (like fir-opt) separately. Changes in comparison to a8ac93: * added information about required targets for test flang/test/Driver/driver-help.f90
-
Adrian Kuegel authored
This reverts commit 4b8964df.
-
Qiu Chaofan authored
-
Danny Mösch authored
-
David Green authored
This includes a couple of fixes after #71908 for bundles and some cleanup for the debug output. One was an iterator type that asserted on bundles, the second a rather subtle issue where forAllMIsUntilDef would hit the LdStLimit when renaming registers, meaning the last instruction was not updated leaving an invalid `ldp x6, x6` instruction.
-
Craig Topper authored
We have to force the register bank to FPRB if the type is s64 and the GPR is 32 bits.
-
wanglei authored
``` when a=c=-0.0, b=0.0: -(a * b + (-c)) = -0.0 -a * b + c = 0.0 (fneg (fma a, b (-c))) != (fma (fneg a), b ,c) ``` See https://reviews.llvm.org/D90901 for a similar discussion on X86.
-