riscv-gnu-toolchain/llvm.git - Unnamed repository; edit this file 'description' to name the repository.

diff options

author	Fabian Ritter <fabian.ritter@amd.com>	2024-07-03 08:32:35 +0200
committer	GitHub <noreply@github.com>	2024-07-03 08:32:35 +0200
commit	e1094dd889c516da0c3181bf2be44ad631a84255 (patch)
tree	807ce7a22eb731a762f1eb96769e600b4d39bbef /lldb/source/Plugins/ScriptInterpreter/Python/ScriptInterpreterPythonImpl.h
parent	4e78d3a6b1560fb5debf1b518b4bd62924e900e5 (diff)
download	llvm-e1094dd889c516da0c3181bf2be44ad631a84255.zip llvm-e1094dd889c516da0c3181bf2be44ad631a84255.tar.gz llvm-e1094dd889c516da0c3181bf2be44ad631a84255.tar.bz2

[AMDGPU][DAG] Enable ganging up of memcpy loads/stores for AMDGPU (#96185)

In the SelectionDAG lowering of the memcpy intrinsic, this optimization introduces additional chains between fixed-size groups of loads and the corresponding stores. While initially introduced to ensure that wider load/store-pair instructions are generated on AArch64, this optimization also improves code generation for AMDGPU: Ganged loads are scheduled into a clause; stores only await completion of their corresponding load. The chosen value of 16 performed good in microbenchmarks, values of 8, 32, or 64 would perform similarly. The testcase updates are autogenerated by utils/update_llc_test_checks.py. See also: - PR introducing this optimization: https://reviews.llvm.org/D46477 Part of SWDEV-455845.

Diffstat (limited to 'lldb/source/Plugins/ScriptInterpreter/Python/ScriptInterpreterPythonImpl.h')

0 files changed, 0 insertions, 0 deletions


context:
space:
mode: