aboutsummaryrefslogtreecommitdiff
path: root/lldb/source/Plugins/ScriptInterpreter/Python/ScriptInterpreterPythonImpl.h
diff options
context:
space:
mode:
authorFabian Ritter <fabian.ritter@amd.com>2024-07-03 08:32:35 +0200
committerGitHub <noreply@github.com>2024-07-03 08:32:35 +0200
commite1094dd889c516da0c3181bf2be44ad631a84255 (patch)
tree807ce7a22eb731a762f1eb96769e600b4d39bbef /lldb/source/Plugins/ScriptInterpreter/Python/ScriptInterpreterPythonImpl.h
parent4e78d3a6b1560fb5debf1b518b4bd62924e900e5 (diff)
downloadllvm-e1094dd889c516da0c3181bf2be44ad631a84255.zip
llvm-e1094dd889c516da0c3181bf2be44ad631a84255.tar.gz
llvm-e1094dd889c516da0c3181bf2be44ad631a84255.tar.bz2
[AMDGPU][DAG] Enable ganging up of memcpy loads/stores for AMDGPU (#96185)
In the SelectionDAG lowering of the memcpy intrinsic, this optimization introduces additional chains between fixed-size groups of loads and the corresponding stores. While initially introduced to ensure that wider load/store-pair instructions are generated on AArch64, this optimization also improves code generation for AMDGPU: Ganged loads are scheduled into a clause; stores only await completion of their corresponding load. The chosen value of 16 performed good in microbenchmarks, values of 8, 32, or 64 would perform similarly. The testcase updates are autogenerated by utils/update_llc_test_checks.py. See also: - PR introducing this optimization: https://reviews.llvm.org/D46477 Part of SWDEV-455845.
Diffstat (limited to 'lldb/source/Plugins/ScriptInterpreter/Python/ScriptInterpreterPythonImpl.h')
0 files changed, 0 insertions, 0 deletions