[flang] add fused matmul-transpose to the runtime
This fused operation should run a lot faster than first transposing the lhs array and then multiplying the matrices separately. Based on flang/runtime/matmul.cpp Depends on D145959 Reviewed By: klausler Differential Revision: https://reviews.llvm.org/D145960
parent
a351a60e
Please register or sign in to comment