[llvm] [AArch64][TTI] Allow mixed-extension partial reductions with +dotprod (PR #199762)
Sander de Smalen via llvm-commits
llvm-commits at lists.llvm.org
Mon Jun 8 00:43:43 PDT 2026
================
@@ -6100,6 +6102,10 @@ InstructionCost AArch64TTIImpl::getPartialReductionCost(
// i8 -> i32 usdot requires +i8mm
if (IsUSDot && IsSupported(ST->hasMatMulInt8(), ST->hasMatMulInt8()))
return Cost + INegCost;
+ // Without +i8mm, lower SUMLA via two udots plus an eor and a sub on plain
+ // +dotprod targets. Charge an extra factor for the expansion.
+ if (IsUSDot && IsSupported(false, ST->hasDotProd()))
----------------
sdesmalen-arm wrote:
nit: perhaps add a mention that this is only implemented for NEON because +i8mm is available on all modern cores with SVE? (that would clarify the `IsSupported(false, ` part of the condition)
https://github.com/llvm/llvm-project/pull/199762
More information about the llvm-commits
mailing list