[llvm] [AArch64] Fold four and eight way partial reductions with [SU]ADDLP (PR #214636)

Paul Walker via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 07:10:38 PDT 2026


================
@@ -13984,6 +13984,46 @@ SDValue TargetLowering::expandPartialReduceMLA(SDNode *N,
     break;
   }
 
+  // A wide partial reduction is built from a ladder of narrower ones, a rung
+  // at a time, each halving the element count and doubling the width.
+  unsigned Opc = N->getOpcode();
+  if (Opc != ISD::PARTIAL_REDUCE_FMLA &&
+      MulOpVT.getVectorMinNumElements() > 2 * AccVT.getVectorMinNumElements() &&
+      getPartialReduceMLAAction(Opc, AccVT, MulOpVT) == Custom) {
----------------
paulwalker-arm wrote:

It should not matter how we come to expand the node. My guess is this is hiding another sin?

The following code uses `widenIntegerVectorElementType`, which when combined with the `PARTIAL_REDUCE_#MLA` requirement that the mul element type must not be bigger than the accumulator type, I think you're missing a test for `AccVT.getScalarSizeInBits() / MulOpVT.getScalarSizeInBits() > 2`?

https://github.com/llvm/llvm-project/pull/214636


More information about the llvm-commits mailing list