[llvm] [AArch64] Keep v16i8 -> v2i64 partial_reduce fixed-length for VL > 128 (PR #204938)

Benjamin Maxwell via llvm-commits llvm-commits at lists.llvm.org
Mon Jun 22 03:08:07 PDT 2026


================
@@ -33517,29 +33469,42 @@ AArch64TargetLowering::LowerPARTIAL_REDUCE_MLA(SDValue Op,
   SDValue DotNode = DAG.getNode(Op.getOpcode(), DL, DotVT,
                                 DAG.getConstant(0, DL, DotVT), LHS, RHS);
 
-  SDValue Res;
   bool IsUnsigned = Op.getOpcode() == ISD::PARTIAL_REDUCE_UMLA;
-  if (Subtarget->hasSVE2() || Subtarget->isStreamingSVEAvailable()) {
+
+  // UADDW{B,T}/SADDW{B,T} fold the dot in the scalable domain, spreading the
+  // sums across all VL/64 lanes. That is only valid for a genuinely scalable
+  // result; a fixed-length result must convert from the scalable container
+  // before splitting (below), else the trailing extract drops the high lanes
+  // for any VL > the fixed width.
----------------
MacDue wrote:

I don't think this is true. `UADDW{B,T}/SADDW{B,T}` add even/odd lanes. The result should be correct for both fixed/scalable vectors. It's only splitting the vectors before converting back to fixed vectors that's the issue.

https://github.com/llvm/llvm-project/pull/204938


More information about the llvm-commits mailing list