[llvm] [AArch64] Keep v16i8 -> v2i64 partial_reduce fixed-length for VL > 128 (PR #204938)
Benjamin Maxwell via llvm-commits
llvm-commits at lists.llvm.org
Mon Jun 22 03:08:07 PDT 2026
================
@@ -33517,29 +33469,42 @@ AArch64TargetLowering::LowerPARTIAL_REDUCE_MLA(SDValue Op,
SDValue DotNode = DAG.getNode(Op.getOpcode(), DL, DotVT,
DAG.getConstant(0, DL, DotVT), LHS, RHS);
- SDValue Res;
bool IsUnsigned = Op.getOpcode() == ISD::PARTIAL_REDUCE_UMLA;
- if (Subtarget->hasSVE2() || Subtarget->isStreamingSVEAvailable()) {
+
+ // UADDW{B,T}/SADDW{B,T} fold the dot in the scalable domain, spreading the
+ // sums across all VL/64 lanes. That is only valid for a genuinely scalable
+ // result; a fixed-length result must convert from the scalable container
+ // before splitting (below), else the trailing extract drops the high lanes
+ // for any VL > the fixed width.
----------------
MacDue wrote:
I don't think this is true. `UADDW{B,T}/SADDW{B,T}` add even/odd lanes. The result should be correct for both fixed/scalable vectors. It's only splitting the vectors before converting back to fixed vectors that's the issue.
https://github.com/llvm/llvm-project/pull/204938
More information about the llvm-commits
mailing list