[llvm] [LV][AArch64] Support partial reductions of extended compares (PR #212190)

Adam Scott via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 00:11:22 PDT 2026


================
@@ -776,6 +776,43 @@ VPInstruction *vputils::findComputeReductionResult(VPReductionPHIRecipe *PhiR) {
       cast<VPSingleDefRecipe>(SelR));
 }
 
+/// Returns the type whose width \p Mask occupies or nullptr if unknown.
+static Type *getMaskMaterializationType(VPValue *Mask, unsigned Depth = 0) {
+  // Give up rather than recurse further. Falling back to i1 is always safe.
+  constexpr unsigned MaxMaskDepth = 3;
+  if (Depth > MaxMaskDepth)
+    return nullptr;
+
+  VPValue *A, *B;
+  if (match(Mask, m_Cmp(m_VPValue(A), m_VPValue()))) {
+    Type *CmpTy = A->getScalarType();
+    unsigned Bits = CmpTy->getPrimitiveSizeInBits();
+    // A pointer reports no size and an i1 is the width we came here to replace.
+    if (Bits <= 1)
+      return nullptr;
+    return IntegerType::get(CmpTy->getContext(), Bits);
+  }
+
+  if (match(Mask, m_Binary<Instruction::And>(m_VPValue(A), m_VPValue(B))) ||
+      match(Mask, m_Binary<Instruction::Or>(m_VPValue(A), m_VPValue(B))) ||
+      match(Mask, m_Binary<Instruction::Xor>(m_VPValue(A), m_VPValue(B)))) {
+    Type *TyA = getMaskMaterializationType(A, Depth + 1);
+    Type *TyB = getMaskMaterializationType(B, Depth + 1);
+    return TyA == TyB ? TyA : nullptr;
+  }
+
+  return nullptr;
+}
+
+Type *vputils::getExtendSrcTypeForPartialReduction(VPValue *ExtSrc) {
----------------
as4230 wrote:

My concern is that on NEON a scalar accumulator will be slower than the vector one I was lining this up for. It sucks the targets can't just say what scale factor they want, or that getScaledReductions can't emit both and let the profitability comparison pick.

With i1 the scale factor comes out as the whole VF, which makes the accumulator a single lane, so I'd expect anywhere that divides the VF down by the scale to need some treatment too. Also worth checking a one lane accumulator doesn't push it toward a fixed VF plan.

I'll have a go at the i1 idea but I'm fine with whichever you prefer.

https://github.com/llvm/llvm-project/pull/212190


More information about the llvm-commits mailing list