[llvm] [LV][AArch64] Provide option to use partial reductions by default (PR #216001)
Benjamin Maxwell via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 06:18:09 PDT 2026
================
@@ -5230,13 +5240,56 @@ void VPlanTransforms::createPartialReductions(VPlan &Plan,
MapVector<VPReductionPHIRecipe *, SmallVector<VPPartialReductionChain>>
ChainsByPhi;
VPBasicBlock *HeaderVPBB = Plan.getVectorLoopRegion()->getEntryBasicBlock();
+ SmallVector<VPReductionPHIRecipe *, 4> UnorderedReductions;
for (VPRecipeBase &R : HeaderVPBB->phis()) {
auto *RedPhiR = dyn_cast<VPReductionPHIRecipe>(&R);
if (!RedPhiR)
continue;
if (auto Chains = getScaledReductions(RedPhiR))
ChainsByPhi.try_emplace(RedPhiR, std::move(*Chains));
+ else if (UsePartialReduceByDefault &&
+ (RedPhiR->getRecurrenceKind() == RecurKind::Add ||
+ (RedPhiR->getRecurrenceKind() == RecurKind::FAdd &&
+ !RedPhiR->isOrdered() && !RedPhiR->isInLoop())))
+ UnorderedReductions.push_back(RedPhiR);
+ }
+
+ // For general unordered reductions which aren't part of a candidate
+ // chain for a scaled partial reduction, we can potentially still use
+ // the intrinsic to allow for more optimization later on.
+ for (auto *Rdx : UnorderedReductions) {
+ auto *Backedge = dyn_cast<VPWidenRecipe>(Rdx->getBackedgeValue());
+ VPValue *OtherOp;
+ if (!Backedge ||
+ !match(Backedge,
+ m_CombineOr(m_c_FAdd(m_Specific(Rdx), m_VPValue(OtherOp)),
+ m_c_Add(m_Specific(Rdx), m_VPValue(OtherOp)))))
+ continue;
+
+ // If the target indicates that the intrinsic is as cheap as (or cheaper
+ // than) the add, then prefer the intrinsic.
+ if (!LoopVectorizationPlanner::getDecisionAndClampRange(
+ [&CostCtx, Rdx, Backedge](ElementCount VF) {
+ InstructionCost CurrentCost = Backedge->computeCost(VF, CostCtx);
+ Type *ScalarTy = Backedge->getScalarType();
+ InstructionCost PRCost = CostCtx.TTI.getPartialReductionCost(
+ Backedge->getOpcode(), ScalarTy, ScalarTy, ScalarTy, VF,
----------------
MacDue wrote:
Based on the existing use of `getPartialReductionCost()` (and docs), we only pass the `InputTypeB` when we have a binary operation. So I think this should just pass `ScalarTy` for the accumulator and `InputTypeA` (then nullptr for `InputTypeB`).
https://github.com/llvm/llvm-project/pull/216001
More information about the llvm-commits
mailing list