[llvm] [VPlan] Narrow interleave groups with factor > VF to one wide op. (PR #228252)

David Sherwood via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 01:54:27 PDT 2026


================
@@ -4490,6 +4538,30 @@ VPlanTransforms::narrowInterleaveGroups(VPlan &Plan,
   if (MiddleVPBB->getNumSuccessors() != 2 && !RequiresScalarEpilogue)
     return nullptr;
 
+  // WideVF is zero if the narrowed recipes stay at the plan's VF.
+  unsigned Factor = AllGroups.front()->getInterleaveGroup()->getFactor();
+  unsigned NumIters = VFToOptimize->getKnownMinValue();
+  ElementCount WideVF =
+      Factor > NumIters ? ElementCount::getFixed(Factor) : ElementCount();
+
+  // The wide ops span multiple registers and process a single iteration, so
+  // only narrow if that is strictly cheaper than the interleave groups.
+  if (!WideVF.isZero()) {
+    InstructionCost InterleaveCost = 0;
+    InstructionCost NarrowedCost = 0;
+    for (VPInterleaveRecipe *G : AllGroups) {
+      InterleaveCost += G->cost(*VFToOptimize, Ctx);
+      Instruction *InsertPos = G->getInsertPos();
+      NarrowedCost += Ctx.TTI.getMemoryOpCost(
----------------
david-arm wrote:

This seems sensible, but I'd better go off and make sure that we actually have sensible costs for all non-power-of-2 NEON/SVE loads and stores. :)

https://github.com/llvm/llvm-project/pull/228252


More information about the llvm-commits mailing list