[llvm] [LV] Cost EVL width-adjusting zext/trunc as free (PR #225016)

Pengcheng Wang via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 20 23:57:45 PDT 2026


https://github.com/wangpc-pp created https://github.com/llvm/llvm-project/pull/225016

With EVL tail folding, the scalar zext/trunc that adjusts the i32
ExplicitVectorLength to the canonical IV type was priced as a regular
cast (cost 1). It never lowers to an instruction, only feeding the
(free) IV increment and AVL decrement, so the phantom cost dropped
borderline loops such as TSVC s351 on RV64 to VF 1.

Here we return 0 for such casts so that we don't over-estimate the
cost.

Fixes #224987

Assisted-by: TRAE CLI (Opus 4.8)


>From a49315ac5e1c6ed80a9e588815ba26d65be30245 Mon Sep 17 00:00:00 2001
From: Pengcheng Wang <wangpengcheng.pp at bytedance.com>
Date: Mon, 21 Sep 2026 12:36:35 +0800
Subject: [PATCH] [LV] Cost EVL width-adjusting zext/trunc as free

With EVL tail folding, the scalar zext/trunc that adjusts the i32
ExplicitVectorLength to the canonical IV type was priced as a regular
cast (cost 1). It never lowers to an instruction, only feeding the
(free) IV increment and AVL decrement, so the phantom cost dropped
borderline loops such as TSVC s351 on RV64 to VF 1.

Here we return 0 for such casts so that we don't over-estimate the
cost.

Fixes #224987

Assisted-by: TRAE CLI (Opus 4.8)
---
 llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp         | 10 +++++++++-
 .../Transforms/LoopVectorize/RISCV/force-vect-msg.ll   |  4 ++--
 .../LoopVectorize/RISCV/tail-folding-cost.ll           |  3 +++
 3 files changed, 14 insertions(+), 3 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index 452589bafd534..a50bb7caca503 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -1360,9 +1360,17 @@ InstructionCost VPInstruction::computeCost(ElementCount VF,
   // NOTE: At the moment it seems only possible to expose this path for
   // the trunc, zext and sext opcodes.
   // TODO: Update VF arg to use onlyFirstLaneUsed once WidenCast is unified.
-  if (Instruction::isCast(getOpcode()))
+  if (Instruction::isCast(getOpcode())) {
+    // A scalar zext/trunc that only adjusts the width of an
+    // ExplicitVectorLength to the canonical IV type is free: it feeds only
+    // the IV increment and AVL decrement, which are modeled as free below.
+    if ((getOpcode() == Instruction::ZExt ||
+         getOpcode() == Instruction::Trunc) &&
+        match(getOperand(0), m_EVL(m_VPValue())))
+      return 0;
     return getCostForRecipeWithOpcode(getOpcode(), ElementCount::getFixed(1),
                                       Ctx);
+  }
 
   if (Instruction::isBinaryOp(getOpcode())) {
     if (!getUnderlyingValue() && getOpcode() != Instruction::FMul) {
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/force-vect-msg.ll b/llvm/test/Transforms/LoopVectorize/RISCV/force-vect-msg.ll
index 22560793e4504..4292317345dd4 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/force-vect-msg.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/force-vect-msg.ll
@@ -3,8 +3,8 @@
 
 ; CHECK: LV: Loop hints: force=enabled
 ; CHECK: LV: Scalar loop costs: 4.
-; ChosenFactor.Cost is 11, but the real cost will be divided by the width, which is 2.8
-; CHECK: Cost for VF vscale x 2: 9
+; ChosenFactor.Cost is 7, but the real cost will be divided by the width, which is 1.75
+; CHECK: Cost for VF vscale x 2: 7
 ; Regardless of force vectorization or not, this loop will eventually be vectorized because of the cost model.
 ; Therefore, the following message does not need to be printed even if vectorization is explicitly forced in the metadata.
 ; CHECK-NOT: LV: Vectorization seems to be not beneficial, but was forced by a user.
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/tail-folding-cost.ll b/llvm/test/Transforms/LoopVectorize/RISCV/tail-folding-cost.ll
index 4bd2072a6d680..bb818dc4e988d 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/tail-folding-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/tail-folding-cost.ll
@@ -14,8 +14,11 @@
 ; DATA: Cost of 8 for VF vscale x 4: EMIT{{.*}} = active lane mask
 
 ; EVL: Cost of 1 for VF vscale x 1: EMIT{{.*}} = EXPLICIT-VECTOR-LENGTH
+; EVL: Cost of 0 for VF vscale x 1: EMIT-SCALAR vp<{{.*}}> = zext vp<%evl> to i64
 ; EVL: Cost of 1 for VF vscale x 2: EMIT{{.*}} = EXPLICIT-VECTOR-LENGTH
+; EVL: Cost of 0 for VF vscale x 2: EMIT-SCALAR vp<{{.*}}> = zext vp<%evl> to i64
 ; EVL: Cost of 1 for VF vscale x 4: EMIT{{.*}} = EXPLICIT-VECTOR-LENGTH
+; EVL: Cost of 0 for VF vscale x 4: EMIT-SCALAR vp<{{.*}}> = zext vp<%evl> to i64
 
 define void @simple_memset(i32 %val, ptr %ptr, i64 %n) #0 {
 entry:



More information about the llvm-commits mailing list