[llvm] [LoopVectorize] Don't assert in getVectorCallCost for vector library variants (PR #202085)

via llvm-commits llvm-commits at lists.llvm.org
Tue Jun 23 10:03:49 PDT 2026


https://github.com/seantalts updated https://github.com/llvm/llvm-project/pull/202085

>From c33e3091de1413065e9b53c05077ad848cc58368 Mon Sep 17 00:00:00 2001
From: Sean Talts <sean.talts at gmail.com>
Date: Fri, 5 Jun 2026 18:32:33 +0000
Subject: [PATCH 1/5] [LoopVectorize] Don't assert in getVectorCallCost for
 vector library variants

During loop vectorization, computePredInstDiscount queries the cost of
instructions at vector VF using getInstructionCost. If an instruction is a
CallInst that has a vector library variant, getInstructionCost delegates to
getVectorCallCost, which asserted that vector library variants should not reach
it.

Such a call can reach getVectorCallCost via computePredInstDiscount, before the
call's widening decision is made, when a predicated user (e.g. a scatter store)
is being considered for scalarization. Remove the assert and fall through to the
existing scalarization cost, which is the cost relevant to that analysis.
---
 llvm/lib/Transforms/Vectorize/LoopVectorize.cpp | 11 +++++------
 1 file changed, 5 insertions(+), 6 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 7ad454d4a1797..5e7bbd71b5260 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -2116,12 +2116,11 @@ static bool hasVectorLibraryVariantFor(const CallInst &CI, ElementCount VF,
 InstructionCost
 LoopVectorizationCostModel::getVectorCallCost(CallInst *CI,
                                               ElementCount VF) const {
-  // Vector library variants are priced by VPWidenCallRecipe::computeCost and
-  // should not reach this function.
-  assert((VF.isScalar() ||
-          !hasVectorLibraryVariantFor(*CI, VF, isMaskRequired(CI), TLI)) &&
-         "getVectorCallCost does not price vector library variants");
-
+  // A call with a vector library variant is normally priced by
+  // VPWidenCallRecipe::computeCost. It can still reach here via
+  // computePredInstDiscount, which queries the cost before the call's widening
+  // decision is made; in that case the predicated call is being considered for
+  // scalarization, so fall through to the scalarization cost below.
   Type *RetTy = CI->getType();
   SmallVector<Type *, 4> Tys;
   for (auto &ArgOp : CI->args())

>From cde5808f0d7ead29d5ae11f3a34d6e26ff698a16 Mon Sep 17 00:00:00 2001
From: Sean Talts <sean.talts at gmail.com>
Date: Fri, 12 Jun 2026 08:07:20 -0400
Subject: [PATCH 2/5] [LoopVectorize] Add test for predication-discount cost of
 a library-variant call

When a conditionally-executed call has a vector library variant matching the
VF, the call is a widen-with-mask candidate rather than scalar-with-predication.
If its result feeds an instruction that is scalar-with-predication (here a
scatter store), computePredInstDiscount walks the operand chain and queries the
call's cost at the vector VF via getVectorCallCost.

This used to hit an assert ("getVectorCallCost does not price vector library
variants"); getVectorCallCost now falls through to the scalarization cost. The
test exercises that path and checks the loop vectorizes without crashing.
---
 .../pred-inst-discount-vector-library-call.ll | 114 ++++++++++++++++++
 1 file changed, 114 insertions(+)
 create mode 100644 llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll

diff --git a/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll b/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll
new file mode 100644
index 0000000000000..5acf7a3a95102
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll
@@ -0,0 +1,114 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt < %s -passes=loop-vectorize -force-vector-interleave=1 -force-vector-width=2 -S 2>&1 | FileCheck %s
+
+; Regression test for the cost model. A conditionally-executed call has a vector
+; library variant matching the VF, so the call itself is a widen-with-mask
+; candidate (not scalar-with-predication). Its result feeds a scatter store that
+; *is* scalar-with-predication. While computing the predication discount for that
+; store, computePredInstDiscount walks the operand chain and queries the cost of
+; the call at the vector VF via getVectorCallCost. This used to hit an assert
+; ("getVectorCallCost does not price vector library variants"); getVectorCallCost
+; now falls through to the scalarization cost instead of asserting.
+
+target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
+
+define void @pred_call_with_variant(ptr readonly %src, ptr noalias %dest, i64 %N) {
+; CHECK-LABEL: define void @pred_call_with_variant(
+; CHECK-SAME: ptr readonly [[SRC:%.*]], ptr noalias [[DEST:%.*]], i64 [[N:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*]]:
+; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 2
+; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], 2
+; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE2:.*]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <2 x i64>, ptr [[TMP0]], align 8
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp ult <2 x i64> [[WIDE_LOAD]], splat (i64 5)
+; CHECK-NEXT:    [[TMP2:%.*]] = extractelement <2 x i1> [[TMP1]], i64 0
+; CHECK-NEXT:    br i1 [[TMP2]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
+; CHECK:       [[PRED_STORE_IF]]:
+; CHECK-NEXT:    [[TMP3:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 0
+; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @foo(i64 [[TMP3]]) #[[ATTR0:[0-9]+]]
+; CHECK-NEXT:    [[TMP5:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP3]]
+; CHECK-NEXT:    store i64 [[TMP4]], ptr [[TMP5]], align 8
+; CHECK-NEXT:    br label %[[PRED_STORE_CONTINUE]]
+; CHECK:       [[PRED_STORE_CONTINUE]]:
+; CHECK-NEXT:    [[TMP6:%.*]] = extractelement <2 x i1> [[TMP1]], i64 1
+; CHECK-NEXT:    br i1 [[TMP6]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2]]
+; CHECK:       [[PRED_STORE_IF1]]:
+; CHECK-NEXT:    [[TMP7:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 1
+; CHECK-NEXT:    [[TMP8:%.*]] = call i64 @foo(i64 [[TMP7]]) #[[ATTR0]]
+; CHECK-NEXT:    [[TMP9:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP7]]
+; CHECK-NEXT:    store i64 [[TMP8]], ptr [[TMP9]], align 8
+; CHECK-NEXT:    br label %[[PRED_STORE_CONTINUE2]]
+; CHECK:       [[PRED_STORE_CONTINUE2]]:
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 2
+; CHECK-NEXT:    [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[CMP_N]], label %[[END:.*]], label %[[SCALAR_PH]]
+; CHECK:       [[SCALAR_PH]]:
+; CHECK-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
+; CHECK:       [[FOR_BODY]]:
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_LOOP:.*]] ]
+; CHECK-NEXT:    [[LD_ADDR:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[IV]]
+; CHECK-NEXT:    [[IDX:%.*]] = load i64, ptr [[LD_ADDR]], align 8
+; CHECK-NEXT:    [[IFCOND:%.*]] = icmp ult i64 [[IDX]], 5
+; CHECK-NEXT:    br i1 [[IFCOND]], label %[[IF_THEN:.*]], label %[[FOR_LOOP]]
+; CHECK:       [[IF_THEN]]:
+; CHECK-NEXT:    [[FOO_RET:%.*]] = call i64 @foo(i64 [[IDX]]) #[[ATTR0]]
+; CHECK-NEXT:    [[ST_ADDR:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[IDX]]
+; CHECK-NEXT:    store i64 [[FOO_RET]], ptr [[ST_ADDR]], align 8
+; CHECK-NEXT:    br label %[[FOR_LOOP]]
+; CHECK:       [[FOR_LOOP]]:
+; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT:    [[LOOPCOND:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; CHECK-NEXT:    br i1 [[LOOPCOND]], label %[[END]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK:       [[END]]:
+; CHECK-NEXT:    ret void
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.loop ]
+  %ld.addr = getelementptr inbounds i64, ptr %src, i64 %iv
+  %idx = load i64, ptr %ld.addr, align 8
+  %ifcond = icmp ult i64 %idx, 5
+  br i1 %ifcond, label %if.then, label %for.loop
+
+if.then:
+  %foo.ret = call i64 @foo(i64 %idx) #0
+  ; A scatter to a loaded index: not consecutive, so the store is
+  ; scalar-with-predication and triggers the predication-discount analysis.
+  %st.addr = getelementptr inbounds i64, ptr %dest, i64 %idx
+  store i64 %foo.ret, ptr %st.addr, align 8
+  br label %for.loop
+
+for.loop:
+  %iv.next = add nsw nuw i64 %iv, 1
+  %loopcond = icmp eq i64 %iv.next, %N
+  br i1 %loopcond, label %end, label %for.body
+
+end:
+  ret void
+}
+
+declare i64 @foo(i64) #0
+declare <2 x i64> @vector_foo(<2 x i64>, <2 x i1>)
+
+; A masked vector variant matching the forced VF of 2, so the call is a valid
+; widen-with-mask candidate (not itself scalar-with-predication) while feeding
+; the scalar-with-predication scatter store.
+attributes #0 = { readonly nounwind "vector-function-abi-variant"="_ZGV_LLVM_M2v_foo(vector_foo)" }
+;.
+; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
+; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
+; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
+; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META2]], [[META1]]}
+;.

>From 75c6d8816d2be679fc5cdb8a6c10e9182aeb1479 Mon Sep 17 00:00:00 2001
From: Sean Talts <sean.talts at gmail.com>
Date: Thu, 18 Jun 2026 13:04:07 -0400
Subject: [PATCH 3/5] [LoopVectorize] Return cheapest call lowering and extend
 test

Address review feedback. getVectorCallCost now returns the minimum of the
scalarization, vector-intrinsic, and vector-library-variant costs, so
computePredInstDiscount compares against the most profitable lowering of the
call at the VF (rather than only scalarization).

Extend the test to show the cost driving the decision: at VF=2 the call is
scalarized, while at VF=8 the wide variant is cheaper so the call stays wide
and only the scatter store is scalarized. Generalize the comment, drop the
unneeded datalayout, and regenerate with --check-globals=none.
---
 .../Transforms/Vectorize/LoopVectorize.cpp    |  52 ++--
 .../pred-inst-discount-vector-library-call.ll | 269 ++++++++++++------
 2 files changed, 221 insertions(+), 100 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 96953c096c9db..9268d8a0b8cf9 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -2098,29 +2098,32 @@ static unsigned estimateElementCount(ElementCount VF,
   return EstimatedVF;
 }
 
-/// Returns true iff \p CI has a library vector variant usable at \p VF: a
-/// mapping with matching VF, masked if required, whose vector function is
-/// declared in the module. Such variants are priced by
-/// VPWidenCallRecipe::computeCost rather than by scalarization.
+/// Returns the vector library variant function of \p CI usable at \p VF,
+/// respecting \p MaskRequired, or nullptr if none is found: a mapping with
+/// matching VF, masked if required, whose vector function is declared in the
+/// module.
+static Function *getVectorLibraryVariantFor(const CallInst &CI, ElementCount VF,
+                                            bool MaskRequired,
+                                            const TargetLibraryInfo *TLI) {
+  if (!TLI || CI.isNoBuiltin())
+    return nullptr;
+  for (const VFInfo &Info : VFDatabase::getMappings(CI))
+    if (Info.Shape.VF == VF && (!MaskRequired || Info.isMasked()))
+      if (Function *F = CI.getModule()->getFunction(Info.VectorName))
+        return F;
+  return nullptr;
+}
+
+/// Returns true iff \p CI has a library vector variant usable at \p VF.
 static bool hasVectorLibraryVariantFor(const CallInst &CI, ElementCount VF,
                                        bool MaskRequired,
                                        const TargetLibraryInfo *TLI) {
-  if (!TLI || CI.isNoBuiltin())
-    return false;
-  return any_of(VFDatabase::getMappings(CI), [&](const VFInfo &Info) {
-    return Info.Shape.VF == VF && (!MaskRequired || Info.isMasked()) &&
-           CI.getModule()->getFunction(Info.VectorName);
-  });
+  return getVectorLibraryVariantFor(CI, VF, MaskRequired, TLI) != nullptr;
 }
 
 InstructionCost
 LoopVectorizationCostModel::getVectorCallCost(CallInst *CI,
                                               ElementCount VF) const {
-  // A call with a vector library variant is normally priced by
-  // VPWidenCallRecipe::computeCost. It can still reach here via
-  // computePredInstDiscount, which queries the cost before the call's widening
-  // decision is made; in that case the predicated call is being considered for
-  // scalarization, so fall through to the scalarization cost below.
   Type *RetTy = CI->getType();
   SmallVector<Type *, 4> Tys;
   for (auto &ArgOp : CI->args())
@@ -2136,10 +2139,21 @@ LoopVectorizationCostModel::getVectorCallCost(CallInst *CI,
                              : ScalarCallCost * VF.getKnownMinValue() +
                                    getScalarizationOverhead(CI, VF);
 
-  if (getVectorIntrinsicIDForCall(CI, TLI)) {
-    InstructionCost IntrinsicCost = getVectorIntrinsicCost(CI, VF);
-    return std::min(Cost, IntrinsicCost);
-  }
+  // The call may also have a non-scalarized lowering at this VF, via a vector
+  // intrinsic or a vector library variant. computePredInstDiscount queries this
+  // cost (before the call's widening decision is made) to decide whether
+  // forcing a predicated tree of operations to scalar is profitable, so return
+  // the cheapest available lowering.
+  if (getVectorIntrinsicIDForCall(CI, TLI))
+    Cost = std::min(Cost, getVectorIntrinsicCost(CI, VF));
+
+  if (Function *Variant =
+          getVectorLibraryVariantFor(*CI, VF, isMaskRequired(CI), TLI))
+    Cost = std::min(Cost, TTI.getCallInstrCost(
+                              /*F=*/nullptr, Variant->getReturnType(),
+                              Variant->getFunctionType()->params(),
+                              Config.CostKind));
+
   return Cost;
 }
 
diff --git a/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll b/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll
index 5acf7a3a95102..1d7f96c35949e 100644
--- a/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll
+++ b/llvm/test/Transforms/LoopVectorize/pred-inst-discount-vector-library-call.ll
@@ -1,76 +1,189 @@
-; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
-; RUN: opt < %s -passes=loop-vectorize -force-vector-interleave=1 -force-vector-width=2 -S 2>&1 | FileCheck %s
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --version 6
+; RUN: opt < %s -passes=loop-vectorize -force-vector-interleave=1 -force-vector-width=2 -S | FileCheck %s --check-prefixes=VF2
+; RUN: opt < %s -passes=loop-vectorize -force-vector-interleave=1 -force-vector-width=8 -S | FileCheck %s --check-prefixes=VF8
 
-; Regression test for the cost model. A conditionally-executed call has a vector
-; library variant matching the VF, so the call itself is a widen-with-mask
-; candidate (not scalar-with-predication). Its result feeds a scatter store that
-; *is* scalar-with-predication. While computing the predication discount for that
-; store, computePredInstDiscount walks the operand chain and queries the cost of
-; the call at the vector VF via getVectorCallCost. This used to hit an assert
-; ("getVectorCallCost does not price vector library variants"); getVectorCallCost
-; now falls through to the scalarization cost instead of asserting.
-
-target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
+; Check that a conditionally-executed call with a vector library variant is
+; costed correctly when deciding whether scalarizing a predicated tree of
+; operations is profitable. Its result feeds a scatter store that must be
+; scalarized; the cost model uses the wide-call cost to decide whether to also
+; scalarize the call. At VF=2 scalarizing the call is cheapest, so it is
+; scalarized; at VF=8 the wide variant is cheaper, so the call stays wide and
+; only the store is scalarized. Querying the wide-call cost on this path
+; previously crashed.
 
 define void @pred_call_with_variant(ptr readonly %src, ptr noalias %dest, i64 %N) {
-; CHECK-LABEL: define void @pred_call_with_variant(
-; CHECK-SAME: ptr readonly [[SRC:%.*]], ptr noalias [[DEST:%.*]], i64 [[N:%.*]]) {
-; CHECK-NEXT:  [[ENTRY:.*]]:
-; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 2
-; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
-; CHECK:       [[VECTOR_PH]]:
-; CHECK-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], 2
-; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
-; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK:       [[VECTOR_BODY]]:
-; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE2:.*]] ]
-; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <2 x i64>, ptr [[TMP0]], align 8
-; CHECK-NEXT:    [[TMP1:%.*]] = icmp ult <2 x i64> [[WIDE_LOAD]], splat (i64 5)
-; CHECK-NEXT:    [[TMP2:%.*]] = extractelement <2 x i1> [[TMP1]], i64 0
-; CHECK-NEXT:    br i1 [[TMP2]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
-; CHECK:       [[PRED_STORE_IF]]:
-; CHECK-NEXT:    [[TMP3:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 0
-; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @foo(i64 [[TMP3]]) #[[ATTR0:[0-9]+]]
-; CHECK-NEXT:    [[TMP5:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP3]]
-; CHECK-NEXT:    store i64 [[TMP4]], ptr [[TMP5]], align 8
-; CHECK-NEXT:    br label %[[PRED_STORE_CONTINUE]]
-; CHECK:       [[PRED_STORE_CONTINUE]]:
-; CHECK-NEXT:    [[TMP6:%.*]] = extractelement <2 x i1> [[TMP1]], i64 1
-; CHECK-NEXT:    br i1 [[TMP6]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2]]
-; CHECK:       [[PRED_STORE_IF1]]:
-; CHECK-NEXT:    [[TMP7:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 1
-; CHECK-NEXT:    [[TMP8:%.*]] = call i64 @foo(i64 [[TMP7]]) #[[ATTR0]]
-; CHECK-NEXT:    [[TMP9:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP7]]
-; CHECK-NEXT:    store i64 [[TMP8]], ptr [[TMP9]], align 8
-; CHECK-NEXT:    br label %[[PRED_STORE_CONTINUE2]]
-; CHECK:       [[PRED_STORE_CONTINUE2]]:
-; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 2
-; CHECK-NEXT:    [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-NEXT:    br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
-; CHECK:       [[MIDDLE_BLOCK]]:
-; CHECK-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
-; CHECK-NEXT:    br i1 [[CMP_N]], label %[[END:.*]], label %[[SCALAR_PH]]
-; CHECK:       [[SCALAR_PH]]:
-; CHECK-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
-; CHECK:       [[FOR_BODY]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_LOOP:.*]] ]
-; CHECK-NEXT:    [[LD_ADDR:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[IV]]
-; CHECK-NEXT:    [[IDX:%.*]] = load i64, ptr [[LD_ADDR]], align 8
-; CHECK-NEXT:    [[IFCOND:%.*]] = icmp ult i64 [[IDX]], 5
-; CHECK-NEXT:    br i1 [[IFCOND]], label %[[IF_THEN:.*]], label %[[FOR_LOOP]]
-; CHECK:       [[IF_THEN]]:
-; CHECK-NEXT:    [[FOO_RET:%.*]] = call i64 @foo(i64 [[IDX]]) #[[ATTR0]]
-; CHECK-NEXT:    [[ST_ADDR:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[IDX]]
-; CHECK-NEXT:    store i64 [[FOO_RET]], ptr [[ST_ADDR]], align 8
-; CHECK-NEXT:    br label %[[FOR_LOOP]]
-; CHECK:       [[FOR_LOOP]]:
-; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-NEXT:    [[LOOPCOND:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-NEXT:    br i1 [[LOOPCOND]], label %[[END]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
-; CHECK:       [[END]]:
-; CHECK-NEXT:    ret void
+; VF2-LABEL: define void @pred_call_with_variant(
+; VF2-SAME: ptr readonly [[SRC:%.*]], ptr noalias [[DEST:%.*]], i64 [[N:%.*]]) {
+; VF2-NEXT:  [[ENTRY:.*]]:
+; VF2-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 2
+; VF2-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; VF2:       [[VECTOR_PH]]:
+; VF2-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], 2
+; VF2-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; VF2-NEXT:    br label %[[VECTOR_BODY:.*]]
+; VF2:       [[VECTOR_BODY]]:
+; VF2-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE2:.*]] ]
+; VF2-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[INDEX]]
+; VF2-NEXT:    [[WIDE_LOAD:%.*]] = load <2 x i64>, ptr [[TMP0]], align 8
+; VF2-NEXT:    [[TMP1:%.*]] = icmp ult <2 x i64> [[WIDE_LOAD]], splat (i64 5)
+; VF2-NEXT:    [[TMP2:%.*]] = extractelement <2 x i1> [[TMP1]], i64 0
+; VF2-NEXT:    br i1 [[TMP2]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
+; VF2:       [[PRED_STORE_IF]]:
+; VF2-NEXT:    [[TMP3:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 0
+; VF2-NEXT:    [[TMP4:%.*]] = call i64 @foo(i64 [[TMP3]]) #[[ATTR0:[0-9]+]]
+; VF2-NEXT:    [[TMP5:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP3]]
+; VF2-NEXT:    store i64 [[TMP4]], ptr [[TMP5]], align 8
+; VF2-NEXT:    br label %[[PRED_STORE_CONTINUE]]
+; VF2:       [[PRED_STORE_CONTINUE]]:
+; VF2-NEXT:    [[TMP6:%.*]] = extractelement <2 x i1> [[TMP1]], i64 1
+; VF2-NEXT:    br i1 [[TMP6]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2]]
+; VF2:       [[PRED_STORE_IF1]]:
+; VF2-NEXT:    [[TMP7:%.*]] = extractelement <2 x i64> [[WIDE_LOAD]], i64 1
+; VF2-NEXT:    [[TMP8:%.*]] = call i64 @foo(i64 [[TMP7]]) #[[ATTR0]]
+; VF2-NEXT:    [[TMP9:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP7]]
+; VF2-NEXT:    store i64 [[TMP8]], ptr [[TMP9]], align 8
+; VF2-NEXT:    br label %[[PRED_STORE_CONTINUE2]]
+; VF2:       [[PRED_STORE_CONTINUE2]]:
+; VF2-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 2
+; VF2-NEXT:    [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; VF2-NEXT:    br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; VF2:       [[MIDDLE_BLOCK]]:
+; VF2-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; VF2-NEXT:    br i1 [[CMP_N]], label %[[END:.*]], label %[[SCALAR_PH]]
+; VF2:       [[SCALAR_PH]]:
+; VF2-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; VF2-NEXT:    br label %[[FOR_BODY:.*]]
+; VF2:       [[FOR_BODY]]:
+; VF2-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_LOOP:.*]] ]
+; VF2-NEXT:    [[LD_ADDR:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[IV]]
+; VF2-NEXT:    [[IDX:%.*]] = load i64, ptr [[LD_ADDR]], align 8
+; VF2-NEXT:    [[IFCOND:%.*]] = icmp ult i64 [[IDX]], 5
+; VF2-NEXT:    br i1 [[IFCOND]], label %[[IF_THEN:.*]], label %[[FOR_LOOP]]
+; VF2:       [[IF_THEN]]:
+; VF2-NEXT:    [[FOO_RET:%.*]] = call i64 @foo(i64 [[IDX]]) #[[ATTR0]]
+; VF2-NEXT:    [[ST_ADDR:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[IDX]]
+; VF2-NEXT:    store i64 [[FOO_RET]], ptr [[ST_ADDR]], align 8
+; VF2-NEXT:    br label %[[FOR_LOOP]]
+; VF2:       [[FOR_LOOP]]:
+; VF2-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; VF2-NEXT:    [[LOOPCOND:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; VF2-NEXT:    br i1 [[LOOPCOND]], label %[[END]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; VF2:       [[END]]:
+; VF2-NEXT:    ret void
+;
+; VF8-LABEL: define void @pred_call_with_variant(
+; VF8-SAME: ptr readonly [[SRC:%.*]], ptr noalias [[DEST:%.*]], i64 [[N:%.*]]) {
+; VF8-NEXT:  [[ENTRY:.*]]:
+; VF8-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; VF8-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; VF8:       [[VECTOR_PH]]:
+; VF8-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], 8
+; VF8-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; VF8-NEXT:    br label %[[VECTOR_BODY:.*]]
+; VF8:       [[VECTOR_BODY]]:
+; VF8-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE14:.*]] ]
+; VF8-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[INDEX]]
+; VF8-NEXT:    [[WIDE_LOAD:%.*]] = load <8 x i64>, ptr [[TMP0]], align 8
+; VF8-NEXT:    [[TMP1:%.*]] = icmp ult <8 x i64> [[WIDE_LOAD]], splat (i64 5)
+; VF8-NEXT:    [[TMP2:%.*]] = call <8 x i64> @vector_foo_8(<8 x i64> [[WIDE_LOAD]], <8 x i1> [[TMP1]])
+; VF8-NEXT:    [[TMP3:%.*]] = extractelement <8 x i1> [[TMP1]], i64 0
+; VF8-NEXT:    br i1 [[TMP3]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
+; VF8:       [[PRED_STORE_IF]]:
+; VF8-NEXT:    [[TMP4:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 0
+; VF8-NEXT:    [[TMP5:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP4]]
+; VF8-NEXT:    [[TMP6:%.*]] = extractelement <8 x i64> [[TMP2]], i64 0
+; VF8-NEXT:    store i64 [[TMP6]], ptr [[TMP5]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE]]
+; VF8:       [[PRED_STORE_CONTINUE]]:
+; VF8-NEXT:    [[TMP7:%.*]] = extractelement <8 x i1> [[TMP1]], i64 1
+; VF8-NEXT:    br i1 [[TMP7]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2:.*]]
+; VF8:       [[PRED_STORE_IF1]]:
+; VF8-NEXT:    [[TMP8:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 1
+; VF8-NEXT:    [[TMP9:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP8]]
+; VF8-NEXT:    [[TMP10:%.*]] = extractelement <8 x i64> [[TMP2]], i64 1
+; VF8-NEXT:    store i64 [[TMP10]], ptr [[TMP9]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE2]]
+; VF8:       [[PRED_STORE_CONTINUE2]]:
+; VF8-NEXT:    [[TMP11:%.*]] = extractelement <8 x i1> [[TMP1]], i64 2
+; VF8-NEXT:    br i1 [[TMP11]], label %[[PRED_STORE_IF3:.*]], label %[[PRED_STORE_CONTINUE4:.*]]
+; VF8:       [[PRED_STORE_IF3]]:
+; VF8-NEXT:    [[TMP12:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 2
+; VF8-NEXT:    [[TMP13:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP12]]
+; VF8-NEXT:    [[TMP14:%.*]] = extractelement <8 x i64> [[TMP2]], i64 2
+; VF8-NEXT:    store i64 [[TMP14]], ptr [[TMP13]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE4]]
+; VF8:       [[PRED_STORE_CONTINUE4]]:
+; VF8-NEXT:    [[TMP15:%.*]] = extractelement <8 x i1> [[TMP1]], i64 3
+; VF8-NEXT:    br i1 [[TMP15]], label %[[PRED_STORE_IF5:.*]], label %[[PRED_STORE_CONTINUE6:.*]]
+; VF8:       [[PRED_STORE_IF5]]:
+; VF8-NEXT:    [[TMP16:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 3
+; VF8-NEXT:    [[TMP17:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP16]]
+; VF8-NEXT:    [[TMP18:%.*]] = extractelement <8 x i64> [[TMP2]], i64 3
+; VF8-NEXT:    store i64 [[TMP18]], ptr [[TMP17]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE6]]
+; VF8:       [[PRED_STORE_CONTINUE6]]:
+; VF8-NEXT:    [[TMP19:%.*]] = extractelement <8 x i1> [[TMP1]], i64 4
+; VF8-NEXT:    br i1 [[TMP19]], label %[[PRED_STORE_IF7:.*]], label %[[PRED_STORE_CONTINUE8:.*]]
+; VF8:       [[PRED_STORE_IF7]]:
+; VF8-NEXT:    [[TMP20:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 4
+; VF8-NEXT:    [[TMP21:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP20]]
+; VF8-NEXT:    [[TMP22:%.*]] = extractelement <8 x i64> [[TMP2]], i64 4
+; VF8-NEXT:    store i64 [[TMP22]], ptr [[TMP21]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE8]]
+; VF8:       [[PRED_STORE_CONTINUE8]]:
+; VF8-NEXT:    [[TMP23:%.*]] = extractelement <8 x i1> [[TMP1]], i64 5
+; VF8-NEXT:    br i1 [[TMP23]], label %[[PRED_STORE_IF9:.*]], label %[[PRED_STORE_CONTINUE10:.*]]
+; VF8:       [[PRED_STORE_IF9]]:
+; VF8-NEXT:    [[TMP24:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 5
+; VF8-NEXT:    [[TMP25:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP24]]
+; VF8-NEXT:    [[TMP26:%.*]] = extractelement <8 x i64> [[TMP2]], i64 5
+; VF8-NEXT:    store i64 [[TMP26]], ptr [[TMP25]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE10]]
+; VF8:       [[PRED_STORE_CONTINUE10]]:
+; VF8-NEXT:    [[TMP27:%.*]] = extractelement <8 x i1> [[TMP1]], i64 6
+; VF8-NEXT:    br i1 [[TMP27]], label %[[PRED_STORE_IF11:.*]], label %[[PRED_STORE_CONTINUE12:.*]]
+; VF8:       [[PRED_STORE_IF11]]:
+; VF8-NEXT:    [[TMP28:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 6
+; VF8-NEXT:    [[TMP29:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP28]]
+; VF8-NEXT:    [[TMP30:%.*]] = extractelement <8 x i64> [[TMP2]], i64 6
+; VF8-NEXT:    store i64 [[TMP30]], ptr [[TMP29]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE12]]
+; VF8:       [[PRED_STORE_CONTINUE12]]:
+; VF8-NEXT:    [[TMP31:%.*]] = extractelement <8 x i1> [[TMP1]], i64 7
+; VF8-NEXT:    br i1 [[TMP31]], label %[[PRED_STORE_IF13:.*]], label %[[PRED_STORE_CONTINUE14]]
+; VF8:       [[PRED_STORE_IF13]]:
+; VF8-NEXT:    [[TMP32:%.*]] = extractelement <8 x i64> [[WIDE_LOAD]], i64 7
+; VF8-NEXT:    [[TMP33:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[TMP32]]
+; VF8-NEXT:    [[TMP34:%.*]] = extractelement <8 x i64> [[TMP2]], i64 7
+; VF8-NEXT:    store i64 [[TMP34]], ptr [[TMP33]], align 8
+; VF8-NEXT:    br label %[[PRED_STORE_CONTINUE14]]
+; VF8:       [[PRED_STORE_CONTINUE14]]:
+; VF8-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; VF8-NEXT:    [[TMP35:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; VF8-NEXT:    br i1 [[TMP35]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; VF8:       [[MIDDLE_BLOCK]]:
+; VF8-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; VF8-NEXT:    br i1 [[CMP_N]], label %[[END:.*]], label %[[SCALAR_PH]]
+; VF8:       [[SCALAR_PH]]:
+; VF8-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; VF8-NEXT:    br label %[[FOR_BODY:.*]]
+; VF8:       [[FOR_BODY]]:
+; VF8-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_LOOP:.*]] ]
+; VF8-NEXT:    [[LD_ADDR:%.*]] = getelementptr inbounds i64, ptr [[SRC]], i64 [[IV]]
+; VF8-NEXT:    [[IDX:%.*]] = load i64, ptr [[LD_ADDR]], align 8
+; VF8-NEXT:    [[IFCOND:%.*]] = icmp ult i64 [[IDX]], 5
+; VF8-NEXT:    br i1 [[IFCOND]], label %[[IF_THEN:.*]], label %[[FOR_LOOP]]
+; VF8:       [[IF_THEN]]:
+; VF8-NEXT:    [[FOO_RET:%.*]] = call i64 @foo(i64 [[IDX]]) #[[ATTR0:[0-9]+]]
+; VF8-NEXT:    [[ST_ADDR:%.*]] = getelementptr inbounds i64, ptr [[DEST]], i64 [[IDX]]
+; VF8-NEXT:    store i64 [[FOO_RET]], ptr [[ST_ADDR]], align 8
+; VF8-NEXT:    br label %[[FOR_LOOP]]
+; VF8:       [[FOR_LOOP]]:
+; VF8-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; VF8-NEXT:    [[LOOPCOND:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; VF8-NEXT:    br i1 [[LOOPCOND]], label %[[END]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; VF8:       [[END]]:
+; VF8-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -100,15 +213,9 @@ end:
 }
 
 declare i64 @foo(i64) #0
-declare <2 x i64> @vector_foo(<2 x i64>, <2 x i1>)
+declare <2 x i64> @vector_foo_2(<2 x i64>, <2 x i1>)
+declare <8 x i64> @vector_foo_8(<8 x i64>, <8 x i1>)
 
-; A masked vector variant matching the forced VF of 2, so the call is a valid
-; widen-with-mask candidate (not itself scalar-with-predication) while feeding
-; the scalar-with-predication scatter store.
-attributes #0 = { readonly nounwind "vector-function-abi-variant"="_ZGV_LLVM_M2v_foo(vector_foo)" }
-;.
-; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
-; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
-; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
-; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META2]], [[META1]]}
-;.
+; Masked vector variants for VF=2 and VF=8, so the call is a widen-with-mask
+; candidate at both VFs while feeding the scalar-with-predication scatter store.
+attributes #0 = { readonly nounwind "vector-function-abi-variant"="_ZGV_LLVM_M2v_foo(vector_foo_2),_ZGV_LLVM_M8v_foo(vector_foo_8)" }

>From 1abfcda59e1286853e870721c878c8e93f6215b1 Mon Sep 17 00:00:00 2001
From: seantalts <me at seantalts.com>
Date: Tue, 23 Jun 2026 09:58:02 -0700
Subject: [PATCH 4/5] Update llvm/lib/Transforms/Vectorize/LoopVectorize.cpp

Co-authored-by: Florian Hahn <flo at fhahn.com>
---
 llvm/lib/Transforms/Vectorize/LoopVectorize.cpp | 6 +-----
 1 file changed, 1 insertion(+), 5 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 9268d8a0b8cf9..c147d4bd8cf15 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -2139,11 +2139,7 @@ LoopVectorizationCostModel::getVectorCallCost(CallInst *CI,
                              : ScalarCallCost * VF.getKnownMinValue() +
                                    getScalarizationOverhead(CI, VF);
 
-  // The call may also have a non-scalarized lowering at this VF, via a vector
-  // intrinsic or a vector library variant. computePredInstDiscount queries this
-  // cost (before the call's widening decision is made) to decide whether
-  // forcing a predicated tree of operations to scalar is profitable, so return
-  // the cheapest available lowering.
+  // The call may be vectorized at this VF, via a vector intrinsic or a vector library variant.
   if (getVectorIntrinsicIDForCall(CI, TLI))
     Cost = std::min(Cost, getVectorIntrinsicCost(CI, VF));
 

>From bd05c91fcbbfc7b27be2f9625f616caacb6587fa Mon Sep 17 00:00:00 2001
From: Sean Talts <sean.talts at gmail.com>
Date: Tue, 23 Jun 2026 13:03:36 -0400
Subject: [PATCH 5/5] [LoopVectorize] clang-format getVectorCallCost

---
 llvm/lib/Transforms/Vectorize/LoopVectorize.cpp | 11 ++++++-----
 1 file changed, 6 insertions(+), 5 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index c147d4bd8cf15..f9423017ed834 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -2139,16 +2139,17 @@ LoopVectorizationCostModel::getVectorCallCost(CallInst *CI,
                              : ScalarCallCost * VF.getKnownMinValue() +
                                    getScalarizationOverhead(CI, VF);
 
-  // The call may be vectorized at this VF, via a vector intrinsic or a vector library variant.
+  // The call may be vectorized at this VF, via a vector intrinsic or a vector
+  // library variant.
   if (getVectorIntrinsicIDForCall(CI, TLI))
     Cost = std::min(Cost, getVectorIntrinsicCost(CI, VF));
 
   if (Function *Variant =
           getVectorLibraryVariantFor(*CI, VF, isMaskRequired(CI), TLI))
-    Cost = std::min(Cost, TTI.getCallInstrCost(
-                              /*F=*/nullptr, Variant->getReturnType(),
-                              Variant->getFunctionType()->params(),
-                              Config.CostKind));
+    Cost = std::min(Cost,
+                    TTI.getCallInstrCost(
+                        /*F=*/nullptr, Variant->getReturnType(),
+                        Variant->getFunctionType()->params(), Config.CostKind));
 
   return Cost;
 }



More information about the llvm-commits mailing list