[llvm] [VPlan] Use VPlan-based costs for all integer IVs. (PR #217766)
Florian Hahn via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 24 04:58:21 PDT 2026
https://github.com/fhahn updated https://github.com/llvm/llvm-project/pull/217766
>From 1e887cf7203d02a5a668734907e931feef7bcad0 Mon Sep 17 00:00:00 2001
From: Florian Hahn <flo at fhahn.com>
Date: Fri, 14 Aug 2026 06:57:17 +0100
Subject: [PATCH] [VPlan] Use VPlan-based costs for all IVs, except scalar
pointer IVs.
Continue the migration of induction costs to be completely VPlan-based.
VPScalarIVStepsRecipe now supports compute its cost for integer
inductions. Skip the legacy precomputeCost code for them.
The main change is that previously we never explicitly accounted for the
canonical IV, which we always generate. It was instead folded into
computing cost of the existing integer IVs with only scalar uses, which in
most cases are actually served by the canonical IV.
The new code accounts for the cost explicitly, as
VPInstruction::computeCost returns 0 for all recipes w/o underlying
instruction.
Floating point and pointer induction will be handled as follow-ups to
keep the diffs managable and make it easier to track down regressions.
There are a few tests where cost decisions change, but overall they
should be more accurate:
- llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll:
- partial-reduce-with-invariant-stores.ll
via the legacy cost model, we computed the cost of i64 vector
induction (which has cost of 2/4 for VF 4 and 8), while now we compute
the cost of the generated narrower wide IV and the canonical IV increment.
- epilogue-vectorization-fix-scalar-resume-values.ll
- predicated-costs.ll
- pr81872.ll
- ARM/optsize_minsize.ll
- X86/conversion-cost.ll
Now also account for the cost of the canonical IV increment
- induction-costs-sve.ll
legacy accounted for scalar increment and extend, now we only account
for the actual canonical IV increment (w/o extend)
---
.../Transforms/Vectorize/LoopVectorize.cpp | 34 ++-
.../LoopVectorize/AArch64/cmp_cost.ll | 15 +-
.../AArch64/conditional-branches-cost.ll | 58 ++++-
...-vectorization-fix-scalar-resume-values.ll | 72 +++---
.../AArch64/fully-unrolled-cost.ll | 25 +-
.../AArch64/induction-costs-sve.ll | 44 ++--
.../partial-reduce-with-invariant-stores.ll | 14 +-
.../LoopVectorize/AArch64/predicated-costs.ll | 60 +++--
.../AArch64/replicating-load-store-costs.ll | 71 +++---
.../scalable-vectorization-cost-tuning.ll | 24 +-
.../LoopVectorize/ARM/mve-icmpcost.ll | 27 +-
.../LoopVectorize/ARM/optsize_minsize.ll | 149 ++---------
.../LoopVectorize/PowerPC/reg-usage.ll | 3 +-
.../RISCV/early-exit-live-out.ll | 3 +-
.../Transforms/LoopVectorize/RISCV/pr88802.ll | 40 +--
.../LoopVectorize/RISCV/uniform-load-store.ll | 21 +-
.../WebAssembly/memory-interleave.ll | 234 +++++++++---------
.../X86/CostModel/vpinstruction-cost.ll | 14 +-
.../LoopVectorize/X86/conversion-cost.ll | 48 ++--
.../X86/pr131359-dead-for-splice.ll | 88 +++++--
.../Transforms/LoopVectorize/X86/pr81872.ll | 44 ++--
.../LoopVectorize/X86/reduction-small-size.ll | 3 +-
.../LoopVectorize/induction-cost.ll | 4 +-
23 files changed, 535 insertions(+), 560 deletions(-)
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 5e8a43eb871c5..441a3cce5b7dd 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -2006,7 +2006,11 @@ static void addFullyUnrolledInstructionsToIgnore(
for (const auto &KV : IL) {
// Extract the key by hand so that it can be used in the lambda below. Note
// that captured structured bindings are a C++20 extension.
- const PHINode *IV = KV.first;
+ PHINode *IV = KV.first;
+
+ // The induction is free: a widened induction generates a vector phi with
+ // its start value and an increment that is dead without a backedge.
+ InstsToIgnore.insert(IV);
// Get next iteration value of the induction variable.
Instruction *IVInst =
@@ -5579,8 +5583,9 @@ LoopVectorizationPlanner::precomputeCosts(VPlan &Plan, ElementCount VF,
// TODO: Remove this code after stepping away from the legacy cost model and
// adding code to simplify VPlans before calculating their costs.
auto TC = getSmallConstantTripCount(PSE.getSE(), OrigLoop);
+ bool IsFullyUnrolled = TC == VF && !Plan.hasTailFolded();
SmallPtrSet<const Value *, 4> WidenedIVs;
- if (TC == VF && !Plan.hasTailFolded()) {
+ if (IsFullyUnrolled) {
addFullyUnrolledInstructionsToIgnore(OrigLoop, Legal->getInductionVars(),
CostCtx.SkipCostComputation);
} else {
@@ -5595,7 +5600,16 @@ LoopVectorizationPlanner::precomputeCosts(VPlan &Plan, ElementCount VF,
}
}
+ bool ChargeCanonicalIVIncrement = false;
for (const auto &[IV, IndDesc] : Legal->getInductionVars()) {
+ // Integer inductions are always costed via the VPlan-based cost model.
+ // TODO: Also migrate FP and pointer inductions.
+ if (IndDesc.getKind() == InductionDescriptor::IK_IntInduction) {
+ // If the vector loop is executed exactly once, the increment is
+ // simplified away.
+ ChargeCanonicalIVIncrement |= !IsFullyUnrolled;
+ continue;
+ }
if (WidenedIVs.contains(IV))
continue;
Instruction *IVInc = cast<Instruction>(
@@ -5630,6 +5644,22 @@ LoopVectorizationPlanner::precomputeCosts(VPlan &Plan, ElementCount VF,
}
}
+ // Add the cost for incrementing the canonical IV (or current iteration phi
+ // for EVL) explicitly, as VPInstruction::computeCost returns 0 for recipes
+ // w/o underlying value.
+ if (ChargeCanonicalIVIncrement) {
+ InstructionCost CanIVIncCost =
+ ForceTargetInstructionCost.getNumOccurrences()
+ ? InstructionCost(ForceTargetInstructionCost)
+ : CostCtx.TTI.getArithmeticInstrCost(
+ Instruction::Add,
+ Plan.getVectorLoopRegion()->getCanonicalIVType(),
+ CostCtx.CostKind);
+ LLVM_DEBUG(dbgs() << "Cost of " << CanIVIncCost << " for VF " << VF
+ << ": canonical IV increment\n");
+ Cost += CanIVIncCost;
+ }
+
// Pre-compute the costs for branches except for the backedge, as the number
// of replicate regions in a VPlan may not directly match the number of
// branches, which would lead to different decisions.
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/cmp_cost.ll b/llvm/test/Transforms/LoopVectorize/AArch64/cmp_cost.ll
index 5bc97910eb2f0..99cc88cc65ba0 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/cmp_cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/cmp_cost.ll
@@ -11,6 +11,7 @@ target triple = "aarch64-none-elf"
define float @fmaxnum_reduction_f32(float %base, i32 %n) {
; CHECK-LABEL: 'fmaxnum_reduction_f32'
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 1 for VF 2: ir<%iv> = WIDEN-INDUCTION nuw nsw ir<0>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: WIDEN-REDUCTION-PHI ir<%max> = phi (fmaxnum) ir<-1.000000e+07>, ir<%max.next>
; CHECK: Cost of 1 for VF 2: WIDEN-CAST ir<%iv.f> = sitofp ir<%iv> to float
@@ -40,6 +41,7 @@ define float @fmaxnum_reduction_f32(float %base, i32 %n) {
; CHECK: Cost of 0 for VF 2: EMIT vp<[[VP13:%[0-9]+]]> = and vp<%cmp.n>, vp<[[VP12]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<[[VP13]]>
; CHECK: Cost of 0 for VF 2: IR %max.next.lcssa = phi float [ %max.next, %loop ] (extra operand: vp<[[VP11]]> from middle.block)
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 1 for VF 4: ir<%iv> = WIDEN-INDUCTION nuw nsw ir<0>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: WIDEN-REDUCTION-PHI ir<%max> = phi (fmaxnum) ir<-1.000000e+07>, ir<%max.next>
; CHECK: Cost of 1 for VF 4: WIDEN-CAST ir<%iv.f> = sitofp ir<%iv> to float
@@ -90,6 +92,7 @@ exit:
define double @fmaxnum_reduction_f64(double %base, i64 %n) {
; CHECK-LABEL: 'fmaxnum_reduction_f64'
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 1 for VF 2: ir<%iv> = WIDEN-INDUCTION nuw nsw ir<0>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: WIDEN-REDUCTION-PHI ir<%max> = phi (fmaxnum) ir<-1.000000e+07>, ir<%max.next>
; CHECK: Cost of 1 for VF 2: WIDEN-CAST ir<%iv.f> = sitofp ir<%iv> to double
@@ -142,8 +145,7 @@ exit:
define i32 @switch_to_cmp(ptr %s, ptr %dst, i64 %n) {
; CHECK-LABEL: 'switch_to_cmp'
-; CHECK: Cost of 1 for VF 2: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 2: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop.latch ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 0 for VF 2: forced scalar %gep = getelementptr i8, ptr %s, i64 %iv
; CHECK: Cost of 0 for VF 2: forced scalar %dst.gep = getelementptr i8, ptr %dst, i64 %iv
; CHECK: Cost of 0 for VF 2: WIDEN-REDUCTION-PHI ir<%c> = phi (add) vp<[[VP3:%[0-9]+]]>, ir<%c.next>
@@ -201,8 +203,7 @@ define i32 @switch_to_cmp(ptr %s, ptr %dst, i64 %n) {
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost of 0 for VF 2: IR %c.next.lcssa = phi i32 [ %c.next, %loop.latch ] (extra operand: vp<[[VP39]]> from middle.block)
-; CHECK: Cost of 1 for VF 4: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 4: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop.latch ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 0 for VF 4: forced scalar %gep = getelementptr i8, ptr %s, i64 %iv
; CHECK: Cost of 0 for VF 4: forced scalar %dst.gep = getelementptr i8, ptr %dst, i64 %iv
; CHECK: Cost of 0 for VF 4: WIDEN-REDUCTION-PHI ir<%c> = phi (add) vp<[[VP3]]>, ir<%c.next>
@@ -260,8 +261,7 @@ define i32 @switch_to_cmp(ptr %s, ptr %dst, i64 %n) {
; CHECK: Cost of 1 for VF 4: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 4: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost of 0 for VF 4: IR %c.next.lcssa = phi i32 [ %c.next, %loop.latch ] (extra operand: vp<[[VP39]]> from middle.block)
-; CHECK: Cost of 1 for VF 8: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 8: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop.latch ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 0 for VF 8: forced scalar %gep = getelementptr i8, ptr %s, i64 %iv
; CHECK: Cost of 0 for VF 8: forced scalar %dst.gep = getelementptr i8, ptr %dst, i64 %iv
; CHECK: Cost of 0 for VF 8: WIDEN-REDUCTION-PHI ir<%c> = phi (add) vp<[[VP3]]>, ir<%c.next>
@@ -319,8 +319,7 @@ define i32 @switch_to_cmp(ptr %s, ptr %dst, i64 %n) {
; CHECK: Cost of 1 for VF 8: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 8: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost of 0 for VF 8: IR %c.next.lcssa = phi i32 [ %c.next, %loop.latch ] (extra operand: vp<[[VP39]]> from middle.block)
-; CHECK: Cost of 1 for VF 16: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 16: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop.latch ]
+; CHECK: Cost of 1 for VF 16: canonical IV increment
; CHECK: Cost of 0 for VF 16: forced scalar %gep = getelementptr i8, ptr %s, i64 %iv
; CHECK: Cost of 0 for VF 16: forced scalar %dst.gep = getelementptr i8, ptr %dst, i64 %iv
; CHECK: Cost of 0 for VF 16: WIDEN-REDUCTION-PHI ir<%c> = phi (add) vp<[[VP3]]>, ir<%c.next>
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll
index 87d6325d640e5..0c113484bb31e 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/conditional-branches-cost.ll
@@ -289,17 +289,17 @@ define void @latch_branch_cost(ptr %dst) {
; PRED: [[VECTOR_PH]]:
; PRED-NEXT: br label %[[VECTOR_BODY:.*]]
; PRED: [[VECTOR_BODY]]:
-; PRED-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE6:.*]] ]
-; PRED-NEXT: [[VEC_IND:%.*]] = phi <4 x i8> [ <i8 0, i8 1, i8 2, i8 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE6]] ]
-; PRED-NEXT: [[TMP0:%.*]] = icmp ule <4 x i8> [[VEC_IND]], splat (i8 99)
-; PRED-NEXT: [[TMP1:%.*]] = extractelement <4 x i1> [[TMP0]], i64 0
+; PRED-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE14:.*]] ]
+; PRED-NEXT: [[VEC_IND:%.*]] = phi <8 x i8> [ <i8 0, i8 1, i8 2, i8 3, i8 4, i8 5, i8 6, i8 7>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE14]] ]
+; PRED-NEXT: [[TMP0:%.*]] = icmp ule <8 x i8> [[VEC_IND]], splat (i8 99)
+; PRED-NEXT: [[TMP1:%.*]] = extractelement <8 x i1> [[TMP0]], i64 0
; PRED-NEXT: br i1 [[TMP1]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
; PRED: [[PRED_STORE_IF]]:
; PRED-NEXT: [[TMP3:%.*]] = getelementptr i8, ptr [[DST]], i64 [[INDEX]]
; PRED-NEXT: store i8 0, ptr [[TMP3]], align 1
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE]]
; PRED: [[PRED_STORE_CONTINUE]]:
-; PRED-NEXT: [[TMP4:%.*]] = extractelement <4 x i1> [[TMP0]], i64 1
+; PRED-NEXT: [[TMP4:%.*]] = extractelement <8 x i1> [[TMP0]], i64 1
; PRED-NEXT: br i1 [[TMP4]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2:.*]]
; PRED: [[PRED_STORE_IF1]]:
; PRED-NEXT: [[TMP5:%.*]] = add i64 [[INDEX]], 1
@@ -307,7 +307,7 @@ define void @latch_branch_cost(ptr %dst) {
; PRED-NEXT: store i8 0, ptr [[TMP6]], align 1
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE2]]
; PRED: [[PRED_STORE_CONTINUE2]]:
-; PRED-NEXT: [[TMP7:%.*]] = extractelement <4 x i1> [[TMP0]], i64 2
+; PRED-NEXT: [[TMP7:%.*]] = extractelement <8 x i1> [[TMP0]], i64 2
; PRED-NEXT: br i1 [[TMP7]], label %[[PRED_STORE_IF3:.*]], label %[[PRED_STORE_CONTINUE4:.*]]
; PRED: [[PRED_STORE_IF3]]:
; PRED-NEXT: [[TMP8:%.*]] = add i64 [[INDEX]], 2
@@ -315,21 +315,53 @@ define void @latch_branch_cost(ptr %dst) {
; PRED-NEXT: store i8 0, ptr [[TMP9]], align 1
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE4]]
; PRED: [[PRED_STORE_CONTINUE4]]:
-; PRED-NEXT: [[TMP10:%.*]] = extractelement <4 x i1> [[TMP0]], i64 3
-; PRED-NEXT: br i1 [[TMP10]], label %[[PRED_STORE_IF5:.*]], label %[[PRED_STORE_CONTINUE6]]
+; PRED-NEXT: [[TMP10:%.*]] = extractelement <8 x i1> [[TMP0]], i64 3
+; PRED-NEXT: br i1 [[TMP10]], label %[[PRED_STORE_IF5:.*]], label %[[PRED_STORE_CONTINUE6:.*]]
; PRED: [[PRED_STORE_IF5]]:
; PRED-NEXT: [[TMP11:%.*]] = add i64 [[INDEX]], 3
; PRED-NEXT: [[TMP24:%.*]] = getelementptr i8, ptr [[DST]], i64 [[TMP11]]
; PRED-NEXT: store i8 0, ptr [[TMP24]], align 1
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE6]]
; PRED: [[PRED_STORE_CONTINUE6]]:
-; PRED-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; PRED-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i8> [[VEC_IND]], splat (i8 4)
-; PRED-NEXT: [[TMP12:%.*]] = icmp eq i64 [[INDEX_NEXT]], 100
-; PRED-NEXT: br i1 [[TMP12]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; PRED-NEXT: [[TMP12:%.*]] = extractelement <8 x i1> [[TMP0]], i64 4
+; PRED-NEXT: br i1 [[TMP12]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT:.*]]
; PRED: [[MIDDLE_BLOCK]]:
-; PRED-NEXT: br label %[[EXIT:.*]]
+; PRED-NEXT: [[TMP13:%.*]] = add i64 [[INDEX]], 4
+; PRED-NEXT: [[TMP14:%.*]] = getelementptr i8, ptr [[DST]], i64 [[TMP13]]
+; PRED-NEXT: store i8 0, ptr [[TMP14]], align 1
+; PRED-NEXT: br label %[[EXIT]]
; PRED: [[EXIT]]:
+; PRED-NEXT: [[TMP15:%.*]] = extractelement <8 x i1> [[TMP0]], i64 5
+; PRED-NEXT: br i1 [[TMP15]], label %[[PRED_STORE_IF9:.*]], label %[[PRED_STORE_CONTINUE10:.*]]
+; PRED: [[PRED_STORE_IF9]]:
+; PRED-NEXT: [[TMP16:%.*]] = add i64 [[INDEX]], 5
+; PRED-NEXT: [[TMP17:%.*]] = getelementptr i8, ptr [[DST]], i64 [[TMP16]]
+; PRED-NEXT: store i8 0, ptr [[TMP17]], align 1
+; PRED-NEXT: br label %[[PRED_STORE_CONTINUE10]]
+; PRED: [[PRED_STORE_CONTINUE10]]:
+; PRED-NEXT: [[TMP18:%.*]] = extractelement <8 x i1> [[TMP0]], i64 6
+; PRED-NEXT: br i1 [[TMP18]], label %[[PRED_STORE_IF11:.*]], label %[[PRED_STORE_CONTINUE12:.*]]
+; PRED: [[PRED_STORE_IF11]]:
+; PRED-NEXT: [[TMP19:%.*]] = add i64 [[INDEX]], 6
+; PRED-NEXT: [[TMP20:%.*]] = getelementptr i8, ptr [[DST]], i64 [[TMP19]]
+; PRED-NEXT: store i8 0, ptr [[TMP20]], align 1
+; PRED-NEXT: br label %[[PRED_STORE_CONTINUE12]]
+; PRED: [[PRED_STORE_CONTINUE12]]:
+; PRED-NEXT: [[TMP21:%.*]] = extractelement <8 x i1> [[TMP0]], i64 7
+; PRED-NEXT: br i1 [[TMP21]], label %[[PRED_STORE_IF13:.*]], label %[[PRED_STORE_CONTINUE14]]
+; PRED: [[PRED_STORE_IF13]]:
+; PRED-NEXT: [[TMP22:%.*]] = add i64 [[INDEX]], 7
+; PRED-NEXT: [[TMP23:%.*]] = getelementptr i8, ptr [[DST]], i64 [[TMP22]]
+; PRED-NEXT: store i8 0, ptr [[TMP23]], align 1
+; PRED-NEXT: br label %[[PRED_STORE_CONTINUE14]]
+; PRED: [[PRED_STORE_CONTINUE14]]:
+; PRED-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; PRED-NEXT: [[VEC_IND_NEXT]] = add nuw <8 x i8> [[VEC_IND]], splat (i8 8)
+; PRED-NEXT: [[TMP25:%.*]] = icmp eq i64 [[INDEX_NEXT]], 104
+; PRED-NEXT: br i1 [[TMP25]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; PRED: [[MIDDLE_BLOCK1]]:
+; PRED-NEXT: br label %[[EXIT1:.*]]
+; PRED: [[EXIT1]]:
; PRED-NEXT: ret void
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/epilogue-vectorization-fix-scalar-resume-values.ll b/llvm/test/Transforms/LoopVectorize/AArch64/epilogue-vectorization-fix-scalar-resume-values.ll
index 20c92e77a9d53..16d8e5751e289 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/epilogue-vectorization-fix-scalar-resume-values.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/epilogue-vectorization-fix-scalar-resume-values.ll
@@ -169,7 +169,7 @@ define i64 @find_last_and_any_of(ptr %Dst, ptr %Src) {
; CHECK: [[VECTOR_MEMCHECK]]:
; CHECK-NEXT: [[TMP0:%.*]] = sub i64 [[DST1]], [[SRC2]]
; CHECK-NEXT: [[TMP38:%.*]] = sub i64 [[TMP0]], 1
-; CHECK-NEXT: [[DIFF_CHECK:%.*]] = icmp ult i64 [[TMP38]], 63
+; CHECK-NEXT: [[DIFF_CHECK:%.*]] = icmp ult i64 [[TMP38]], 127
; CHECK-NEXT: br i1 [[DIFF_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
; CHECK: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
; CHECK-NEXT: br i1 false, label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
@@ -177,47 +177,47 @@ define i64 @find_last_and_any_of(ptr %Dst, ptr %Src) {
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <8 x i32> [ splat (i32 -2147483648), %[[VECTOR_PH]] ], [ [[TMP5:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_PHI3:%.*]] = phi <8 x i32> [ splat (i32 -2147483648), %[[VECTOR_PH]] ], [ [[TMP6:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_PHI4:%.*]] = phi <8 x i1> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP9:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_PHI5:%.*]] = phi <8 x i1> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP10:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[STEP_ADD:%.*]] = add <8 x i32> [[VEC_IND]], splat (i32 8)
+; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <16 x i32> [ splat (i32 -2147483648), %[[VECTOR_PH]] ], [ [[TMP6:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI3:%.*]] = phi <16 x i32> [ splat (i32 -2147483648), %[[VECTOR_PH]] ], [ [[TMP7:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI4:%.*]] = phi <16 x i1> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP10:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI5:%.*]] = phi <16 x i1> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP11:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <16 x i32> [ <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 14, i32 15>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[STEP_ADD:%.*]] = add <16 x i32> [[VEC_IND]], splat (i32 16)
; CHECK-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[TMP1]], i64 8
-; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <8 x i32>, ptr [[TMP1]], align 4
-; CHECK-NEXT: [[WIDE_LOAD6:%.*]] = load <8 x i32>, ptr [[TMP2]], align 4
-; CHECK-NEXT: [[TMP3:%.*]] = icmp sgt <8 x i32> [[WIDE_LOAD]], zeroinitializer
-; CHECK-NEXT: [[TMP4:%.*]] = icmp sgt <8 x i32> [[WIDE_LOAD6]], zeroinitializer
-; CHECK-NEXT: [[TMP5]] = select <8 x i1> [[TMP3]], <8 x i32> [[VEC_IND]], <8 x i32> [[VEC_PHI]]
-; CHECK-NEXT: [[TMP6]] = select <8 x i1> [[TMP4]], <8 x i32> [[STEP_ADD]], <8 x i32> [[VEC_PHI3]]
-; CHECK-NEXT: [[TMP7:%.*]] = icmp sgt <8 x i32> [[WIDE_LOAD]], splat (i32 100)
-; CHECK-NEXT: [[TMP8:%.*]] = icmp sgt <8 x i32> [[WIDE_LOAD6]], splat (i32 100)
-; CHECK-NEXT: [[TMP9]] = or <8 x i1> [[VEC_PHI4]], [[TMP7]]
-; CHECK-NEXT: [[TMP10]] = or <8 x i1> [[VEC_PHI5]], [[TMP8]]
-; CHECK-NEXT: [[TMP11:%.*]] = add nsw <8 x i32> [[WIDE_LOAD]], splat (i32 7)
-; CHECK-NEXT: [[TMP12:%.*]] = add nsw <8 x i32> [[WIDE_LOAD6]], splat (i32 7)
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[TMP1]], i64 16
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-NEXT: [[WIDE_LOAD6:%.*]] = load <16 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP4:%.*]] = icmp sgt <16 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT: [[TMP5:%.*]] = icmp sgt <16 x i32> [[WIDE_LOAD6]], zeroinitializer
+; CHECK-NEXT: [[TMP6]] = select <16 x i1> [[TMP4]], <16 x i32> [[VEC_IND]], <16 x i32> [[VEC_PHI]]
+; CHECK-NEXT: [[TMP7]] = select <16 x i1> [[TMP5]], <16 x i32> [[STEP_ADD]], <16 x i32> [[VEC_PHI3]]
+; CHECK-NEXT: [[TMP8:%.*]] = icmp sgt <16 x i32> [[WIDE_LOAD]], splat (i32 100)
+; CHECK-NEXT: [[TMP9:%.*]] = icmp sgt <16 x i32> [[WIDE_LOAD6]], splat (i32 100)
+; CHECK-NEXT: [[TMP10]] = or <16 x i1> [[VEC_PHI4]], [[TMP8]]
+; CHECK-NEXT: [[TMP11]] = or <16 x i1> [[VEC_PHI5]], [[TMP9]]
+; CHECK-NEXT: [[TMP12:%.*]] = add nsw <16 x i32> [[WIDE_LOAD]], splat (i32 7)
+; CHECK-NEXT: [[TMP14:%.*]] = add nsw <16 x i32> [[WIDE_LOAD6]], splat (i32 7)
; CHECK-NEXT: [[TMP13:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[INDEX]]
-; CHECK-NEXT: [[TMP14:%.*]] = getelementptr inbounds i32, ptr [[TMP13]], i64 8
-; CHECK-NEXT: store <8 x i32> [[TMP11]], ptr [[TMP13]], align 4
-; CHECK-NEXT: store <8 x i32> [[TMP12]], ptr [[TMP14]], align 4
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <8 x i32> [[STEP_ADD]], splat (i32 8)
-; CHECK-NEXT: [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1008
-; CHECK-NEXT: br i1 [[TMP15]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
+; CHECK-NEXT: [[TMP15:%.*]] = getelementptr inbounds i32, ptr [[TMP13]], i64 16
+; CHECK-NEXT: store <16 x i32> [[TMP12]], ptr [[TMP13]], align 4
+; CHECK-NEXT: store <16 x i32> [[TMP14]], ptr [[TMP15]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 32
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <16 x i32> [[STEP_ADD]], splat (i32 16)
+; CHECK-NEXT: [[TMP39:%.*]] = icmp eq i64 [[INDEX_NEXT]], 992
+; CHECK-NEXT: br i1 [[TMP39]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
-; CHECK-NEXT: [[RDX_MINMAX:%.*]] = call <8 x i32> @llvm.smax.v8i32(<8 x i32> [[TMP5]], <8 x i32> [[TMP6]])
-; CHECK-NEXT: [[TMP16:%.*]] = call i32 @llvm.vector.reduce.smax.v8i32(<8 x i32> [[RDX_MINMAX]])
+; CHECK-NEXT: [[RDX_MINMAX:%.*]] = call <16 x i32> @llvm.smax.v16i32(<16 x i32> [[TMP6]], <16 x i32> [[TMP7]])
+; CHECK-NEXT: [[TMP16:%.*]] = call i32 @llvm.vector.reduce.smax.v16i32(<16 x i32> [[RDX_MINMAX]])
; CHECK-NEXT: [[TMP17:%.*]] = icmp ne i32 [[TMP16]], -2147483648
; CHECK-NEXT: [[TMP18:%.*]] = select i1 [[TMP17]], i32 [[TMP16]], i32 -1
-; CHECK-NEXT: [[BIN_RDX:%.*]] = or <8 x i1> [[TMP10]], [[TMP9]]
-; CHECK-NEXT: [[TMP19:%.*]] = call i1 @llvm.vector.reduce.or.v8i1(<8 x i1> [[BIN_RDX]])
+; CHECK-NEXT: [[BIN_RDX:%.*]] = or <16 x i1> [[TMP11]], [[TMP10]]
+; CHECK-NEXT: [[TMP19:%.*]] = call i1 @llvm.vector.reduce.or.v16i1(<16 x i1> [[BIN_RDX]])
; CHECK-NEXT: [[TMP20:%.*]] = freeze i1 [[TMP19]]
; CHECK-NEXT: br i1 false, label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
; CHECK: [[VEC_EPILOG_ITER_CHECK]]:
-; CHECK-NEXT: br i1 false, label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF13:![0-9]+]]
+; CHECK-NEXT: br i1 false, label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF9]]
; CHECK: [[VEC_EPILOG_PH]]:
-; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ 1008, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ 992, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i32 [ [[TMP18]], %[[VEC_EPILOG_ITER_CHECK]] ], [ -1, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
; CHECK-NEXT: [[BC_MERGE_RDX7:%.*]] = phi i1 [ [[TMP20]], %[[VEC_EPILOG_ITER_CHECK]] ], [ false, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
; CHECK-NEXT: [[TMP21:%.*]] = icmp eq i32 [[BC_MERGE_RDX]], -1
@@ -249,7 +249,7 @@ define i64 @find_last_and_any_of(ptr %Dst, ptr %Src) {
; CHECK-NEXT: [[INDEX_NEXT17]] = add nuw i64 [[INDEX12]], 4
; CHECK-NEXT: [[VEC_IND_NEXT18]] = add <4 x i32> [[VEC_IND15]], splat (i32 4)
; CHECK-NEXT: [[TMP32:%.*]] = icmp eq i64 [[INDEX_NEXT17]], 1020
-; CHECK-NEXT: br i1 [[TMP32]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP14:![0-9]+]]
+; CHECK-NEXT: br i1 [[TMP32]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP13:![0-9]+]]
; CHECK: [[VEC_EPILOG_MIDDLE_BLOCK]]:
; CHECK-NEXT: [[TMP33:%.*]] = call i32 @llvm.vector.reduce.smax.v4i32(<4 x i32> [[TMP27]])
; CHECK-NEXT: [[TMP34:%.*]] = icmp ne i32 [[TMP33]], -2147483648
@@ -258,7 +258,7 @@ define i64 @find_last_and_any_of(ptr %Dst, ptr %Src) {
; CHECK-NEXT: [[TMP37:%.*]] = freeze i1 [[TMP36]]
; CHECK-NEXT: br i1 true, label %[[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
; CHECK: [[VEC_EPILOG_SCALAR_PH]]:
-; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ 1020, %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ 1008, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MEMCHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ 1020, %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ 992, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MEMCHECK]] ], [ 0, %[[ITER_CHECK]] ]
; CHECK-NEXT: [[BC_MERGE_RDX19:%.*]] = phi i32 [ [[TMP35]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[TMP18]], %[[VEC_EPILOG_ITER_CHECK]] ], [ -1, %[[VECTOR_MEMCHECK]] ], [ -1, %[[ITER_CHECK]] ]
; CHECK-NEXT: [[BC_MERGE_RDX20:%.*]] = phi i1 [ [[TMP37]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[TMP20]], %[[VEC_EPILOG_ITER_CHECK]] ], [ false, %[[VECTOR_MEMCHECK]] ], [ false, %[[ITER_CHECK]] ]
; CHECK-NEXT: br label %[[LOOP:.*]]
@@ -278,7 +278,7 @@ define i64 @find_last_and_any_of(ptr %Dst, ptr %Src) {
; CHECK-NEXT: store i32 [[ADD]], ptr [[GEP_DST]], align 4
; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
; CHECK-NEXT: [[EC:%.*]] = icmp eq i64 [[IV_NEXT]], 1020
-; CHECK-NEXT: br i1 [[EC]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP15:![0-9]+]]
+; CHECK-NEXT: br i1 [[EC]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP14:![0-9]+]]
; CHECK: [[EXIT]]:
; CHECK-NEXT: [[RES_LCSSA:%.*]] = phi i32 [ [[RES_NEXT]], %[[LOOP]] ], [ [[TMP18]], %[[MIDDLE_BLOCK]] ], [ [[TMP35]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ]
; CHECK-NEXT: [[ANY_LCSSA:%.*]] = phi i1 [ [[ANY_NEXT]], %[[LOOP]] ], [ [[TMP20]], %[[MIDDLE_BLOCK]] ], [ [[TMP37]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/fully-unrolled-cost.ll b/llvm/test/Transforms/LoopVectorize/AArch64/fully-unrolled-cost.ll
index 9b051605a874a..532a4a18a5133 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/fully-unrolled-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/fully-unrolled-cost.ll
@@ -7,11 +7,10 @@ target triple="aarch64--linux-gnu"
; vector loop gets executed exactly once with the given VF.
define i64 @test(ptr %a, ptr %b) #0 {
; CHECK-LABEL: LV: Checking a loop in 'test'
-; CHECK: Cost of 1 for VF 8: induction instruction %i.iv.next = add nuw nsw i64 %i.iv, 1
-; CHECK-NEXT: Cost of 0 for VF 8: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 1 for VF 8: EMIT branch-on-count vp<%index.next>, vp<{{.+}}>
; CHECK: Cost for VF 8: 30
-; CHECK-NEXT: Cost of 0 for VF 16: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
+; CHECK-NOT: canonical IV increment
; CHECK: Cost of 0 for VF 16: EMIT branch-on-count vp<%index.next>, vp<{{.+}}>
; CHECK: Cost for VF 16: 56
; CHECK: LV: Selecting VF: 16
@@ -40,12 +39,10 @@ for.body:
; Same as above, but in the next iteration IV has extra users, and thus, the cost is not zero.
define i64 @test_external_iv_user(ptr %a, ptr %b) #0 {
; CHECK-LABEL: LV: Checking a loop in 'test_external_iv_user'
-; CHECK: Cost of 1 for VF 8: induction instruction %i.iv.next = add nuw nsw i64 %i.iv, 1
-; CHECK-NEXT: Cost of 0 for VF 8: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 1 for VF 8: EMIT branch-on-count vp<%index.next>, vp<{{.+}}>
-; CHECK: Cost for VF 8: 30
-; CHECK-NEXT: Cost of 1 for VF 16: induction instruction %i.iv.next = add nuw nsw i64 %i.iv, 1
-; CHECK-NEXT: Cost of 0 for VF 16: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
+; CHECK: Cost for VF 8: 31
+; CHECK-NOT: canonical IV increment
; CHECK: Cost of 0 for VF 16: EMIT branch-on-count vp<%index.next>, vp<{{.+}}>
; CHECK: Cost for VF 16: 57
; CHECK: LV: Selecting VF: 16
@@ -74,14 +71,9 @@ exit:
; Same as above but with two IVs without extra users. They all have zero cost when VF equals the number of iterations.
define i64 @test_two_ivs(ptr %a, ptr %b, i64 %start) #0 {
; CHECK-LABEL: LV: Checking a loop in 'test_two_ivs'
-; CHECK: Cost of 1 for VF 8: induction instruction %i.iv.next = add nuw nsw i64 %i.iv, 1
-; CHECK-NEXT: Cost of 0 for VF 8: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
-; CHECK-NEXT: Cost of 1 for VF 8: induction instruction %j.iv.next = add nuw nsw i64 %j.iv, 1
-; CHECK-NEXT: Cost of 0 for VF 8: induction instruction %j.iv = phi i64 [ %start, %entry ], [ %j.iv.next, %for.body ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 1 for VF 8: EMIT branch-on-count vp<%index.next>, vp<{{.+}}>
-; CHECK: Cost for VF 8: 17
-; CHECK-NEXT: Cost of 0 for VF 16: induction instruction %i.iv = phi i64 [ 0, %entry ], [ %i.iv.next, %for.body ]
-; CHECK-NEXT: Cost of 0 for VF 16: induction instruction %j.iv = phi i64 [ %start, %entry ], [ %j.iv.next, %for.body ]
+; CHECK: Cost for VF 8: 16
; CHECK: Cost of 1 for VF 16: EXPRESSION vp<%11> = ir<%sum> + partial.reduce.add (mul nuw nsw (ir<%1> zext to i64), (ir<%0> zext to i64))
; CHECK: Cost for VF 16: 4
; CHECK: LV: Selecting VF: 16
@@ -115,8 +107,7 @@ define i1 @test_extra_cmp_user(ptr nocapture noundef %dst, ptr nocapture noundef
; CHECK: Cost of 4 for VF 8: WIDEN ir<%exitcond.not> = icmp eq ir<%indvars.iv.next>, ir<16>
; CHECK: Cost of 1 for VF 8: EMIT vp<%cmp.n> = icmp eq ir<16>, vp<%2>
; CHECK-NEXT: Cost of 0 for VF 8: EMIT branch-on-cond vp<%cmp.n>
-; CHECK: Cost for VF 8: 9
-; CHECK: Cost of 0 for VF 16: induction instruction %indvars.iv = phi i64 [ 0, %entry ], [ %indvars.iv.next, %for.body ]
+; CHECK: Cost for VF 8: 10
; CHECK: Cost of 0 for VF 16: WIDEN ir<%exitcond.not> = icmp eq ir<%indvars.iv.next>, ir<16>
; CHECK: Cost of 1 for VF 16: EMIT vp<%cmp.n> = icmp eq ir<16>, vp<%2>
; CHECK-NEXT: Cost of 0 for VF 16: EMIT branch-on-cond vp<%cmp.n>
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/induction-costs-sve.ll b/llvm/test/Transforms/LoopVectorize/AArch64/induction-costs-sve.ll
index 68968f4957388..fc43405404b4d 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/induction-costs-sve.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/induction-costs-sve.ll
@@ -619,37 +619,53 @@ define void @exit_cond_zext_iv_store32(ptr %dst, i64 %N) {
; PRED-NEXT: [[TMP6:%.*]] = or i1 [[TMP4]], [[TMP5]]
; PRED-NEXT: br i1 [[TMP6]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; PRED: [[VECTOR_PH]]:
-; PRED-NEXT: [[N_RND_UP:%.*]] = add i64 [[TMP0]], 1
-; PRED-NEXT: [[TMP7:%.*]] = and i64 [[N_RND_UP]], 1
+; PRED-NEXT: [[N_RND_UP:%.*]] = add i64 [[TMP0]], 3
+; PRED-NEXT: [[TMP7:%.*]] = and i64 [[N_RND_UP]], 3
; PRED-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP7]]
; PRED-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[TMP0]], 1
-; PRED-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <2 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
-; PRED-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <2 x i64> [[BROADCAST_SPLATINSERT]], <2 x i64> poison, <2 x i32> zeroinitializer
+; PRED-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; PRED-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
; PRED-NEXT: br label %[[VECTOR_BODY:.*]]
; PRED: [[VECTOR_BODY]]:
-; PRED-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE2:.*]] ]
-; PRED-NEXT: [[VEC_IND:%.*]] = phi <2 x i64> [ <i64 0, i64 1>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE2]] ]
-; PRED-NEXT: [[TMP8:%.*]] = icmp ule <2 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
-; PRED-NEXT: [[TMP9:%.*]] = extractelement <2 x i1> [[TMP8]], i64 0
+; PRED-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT1:%.*]], %[[PRED_STORE_CONTINUE6:.*]] ]
+; PRED-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE6]] ]
+; PRED-NEXT: [[TMP8:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; PRED-NEXT: [[TMP9:%.*]] = extractelement <4 x i1> [[TMP8]], i64 0
; PRED-NEXT: br i1 [[TMP9]], label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
; PRED: [[PRED_STORE_IF]]:
; PRED-NEXT: [[TMP10:%.*]] = getelementptr { [100 x i32], i32, i32 }, ptr [[DST]], i64 [[INDEX]], i32 2
; PRED-NEXT: store i32 0, ptr [[TMP10]], align 8
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE]]
; PRED: [[PRED_STORE_CONTINUE]]:
-; PRED-NEXT: [[TMP11:%.*]] = extractelement <2 x i1> [[TMP8]], i64 1
-; PRED-NEXT: br i1 [[TMP11]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2]]
+; PRED-NEXT: [[TMP11:%.*]] = extractelement <4 x i1> [[TMP8]], i64 1
+; PRED-NEXT: br i1 [[TMP11]], label %[[PRED_STORE_IF1:.*]], label %[[PRED_STORE_CONTINUE2:.*]]
; PRED: [[PRED_STORE_IF1]]:
; PRED-NEXT: [[TMP12:%.*]] = add i64 [[INDEX]], 1
; PRED-NEXT: [[TMP13:%.*]] = getelementptr { [100 x i32], i32, i32 }, ptr [[DST]], i64 [[TMP12]], i32 2
; PRED-NEXT: store i32 0, ptr [[TMP13]], align 8
; PRED-NEXT: br label %[[PRED_STORE_CONTINUE2]]
; PRED: [[PRED_STORE_CONTINUE2]]:
-; PRED-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 2
-; PRED-NEXT: [[VEC_IND_NEXT]] = add nuw <2 x i64> [[VEC_IND]], splat (i64 2)
-; PRED-NEXT: [[TMP14:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; PRED-NEXT: br i1 [[TMP14]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; PRED-NEXT: [[TMP14:%.*]] = extractelement <4 x i1> [[TMP8]], i64 2
+; PRED-NEXT: br i1 [[TMP14]], label %[[PRED_STORE_IF3:.*]], label %[[MIDDLE_BLOCK:.*]]
+; PRED: [[PRED_STORE_IF3]]:
+; PRED-NEXT: [[INDEX_NEXT:%.*]] = add i64 [[INDEX]], 2
+; PRED-NEXT: [[TMP16:%.*]] = getelementptr { [100 x i32], i32, i32 }, ptr [[DST]], i64 [[INDEX_NEXT]], i32 2
+; PRED-NEXT: store i32 0, ptr [[TMP16]], align 8
+; PRED-NEXT: br label %[[MIDDLE_BLOCK]]
; PRED: [[MIDDLE_BLOCK]]:
+; PRED-NEXT: [[TMP17:%.*]] = extractelement <4 x i1> [[TMP8]], i64 3
+; PRED-NEXT: br i1 [[TMP17]], label %[[PRED_STORE_IF5:.*]], label %[[PRED_STORE_CONTINUE6]]
+; PRED: [[PRED_STORE_IF5]]:
+; PRED-NEXT: [[TMP18:%.*]] = add i64 [[INDEX]], 3
+; PRED-NEXT: [[TMP19:%.*]] = getelementptr { [100 x i32], i32, i32 }, ptr [[DST]], i64 [[TMP18]], i32 2
+; PRED-NEXT: store i32 0, ptr [[TMP19]], align 8
+; PRED-NEXT: br label %[[PRED_STORE_CONTINUE6]]
+; PRED: [[PRED_STORE_CONTINUE6]]:
+; PRED-NEXT: [[INDEX_NEXT1]] = add i64 [[INDEX]], 4
+; PRED-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; PRED-NEXT: [[TMP20:%.*]] = icmp eq i64 [[INDEX_NEXT1]], [[N_VEC]]
+; PRED-NEXT: br i1 [[TMP20]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; PRED: [[MIDDLE_BLOCK1]]:
; PRED-NEXT: br label %[[EXIT:.*]]
; PRED: [[SCALAR_PH]]:
; PRED-NEXT: br label %[[LOOP:.*]]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/partial-reduce-with-invariant-stores.ll b/llvm/test/Transforms/LoopVectorize/AArch64/partial-reduce-with-invariant-stores.ll
index 10cbc67177bb8..952be051e3476 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/partial-reduce-with-invariant-stores.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/partial-reduce-with-invariant-stores.ll
@@ -74,7 +74,7 @@ define void @chained_sext_adds(ptr noalias %src, ptr noalias %dst) #0 {
; CHECK-SAME: ptr noalias [[SRC:%.*]], ptr noalias [[DST:%.*]]) #[[ATTR1:[0-9]+]] {
; CHECK-NEXT: [[ENTRY:.*]]:
; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
-; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 3
; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 1000, [[TMP1]]
; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
@@ -83,17 +83,17 @@ define void @chained_sext_adds(ptr noalias %src, ptr noalias %dst) #0 {
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <vscale x 4 x i32> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP7:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_PHI:%.*]] = phi <vscale x 8 x i32> [ zeroinitializer, %[[VECTOR_PH]] ], [ [[TMP5:%.*]], %[[VECTOR_BODY]] ]
; CHECK-NEXT: [[TMP4:%.*]] = getelementptr i8, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i8>, ptr [[TMP4]], align 1
-; CHECK-NEXT: [[TMP5:%.*]] = sext <vscale x 4 x i8> [[WIDE_LOAD]] to <vscale x 4 x i32>
-; CHECK-NEXT: [[TMP6:%.*]] = add <vscale x 4 x i32> [[VEC_PHI]], [[TMP5]]
-; CHECK-NEXT: [[TMP7]] = add <vscale x 4 x i32> [[TMP6]], [[TMP5]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 8 x i8>, ptr [[TMP4]], align 1
+; CHECK-NEXT: [[TMP3:%.*]] = sext <vscale x 8 x i8> [[WIDE_LOAD]] to <vscale x 8 x i32>
+; CHECK-NEXT: [[TMP6:%.*]] = add <vscale x 8 x i32> [[VEC_PHI]], [[TMP3]]
+; CHECK-NEXT: [[TMP5]] = add <vscale x 8 x i32> [[TMP6]], [[TMP3]]
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
; CHECK-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
-; CHECK-NEXT: [[TMP9:%.*]] = call i32 @llvm.vector.reduce.add.nxv4i32(<vscale x 4 x i32> [[TMP7]])
+; CHECK-NEXT: [[TMP9:%.*]] = call i32 @llvm.vector.reduce.add.nxv8i32(<vscale x 8 x i32> [[TMP5]])
; CHECK-NEXT: store i32 [[TMP9]], ptr [[DST]], align 4
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 1000, [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/predicated-costs.ll b/llvm/test/Transforms/LoopVectorize/AArch64/predicated-costs.ll
index 533240f48967f..285e978a5b9a8 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/predicated-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/predicated-costs.ll
@@ -66,42 +66,62 @@ define void @test_predicated_load_cast_hint(ptr %dst.1, ptr %dst.2, ptr %src, i8
; CHECK-NEXT: [[CONFLICT_RDX15:%.*]] = or i1 [[CONFLICT_RDX]], [[FOUND_CONFLICT14]]
; CHECK-NEXT: br i1 [[CONFLICT_RDX15]], label %[[SCALAR_PH]], label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
-; CHECK-NEXT: [[TMP27:%.*]] = load i8, ptr [[SRC]], align 1, !alias.scope [[META0:![0-9]+]], !noalias [[META3:![0-9]+]]
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <2 x i8> poison, i8 [[TMP27]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <2 x i8> [[BROADCAST_SPLATINSERT]], <2 x i8> poison, <2 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP25:%.*]] = zext <2 x i8> [[BROADCAST_SPLAT]] to <2 x i64>
-; CHECK-NEXT: [[ACTIVE_LANE_MASK_ENTRY:%.*]] = call <2 x i1> @llvm.get.active.lane.mask.v2i1.i32(i32 0, i32 [[TMP2]])
+; CHECK-NEXT: [[TMP25:%.*]] = load i8, ptr [[SRC]], align 1, !alias.scope [[META0:![0-9]+]], !noalias [[META3:![0-9]+]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i8> poison, i8 [[TMP25]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i8> [[BROADCAST_SPLATINSERT]], <4 x i8> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: [[TMP26:%.*]] = zext <4 x i8> [[BROADCAST_SPLAT]] to <4 x i64>
+; CHECK-NEXT: [[ACTIVE_LANE_MASK_ENTRY:%.*]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i32(i32 0, i32 [[TMP2]])
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE22:.*]] ]
-; CHECK-NEXT: [[ACTIVE_LANE_MASK:%.*]] = phi <2 x i1> [ [[ACTIVE_LANE_MASK_ENTRY]], %[[VECTOR_PH]] ], [ [[ACTIVE_LANE_MASK_NEXT:%.*]], %[[PRED_STORE_CONTINUE22]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <2 x i8> [ <i8 0, i8 4>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE22]] ]
-; CHECK-NEXT: [[TMP26:%.*]] = zext <2 x i8> [[VEC_IND]] to <2 x i64>
-; CHECK-NEXT: [[TMP37:%.*]] = extractelement <2 x i1> [[ACTIVE_LANE_MASK]], i64 0
+; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[PRED_STORE_CONTINUE21:.*]] ]
+; CHECK-NEXT: [[ACTIVE_LANE_MASK:%.*]] = phi <4 x i1> [ [[ACTIVE_LANE_MASK_ENTRY]], %[[VECTOR_PH]] ], [ [[ACTIVE_LANE_MASK_NEXT:%.*]], %[[PRED_STORE_CONTINUE21]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i8> [ <i8 0, i8 4, i8 8, i8 12>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[PRED_STORE_CONTINUE21]] ]
+; CHECK-NEXT: [[TMP27:%.*]] = zext <4 x i8> [[VEC_IND]] to <4 x i64>
+; CHECK-NEXT: [[TMP37:%.*]] = extractelement <4 x i1> [[ACTIVE_LANE_MASK]], i64 0
; CHECK-NEXT: br i1 [[TMP37]], label %[[PRED_STORE_IF19:.*]], label %[[PRED_STORE_CONTINUE20:.*]]
; CHECK: [[PRED_STORE_IF19]]:
-; CHECK-NEXT: [[TMP114:%.*]] = extractelement <2 x i64> [[TMP26]], i64 0
+; CHECK-NEXT: [[TMP114:%.*]] = extractelement <4 x i64> [[TMP27]], i64 0
; CHECK-NEXT: [[TMP115:%.*]] = getelementptr [16 x i64], ptr [[DST_1]], i64 [[TMP114]], i64 [[OFF]]
-; CHECK-NEXT: [[TMP116:%.*]] = extractelement <2 x i64> [[TMP25]], i64 0
+; CHECK-NEXT: [[TMP116:%.*]] = extractelement <4 x i64> [[TMP26]], i64 0
; CHECK-NEXT: [[TMP117:%.*]] = or i64 [[TMP116]], 1
; CHECK-NEXT: store i64 [[TMP117]], ptr [[TMP115]], align 8, !alias.scope [[META3]]
; CHECK-NEXT: br label %[[PRED_STORE_CONTINUE20]]
; CHECK: [[PRED_STORE_CONTINUE20]]:
-; CHECK-NEXT: [[TMP42:%.*]] = extractelement <2 x i1> [[ACTIVE_LANE_MASK]], i64 1
-; CHECK-NEXT: br i1 [[TMP42]], label %[[PRED_STORE_IF21:.*]], label %[[PRED_STORE_CONTINUE22]]
+; CHECK-NEXT: [[TMP42:%.*]] = extractelement <4 x i1> [[ACTIVE_LANE_MASK]], i64 1
+; CHECK-NEXT: br i1 [[TMP42]], label %[[PRED_STORE_IF21:.*]], label %[[PRED_STORE_CONTINUE22:.*]]
; CHECK: [[PRED_STORE_IF21]]:
-; CHECK-NEXT: [[TMP120:%.*]] = extractelement <2 x i64> [[TMP26]], i64 1
+; CHECK-NEXT: [[TMP120:%.*]] = extractelement <4 x i64> [[TMP27]], i64 1
; CHECK-NEXT: [[TMP121:%.*]] = getelementptr [16 x i64], ptr [[DST_1]], i64 [[TMP120]], i64 [[OFF]]
-; CHECK-NEXT: [[TMP122:%.*]] = extractelement <2 x i64> [[TMP25]], i64 1
+; CHECK-NEXT: [[TMP122:%.*]] = extractelement <4 x i64> [[TMP26]], i64 1
; CHECK-NEXT: [[TMP123:%.*]] = or i64 [[TMP122]], 1
; CHECK-NEXT: store i64 [[TMP123]], ptr [[TMP121]], align 8, !alias.scope [[META3]]
; CHECK-NEXT: br label %[[PRED_STORE_CONTINUE22]]
; CHECK: [[PRED_STORE_CONTINUE22]]:
-; CHECK-NEXT: [[INDEX_NEXT]] = add i32 [[INDEX]], 2
-; CHECK-NEXT: [[ACTIVE_LANE_MASK_NEXT]] = call <2 x i1> @llvm.get.active.lane.mask.v2i1.i32(i32 [[INDEX_NEXT]], i32 [[TMP2]])
-; CHECK-NEXT: [[TMP47:%.*]] = extractelement <2 x i1> [[ACTIVE_LANE_MASK_NEXT]], i64 0
+; CHECK-NEXT: [[TMP38:%.*]] = extractelement <4 x i1> [[ACTIVE_LANE_MASK]], i64 2
+; CHECK-NEXT: br i1 [[TMP38]], label %[[PRED_STORE_IF18:.*]], label %[[PRED_STORE_CONTINUE19:.*]]
+; CHECK: [[PRED_STORE_IF18]]:
+; CHECK-NEXT: [[TMP39:%.*]] = extractelement <4 x i64> [[TMP27]], i64 2
+; CHECK-NEXT: [[TMP40:%.*]] = getelementptr [16 x i64], ptr [[DST_1]], i64 [[TMP39]], i64 [[OFF]]
+; CHECK-NEXT: [[TMP41:%.*]] = extractelement <4 x i64> [[TMP26]], i64 2
+; CHECK-NEXT: [[TMP49:%.*]] = or i64 [[TMP41]], 1
+; CHECK-NEXT: store i64 [[TMP49]], ptr [[TMP40]], align 8, !alias.scope [[META3]]
+; CHECK-NEXT: br label %[[PRED_STORE_CONTINUE19]]
+; CHECK: [[PRED_STORE_CONTINUE19]]:
+; CHECK-NEXT: [[TMP43:%.*]] = extractelement <4 x i1> [[ACTIVE_LANE_MASK]], i64 3
+; CHECK-NEXT: br i1 [[TMP43]], label %[[PRED_STORE_IF20:.*]], label %[[PRED_STORE_CONTINUE21]]
+; CHECK: [[PRED_STORE_IF20]]:
+; CHECK-NEXT: [[TMP44:%.*]] = extractelement <4 x i64> [[TMP27]], i64 3
+; CHECK-NEXT: [[TMP45:%.*]] = getelementptr [16 x i64], ptr [[DST_1]], i64 [[TMP44]], i64 [[OFF]]
+; CHECK-NEXT: [[TMP46:%.*]] = extractelement <4 x i64> [[TMP26]], i64 3
+; CHECK-NEXT: [[TMP50:%.*]] = or i64 [[TMP46]], 1
+; CHECK-NEXT: store i64 [[TMP50]], ptr [[TMP45]], align 8, !alias.scope [[META3]]
+; CHECK-NEXT: br label %[[PRED_STORE_CONTINUE21]]
+; CHECK: [[PRED_STORE_CONTINUE21]]:
+; CHECK-NEXT: [[INDEX_NEXT]] = add i32 [[INDEX]], 4
+; CHECK-NEXT: [[ACTIVE_LANE_MASK_NEXT]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i32(i32 [[INDEX_NEXT]], i32 [[TMP2]])
+; CHECK-NEXT: [[TMP47:%.*]] = extractelement <4 x i1> [[ACTIVE_LANE_MASK_NEXT]], i64 0
; CHECK-NEXT: [[TMP48:%.*]] = xor i1 [[TMP47]], true
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <2 x i8> [[VEC_IND]], splat (i8 8)
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i8> [[VEC_IND]], splat (i8 16)
; CHECK-NEXT: br i1 [[TMP48]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: store i8 0, ptr [[DST_2]], align 1, !alias.scope [[META9:![0-9]+]], !noalias [[META11:![0-9]+]]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/replicating-load-store-costs.ll b/llvm/test/Transforms/LoopVectorize/AArch64/replicating-load-store-costs.ll
index ee038890245ae..8d492f5d455cc 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/replicating-load-store-costs.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/replicating-load-store-costs.ll
@@ -216,45 +216,32 @@ define void @test_load_gep_widen_induction(ptr noalias %dst, ptr noalias %dst2)
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[OFFSET_IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <2 x i64> [ <i64 0, i64 1>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[STEP_ADD:%.*]] = add nuw <2 x i64> [[VEC_IND]], splat (i64 2)
-; CHECK-NEXT: [[STEP_ADD_2:%.*]] = add nuw <2 x i64> [[STEP_ADD]], splat (i64 2)
-; CHECK-NEXT: [[STEP_ADD_3:%.*]] = add nuw <2 x i64> [[STEP_ADD_2]], splat (i64 2)
-; CHECK-NEXT: [[TMP0:%.*]] = getelementptr i128, ptr [[DST]], <2 x i64> [[VEC_IND]]
-; CHECK-NEXT: [[TMP1:%.*]] = getelementptr i128, ptr [[DST]], <2 x i64> [[STEP_ADD]]
-; CHECK-NEXT: [[TMP2:%.*]] = getelementptr i128, ptr [[DST]], <2 x i64> [[STEP_ADD_2]]
-; CHECK-NEXT: [[TMP3:%.*]] = getelementptr i128, ptr [[DST]], <2 x i64> [[STEP_ADD_3]]
-; CHECK-NEXT: [[TMP5:%.*]] = extractelement <2 x ptr> [[TMP0]], i64 0
-; CHECK-NEXT: store ptr null, ptr [[TMP5]], align 8
-; CHECK-NEXT: [[TMP6:%.*]] = extractelement <2 x ptr> [[TMP0]], i64 1
-; CHECK-NEXT: store ptr null, ptr [[TMP6]], align 8
-; CHECK-NEXT: [[TMP7:%.*]] = extractelement <2 x ptr> [[TMP1]], i64 0
-; CHECK-NEXT: store ptr null, ptr [[TMP7]], align 8
-; CHECK-NEXT: [[TMP8:%.*]] = extractelement <2 x ptr> [[TMP1]], i64 1
-; CHECK-NEXT: store ptr null, ptr [[TMP8]], align 8
-; CHECK-NEXT: [[TMP9:%.*]] = extractelement <2 x ptr> [[TMP2]], i64 0
+; CHECK-NEXT: [[TMP0:%.*]] = add i64 [[OFFSET_IDX]], 1
+; CHECK-NEXT: [[TMP1:%.*]] = add i64 [[OFFSET_IDX]], 2
+; CHECK-NEXT: [[TMP2:%.*]] = add i64 [[OFFSET_IDX]], 3
+; CHECK-NEXT: [[TMP9:%.*]] = getelementptr i128, ptr [[DST]], i64 [[OFFSET_IDX]]
+; CHECK-NEXT: [[TMP10:%.*]] = getelementptr i128, ptr [[DST]], i64 [[TMP0]]
+; CHECK-NEXT: [[TMP11:%.*]] = getelementptr i128, ptr [[DST]], i64 [[TMP1]]
+; CHECK-NEXT: [[TMP17:%.*]] = getelementptr i128, ptr [[DST]], i64 [[TMP2]]
; CHECK-NEXT: store ptr null, ptr [[TMP9]], align 8
-; CHECK-NEXT: [[TMP10:%.*]] = extractelement <2 x ptr> [[TMP2]], i64 1
; CHECK-NEXT: store ptr null, ptr [[TMP10]], align 8
-; CHECK-NEXT: [[TMP11:%.*]] = extractelement <2 x ptr> [[TMP3]], i64 0
; CHECK-NEXT: store ptr null, ptr [[TMP11]], align 8
-; CHECK-NEXT: [[TMP17:%.*]] = extractelement <2 x ptr> [[TMP3]], i64 1
; CHECK-NEXT: store ptr null, ptr [[TMP17]], align 8
; CHECK-NEXT: [[TMP12:%.*]] = getelementptr ptr, ptr [[DST2]], i64 [[OFFSET_IDX]]
-; CHECK-NEXT: [[TMP13:%.*]] = getelementptr ptr, ptr [[TMP12]], i64 2
-; CHECK-NEXT: [[TMP14:%.*]] = getelementptr ptr, ptr [[TMP12]], i64 4
-; CHECK-NEXT: [[TMP15:%.*]] = getelementptr ptr, ptr [[TMP12]], i64 6
-; CHECK-NEXT: store <2 x ptr> [[TMP0]], ptr [[TMP12]], align 8
-; CHECK-NEXT: store <2 x ptr> [[TMP1]], ptr [[TMP13]], align 8
-; CHECK-NEXT: store <2 x ptr> [[TMP2]], ptr [[TMP14]], align 8
-; CHECK-NEXT: store <2 x ptr> [[TMP3]], ptr [[TMP15]], align 8
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[OFFSET_IDX]], 8
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <2 x i64> [[STEP_ADD_3]], splat (i64 2)
-; CHECK-NEXT: [[TMP16:%.*]] = icmp eq i64 [[INDEX_NEXT]], 96
-; CHECK-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK-NEXT: [[TMP8:%.*]] = getelementptr ptr, ptr [[DST2]], i64 [[TMP0]]
+; CHECK-NEXT: [[TMP13:%.*]] = getelementptr ptr, ptr [[DST2]], i64 [[TMP1]]
+; CHECK-NEXT: [[TMP14:%.*]] = getelementptr ptr, ptr [[DST2]], i64 [[TMP2]]
+; CHECK-NEXT: store ptr [[TMP9]], ptr [[TMP12]], align 8
+; CHECK-NEXT: store ptr [[TMP10]], ptr [[TMP8]], align 8
+; CHECK-NEXT: store ptr [[TMP11]], ptr [[TMP13]], align 8
+; CHECK-NEXT: store ptr [[TMP17]], ptr [[TMP14]], align 8
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[OFFSET_IDX]], 4
+; CHECK-NEXT: [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT]], 100
+; CHECK-NEXT: br i1 [[TMP15]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
; CHECK: [[SCALAR_PH]]:
+; CHECK-NEXT: ret void
;
entry:
br label %loop
@@ -314,7 +301,7 @@ define ptr @replicating_store_in_conditional_latch(ptr %p, i32 %n) #0 {
; CHECK-NEXT: store ptr null, ptr [[TMP15]], align 8
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-NEXT: [[TMP16:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
+; CHECK-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: br label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
@@ -478,21 +465,21 @@ define void @test_prefer_vector_addressing(ptr %start, ptr %ms, ptr noalias %src
; CHECK-NEXT: [[NEXT_GEP3:%.*]] = getelementptr i8, ptr [[START]], i64 [[TMP11]]
; CHECK-NEXT: [[NEXT_GEP4:%.*]] = getelementptr i8, ptr [[START]], i64 [[TMP12]]
; CHECK-NEXT: [[NEXT_GEP5:%.*]] = getelementptr i8, ptr [[START]], i64 [[TMP13]]
-; CHECK-NEXT: [[TMP14:%.*]] = load i64, ptr [[NEXT_GEP]], align 1, !tbaa [[LONG_LONG_TBAA12:![0-9]+]]
-; CHECK-NEXT: [[TMP15:%.*]] = load i64, ptr [[NEXT_GEP3]], align 1, !tbaa [[LONG_LONG_TBAA12]]
-; CHECK-NEXT: [[TMP16:%.*]] = load i64, ptr [[NEXT_GEP4]], align 1, !tbaa [[LONG_LONG_TBAA12]]
-; CHECK-NEXT: [[TMP17:%.*]] = load i64, ptr [[NEXT_GEP5]], align 1, !tbaa [[LONG_LONG_TBAA12]]
+; CHECK-NEXT: [[TMP14:%.*]] = load i64, ptr [[NEXT_GEP]], align 1, !tbaa [[LONG_LONG_TBAA11:![0-9]+]]
+; CHECK-NEXT: [[TMP15:%.*]] = load i64, ptr [[NEXT_GEP3]], align 1, !tbaa [[LONG_LONG_TBAA11]]
+; CHECK-NEXT: [[TMP16:%.*]] = load i64, ptr [[NEXT_GEP4]], align 1, !tbaa [[LONG_LONG_TBAA11]]
+; CHECK-NEXT: [[TMP17:%.*]] = load i64, ptr [[NEXT_GEP5]], align 1, !tbaa [[LONG_LONG_TBAA11]]
; CHECK-NEXT: [[TMP18:%.*]] = getelementptr i8, ptr [[SRC]], i64 [[TMP14]]
; CHECK-NEXT: [[TMP19:%.*]] = getelementptr i8, ptr [[SRC]], i64 [[TMP15]]
; CHECK-NEXT: [[TMP20:%.*]] = getelementptr i8, ptr [[SRC]], i64 [[TMP16]]
; CHECK-NEXT: [[TMP21:%.*]] = getelementptr i8, ptr [[SRC]], i64 [[TMP17]]
-; CHECK-NEXT: store i32 0, ptr [[TMP18]], align 4, !tbaa [[INT_TBAA17:![0-9]+]]
-; CHECK-NEXT: store i32 0, ptr [[TMP19]], align 4, !tbaa [[INT_TBAA17]]
-; CHECK-NEXT: store i32 0, ptr [[TMP20]], align 4, !tbaa [[INT_TBAA17]]
-; CHECK-NEXT: store i32 0, ptr [[TMP21]], align 4, !tbaa [[INT_TBAA17]]
+; CHECK-NEXT: store i32 0, ptr [[TMP18]], align 4, !tbaa [[INT_TBAA16:![0-9]+]]
+; CHECK-NEXT: store i32 0, ptr [[TMP19]], align 4, !tbaa [[INT_TBAA16]]
+; CHECK-NEXT: store i32 0, ptr [[TMP20]], align 4, !tbaa [[INT_TBAA16]]
+; CHECK-NEXT: store i32 0, ptr [[TMP21]], align 4, !tbaa [[INT_TBAA16]]
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-NEXT: [[TMP22:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-NEXT: br i1 [[TMP22]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP19:![0-9]+]]
+; CHECK-NEXT: br i1 [[TMP22]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP18:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[TMP6]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
@@ -745,7 +732,7 @@ define i32 @test_or_reduction_with_stride_2(i32 %scale, ptr %src) {
; CHECK-NEXT: [[TMP66]] = or <16 x i32> [[TMP65]], [[VEC_PHI]]
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
; CHECK-NEXT: [[TMP67:%.*]] = icmp eq i64 [[INDEX_NEXT]], 48
-; CHECK-NEXT: br i1 [[TMP67]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP21:![0-9]+]]
+; CHECK-NEXT: br i1 [[TMP67]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP20:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[TMP68:%.*]] = call i32 @llvm.vector.reduce.or.v16i32(<16 x i32> [[TMP66]])
; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/scalable-vectorization-cost-tuning.ll b/llvm/test/Transforms/LoopVectorize/AArch64/scalable-vectorization-cost-tuning.ll
index 33b8f22bd048e..5d5c99ef39f9a 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/scalable-vectorization-cost-tuning.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/scalable-vectorization-cost-tuning.ll
@@ -27,26 +27,26 @@
define void @test0(ptr %a, ptr %b, ptr %c) #0 {
; VSCALEFORTUNING1-LABEL: 'test0'
; VSCALEFORTUNING1: Cost for VF vscale x 1: Invalid (Estimated cost per lane: Invalid)
-; VSCALEFORTUNING1: Cost for VF vscale x 2: 10 (Estimated cost per lane: 5)
-; VSCALEFORTUNING1: Cost for VF vscale x 4: 10 (Estimated cost per lane: 2.5)
-; VSCALEFORTUNING1: Cost for VF vscale x 8: 10 (Estimated cost per lane: 1.25)
-; VSCALEFORTUNING1: Cost for VF vscale x 16: 10 (Estimated cost per lane: 0.625)
+; VSCALEFORTUNING1: Cost for VF vscale x 2: 9 (Estimated cost per lane: 4.5)
+; VSCALEFORTUNING1: Cost for VF vscale x 4: 9 (Estimated cost per lane: 2.25)
+; VSCALEFORTUNING1: Cost for VF vscale x 8: 9 (Estimated cost per lane: 1.13)
+; VSCALEFORTUNING1: Cost for VF vscale x 16: 9 (Estimated cost per lane: 0.563)
; VSCALEFORTUNING1: LV: Selecting VF: vscale x 16.
;
; VSCALEFORTUNING2-LABEL: 'test0'
; VSCALEFORTUNING2: Cost for VF vscale x 1: Invalid (Estimated cost per lane: Invalid)
-; VSCALEFORTUNING2: Cost for VF vscale x 2: 10 (Estimated cost per lane: 2.5)
-; VSCALEFORTUNING2: Cost for VF vscale x 4: 10 (Estimated cost per lane: 1.25)
-; VSCALEFORTUNING2: Cost for VF vscale x 8: 10 (Estimated cost per lane: 0.625)
-; VSCALEFORTUNING2: Cost for VF vscale x 16: 10 (Estimated cost per lane: 0.313)
+; VSCALEFORTUNING2: Cost for VF vscale x 2: 9 (Estimated cost per lane: 2.25)
+; VSCALEFORTUNING2: Cost for VF vscale x 4: 9 (Estimated cost per lane: 1.13)
+; VSCALEFORTUNING2: Cost for VF vscale x 8: 9 (Estimated cost per lane: 0.563)
+; VSCALEFORTUNING2: Cost for VF vscale x 16: 9 (Estimated cost per lane: 0.281)
; VSCALEFORTUNING2: LV: Selecting VF: vscale x 16.
;
; VSCALEFORTUNING1-PREFER-FIXED-LABEL: 'test0'
; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 1: Invalid (Estimated cost per lane: Invalid)
-; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 2: 10 (Estimated cost per lane: 5)
-; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 4: 10 (Estimated cost per lane: 2.5)
-; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 8: 10 (Estimated cost per lane: 1.25)
-; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 16: 10 (Estimated cost per lane: 0.625)
+; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 2: 9 (Estimated cost per lane: 4.5)
+; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 4: 9 (Estimated cost per lane: 2.25)
+; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 8: 9 (Estimated cost per lane: 1.13)
+; VSCALEFORTUNING1-PREFER-FIXED: Cost for VF vscale x 16: 9 (Estimated cost per lane: 0.563)
; VSCALEFORTUNING1-PREFER-FIXED: LV: Selecting VF: 16.
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/ARM/mve-icmpcost.ll b/llvm/test/Transforms/LoopVectorize/ARM/mve-icmpcost.ll
index 44eb3626c79e9..87ef35b7e90fa 100644
--- a/llvm/test/Transforms/LoopVectorize/ARM/mve-icmpcost.ll
+++ b/llvm/test/Transforms/LoopVectorize/ARM/mve-icmpcost.ll
@@ -20,8 +20,7 @@ define void @expensive_icmp(ptr noalias nocapture %d, ptr nocapture readonly %s,
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %inc = add nuw nsw i32 %i.016, 1
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %exitcond.not = icmp eq i32 %inc, %n
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %exitcond.not, label %for.cond.cleanup.loopexit, label %for.body
-; CHECK: Cost of 1 for VF 2: induction instruction %inc = add nuw nsw i32 %i.016, 1
-; CHECK: Cost of 0 for VF 2: induction instruction %i.016 = phi i32 [ 0, %for.body.lr.ph ], [ %inc, %for.inc ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 0 for VF 2: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3:%[0-9]+]]>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: CLONE ir<%arrayidx> = getelementptr inbounds ir<%s>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 2: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i16, ir<%arrayidx>, ir<1>
@@ -46,8 +45,7 @@ define void @expensive_icmp(ptr noalias nocapture %d, ptr nocapture readonly %s,
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost for VF 2: 86 (Estimated cost per lane: 43)
-; CHECK: Cost of 1 for VF 4: induction instruction %inc = add nuw nsw i32 %i.016, 1
-; CHECK: Cost of 0 for VF 4: induction instruction %i.016 = phi i32 [ 0, %for.body.lr.ph ], [ %inc, %for.inc ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 0 for VF 4: vp<[[VP4]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: CLONE ir<%arrayidx> = getelementptr inbounds ir<%s>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 4: vp<[[VP5]]> = vector-pointer inbounds i16, ir<%arrayidx>, ir<1>
@@ -72,8 +70,7 @@ define void @expensive_icmp(ptr noalias nocapture %d, ptr nocapture readonly %s,
; CHECK: Cost of 1 for VF 4: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 4: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost for VF 4: 10 (Estimated cost per lane: 2.5)
-; CHECK: Cost of 1 for VF 8: induction instruction %inc = add nuw nsw i32 %i.016, 1
-; CHECK: Cost of 0 for VF 8: induction instruction %i.016 = phi i32 [ 0, %for.body.lr.ph ], [ %inc, %for.inc ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 0 for VF 8: vp<[[VP4]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 8: CLONE ir<%arrayidx> = getelementptr inbounds ir<%s>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 8: vp<[[VP5]]> = vector-pointer inbounds i16, ir<%arrayidx>, ir<1>
@@ -156,14 +153,13 @@ define void @cheap_icmp(ptr nocapture readonly %pSrcA, ptr nocapture readonly %p
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %dec = add i32 %blkCnt.012, -1
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %cmp.not = icmp eq i32 %dec, 0
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %cmp.not, label %while.end.loopexit, label %while.body
-; CHECK: Cost of 1 for VF 2: induction instruction %dec = add i32 %blkCnt.012, -1
-; CHECK: Cost of 0 for VF 2: induction instruction %blkCnt.012 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 2: induction instruction %incdec.ptr = getelementptr inbounds i8, ptr %pSrcA.addr.011, i32 1
; CHECK: Cost of 0 for VF 2: induction instruction %pSrcA.addr.011 = phi ptr [ %incdec.ptr, %while.body ], [ %pSrcA, %while.body.preheader ]
; CHECK: Cost of 0 for VF 2: induction instruction %incdec.ptr5 = getelementptr inbounds i8, ptr %pDst.addr.010, i32 1
; CHECK: Cost of 0 for VF 2: induction instruction %pDst.addr.010 = phi ptr [ %incdec.ptr5, %while.body ], [ %pDst, %while.body.preheader ]
; CHECK: Cost of 0 for VF 2: induction instruction %incdec.ptr2 = getelementptr inbounds i8, ptr %pSrcB.addr.09, i32 1
; CHECK: Cost of 0 for VF 2: induction instruction %pSrcB.addr.09 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrcB, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 0 for VF 2: vp<[[VP8:%[0-9]+]]> = SCALAR-STEPS vp<[[VP7:%[0-9]+]]>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: EMIT vp<%next.gep> = ptradd ir<%pSrcA>, vp<[[VP8]]>
; CHECK: Cost of 0 for VF 2: vp<[[VP9:%[0-9]+]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
@@ -216,14 +212,13 @@ define void @cheap_icmp(ptr nocapture readonly %pSrcA, ptr nocapture readonly %p
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<%blockSize>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost for VF 2: 130 (Estimated cost per lane: 65)
-; CHECK: Cost of 1 for VF 4: induction instruction %dec = add i32 %blkCnt.012, -1
-; CHECK: Cost of 0 for VF 4: induction instruction %blkCnt.012 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 4: induction instruction %incdec.ptr = getelementptr inbounds i8, ptr %pSrcA.addr.011, i32 1
; CHECK: Cost of 0 for VF 4: induction instruction %pSrcA.addr.011 = phi ptr [ %incdec.ptr, %while.body ], [ %pSrcA, %while.body.preheader ]
; CHECK: Cost of 0 for VF 4: induction instruction %incdec.ptr5 = getelementptr inbounds i8, ptr %pDst.addr.010, i32 1
; CHECK: Cost of 0 for VF 4: induction instruction %pDst.addr.010 = phi ptr [ %incdec.ptr5, %while.body ], [ %pDst, %while.body.preheader ]
; CHECK: Cost of 0 for VF 4: induction instruction %incdec.ptr2 = getelementptr inbounds i8, ptr %pSrcB.addr.09, i32 1
; CHECK: Cost of 0 for VF 4: induction instruction %pSrcB.addr.09 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrcB, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 0 for VF 4: vp<[[VP8]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: EMIT vp<%next.gep> = ptradd ir<%pSrcA>, vp<[[VP8]]>
; CHECK: Cost of 0 for VF 4: vp<[[VP9]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
@@ -276,14 +271,13 @@ define void @cheap_icmp(ptr nocapture readonly %pSrcA, ptr nocapture readonly %p
; CHECK: Cost of 1 for VF 4: EMIT vp<%cmp.n> = icmp eq ir<%blockSize>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 4: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost for VF 4: 14 (Estimated cost per lane: 3.5)
-; CHECK: Cost of 1 for VF 8: induction instruction %dec = add i32 %blkCnt.012, -1
-; CHECK: Cost of 0 for VF 8: induction instruction %blkCnt.012 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 8: induction instruction %incdec.ptr = getelementptr inbounds i8, ptr %pSrcA.addr.011, i32 1
; CHECK: Cost of 0 for VF 8: induction instruction %pSrcA.addr.011 = phi ptr [ %incdec.ptr, %while.body ], [ %pSrcA, %while.body.preheader ]
; CHECK: Cost of 0 for VF 8: induction instruction %incdec.ptr5 = getelementptr inbounds i8, ptr %pDst.addr.010, i32 1
; CHECK: Cost of 0 for VF 8: induction instruction %pDst.addr.010 = phi ptr [ %incdec.ptr5, %while.body ], [ %pDst, %while.body.preheader ]
; CHECK: Cost of 0 for VF 8: induction instruction %incdec.ptr2 = getelementptr inbounds i8, ptr %pSrcB.addr.09, i32 1
; CHECK: Cost of 0 for VF 8: induction instruction %pSrcB.addr.09 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrcB, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 8: canonical IV increment
; CHECK: Cost of 0 for VF 8: vp<[[VP8]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 8: EMIT vp<%next.gep> = ptradd ir<%pSrcA>, vp<[[VP8]]>
; CHECK: Cost of 0 for VF 8: vp<[[VP9]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
@@ -336,14 +330,13 @@ define void @cheap_icmp(ptr nocapture readonly %pSrcA, ptr nocapture readonly %p
; CHECK: Cost of 1 for VF 8: EMIT vp<%cmp.n> = icmp eq ir<%blockSize>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 8: EMIT branch-on-cond vp<%cmp.n>
; CHECK: Cost for VF 8: 26 (Estimated cost per lane: 3.25)
-; CHECK: Cost of 1 for VF 16: induction instruction %dec = add i32 %blkCnt.012, -1
-; CHECK: Cost of 0 for VF 16: induction instruction %blkCnt.012 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 16: induction instruction %incdec.ptr = getelementptr inbounds i8, ptr %pSrcA.addr.011, i32 1
; CHECK: Cost of 0 for VF 16: induction instruction %pSrcA.addr.011 = phi ptr [ %incdec.ptr, %while.body ], [ %pSrcA, %while.body.preheader ]
; CHECK: Cost of 0 for VF 16: induction instruction %incdec.ptr5 = getelementptr inbounds i8, ptr %pDst.addr.010, i32 1
; CHECK: Cost of 0 for VF 16: induction instruction %pDst.addr.010 = phi ptr [ %incdec.ptr5, %while.body ], [ %pDst, %while.body.preheader ]
; CHECK: Cost of 0 for VF 16: induction instruction %incdec.ptr2 = getelementptr inbounds i8, ptr %pSrcB.addr.09, i32 1
; CHECK: Cost of 0 for VF 16: induction instruction %pSrcB.addr.09 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrcB, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 16: canonical IV increment
; CHECK: Cost of 0 for VF 16: vp<[[VP8]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 16: EMIT vp<%next.gep> = ptradd ir<%pSrcA>, vp<[[VP8]]>
; CHECK: Cost of 0 for VF 16: vp<[[VP9]]> = SCALAR-STEPS vp<[[VP7]]>, ir<1>, vp<[[VP0]]>
@@ -453,10 +446,9 @@ define void @floatcmp(ptr nocapture readonly %pSrc, ptr nocapture %pDst, i32 %bl
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %cmp.not, label %while.end.loopexit, label %while.body
; CHECK: Cost of 0 for VF 2: induction instruction %incdec.ptr2 = getelementptr inbounds float, ptr %pSrc.addr.010, i32 1
; CHECK: Cost of 0 for VF 2: induction instruction %pSrc.addr.010 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrc, %while.body.preheader ]
-; CHECK: Cost of 1 for VF 2: induction instruction %dec = add i32 %blockSize.addr.09, -1
-; CHECK: Cost of 0 for VF 2: induction instruction %blockSize.addr.09 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 2: induction instruction %incdec.ptr = getelementptr inbounds i32, ptr %pDst.addr.08, i32 1
; CHECK: Cost of 0 for VF 2: induction instruction %pDst.addr.08 = phi ptr [ %incdec.ptr, %while.body ], [ %pDst, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 1 for VF 2: vp<[[VP7:%[0-9]+]]> = DERIVED-IV ir<0> + vp<[[VP6:%[0-9]+]]> * ir<4>
; CHECK: Cost of 0 for VF 2: vp<[[VP8:%[0-9]+]]> = SCALAR-STEPS vp<[[VP7]]>, ir<4>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: EMIT vp<%next.gep> = ptradd ir<%pSrc>, vp<[[VP8]]>
@@ -496,10 +488,9 @@ define void @floatcmp(ptr nocapture readonly %pSrc, ptr nocapture %pDst, i32 %bl
; CHECK: Cost for VF 2: 84 (Estimated cost per lane: 42)
; CHECK: Cost of 0 for VF 4: induction instruction %incdec.ptr2 = getelementptr inbounds float, ptr %pSrc.addr.010, i32 1
; CHECK: Cost of 0 for VF 4: induction instruction %pSrc.addr.010 = phi ptr [ %incdec.ptr2, %while.body ], [ %pSrc, %while.body.preheader ]
-; CHECK: Cost of 1 for VF 4: induction instruction %dec = add i32 %blockSize.addr.09, -1
-; CHECK: Cost of 0 for VF 4: induction instruction %blockSize.addr.09 = phi i32 [ %dec, %while.body ], [ %blockSize, %while.body.preheader ]
; CHECK: Cost of 0 for VF 4: induction instruction %incdec.ptr = getelementptr inbounds i32, ptr %pDst.addr.08, i32 1
; CHECK: Cost of 0 for VF 4: induction instruction %pDst.addr.08 = phi ptr [ %incdec.ptr, %while.body ], [ %pDst, %while.body.preheader ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 1 for VF 4: vp<[[VP7]]> = DERIVED-IV ir<0> + vp<[[VP6]]> * ir<4>
; CHECK: Cost of 0 for VF 4: vp<[[VP8]]> = SCALAR-STEPS vp<[[VP7]]>, ir<4>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: EMIT vp<%next.gep> = ptradd ir<%pSrc>, vp<[[VP8]]>
diff --git a/llvm/test/Transforms/LoopVectorize/ARM/optsize_minsize.ll b/llvm/test/Transforms/LoopVectorize/ARM/optsize_minsize.ll
index e7683c6423575..13c6ea765568b 100644
--- a/llvm/test/Transforms/LoopVectorize/ARM/optsize_minsize.ll
+++ b/llvm/test/Transforms/LoopVectorize/ARM/optsize_minsize.ll
@@ -181,137 +181,24 @@ for.cond.cleanup:
define void @tail_predicate_without_optsize(ptr %p, i8 %a, i8 %b, i8 %c, i32 %n) {
; DEFAULT-LABEL: define void @tail_predicate_without_optsize(
; DEFAULT-SAME: ptr [[P:%.*]], i8 [[A:%.*]], i8 [[B:%.*]], i8 [[C:%.*]], i32 [[N:%.*]]) {
-; DEFAULT-NEXT: [[ENTRY:.*:]]
-; DEFAULT-NEXT: br label %[[VECTOR_PH:.*]]
-; DEFAULT: [[VECTOR_PH]]:
-; DEFAULT-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i8> poison, i8 [[A]], i64 0
-; DEFAULT-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i8> [[BROADCAST_SPLATINSERT]], <16 x i8> poison, <16 x i32> zeroinitializer
-; DEFAULT-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <16 x i8> poison, i8 [[B]], i64 0
-; DEFAULT-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <16 x i8> [[BROADCAST_SPLATINSERT1]], <16 x i8> poison, <16 x i32> zeroinitializer
-; DEFAULT-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <16 x i8> poison, i8 [[C]], i64 0
-; DEFAULT-NEXT: [[BROADCAST_SPLAT4:%.*]] = shufflevector <16 x i8> [[BROADCAST_SPLATINSERT3]], <16 x i8> poison, <16 x i32> zeroinitializer
-; DEFAULT-NEXT: br label %[[VECTOR_BODY:.*]]
-; DEFAULT: [[VECTOR_BODY]]:
-; DEFAULT-NEXT: [[TMP0:%.*]] = mul <16 x i8> [[BROADCAST_SPLAT]], <i8 0, i8 1, i8 2, i8 3, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15>
-; DEFAULT-NEXT: [[TMP1:%.*]] = mul <16 x i8> <i8 0, i8 0, i8 1, i8 1, i8 2, i8 2, i8 3, i8 3, i8 4, i8 4, i8 5, i8 5, i8 6, i8 6, i8 7, i8 7>, [[BROADCAST_SPLAT2]]
-; DEFAULT-NEXT: [[TMP2:%.*]] = add <16 x i8> [[TMP1]], [[TMP0]]
-; DEFAULT-NEXT: [[TMP3:%.*]] = mul <16 x i8> <i8 0, i8 0, i8 0, i8 0, i8 1, i8 1, i8 1, i8 1, i8 2, i8 2, i8 2, i8 2, i8 3, i8 3, i8 3, i8 3>, [[BROADCAST_SPLAT4]]
-; DEFAULT-NEXT: [[TMP4:%.*]] = add <16 x i8> [[TMP2]], [[TMP3]]
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF:.*]], label %[[PRED_STORE_CONTINUE:.*]]
-; DEFAULT: [[PRED_STORE_IF]]:
-; DEFAULT-NEXT: [[TMP5:%.*]] = extractelement <16 x i8> [[TMP4]], i64 0
-; DEFAULT-NEXT: store i8 [[TMP5]], ptr [[P]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE]]
-; DEFAULT: [[PRED_STORE_CONTINUE]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF5:.*]], label %[[PRED_STORE_CONTINUE6:.*]]
-; DEFAULT: [[PRED_STORE_IF5]]:
-; DEFAULT-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 1
-; DEFAULT-NEXT: [[TMP7:%.*]] = extractelement <16 x i8> [[TMP4]], i64 1
-; DEFAULT-NEXT: store i8 [[TMP7]], ptr [[TMP6]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE6]]
-; DEFAULT: [[PRED_STORE_CONTINUE6]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF7:.*]], label %[[PRED_STORE_CONTINUE8:.*]]
-; DEFAULT: [[PRED_STORE_IF7]]:
-; DEFAULT-NEXT: [[TMP8:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 2
-; DEFAULT-NEXT: [[TMP9:%.*]] = extractelement <16 x i8> [[TMP4]], i64 2
-; DEFAULT-NEXT: store i8 [[TMP9]], ptr [[TMP8]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE8]]
-; DEFAULT: [[PRED_STORE_CONTINUE8]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF9:.*]], label %[[PRED_STORE_CONTINUE10:.*]]
-; DEFAULT: [[PRED_STORE_IF9]]:
-; DEFAULT-NEXT: [[TMP10:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 3
-; DEFAULT-NEXT: [[TMP11:%.*]] = extractelement <16 x i8> [[TMP4]], i64 3
-; DEFAULT-NEXT: store i8 [[TMP11]], ptr [[TMP10]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE10]]
-; DEFAULT: [[PRED_STORE_CONTINUE10]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF11:.*]], label %[[PRED_STORE_CONTINUE12:.*]]
-; DEFAULT: [[PRED_STORE_IF11]]:
-; DEFAULT-NEXT: [[TMP12:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 4
-; DEFAULT-NEXT: [[TMP13:%.*]] = extractelement <16 x i8> [[TMP4]], i64 4
-; DEFAULT-NEXT: store i8 [[TMP13]], ptr [[TMP12]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE12]]
-; DEFAULT: [[PRED_STORE_CONTINUE12]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF13:.*]], label %[[PRED_STORE_CONTINUE14:.*]]
-; DEFAULT: [[PRED_STORE_IF13]]:
-; DEFAULT-NEXT: [[TMP14:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 5
-; DEFAULT-NEXT: [[TMP15:%.*]] = extractelement <16 x i8> [[TMP4]], i64 5
-; DEFAULT-NEXT: store i8 [[TMP15]], ptr [[TMP14]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE14]]
-; DEFAULT: [[PRED_STORE_CONTINUE14]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF15:.*]], label %[[PRED_STORE_CONTINUE16:.*]]
-; DEFAULT: [[PRED_STORE_IF15]]:
-; DEFAULT-NEXT: [[TMP16:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 6
-; DEFAULT-NEXT: [[TMP17:%.*]] = extractelement <16 x i8> [[TMP4]], i64 6
-; DEFAULT-NEXT: store i8 [[TMP17]], ptr [[TMP16]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE16]]
-; DEFAULT: [[PRED_STORE_CONTINUE16]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF17:.*]], label %[[PRED_STORE_CONTINUE18:.*]]
-; DEFAULT: [[PRED_STORE_IF17]]:
-; DEFAULT-NEXT: [[TMP18:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 7
-; DEFAULT-NEXT: [[TMP19:%.*]] = extractelement <16 x i8> [[TMP4]], i64 7
-; DEFAULT-NEXT: store i8 [[TMP19]], ptr [[TMP18]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE18]]
-; DEFAULT: [[PRED_STORE_CONTINUE18]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF19:.*]], label %[[PRED_STORE_CONTINUE20:.*]]
-; DEFAULT: [[PRED_STORE_IF19]]:
-; DEFAULT-NEXT: [[TMP20:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 8
-; DEFAULT-NEXT: [[TMP21:%.*]] = extractelement <16 x i8> [[TMP4]], i64 8
-; DEFAULT-NEXT: store i8 [[TMP21]], ptr [[TMP20]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE20]]
-; DEFAULT: [[PRED_STORE_CONTINUE20]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF21:.*]], label %[[PRED_STORE_CONTINUE22:.*]]
-; DEFAULT: [[PRED_STORE_IF21]]:
-; DEFAULT-NEXT: [[TMP22:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 9
-; DEFAULT-NEXT: [[TMP23:%.*]] = extractelement <16 x i8> [[TMP4]], i64 9
-; DEFAULT-NEXT: store i8 [[TMP23]], ptr [[TMP22]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE22]]
-; DEFAULT: [[PRED_STORE_CONTINUE22]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF23:.*]], label %[[PRED_STORE_CONTINUE24:.*]]
-; DEFAULT: [[PRED_STORE_IF23]]:
-; DEFAULT-NEXT: [[TMP24:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 10
-; DEFAULT-NEXT: [[TMP25:%.*]] = extractelement <16 x i8> [[TMP4]], i64 10
-; DEFAULT-NEXT: store i8 [[TMP25]], ptr [[TMP24]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE24]]
-; DEFAULT: [[PRED_STORE_CONTINUE24]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF25:.*]], label %[[PRED_STORE_CONTINUE26:.*]]
-; DEFAULT: [[PRED_STORE_IF25]]:
-; DEFAULT-NEXT: [[TMP26:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 11
-; DEFAULT-NEXT: [[TMP27:%.*]] = extractelement <16 x i8> [[TMP4]], i64 11
-; DEFAULT-NEXT: store i8 [[TMP27]], ptr [[TMP26]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE26]]
-; DEFAULT: [[PRED_STORE_CONTINUE26]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF27:.*]], label %[[PRED_STORE_CONTINUE28:.*]]
-; DEFAULT: [[PRED_STORE_IF27]]:
-; DEFAULT-NEXT: [[TMP28:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 12
-; DEFAULT-NEXT: [[TMP29:%.*]] = extractelement <16 x i8> [[TMP4]], i64 12
-; DEFAULT-NEXT: store i8 [[TMP29]], ptr [[TMP28]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE28]]
-; DEFAULT: [[PRED_STORE_CONTINUE28]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF29:.*]], label %[[PRED_STORE_CONTINUE30:.*]]
-; DEFAULT: [[PRED_STORE_IF29]]:
-; DEFAULT-NEXT: [[TMP30:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 13
-; DEFAULT-NEXT: [[TMP31:%.*]] = extractelement <16 x i8> [[TMP4]], i64 13
-; DEFAULT-NEXT: store i8 [[TMP31]], ptr [[TMP30]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE30]]
-; DEFAULT: [[PRED_STORE_CONTINUE30]]:
-; DEFAULT-NEXT: br i1 true, label %[[PRED_STORE_IF31:.*]], label %[[PRED_STORE_CONTINUE32:.*]]
-; DEFAULT: [[PRED_STORE_IF31]]:
-; DEFAULT-NEXT: [[TMP32:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 14
-; DEFAULT-NEXT: [[TMP33:%.*]] = extractelement <16 x i8> [[TMP4]], i64 14
-; DEFAULT-NEXT: store i8 [[TMP33]], ptr [[TMP32]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE32]]
-; DEFAULT: [[PRED_STORE_CONTINUE32]]:
-; DEFAULT-NEXT: br i1 false, label %[[PRED_STORE_IF33:.*]], label %[[PRED_STORE_CONTINUE34:.*]]
-; DEFAULT: [[PRED_STORE_IF33]]:
-; DEFAULT-NEXT: [[TMP34:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 15
-; DEFAULT-NEXT: [[TMP35:%.*]] = extractelement <16 x i8> [[TMP4]], i64 15
-; DEFAULT-NEXT: store i8 [[TMP35]], ptr [[TMP34]], align 1
-; DEFAULT-NEXT: br label %[[PRED_STORE_CONTINUE34]]
-; DEFAULT: [[PRED_STORE_CONTINUE34]]:
-; DEFAULT-NEXT: br label %[[MIDDLE_BLOCK:.*]]
-; DEFAULT: [[MIDDLE_BLOCK]]:
-; DEFAULT-NEXT: br label %[[FOR_COND_CLEANUP:.*]]
-; DEFAULT: [[FOR_COND_CLEANUP]]:
+; DEFAULT-NEXT: [[PRED_STORE_IF32:.*]]:
+; DEFAULT-NEXT: br label %[[FOR_BODY:.*]]
+; DEFAULT: [[FOR_BODY]]:
+; DEFAULT-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[PRED_STORE_IF32]] ], [ [[IV_NEXT:%.*]], %[[FOR_BODY]] ]
+; DEFAULT-NEXT: [[TMP0:%.*]] = trunc nuw nsw i64 [[IV]] to i8
+; DEFAULT-NEXT: [[MUL:%.*]] = mul i8 [[A]], [[TMP0]]
+; DEFAULT-NEXT: [[SHR:%.*]] = lshr i8 [[TMP0]], 1
+; DEFAULT-NEXT: [[MUL5:%.*]] = mul i8 [[SHR]], [[B]]
+; DEFAULT-NEXT: [[ADD:%.*]] = add i8 [[MUL5]], [[MUL]]
+; DEFAULT-NEXT: [[SHR7:%.*]] = lshr i8 [[TMP0]], 2
+; DEFAULT-NEXT: [[MUL9:%.*]] = mul i8 [[SHR7]], [[C]]
+; DEFAULT-NEXT: [[TMP37:%.*]] = add i8 [[ADD]], [[MUL9]]
+; DEFAULT-NEXT: [[TMP36:%.*]] = getelementptr inbounds i8, ptr [[P]], i64 [[IV]]
+; DEFAULT-NEXT: store i8 [[TMP37]], ptr [[TMP36]], align 1
+; DEFAULT-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; DEFAULT-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 15
+; DEFAULT-NEXT: br i1 [[EXITCOND_NOT]], label %[[FOR_COND_CLEANUP1:.*]], label %[[FOR_BODY]]
+; DEFAULT: [[FOR_COND_CLEANUP1]]:
; DEFAULT-NEXT: ret void
;
; OPTSIZE-LABEL: define void @tail_predicate_without_optsize(
diff --git a/llvm/test/Transforms/LoopVectorize/PowerPC/reg-usage.ll b/llvm/test/Transforms/LoopVectorize/PowerPC/reg-usage.ll
index f8a9f41338489..4a2b91d96cae5 100644
--- a/llvm/test/Transforms/LoopVectorize/PowerPC/reg-usage.ll
+++ b/llvm/test/Transforms/LoopVectorize/PowerPC/reg-usage.ll
@@ -79,7 +79,8 @@ for.body:
define i64 @bar(ptr nocapture %a) {
; CHECK-LABEL: bar
-; CHECK: Executing best plan with VF=2, UF=8
+; CHECK-PWR8: Executing best plan with VF=2, UF=8
+; CHECK-PWR9: Executing best plan with VF=1, UF=4
entry:
br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/early-exit-live-out.ll b/llvm/test/Transforms/LoopVectorize/RISCV/early-exit-live-out.ll
index 97e2ba11f7b01..4ae1a16de9ca0 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/early-exit-live-out.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/early-exit-live-out.ll
@@ -175,8 +175,7 @@ define i64 @strided_search(ptr align 8 dereferenceable(14784) %p) {
; RV64-NEXT: [[ENTRY:.*]]:
; RV64-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
; RV64-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 1
-; RV64-NEXT: [[UMAX:%.*]] = call i64 @llvm.umax.i64(i64 [[TMP1]], i64 3)
-; RV64-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 132, [[UMAX]]
+; RV64-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 132, [[TMP1]]
; RV64-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; RV64: [[VECTOR_PH]]:
; RV64-NEXT: [[N_MOD_VF:%.*]] = urem i64 132, [[TMP1]]
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/pr88802.ll b/llvm/test/Transforms/LoopVectorize/RISCV/pr88802.ll
index 140ce955d4b49..0a8e4133892fd 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/pr88802.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/pr88802.ll
@@ -7,31 +7,31 @@ define void @test(ptr %p, i64 %a, i8 %b) {
; CHECK-NEXT: entry:
; CHECK-NEXT: br label [[VECTOR_PH:%.*]]
; CHECK: vector.ph:
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 2 x i8> poison, i8 [[B]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 2 x i8> [[BROADCAST_SPLATINSERT]], <vscale x 2 x i8> poison, <vscale x 2 x i32> zeroinitializer
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 8 x i8> poison, i8 [[B]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 8 x i8> [[BROADCAST_SPLATINSERT]], <vscale x 8 x i8> poison, <vscale x 8 x i32> zeroinitializer
; CHECK-NEXT: [[TMP0:%.*]] = shl i64 [[A]], 48
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <vscale x 2 x i64> poison, i64 [[TMP0]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <vscale x 2 x i64> [[BROADCAST_SPLATINSERT1]], <vscale x 2 x i64> poison, <vscale x 2 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP6:%.*]] = ashr <vscale x 2 x i64> [[BROADCAST_SPLAT2]], splat (i64 52)
-; CHECK-NEXT: [[TMP7:%.*]] = trunc <vscale x 2 x i64> [[TMP6]] to <vscale x 2 x i32>
-; CHECK-NEXT: [[TMP8:%.*]] = zext <vscale x 2 x i8> [[BROADCAST_SPLAT]] to <vscale x 2 x i32>
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <vscale x 2 x ptr> poison, ptr [[P]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT4:%.*]] = shufflevector <vscale x 2 x ptr> [[BROADCAST_SPLATINSERT3]], <vscale x 2 x ptr> poison, <vscale x 2 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP9:%.*]] = call <vscale x 2 x i32> @llvm.stepvector.nxv2i32()
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <vscale x 8 x i64> poison, i64 [[TMP0]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <vscale x 8 x i64> [[BROADCAST_SPLATINSERT1]], <vscale x 8 x i64> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: [[TMP1:%.*]] = ashr <vscale x 8 x i64> [[BROADCAST_SPLAT2]], splat (i64 52)
+; CHECK-NEXT: [[TMP2:%.*]] = trunc <vscale x 8 x i64> [[TMP1]] to <vscale x 8 x i32>
+; CHECK-NEXT: [[TMP3:%.*]] = zext <vscale x 8 x i8> [[BROADCAST_SPLAT]] to <vscale x 8 x i32>
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <vscale x 8 x ptr> poison, ptr [[P]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT4:%.*]] = shufflevector <vscale x 8 x ptr> [[BROADCAST_SPLATINSERT3]], <vscale x 8 x ptr> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: [[TMP4:%.*]] = call <vscale x 8 x i32> @llvm.stepvector.nxv8i32()
; CHECK-NEXT: br label [[FOR_COND:%.*]]
; CHECK: vector.body:
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <vscale x 2 x i32> [ [[TMP9]], [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[FOR_COND]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <vscale x 8 x i32> [ [[TMP4]], [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[FOR_COND]] ]
; CHECK-NEXT: [[AVL:%.*]] = phi i32 [ 9, [[VECTOR_PH]] ], [ [[AVL_NEXT:%.*]], [[FOR_COND]] ]
-; CHECK-NEXT: [[TMP11:%.*]] = call i32 @llvm.experimental.get.vector.length.i32(i32 [[AVL]], i32 2, i1 true)
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT7:%.*]] = insertelement <vscale x 2 x i32> poison, i32 [[TMP11]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT8:%.*]] = shufflevector <vscale x 2 x i32> [[BROADCAST_SPLATINSERT7]], <vscale x 2 x i32> poison, <vscale x 2 x i32> zeroinitializer
-; CHECK-NEXT: [[TMP12:%.*]] = icmp slt <vscale x 2 x i32> [[VEC_IND]], splat (i32 2)
-; CHECK-NEXT: [[PREDPHI:%.*]] = select <vscale x 2 x i1> [[TMP12]], <vscale x 2 x i32> [[TMP8]], <vscale x 2 x i32> [[TMP7]]
-; CHECK-NEXT: [[TMP16:%.*]] = shl <vscale x 2 x i32> [[PREDPHI]], splat (i32 8)
-; CHECK-NEXT: [[TMP17:%.*]] = trunc <vscale x 2 x i32> [[TMP16]] to <vscale x 2 x i8>
-; CHECK-NEXT: call void @llvm.vp.scatter.nxv2i8.nxv2p0(<vscale x 2 x i8> [[TMP17]], <vscale x 2 x ptr> align 1 [[BROADCAST_SPLAT4]], <vscale x 2 x i1> splat (i1 true), i32 [[TMP11]])
+; CHECK-NEXT: [[TMP11:%.*]] = call i32 @llvm.experimental.get.vector.length.i32(i32 [[AVL]], i32 8, i1 true)
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT5:%.*]] = insertelement <vscale x 8 x i32> poison, i32 [[TMP11]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT6:%.*]] = shufflevector <vscale x 8 x i32> [[BROADCAST_SPLATINSERT5]], <vscale x 8 x i32> poison, <vscale x 8 x i32> zeroinitializer
+; CHECK-NEXT: [[TMP6:%.*]] = icmp slt <vscale x 8 x i32> [[VEC_IND]], splat (i32 2)
+; CHECK-NEXT: [[PREDPHI:%.*]] = select <vscale x 8 x i1> [[TMP6]], <vscale x 8 x i32> [[TMP3]], <vscale x 8 x i32> [[TMP2]]
+; CHECK-NEXT: [[TMP7:%.*]] = shl <vscale x 8 x i32> [[PREDPHI]], splat (i32 8)
+; CHECK-NEXT: [[TMP8:%.*]] = trunc <vscale x 8 x i32> [[TMP7]] to <vscale x 8 x i8>
+; CHECK-NEXT: call void @llvm.vp.scatter.nxv8i8.nxv8p0(<vscale x 8 x i8> [[TMP8]], <vscale x 8 x ptr> align 1 [[BROADCAST_SPLAT4]], <vscale x 8 x i1> splat (i1 true), i32 [[TMP11]])
; CHECK-NEXT: [[AVL_NEXT]] = sub nuw i32 [[AVL]], [[TMP11]]
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <vscale x 2 x i32> [[VEC_IND]], [[BROADCAST_SPLAT8]]
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <vscale x 8 x i32> [[VEC_IND]], [[BROADCAST_SPLAT6]]
; CHECK-NEXT: [[TMP21:%.*]] = icmp eq i32 [[AVL_NEXT]], 0
; CHECK-NEXT: br i1 [[TMP21]], label [[MIDDLE_BLOCK:%.*]], label [[FOR_COND]], !llvm.loop [[LOOP0:![0-9]+]]
; CHECK: middle.block:
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/uniform-load-store.ll b/llvm/test/Transforms/LoopVectorize/RISCV/uniform-load-store.ll
index c7584291766a8..966f78f371b59 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/uniform-load-store.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/uniform-load-store.ll
@@ -581,23 +581,18 @@ define void @uniform_store_of_loop_varying(ptr noalias nocapture %a, ptr noalias
; FIXEDLEN-NEXT: [[ENTRY:.*:]]
; FIXEDLEN-NEXT: br label %[[VECTOR_PH:.*]]
; FIXEDLEN: [[VECTOR_PH]]:
-; FIXEDLEN-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <2 x ptr> poison, ptr [[B]], i64 0
-; FIXEDLEN-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <2 x ptr> [[BROADCAST_SPLATINSERT]], <2 x ptr> poison, <2 x i32> zeroinitializer
-; FIXEDLEN-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <2 x i64> poison, i64 [[V]], i64 0
-; FIXEDLEN-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <2 x i64> [[BROADCAST_SPLATINSERT1]], <2 x i64> poison, <2 x i32> zeroinitializer
+; FIXEDLEN-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[V]], i64 0
+; FIXEDLEN-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
; FIXEDLEN-NEXT: br label %[[VECTOR_BODY:.*]]
; FIXEDLEN: [[VECTOR_BODY]]:
; FIXEDLEN-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; FIXEDLEN-NEXT: [[VEC_IND:%.*]] = phi <2 x i64> [ <i64 0, i64 1>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; FIXEDLEN-NEXT: [[STEP_ADD:%.*]] = add nuw <2 x i64> [[VEC_IND]], splat (i64 2)
-; FIXEDLEN-NEXT: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> [[VEC_IND]], <2 x ptr> align 8 [[BROADCAST_SPLAT]], <2 x i1> splat (i1 true))
-; FIXEDLEN-NEXT: call void @llvm.masked.scatter.v2i64.v2p0(<2 x i64> [[STEP_ADD]], <2 x ptr> align 8 [[BROADCAST_SPLAT]], <2 x i1> splat (i1 true))
+; FIXEDLEN-NEXT: [[TMP0:%.*]] = add i64 [[INDEX]], 7
+; FIXEDLEN-NEXT: store i64 [[TMP0]], ptr [[B]], align 8
; FIXEDLEN-NEXT: [[TMP5:%.*]] = getelementptr inbounds i64, ptr [[A]], i64 [[INDEX]]
-; FIXEDLEN-NEXT: [[TMP1:%.*]] = getelementptr inbounds i64, ptr [[TMP5]], i64 2
-; FIXEDLEN-NEXT: store <2 x i64> [[BROADCAST_SPLAT2]], ptr [[TMP5]], align 8
-; FIXEDLEN-NEXT: store <2 x i64> [[BROADCAST_SPLAT2]], ptr [[TMP1]], align 8
-; FIXEDLEN-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; FIXEDLEN-NEXT: [[VEC_IND_NEXT]] = add nuw nsw <2 x i64> [[STEP_ADD]], splat (i64 2)
+; FIXEDLEN-NEXT: [[TMP2:%.*]] = getelementptr inbounds i64, ptr [[TMP5]], i64 4
+; FIXEDLEN-NEXT: store <4 x i64> [[BROADCAST_SPLAT]], ptr [[TMP5]], align 8
+; FIXEDLEN-NEXT: store <4 x i64> [[BROADCAST_SPLAT]], ptr [[TMP2]], align 8
+; FIXEDLEN-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
; FIXEDLEN-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
; FIXEDLEN-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
; FIXEDLEN: [[MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/WebAssembly/memory-interleave.ll b/llvm/test/Transforms/LoopVectorize/WebAssembly/memory-interleave.ll
index d123a62d443cc..14e3967f0a387 100644
--- a/llvm/test/Transforms/LoopVectorize/WebAssembly/memory-interleave.ll
+++ b/llvm/test/Transforms/LoopVectorize/WebAssembly/memory-interleave.ll
@@ -31,7 +31,7 @@ define hidden void @two_ints_same_op(ptr noalias nocapture noundef writeonly %0,
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 2: 27 (Estimated cost per lane: 13.5)
+; CHECK: Cost for VF 2: 26 (Estimated cost per lane: 13)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -41,7 +41,7 @@ define hidden void @two_ints_same_op(ptr noalias nocapture noundef writeonly %0,
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 4: 24 (Estimated cost per lane: 6)
+; CHECK: Cost for VF 4: 23 (Estimated cost per lane: 5.75)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -83,7 +83,7 @@ define hidden void @two_ints_vary_op(ptr noalias nocapture noundef writeonly %0,
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 2: 27 (Estimated cost per lane: 13.5)
+; CHECK: Cost for VF 2: 26 (Estimated cost per lane: 13)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -93,7 +93,7 @@ define hidden void @two_ints_vary_op(ptr noalias nocapture noundef writeonly %0,
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 4: 24 (Estimated cost per lane: 6)
+; CHECK: Cost for VF 4: 23 (Estimated cost per lane: 5.75)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -135,7 +135,7 @@ define hidden void @three_ints(ptr noalias nocapture noundef writeonly %0, ptr n
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%13>, ir<%14>
@@ -145,7 +145,7 @@ define hidden void @three_ints(ptr noalias nocapture noundef writeonly %0, ptr n
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 4: 116 (Estimated cost per lane: 29)
+; CHECK: Cost for VF 4: 115 (Estimated cost per lane: 28.8)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -194,7 +194,7 @@ define hidden void @three_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%13>, ir<%14>
@@ -204,7 +204,7 @@ define hidden void @three_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 4: 116 (Estimated cost per lane: 29)
+; CHECK: Cost for VF 4: 115 (Estimated cost per lane: 28.8)
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%13>, ir<%14>
@@ -214,7 +214,7 @@ define hidden void @three_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 8: 224 (Estimated cost per lane: 28)
+; CHECK: Cost for VF 8: 223 (Estimated cost per lane: 27.9)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -269,7 +269,7 @@ define hidden void @four_shorts_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -285,7 +285,7 @@ define hidden void @four_shorts_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 62 (Estimated cost per lane: 15.5)
+; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
; CHECK: Cost of 68 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -301,7 +301,7 @@ define hidden void @four_shorts_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 212 (Estimated cost per lane: 26.5)
+; CHECK: Cost for VF 8: 211 (Estimated cost per lane: 26.4)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -363,7 +363,7 @@ define hidden void @four_shorts_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -379,7 +379,7 @@ define hidden void @four_shorts_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 62 (Estimated cost per lane: 15.5)
+; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
; CHECK: Cost of 68 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -395,7 +395,7 @@ define hidden void @four_shorts_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 212 (Estimated cost per lane: 26.5)
+; CHECK: Cost for VF 8: 211 (Estimated cost per lane: 26.4)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -457,7 +457,7 @@ define hidden void @four_shorts_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -473,7 +473,7 @@ define hidden void @four_shorts_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 62 (Estimated cost per lane: 15.5)
+; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
; CHECK: Cost of 68 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -489,7 +489,7 @@ define hidden void @four_shorts_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 212 (Estimated cost per lane: 26.5)
+; CHECK: Cost for VF 8: 211 (Estimated cost per lane: 26.4)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -551,7 +551,7 @@ define hidden void @five_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%34> = load ir<%33>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%36> = load ir<%35>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%37>, ir<%38>
-; CHECK: Cost for VF 2: 100 (Estimated cost per lane: 50)
+; CHECK: Cost for VF 2: 99 (Estimated cost per lane: 49.5)
; CHECK: Cost of 42 for VF 4: INTERLEAVE-GROUP with factor 5, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -570,7 +570,7 @@ define hidden void @five_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
; CHECK: store ir<%37> to index 4
-; CHECK: Cost for VF 4: 135 (Estimated cost per lane: 33.8)
+; CHECK: Cost for VF 4: 134 (Estimated cost per lane: 33.5)
; CHECK: Cost of 84 for VF 8: INTERLEAVE-GROUP with factor 5, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -589,7 +589,7 @@ define hidden void @five_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
; CHECK: store ir<%37> to index 4
-; CHECK: Cost for VF 8: 261 (Estimated cost per lane: 32.6)
+; CHECK: Cost for VF 8: 260 (Estimated cost per lane: 32.5)
; CHECK: Cost of 168 for VF 16: INTERLEAVE-GROUP with factor 5, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -608,7 +608,7 @@ define hidden void @five_shorts(ptr noalias nocapture noundef writeonly %0, ptr
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
; CHECK: store ir<%37> to index 4
-; CHECK: Cost for VF 16: 513 (Estimated cost per lane: 32.1)
+; CHECK: Cost for VF 16: 512 (Estimated cost per lane: 32)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -668,7 +668,7 @@ define hidden void @two_bytes_same_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%16> = load ir<%15>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%18> = load ir<%17>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%19>, ir<%20>
-; CHECK: Cost for VF 2: 53 (Estimated cost per lane: 26.5)
+; CHECK: Cost for VF 2: 52 (Estimated cost per lane: 26)
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -678,7 +678,7 @@ define hidden void @two_bytes_same_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
+; CHECK: Cost for VF 4: 60 (Estimated cost per lane: 15)
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -688,7 +688,7 @@ define hidden void @two_bytes_same_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 8: 33 (Estimated cost per lane: 4.13)
+; CHECK: Cost for VF 8: 32 (Estimated cost per lane: 4)
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -698,7 +698,7 @@ define hidden void @two_bytes_same_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 16: 30 (Estimated cost per lane: 1.88)
+; CHECK: Cost for VF 16: 29 (Estimated cost per lane: 1.81)
; CHECK: LV: Selecting VF: 16.
;
%5 = icmp eq i32 %3, 0
@@ -737,7 +737,7 @@ define hidden void @two_bytes_vary_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%16> = load ir<%15>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%18> = load ir<%17>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%19>, ir<%20>
-; CHECK: Cost for VF 2: 48 (Estimated cost per lane: 24)
+; CHECK: Cost for VF 2: 47 (Estimated cost per lane: 23.5)
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -747,7 +747,7 @@ define hidden void @two_bytes_vary_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 4: 50 (Estimated cost per lane: 12.5)
+; CHECK: Cost for VF 4: 49 (Estimated cost per lane: 12.3)
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -757,7 +757,7 @@ define hidden void @two_bytes_vary_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 8: 30 (Estimated cost per lane: 3.75)
+; CHECK: Cost for VF 8: 29 (Estimated cost per lane: 3.63)
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -767,7 +767,7 @@ define hidden void @two_bytes_vary_op(ptr noalias nocapture noundef writeonly %0
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, ir<%14>
; CHECK: store ir<%13> to index 0
; CHECK: store ir<%19> to index 1
-; CHECK: Cost for VF 16: 27 (Estimated cost per lane: 1.69)
+; CHECK: Cost for VF 16: 26 (Estimated cost per lane: 1.63)
; CHECK: LV: Selecting VF: 16.
;
%5 = icmp eq i32 %3, 0
@@ -809,7 +809,7 @@ define hidden void @three_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%13>, ir<%14>
@@ -819,7 +819,7 @@ define hidden void @three_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 4: 116 (Estimated cost per lane: 29)
+; CHECK: Cost for VF 4: 115 (Estimated cost per lane: 28.8)
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%13>, ir<%14>
@@ -829,7 +829,7 @@ define hidden void @three_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 8: 224 (Estimated cost per lane: 28)
+; CHECK: Cost for VF 8: 223 (Estimated cost per lane: 27.9)
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 48 for VF 16: REPLICATE store ir<%13>, ir<%14>
@@ -839,7 +839,7 @@ define hidden void @three_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 48 for VF 16: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 16: 440 (Estimated cost per lane: 27.5)
+; CHECK: Cost for VF 16: 439 (Estimated cost per lane: 27.4)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -888,7 +888,7 @@ define hidden void @three_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 2: 62 (Estimated cost per lane: 31)
+; CHECK: Cost for VF 2: 61 (Estimated cost per lane: 30.5)
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%13>, ir<%14>
@@ -898,7 +898,7 @@ define hidden void @three_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 12 for VF 4: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 12 for VF 4: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 4: 116 (Estimated cost per lane: 29)
+; CHECK: Cost for VF 4: 115 (Estimated cost per lane: 28.8)
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%13>, ir<%14>
@@ -908,7 +908,7 @@ define hidden void @three_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 24 for VF 8: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 24 for VF 8: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 8: 224 (Estimated cost per lane: 28)
+; CHECK: Cost for VF 8: 223 (Estimated cost per lane: 27.9)
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%10> = load ir<%9>
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%12> = load ir<%11>
; CHECK: Cost of 48 for VF 16: REPLICATE store ir<%13>, ir<%14>
@@ -918,7 +918,7 @@ define hidden void @three_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%22> = load ir<%21>
; CHECK: Cost of 48 for VF 16: REPLICATE ir<%24> = load ir<%23>
; CHECK: Cost of 48 for VF 16: REPLICATE store ir<%25>, ir<%26>
-; CHECK: Cost for VF 16: 440 (Estimated cost per lane: 27.5)
+; CHECK: Cost for VF 16: 439 (Estimated cost per lane: 27.4)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -970,7 +970,7 @@ define hidden void @four_bytes_same_op(ptr noalias nocapture noundef writeonly %
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%28> = load ir<%27>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%30> = load ir<%29>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%31>, ir<%32>
-; CHECK: Cost for VF 2: 81 (Estimated cost per lane: 40.5)
+; CHECK: Cost for VF 2: 80 (Estimated cost per lane: 40)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -986,7 +986,7 @@ define hidden void @four_bytes_same_op(ptr noalias nocapture noundef writeonly %
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 62 (Estimated cost per lane: 15.5)
+; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1002,7 +1002,7 @@ define hidden void @four_bytes_same_op(ptr noalias nocapture noundef writeonly %
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 86 (Estimated cost per lane: 10.8)
+; CHECK: Cost for VF 8: 85 (Estimated cost per lane: 10.6)
; CHECK: Cost of 132 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1018,7 +1018,7 @@ define hidden void @four_bytes_same_op(ptr noalias nocapture noundef writeonly %
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 16: 404 (Estimated cost per lane: 25.3)
+; CHECK: Cost for VF 16: 403 (Estimated cost per lane: 25.2)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp eq i32 %3, 0
@@ -1077,7 +1077,7 @@ define hidden void @four_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%28> = load ir<%27>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%30> = load ir<%29>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%31>, ir<%32>
-; CHECK: Cost for VF 2: 91 (Estimated cost per lane: 45.5)
+; CHECK: Cost for VF 2: 90 (Estimated cost per lane: 45)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1093,7 +1093,7 @@ define hidden void @four_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 84 (Estimated cost per lane: 21)
+; CHECK: Cost for VF 4: 83 (Estimated cost per lane: 20.8)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1109,7 +1109,7 @@ define hidden void @four_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 92 (Estimated cost per lane: 11.5)
+; CHECK: Cost for VF 8: 91 (Estimated cost per lane: 11.4)
; CHECK: Cost of 132 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1125,7 +1125,7 @@ define hidden void @four_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 16: 410 (Estimated cost per lane: 25.6)
+; CHECK: Cost for VF 16: 409 (Estimated cost per lane: 25.6)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp eq i32 %3, 0
@@ -1185,7 +1185,7 @@ define hidden void @four_bytes_interleave_op(ptr noalias nocapture noundef write
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%28> = load ir<%27>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%30> = load ir<%29>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%31>, ir<%32>
-; CHECK: Cost for VF 2: 81 (Estimated cost per lane: 40.5)
+; CHECK: Cost for VF 2: 80 (Estimated cost per lane: 40)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1201,7 +1201,7 @@ define hidden void @four_bytes_interleave_op(ptr noalias nocapture noundef write
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 4: 62 (Estimated cost per lane: 15.5)
+; CHECK: Cost for VF 4: 61 (Estimated cost per lane: 15.3)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1217,7 +1217,7 @@ define hidden void @four_bytes_interleave_op(ptr noalias nocapture noundef write
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 8: 86 (Estimated cost per lane: 10.8)
+; CHECK: Cost for VF 8: 85 (Estimated cost per lane: 10.6)
; CHECK: Cost of 132 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1233,7 +1233,7 @@ define hidden void @four_bytes_interleave_op(ptr noalias nocapture noundef write
; CHECK: store ir<%19> to index 1
; CHECK: store ir<%25> to index 2
; CHECK: store ir<%31> to index 3
-; CHECK: Cost for VF 16: 404 (Estimated cost per lane: 25.3)
+; CHECK: Cost for VF 16: 403 (Estimated cost per lane: 25.2)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp eq i32 %3, 0
@@ -1308,7 +1308,7 @@ define hidden void @eight_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 2: 154 (Estimated cost per lane: 77)
+; CHECK: Cost for VF 2: 153 (Estimated cost per lane: 76.5)
; CHECK: Cost of 66 for VF 4: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1336,7 +1336,7 @@ define hidden void @eight_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 4: 298 (Estimated cost per lane: 74.5)
+; CHECK: Cost for VF 4: 297 (Estimated cost per lane: 74.3)
; CHECK: Cost of 132 for VF 8: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1364,7 +1364,7 @@ define hidden void @eight_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 8: 432 (Estimated cost per lane: 54)
+; CHECK: Cost for VF 8: 431 (Estimated cost per lane: 53.9)
; CHECK: Cost of 264 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1392,7 +1392,7 @@ define hidden void @eight_bytes_same_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 16: 828 (Estimated cost per lane: 51.8)
+; CHECK: Cost for VF 16: 827 (Estimated cost per lane: 51.7)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -1494,7 +1494,7 @@ define hidden void @eight_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 2: 114 (Estimated cost per lane: 57)
+; CHECK: Cost for VF 2: 113 (Estimated cost per lane: 56.5)
; CHECK: Cost of 66 for VF 4: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1522,7 +1522,7 @@ define hidden void @eight_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 4: 210 (Estimated cost per lane: 52.5)
+; CHECK: Cost for VF 4: 209 (Estimated cost per lane: 52.3)
; CHECK: Cost of 132 for VF 8: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1550,7 +1550,7 @@ define hidden void @eight_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 8: 408 (Estimated cost per lane: 51)
+; CHECK: Cost for VF 8: 407 (Estimated cost per lane: 50.9)
; CHECK: Cost of 264 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1578,7 +1578,7 @@ define hidden void @eight_bytes_split_op(ptr noalias nocapture noundef writeonly
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 16: 804 (Estimated cost per lane: 50.3)
+; CHECK: Cost for VF 16: 803 (Estimated cost per lane: 50.2)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -1680,7 +1680,7 @@ define hidden void @eight_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 2: 114 (Estimated cost per lane: 57)
+; CHECK: Cost for VF 2: 113 (Estimated cost per lane: 56.5)
; CHECK: Cost of 66 for VF 4: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1708,7 +1708,7 @@ define hidden void @eight_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 4: 210 (Estimated cost per lane: 52.5)
+; CHECK: Cost for VF 4: 209 (Estimated cost per lane: 52.3)
; CHECK: Cost of 132 for VF 8: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1736,7 +1736,7 @@ define hidden void @eight_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 8: 408 (Estimated cost per lane: 51)
+; CHECK: Cost for VF 8: 407 (Estimated cost per lane: 50.9)
; CHECK: Cost of 264 for VF 16: INTERLEAVE-GROUP with factor 8, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%16> = load from index 1
@@ -1764,7 +1764,7 @@ define hidden void @eight_bytes_interleave_op(ptr noalias nocapture noundef writ
; CHECK: store ir<%43> to index 5
; CHECK: store ir<%49> to index 6
; CHECK: store ir<%55> to index 7
-; CHECK: Cost for VF 16: 804 (Estimated cost per lane: 50.3)
+; CHECK: Cost for VF 16: 803 (Estimated cost per lane: 50.2)
; CHECK: LV: Selecting VF: 1.
;
%5 = icmp eq i32 %3, 0
@@ -1857,7 +1857,7 @@ define hidden void @four_bytes_into_four_ints_same_op(ptr noalias nocapture noun
; CHECK: store ir<%28> to index 1
; CHECK: store ir<%38> to index 2
; CHECK: store ir<%48> to index 3
-; CHECK: Cost for VF 2: 89 (Estimated cost per lane: 44.5)
+; CHECK: Cost for VF 2: 88 (Estimated cost per lane: 44)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%20> = load from index 1
@@ -1878,7 +1878,7 @@ define hidden void @four_bytes_into_four_ints_same_op(ptr noalias nocapture noun
; CHECK: store ir<%28> to index 1
; CHECK: store ir<%38> to index 2
; CHECK: store ir<%48> to index 3
-; CHECK: Cost for VF 4: 104 (Estimated cost per lane: 26)
+; CHECK: Cost for VF 4: 103 (Estimated cost per lane: 25.8)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -1954,7 +1954,7 @@ define hidden void @four_bytes_into_four_ints_vary_op(ptr noalias nocapture noun
; CHECK: store ir<%23> to index 1
; CHECK: store ir<%31> to index 2
; CHECK: store ir<%38> to index 3
-; CHECK: Cost for VF 2: 72 (Estimated cost per lane: 36)
+; CHECK: Cost for VF 2: 71 (Estimated cost per lane: 35.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%9>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%18> = load from index 1
@@ -1970,7 +1970,7 @@ define hidden void @four_bytes_into_four_ints_vary_op(ptr noalias nocapture noun
; CHECK: store ir<%23> to index 1
; CHECK: store ir<%31> to index 2
; CHECK: store ir<%38> to index 3
-; CHECK: Cost for VF 4: 80 (Estimated cost per lane: 20)
+; CHECK: Cost for VF 4: 79 (Estimated cost per lane: 19.8)
; CHECK: LV: Selecting VF: 4.
;
%5 = icmp eq i32 %3, 0
@@ -2030,21 +2030,21 @@ define hidden void @scale_uv_row_down2(ptr nocapture noundef readonly %0, i32 no
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%11> to index 0
; CHECK: store ir<%13> to index 1
-; CHECK: Cost for VF 4: 37 (Estimated cost per lane: 9.25)
+; CHECK: Cost for VF 4: 36 (Estimated cost per lane: 9)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, ir<%10>
; CHECK: ir<%11> = load from index 0
; CHECK: ir<%13> = load from index 1
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%11> to index 0
; CHECK: store ir<%13> to index 1
-; CHECK: Cost for VF 8: 41 (Estimated cost per lane: 5.13)
+; CHECK: Cost for VF 8: 40 (Estimated cost per lane: 5)
; CHECK: Cost of 68 for VF 16: INTERLEAVE-GROUP with factor 4, ir<%10>
; CHECK: ir<%11> = load from index 0
; CHECK: ir<%13> = load from index 1
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%11> to index 0
; CHECK: store ir<%13> to index 1
-; CHECK: Cost for VF 16: 82 (Estimated cost per lane: 5.13)
+; CHECK: Cost for VF 16: 81 (Estimated cost per lane: 5.06)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp sgt i32 %3, 0
@@ -2084,7 +2084,7 @@ define hidden void @scale_uv_row_down2_box(ptr nocapture noundef readonly %0, i3
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%38> = load ir<%37>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%41> = load ir<%40>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%48>, ir<%49>
-; CHECK: Cost for VF 2: 82 (Estimated cost per lane: 41)
+; CHECK: Cost for VF 2: 81 (Estimated cost per lane: 40.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%14> = load from index 0
; CHECK: ir<%32> = load from index 1
@@ -2098,7 +2098,7 @@ define hidden void @scale_uv_row_down2_box(ptr nocapture noundef readonly %0, i3
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%30> to index 0
; CHECK: store ir<%48> to index 1
-; CHECK: Cost for VF 4: 75 (Estimated cost per lane: 18.8)
+; CHECK: Cost for VF 4: 74 (Estimated cost per lane: 18.5)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%14> = load from index 0
; CHECK: ir<%32> = load from index 1
@@ -2112,7 +2112,7 @@ define hidden void @scale_uv_row_down2_box(ptr nocapture noundef readonly %0, i3
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%30> to index 0
; CHECK: store ir<%48> to index 1
-; CHECK: Cost for VF 8: 91 (Estimated cost per lane: 11.4)
+; CHECK: Cost for VF 8: 90 (Estimated cost per lane: 11.3)
; CHECK: Cost of 132 for VF 16: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%14> = load from index 0
; CHECK: ir<%32> = load from index 1
@@ -2126,7 +2126,7 @@ define hidden void @scale_uv_row_down2_box(ptr nocapture noundef readonly %0, i3
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%30> to index 0
; CHECK: store ir<%48> to index 1
-; CHECK: Cost for VF 16: 324 (Estimated cost per lane: 20.3)
+; CHECK: Cost for VF 16: 323 (Estimated cost per lane: 20.2)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp sgt i32 %3, 0
@@ -2199,7 +2199,7 @@ define hidden void @scale_uv_row_down2_linear(ptr nocapture noundef readonly %0,
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%20> = load ir<%19>
; CHECK: Cost of 6 for VF 2: REPLICATE ir<%23> = load ir<%22>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%28>, ir<%29>
-; CHECK: Cost for VF 2: 54 (Estimated cost per lane: 27)
+; CHECK: Cost for VF 2: 53 (Estimated cost per lane: 26.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%20> = load from index 1
@@ -2208,7 +2208,7 @@ define hidden void @scale_uv_row_down2_linear(ptr nocapture noundef readonly %0,
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%18> to index 0
; CHECK: store ir<%28> to index 1
-; CHECK: Cost for VF 4: 49 (Estimated cost per lane: 12.3)
+; CHECK: Cost for VF 4: 48 (Estimated cost per lane: 12)
; CHECK: Cost of 26 for VF 8: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%20> = load from index 1
@@ -2217,7 +2217,7 @@ define hidden void @scale_uv_row_down2_linear(ptr nocapture noundef readonly %0,
; CHECK: Cost of 7 for VF 8: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%18> to index 0
; CHECK: store ir<%28> to index 1
-; CHECK: Cost for VF 8: 57 (Estimated cost per lane: 7.13)
+; CHECK: Cost for VF 8: 56 (Estimated cost per lane: 7)
; CHECK: Cost of 132 for VF 16: INTERLEAVE-GROUP with factor 4, vp<%next.gep>
; CHECK: ir<%10> = load from index 0
; CHECK: ir<%20> = load from index 1
@@ -2226,7 +2226,7 @@ define hidden void @scale_uv_row_down2_linear(ptr nocapture noundef readonly %0,
; CHECK: Cost of 6 for VF 16: INTERLEAVE-GROUP with factor 2, vp<%next.gep>.1
; CHECK: store ir<%18> to index 0
; CHECK: store ir<%28> to index 1
-; CHECK: Cost for VF 16: 176 (Estimated cost per lane: 11)
+; CHECK: Cost for VF 16: 175 (Estimated cost per lane: 10.9)
; CHECK: LV: Selecting VF: 8.
;
%5 = icmp sgt i32 %3, 0
@@ -2280,7 +2280,7 @@ define hidden void @two_floats_same_op(ptr noundef readonly captures(none) %a, p
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul8> to index 1
-; CHECK: Cost for VF 2: 29 (Estimated cost per lane: 14.5)
+; CHECK: Cost for VF 2: 28 (Estimated cost per lane: 14)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2290,7 +2290,7 @@ define hidden void @two_floats_same_op(ptr noundef readonly captures(none) %a, p
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul8> to index 1
-; CHECK: Cost for VF 4: 26 (Estimated cost per lane: 6.5)
+; CHECK: Cost for VF 4: 25 (Estimated cost per lane: 6.25)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2333,7 +2333,7 @@ define hidden void @two_floats_vary_op(ptr noundef readonly captures(none) %a, p
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 2: 29 (Estimated cost per lane: 14.5)
+; CHECK: Cost for VF 2: 28 (Estimated cost per lane: 14)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2343,7 +2343,7 @@ define hidden void @two_floats_vary_op(ptr noundef readonly captures(none) %a, p
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 4: 26 (Estimated cost per lane: 6.5)
+; CHECK: Cost for VF 4: 25 (Estimated cost per lane: 6.25)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2384,7 +2384,7 @@ define hidden void @two_bytes_two_floats_same_op(ptr noundef readonly captures(n
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul11> to index 1
-; CHECK: Cost for VF 2: 52 (Estimated cost per lane: 26)
+; CHECK: Cost for VF 2: 51 (Estimated cost per lane: 25.5)
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2394,7 +2394,7 @@ define hidden void @two_bytes_two_floats_same_op(ptr noundef readonly captures(n
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul11> to index 1
-; CHECK: Cost for VF 4: 48 (Estimated cost per lane: 12)
+; CHECK: Cost for VF 4: 47 (Estimated cost per lane: 11.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2439,7 +2439,7 @@ define hidden void @two_bytes_two_floats_vary_op(ptr noundef readonly captures(n
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 2: 52 (Estimated cost per lane: 26)
+; CHECK: Cost for VF 2: 51 (Estimated cost per lane: 25.5)
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2449,7 +2449,7 @@ define hidden void @two_bytes_two_floats_vary_op(ptr noundef readonly captures(n
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 4: 48 (Estimated cost per lane: 12)
+; CHECK: Cost for VF 4: 47 (Estimated cost per lane: 11.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2495,7 +2495,7 @@ define hidden void @two_floats_two_bytes_same_op(ptr noundef readonly captures(n
; CHECK: ir<%3> = load from index 1
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv>, ir<%arrayidx3>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv9>, ir<%y11>
-; CHECK: Cost for VF 2: 47 (Estimated cost per lane: 23.5)
+; CHECK: Cost for VF 2: 46 (Estimated cost per lane: 23)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2505,7 +2505,7 @@ define hidden void @two_floats_two_bytes_same_op(ptr noundef readonly captures(n
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv9> to index 1
-; CHECK: Cost for VF 4: 43 (Estimated cost per lane: 10.8)
+; CHECK: Cost for VF 4: 42 (Estimated cost per lane: 10.5)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2549,7 +2549,7 @@ define hidden void @two_floats_two_bytes_vary_op(ptr noundef readonly captures(n
; CHECK: ir<%3> = load from index 1
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv>, ir<%arrayidx3>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv8>, ir<%y10>
-; CHECK: Cost for VF 2: 47 (Estimated cost per lane: 23.5)
+; CHECK: Cost for VF 2: 46 (Estimated cost per lane: 23)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2559,7 +2559,7 @@ define hidden void @two_floats_two_bytes_vary_op(ptr noundef readonly captures(n
; CHECK: Cost of 11 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv8> to index 1
-; CHECK: Cost for VF 4: 43 (Estimated cost per lane: 10.8)
+; CHECK: Cost for VF 4: 42 (Estimated cost per lane: 10.5)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2604,7 +2604,7 @@ define hidden void @two_shorts_two_floats_same_op(ptr noundef readonly captures(
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul11> to index 1
-; CHECK: Cost for VF 2: 45 (Estimated cost per lane: 22.5)
+; CHECK: Cost for VF 2: 44 (Estimated cost per lane: 22)
; CHECK: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2614,7 +2614,7 @@ define hidden void @two_shorts_two_floats_same_op(ptr noundef readonly captures(
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%mul> to index 0
; CHECK: store ir<%mul11> to index 1
-; CHECK: Cost for VF 4: 36 (Estimated cost per lane: 9)
+; CHECK: Cost for VF 4: 35 (Estimated cost per lane: 8.75)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2661,7 +2661,7 @@ define hidden void @two_shorts_two_floats_vary_op(ptr noundef readonly captures(
; CHECK: Cost of 7 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 2: 45 (Estimated cost per lane: 22.5)
+; CHECK: Cost for VF 2: 44 (Estimated cost per lane: 22)
; CHECK: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2671,7 +2671,7 @@ define hidden void @two_shorts_two_floats_vary_op(ptr noundef readonly captures(
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx4>
; CHECK: store ir<%add> to index 0
; CHECK: store ir<%sub> to index 1
-; CHECK: Cost for VF 4: 36 (Estimated cost per lane: 9)
+; CHECK: Cost for VF 4: 35 (Estimated cost per lane: 8.75)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2718,7 +2718,7 @@ define hidden void @two_floats_two_shorts_same_op(ptr noundef readonly captures(
; CHECK: Cost of 11 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv9> to index 1
-; CHECK: Cost for VF 2: 41 (Estimated cost per lane: 20.5)
+; CHECK: Cost for VF 2: 40 (Estimated cost per lane: 20)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2728,7 +2728,7 @@ define hidden void @two_floats_two_shorts_same_op(ptr noundef readonly captures(
; CHECK: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv9> to index 1
-; CHECK: Cost for VF 4: 35 (Estimated cost per lane: 8.75)
+; CHECK: Cost for VF 4: 34 (Estimated cost per lane: 8.5)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2773,7 +2773,7 @@ define hidden void @two_floats_two_shorts_vary_op(ptr noundef readonly captures(
; CHECK: Cost of 11 for VF 2: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv8> to index 1
-; CHECK: Cost for VF 2: 41 (Estimated cost per lane: 20.5)
+; CHECK: Cost for VF 2: 40 (Estimated cost per lane: 20)
; CHECK: Cost of 6 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2783,7 +2783,7 @@ define hidden void @two_floats_two_shorts_vary_op(ptr noundef readonly captures(
; CHECK: Cost of 7 for VF 4: INTERLEAVE-GROUP with factor 2, ir<%arrayidx3>
; CHECK: store ir<%conv> to index 0
; CHECK: store ir<%conv8> to index 1
-; CHECK: Cost for VF 4: 35 (Estimated cost per lane: 8.75)
+; CHECK: Cost for VF 4: 34 (Estimated cost per lane: 8.5)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2834,8 +2834,8 @@ define hidden void @four_floats_same_op(ptr noundef readonly captures(none) %a,
; CHECK: store ir<%mul8> to index 1
; CHECK: store ir<%mul14> to index 2
; CHECK: store ir<%mul20> to index 3
-; CHECK: Cost for VF 2: 54 (Estimated cost per lane: 27)
-; CHECK: Cost for VF 4: 12 (Estimated cost per lane: 3)
+; CHECK: Cost for VF 2: 53 (Estimated cost per lane: 26.5)
+; CHECK: Cost for VF 4: 11 (Estimated cost per lane: 2.75)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -2898,7 +2898,7 @@ define hidden void @four_floats_vary_op(ptr noundef readonly captures(none) %a,
; CHECK: store ir<%sub> to index 1
; CHECK: store ir<%mul> to index 2
; CHECK: store ir<%div> to index 3
-; CHECK: Cost for VF 2: 54 (Estimated cost per lane: 27)
+; CHECK: Cost for VF 2: 53 (Estimated cost per lane: 26.5)
; CHECK: Cost of 36 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2914,7 +2914,7 @@ define hidden void @four_floats_vary_op(ptr noundef readonly captures(none) %a,
; CHECK: store ir<%sub> to index 1
; CHECK: store ir<%mul> to index 2
; CHECK: store ir<%div> to index 3
-; CHECK: Cost for VF 4: 120 (Estimated cost per lane: 30)
+; CHECK: Cost for VF 4: 119 (Estimated cost per lane: 29.8)
; CHECK: LV: Selecting VF: 1.
;
entry:
@@ -2975,7 +2975,7 @@ define hidden void @four_bytes_four_floats_same_op(ptr noundef readonly captures
; CHECK: store ir<%mul11> to index 1
; CHECK: store ir<%mul19> to index 2
; CHECK: store ir<%mul27> to index 3
-; CHECK: Cost for VF 2: 99 (Estimated cost per lane: 49.5)
+; CHECK: Cost for VF 2: 98 (Estimated cost per lane: 49)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -2991,7 +2991,7 @@ define hidden void @four_bytes_four_floats_same_op(ptr noundef readonly captures
; CHECK: store ir<%mul11> to index 1
; CHECK: store ir<%mul19> to index 2
; CHECK: store ir<%mul27> to index 3
-; CHECK: Cost for VF 4: 108 (Estimated cost per lane: 27)
+; CHECK: Cost for VF 4: 107 (Estimated cost per lane: 26.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -3060,7 +3060,7 @@ define hidden void @four_bytes_four_floats_vary_op(ptr noundef readonly captures
; CHECK: store ir<%add> to index 1
; CHECK: store ir<%div> to index 2
; CHECK: store ir<%sub> to index 3
-; CHECK: Cost for VF 2: 99 (Estimated cost per lane: 49.5)
+; CHECK: Cost for VF 2: 98 (Estimated cost per lane: 49)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3076,7 +3076,7 @@ define hidden void @four_bytes_four_floats_vary_op(ptr noundef readonly captures
; CHECK: store ir<%add> to index 1
; CHECK: store ir<%div> to index 2
; CHECK: store ir<%sub> to index 3
-; CHECK: Cost for VF 4: 108 (Estimated cost per lane: 27)
+; CHECK: Cost for VF 4: 107 (Estimated cost per lane: 26.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -3146,7 +3146,7 @@ define hidden void @four_floats_four_bytes_same_op(ptr noundef readonly captures
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv9>, ir<%y11>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv16>, ir<%z18>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv23>, ir<%w25>
-; CHECK: Cost for VF 2: 89 (Estimated cost per lane: 44.5)
+; CHECK: Cost for VF 2: 88 (Estimated cost per lane: 44)
; CHECK: Cost of 36 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3162,7 +3162,7 @@ define hidden void @four_floats_four_bytes_same_op(ptr noundef readonly captures
; CHECK: store ir<%conv9> to index 1
; CHECK: store ir<%conv16> to index 2
; CHECK: store ir<%conv23> to index 3
-; CHECK: Cost for VF 4: 126 (Estimated cost per lane: 31.5)
+; CHECK: Cost for VF 4: 125 (Estimated cost per lane: 31.3)
; CHECK: LV: Selecting VF: 1.
;
entry:
@@ -3228,7 +3228,7 @@ define hidden void @four_floats_four_bytes_vary_op(ptr noundef readonly captures
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv8>, ir<%y10>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv14>, ir<%z16>
; CHECK: Cost of 6 for VF 2: REPLICATE store ir<%conv20>, ir<%w22>
-; CHECK: Cost for VF 2: 89 (Estimated cost per lane: 44.5)
+; CHECK: Cost for VF 2: 88 (Estimated cost per lane: 44)
; CHECK: Cost of 36 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3244,7 +3244,7 @@ define hidden void @four_floats_four_bytes_vary_op(ptr noundef readonly captures
; CHECK: store ir<%conv8> to index 1
; CHECK: store ir<%conv14> to index 2
; CHECK: store ir<%conv20> to index 3
-; CHECK: Cost for VF 4: 126 (Estimated cost per lane: 31.5)
+; CHECK: Cost for VF 4: 125 (Estimated cost per lane: 31.3)
; CHECK: LV: Selecting VF: 1.
;
entry:
@@ -3311,7 +3311,7 @@ define hidden void @four_shorts_four_floats_same_op(ptr noundef readonly capture
; CHECK: store ir<%mul11> to index 1
; CHECK: store ir<%mul19> to index 2
; CHECK: store ir<%mul27> to index 3
-; CHECK: Cost for VF 2: 78 (Estimated cost per lane: 39)
+; CHECK: Cost for VF 2: 77 (Estimated cost per lane: 38.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3327,7 +3327,7 @@ define hidden void @four_shorts_four_floats_same_op(ptr noundef readonly capture
; CHECK: store ir<%mul11> to index 1
; CHECK: store ir<%mul19> to index 2
; CHECK: store ir<%mul27> to index 3
-; CHECK: Cost for VF 4: 100 (Estimated cost per lane: 25)
+; CHECK: Cost for VF 4: 99 (Estimated cost per lane: 24.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -3398,7 +3398,7 @@ define hidden void @four_shorts_four_floats_vary_op(ptr noundef readonly capture
; CHECK: store ir<%add> to index 1
; CHECK: store ir<%div> to index 2
; CHECK: store ir<%sub> to index 3
-; CHECK: Cost for VF 2: 78 (Estimated cost per lane: 39)
+; CHECK: Cost for VF 2: 77 (Estimated cost per lane: 38.5)
; CHECK: Cost of 18 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3414,7 +3414,7 @@ define hidden void @four_shorts_four_floats_vary_op(ptr noundef readonly capture
; CHECK: store ir<%add> to index 1
; CHECK: store ir<%div> to index 2
; CHECK: store ir<%sub> to index 3
-; CHECK: Cost for VF 4: 100 (Estimated cost per lane: 25)
+; CHECK: Cost for VF 4: 99 (Estimated cost per lane: 24.8)
; CHECK: LV: Selecting VF: 4.
;
entry:
@@ -3485,7 +3485,7 @@ define hidden void @four_floats_four_shorts_same_op(ptr noundef readonly capture
; CHECK: store ir<%conv9> to index 1
; CHECK: store ir<%conv16> to index 2
; CHECK: store ir<%conv23> to index 3
-; CHECK: Cost for VF 2: 74 (Estimated cost per lane: 37)
+; CHECK: Cost for VF 2: 73 (Estimated cost per lane: 36.5)
; CHECK: Cost of 36 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3501,7 +3501,7 @@ define hidden void @four_floats_four_shorts_same_op(ptr noundef readonly capture
; CHECK: store ir<%conv9> to index 1
; CHECK: store ir<%conv16> to index 2
; CHECK: store ir<%conv23> to index 3
-; CHECK: Cost for VF 4: 118 (Estimated cost per lane: 29.5)
+; CHECK: Cost for VF 4: 117 (Estimated cost per lane: 29.3)
; CHECK: LV: Selecting VF: 1.
;
entry:
@@ -3568,7 +3568,7 @@ define hidden void @four_floats_four_shorts_vary_op(ptr noundef readonly capture
; CHECK: store ir<%conv8> to index 1
; CHECK: store ir<%conv14> to index 2
; CHECK: store ir<%conv20> to index 3
-; CHECK: Cost for VF 2: 74 (Estimated cost per lane: 37)
+; CHECK: Cost for VF 2: 73 (Estimated cost per lane: 36.5)
; CHECK: Cost of 36 for VF 4: INTERLEAVE-GROUP with factor 4, ir<%arrayidx>
; CHECK: ir<%0> = load from index 0
; CHECK: ir<%2> = load from index 1
@@ -3584,7 +3584,7 @@ define hidden void @four_floats_four_shorts_vary_op(ptr noundef readonly capture
; CHECK: store ir<%conv8> to index 1
; CHECK: store ir<%conv14> to index 2
; CHECK: store ir<%conv20> to index 3
-; CHECK: Cost for VF 4: 118 (Estimated cost per lane: 29.5)
+; CHECK: Cost for VF 4: 117 (Estimated cost per lane: 29.3)
; CHECK: LV: Selecting VF: 1.
;
entry:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/CostModel/vpinstruction-cost.ll b/llvm/test/Transforms/LoopVectorize/X86/CostModel/vpinstruction-cost.ll
index 61beec08b0e09..b145e06143a7d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/CostModel/vpinstruction-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/CostModel/vpinstruction-cost.ll
@@ -20,6 +20,7 @@ define void @wide_or_replaced_with_add_vpinstruction(ptr %src, ptr noalias %dst)
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %iv.next = add nuw nsw i64 %iv, 1
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %exitcond = icmp eq i64 %iv.next, 32
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %exitcond, label %exit, label %loop.header
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 1 for VF 2: ir<%iv> = WIDEN-INDUCTION nuw nsw ir<0>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3:%[0-9]+]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 2: CLONE ir<%g.src> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
@@ -42,6 +43,7 @@ define void @wide_or_replaced_with_add_vpinstruction(ptr %src, ptr noalias %dst)
; CHECK: Cost of 0 for VF 2: IR %c = icmp ule i64 %l, 128
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<32>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 1 for VF 4: ir<%iv> = WIDEN-INDUCTION nuw nsw ir<0>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: vp<[[VP4]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: CLONE ir<%g.src> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
@@ -104,8 +106,7 @@ define void @test_vpinstruction_freeze_cost(ptr %src, ptr noalias %dst) {
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %iv.next = add nuw nsw i64 %iv, 1
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %ec = icmp eq i64 %iv.next, 32
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %ec, label %exit, label %loop
-; CHECK: Cost of 1 for VF 2: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 2: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 0 for VF 2: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3:%[0-9]+]]>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: CLONE ir<%g.src> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 2: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i64, ir<%g.src>, ir<1>
@@ -128,8 +129,7 @@ define void @test_vpinstruction_freeze_cost(ptr %src, ptr noalias %dst) {
; CHECK: Cost of 0 for VF 2: IR %ec = icmp eq i64 %iv.next, 32
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<32>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
-; CHECK: Cost of 1 for VF 4: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 4: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 0 for VF 4: vp<[[VP4]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: CLONE ir<%g.src> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 4: vp<[[VP5]]> = vector-pointer inbounds i64, ir<%g.src>, ir<1>
@@ -300,8 +300,7 @@ define void @test_vpinstruction_extractvalue_cost(ptr noalias %dst, {i64, i64} %
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %iv.next = add nuw nsw i64 %iv, 1
; CHECK: LV: Found an estimated cost of 1 for VF 1 For instruction: %ec = icmp eq i64 %iv.next, 1000
; CHECK: LV: Found an estimated cost of 0 for VF 1 For instruction: br i1 %ec, label %exit, label %loop
-; CHECK: Cost of 1 for VF 2: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 2: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 0 for VF 2: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3:%[0-9]+]]>, ir<1>, vp<[[VP0:%[0-9]+]]>
; CHECK: Cost of 0 for VF 2: CLONE ir<%g.dst> = getelementptr inbounds ir<%dst>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 2: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i64, ir<%g.dst>, ir<1>
@@ -323,8 +322,7 @@ define void @test_vpinstruction_extractvalue_cost(ptr noalias %dst, {i64, i64} %
; CHECK: Cost of 1 for VF 2: CLONE ir<%add> = add ir<%a>, ir<%b>
; CHECK: Cost of 1 for VF 2: EMIT vp<%cmp.n> = icmp eq ir<1000>, vp<[[VP2]]>
; CHECK: Cost of 0 for VF 2: EMIT branch-on-cond vp<%cmp.n>
-; CHECK: Cost of 1 for VF 4: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 0 for VF 4: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
+; CHECK: Cost of 1 for VF 4: canonical IV increment
; CHECK: Cost of 0 for VF 4: vp<[[VP4]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK: Cost of 0 for VF 4: CLONE ir<%g.dst> = getelementptr inbounds ir<%dst>, vp<[[VP4]]>
; CHECK: Cost of 0 for VF 4: vp<[[VP5]]> = vector-pointer inbounds i64, ir<%g.dst>, ir<1>
diff --git a/llvm/test/Transforms/LoopVectorize/X86/conversion-cost.ll b/llvm/test/Transforms/LoopVectorize/X86/conversion-cost.ll
index b272d81a9b086..278eb68208b3d 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/conversion-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/conversion-cost.ll
@@ -12,60 +12,60 @@ define void @conversion_cost1(i32 %n, ptr nocapture %A, ptr nocapture %B) {
; CHECK: [[ITER_CHECK]]:
; CHECK-NEXT: [[TMP2:%.*]] = add i32 [[N]], -3
; CHECK-NEXT: [[TMP3:%.*]] = zext i32 [[TMP2]] to i64
-; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP3]], 8
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP3]], 16
; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
; CHECK: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
-; CHECK-NEXT: [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[TMP3]], 64
+; CHECK-NEXT: [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[TMP3]], 128
; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK1]], label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
-; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i64 [[TMP3]], 63
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = and i64 [[TMP3]], 127
; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[TMP3]], [[N_MOD_VF]]
; CHECK-NEXT: [[TMP4:%.*]] = add i64 3, [[N_VEC]]
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <16 x i8> [ <i8 3, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[STEP_ADD:%.*]] = add <16 x i8> [[VEC_IND]], splat (i8 16)
-; CHECK-NEXT: [[STEP_ADD_2:%.*]] = add <16 x i8> [[STEP_ADD]], splat (i8 16)
-; CHECK-NEXT: [[STEP_ADD_3:%.*]] = add <16 x i8> [[STEP_ADD_2]], splat (i8 16)
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <32 x i8> [ <i8 3, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 20, i8 21, i8 22, i8 23, i8 24, i8 25, i8 26, i8 27, i8 28, i8 29, i8 30, i8 31, i8 32, i8 33, i8 34>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[STEP_ADD:%.*]] = add <32 x i8> [[VEC_IND]], splat (i8 32)
+; CHECK-NEXT: [[STEP_ADD_2:%.*]] = add <32 x i8> [[STEP_ADD]], splat (i8 32)
+; CHECK-NEXT: [[STEP_ADD_3:%.*]] = add <32 x i8> [[STEP_ADD_2]], splat (i8 32)
; CHECK-NEXT: [[TMP5:%.*]] = add i64 3, [[INDEX]]
; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[A]], i64 [[TMP5]]
-; CHECK-NEXT: [[TMP15:%.*]] = getelementptr inbounds i8, ptr [[TMP6]], i64 16
; CHECK-NEXT: [[TMP16:%.*]] = getelementptr inbounds i8, ptr [[TMP6]], i64 32
-; CHECK-NEXT: [[TMP17:%.*]] = getelementptr inbounds i8, ptr [[TMP6]], i64 48
-; CHECK-NEXT: store <16 x i8> [[VEC_IND]], ptr [[TMP6]], align 1
-; CHECK-NEXT: store <16 x i8> [[STEP_ADD]], ptr [[TMP15]], align 1
-; CHECK-NEXT: store <16 x i8> [[STEP_ADD_2]], ptr [[TMP16]], align 1
-; CHECK-NEXT: store <16 x i8> [[STEP_ADD_3]], ptr [[TMP17]], align 1
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 64
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add <16 x i8> [[STEP_ADD_3]], splat (i8 16)
+; CHECK-NEXT: [[TMP15:%.*]] = getelementptr inbounds i8, ptr [[TMP6]], i64 64
+; CHECK-NEXT: [[TMP17:%.*]] = getelementptr inbounds i8, ptr [[TMP6]], i64 96
+; CHECK-NEXT: store <32 x i8> [[VEC_IND]], ptr [[TMP6]], align 1
+; CHECK-NEXT: store <32 x i8> [[STEP_ADD]], ptr [[TMP16]], align 1
+; CHECK-NEXT: store <32 x i8> [[STEP_ADD_2]], ptr [[TMP15]], align 1
+; CHECK-NEXT: store <32 x i8> [[STEP_ADD_3]], ptr [[TMP17]], align 1
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 128
+; CHECK-NEXT: [[VEC_IND_NEXT]] = add <32 x i8> [[STEP_ADD_3]], splat (i8 32)
; CHECK-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[TMP3]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], [[DOT_CRIT_EDGE_LOOPEXIT:label %.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
; CHECK: [[VEC_EPILOG_ITER_CHECK]]:
-; CHECK-NEXT: [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[N_MOD_VF]], 8
+; CHECK-NEXT: [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[N_MOD_VF]], 16
; CHECK-NEXT: br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3:![0-9]+]]
; CHECK: [[VEC_EPILOG_PH]]:
; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[TMP4]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 3, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
-; CHECK-NEXT: [[N_MOD_VF2:%.*]] = and i64 [[TMP3]], 7
+; CHECK-NEXT: [[N_MOD_VF2:%.*]] = and i64 [[TMP3]], 15
; CHECK-NEXT: [[N_VEC3:%.*]] = sub i64 [[TMP3]], [[N_MOD_VF2]]
; CHECK-NEXT: [[TMP8:%.*]] = add i64 3, [[N_VEC3]]
; CHECK-NEXT: [[TMP9:%.*]] = trunc i64 [[BC_RESUME_VAL]] to i8
-; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <8 x i8> poison, i8 [[TMP9]], i64 0
-; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <8 x i8> [[BROADCAST_SPLATINSERT]], <8 x i8> poison, <8 x i32> zeroinitializer
-; CHECK-NEXT: [[INDUCTION:%.*]] = add <8 x i8> [[BROADCAST_SPLAT]], <i8 0, i8 1, i8 2, i8 3, i8 4, i8 5, i8 6, i8 7>
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i8> poison, i8 [[TMP9]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i8> [[BROADCAST_SPLATINSERT]], <16 x i8> poison, <16 x i32> zeroinitializer
+; CHECK-NEXT: [[INDUCTION:%.*]] = add <16 x i8> [[BROADCAST_SPLAT]], <i8 0, i8 1, i8 2, i8 3, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15>
; CHECK-NEXT: br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
; CHECK: [[VEC_EPILOG_VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX4:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT6:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND5:%.*]] = phi <8 x i8> [ [[INDUCTION]], %[[VEC_EPILOG_PH]] ], [ [[VEC_IND_NEXT7:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT: [[VEC_IND4:%.*]] = phi <16 x i8> [ [[INDUCTION]], %[[VEC_EPILOG_PH]] ], [ [[VEC_IND_NEXT6:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
; CHECK-NEXT: [[TMP10:%.*]] = add i64 3, [[INDEX4]]
; CHECK-NEXT: [[TMP11:%.*]] = getelementptr inbounds i8, ptr [[A]], i64 [[TMP10]]
-; CHECK-NEXT: store <8 x i8> [[VEC_IND5]], ptr [[TMP11]], align 1
-; CHECK-NEXT: [[INDEX_NEXT6]] = add nuw i64 [[INDEX4]], 8
-; CHECK-NEXT: [[VEC_IND_NEXT7]] = add <8 x i8> [[VEC_IND5]], splat (i8 8)
+; CHECK-NEXT: store <16 x i8> [[VEC_IND4]], ptr [[TMP11]], align 1
+; CHECK-NEXT: [[INDEX_NEXT6]] = add nuw i64 [[INDEX4]], 16
+; CHECK-NEXT: [[VEC_IND_NEXT6]] = add <16 x i8> [[VEC_IND4]], splat (i8 16)
; CHECK-NEXT: [[TMP12:%.*]] = icmp eq i64 [[INDEX_NEXT6]], [[N_VEC3]]
; CHECK-NEXT: br i1 [[TMP12]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK: [[VEC_EPILOG_MIDDLE_BLOCK]]:
diff --git a/llvm/test/Transforms/LoopVectorize/X86/pr131359-dead-for-splice.ll b/llvm/test/Transforms/LoopVectorize/X86/pr131359-dead-for-splice.ll
index 91958b6e74529..5a4acc96c1583 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/pr131359-dead-for-splice.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/pr131359-dead-for-splice.ll
@@ -9,29 +9,51 @@ target triple = "x86_64"
define void @no_use() {
; CHECK-LABEL: define void @no_use() {
-; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br i1 false, label %[[SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-NEXT: br i1 false, label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH1:.*]]
+; CHECK: [[VECTOR_PH1]]:
; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_PH]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7>, %[[VECTOR_PH1]] ], [ [[VEC_IND_NEXT1:%.*]], %[[VECTOR_PH]] ]
+; CHECK-NEXT: [[STEP_ADD1:%.*]] = add nuw <8 x i32> [[VEC_IND]], splat (i32 8)
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 16
+; CHECK-NEXT: [[VEC_IND_NEXT1]] = add <8 x i32> [[STEP_ADD1]], splat (i32 8)
+; CHECK-NEXT: [[TMP0:%.*]] = icmp eq i32 [[INDEX_NEXT]], 32
+; CHECK-NEXT: br i1 [[TMP0]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_PH]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK1]]:
+; CHECK-NEXT: [[TMP1:%.*]] = extractelement <8 x i32> [[STEP_ADD1]], i64 7
+; CHECK-NEXT: br i1 false, label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK: [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-NEXT: br i1 false, label %[[SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3:![0-9]+]]
+; CHECK: [[VEC_EPILOG_PH]]:
+; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i32 [ 32, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[VEC_EPILOG_RESUME_VAL]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: [[INDUCTION:%.*]] = add <4 x i32> [[BROADCAST_SPLAT]], <i32 0, i32 1, i32 2, i32 3>
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 1, i32 2, i32 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[STEP_ADD:%.*]] = add nuw <4 x i32> [[VEC_IND]], splat (i32 4)
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT: [[INDEX1:%.*]] = phi i32 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT3:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[STEP_ADD:%.*]] = phi <4 x i32> [ [[INDUCTION]], %[[VEC_EPILOG_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[INDEX_NEXT3]] = add nuw i32 [[INDEX1]], 4
; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[STEP_ADD]], splat (i32 4)
-; CHECK-NEXT: [[TMP0:%.*]] = icmp eq i32 [[INDEX_NEXT]], 40
-; CHECK-NEXT: br i1 [[TMP0]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-NEXT: [[TMP2:%.*]] = icmp eq i32 [[INDEX_NEXT3]], 44
+; CHECK-NEXT: br i1 [[TMP2]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[VECTOR_RECUR_EXTRACT:%.*]] = extractelement <4 x i32> [[STEP_ADD]], i64 3
-; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
+; CHECK-NEXT: br i1 true, label %[[EXIT]], label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
+; CHECK-NEXT: [[TMP4:%.*]] = phi i32 [ [[VECTOR_RECUR_EXTRACT]], %[[MIDDLE_BLOCK]] ], [ [[TMP1]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-NEXT: [[TMP5:%.*]] = phi i32 [ 44, %[[MIDDLE_BLOCK]] ], [ 32, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ENTRY]] ]
; CHECK-NEXT: br label %[[LOOP:.*]]
; CHECK: [[LOOP]]:
-; CHECK-NEXT: [[FOR:%.*]] = phi i32 [ [[VECTOR_RECUR_EXTRACT]], %[[SCALAR_PH]] ], [ [[E_0_I:%.*]], %[[LOOP]] ]
-; CHECK-NEXT: [[E_0_I]] = phi i32 [ 40, %[[SCALAR_PH]] ], [ [[INC_I:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[FOR:%.*]] = phi i32 [ [[TMP4]], %[[SCALAR_PH]] ], [ [[E_0_I:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[E_0_I]] = phi i32 [ [[TMP5]], %[[SCALAR_PH]] ], [ [[INC_I:%.*]], %[[LOOP]] ]
; CHECK-NEXT: [[INC_I]] = add i32 [[E_0_I]], 1
; CHECK-NEXT: [[EXITCOND_NOT_I:%.*]] = icmp eq i32 [[E_0_I]], 43
-; CHECK-NEXT: br i1 [[EXITCOND_NOT_I]], label %[[EXIT:.*]], label %[[LOOP]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK-NEXT: br i1 [[EXITCOND_NOT_I]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK: [[EXIT]]:
; CHECK-NEXT: ret void
;
@@ -51,30 +73,52 @@ exit:
define void @dead_use() {
; CHECK-LABEL: define void @dead_use() {
-; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[ENTRY:.*]]:
+; CHECK-NEXT: br i1 false, label %[[SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-NEXT: br i1 false, label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH1:.*]]
+; CHECK: [[VECTOR_PH1]]:
; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_PH]] ]
+; CHECK-NEXT: [[VEC_IND:%.*]] = phi <8 x i32> [ <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7>, %[[VECTOR_PH1]] ], [ [[VEC_IND_NEXT1:%.*]], %[[VECTOR_PH]] ]
+; CHECK-NEXT: [[STEP_ADD1:%.*]] = add nuw <8 x i32> [[VEC_IND]], splat (i32 8)
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 16
+; CHECK-NEXT: [[VEC_IND_NEXT1]] = add <8 x i32> [[STEP_ADD1]], splat (i32 8)
+; CHECK-NEXT: [[TMP0:%.*]] = icmp eq i32 [[INDEX_NEXT]], 32
+; CHECK-NEXT: br i1 [[TMP0]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_PH]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK1]]:
+; CHECK-NEXT: [[TMP1:%.*]] = extractelement <8 x i32> [[STEP_ADD1]], i64 7
+; CHECK-NEXT: br i1 false, label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK: [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-NEXT: br i1 false, label %[[SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3]]
+; CHECK: [[VEC_EPILOG_PH]]:
+; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i32 [ 32, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[VEC_EPILOG_RESUME_VAL]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: [[INDUCTION:%.*]] = add <4 x i32> [[BROADCAST_SPLAT]], <i32 0, i32 1, i32 2, i32 3>
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i32> [ <i32 0, i32 1, i32 2, i32 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[STEP_ADD:%.*]] = add nuw <4 x i32> [[VEC_IND]], splat (i32 4)
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i32 [[INDEX]], 8
+; CHECK-NEXT: [[INDEX1:%.*]] = phi i32 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT3:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[STEP_ADD:%.*]] = phi <4 x i32> [ [[INDUCTION]], %[[VEC_EPILOG_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[INDEX_NEXT3]] = add nuw i32 [[INDEX1]], 4
; CHECK-NEXT: [[VEC_IND_NEXT]] = add <4 x i32> [[STEP_ADD]], splat (i32 4)
-; CHECK-NEXT: [[TMP0:%.*]] = icmp eq i32 [[INDEX_NEXT]], 40
-; CHECK-NEXT: br i1 [[TMP0]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-NEXT: [[TMP2:%.*]] = icmp eq i32 [[INDEX_NEXT3]], 44
+; CHECK-NEXT: br i1 [[TMP2]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP7:![0-9]+]]
; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[VECTOR_RECUR_EXTRACT:%.*]] = extractelement <4 x i32> [[STEP_ADD]], i64 3
-; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
+; CHECK-NEXT: br i1 true, label %[[EXIT]], label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
+; CHECK-NEXT: [[TMP4:%.*]] = phi i32 [ [[VECTOR_RECUR_EXTRACT]], %[[MIDDLE_BLOCK]] ], [ [[TMP1]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-NEXT: [[TMP5:%.*]] = phi i32 [ 44, %[[MIDDLE_BLOCK]] ], [ 32, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ENTRY]] ]
; CHECK-NEXT: br label %[[LOOP:.*]]
; CHECK: [[LOOP]]:
-; CHECK-NEXT: [[D_0_I:%.*]] = phi i32 [ [[VECTOR_RECUR_EXTRACT]], %[[SCALAR_PH]] ], [ [[E_0_I:%.*]], %[[LOOP]] ]
-; CHECK-NEXT: [[E_0_I]] = phi i32 [ 40, %[[SCALAR_PH]] ], [ [[INC_I:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[D_0_I:%.*]] = phi i32 [ [[TMP4]], %[[SCALAR_PH]] ], [ [[E_0_I:%.*]], %[[LOOP]] ]
+; CHECK-NEXT: [[E_0_I]] = phi i32 [ [[TMP5]], %[[SCALAR_PH]] ], [ [[INC_I:%.*]], %[[LOOP]] ]
; CHECK-NEXT: [[DEAD:%.*]] = add i32 [[D_0_I]], 1
; CHECK-NEXT: [[INC_I]] = add i32 [[E_0_I]], 1
; CHECK-NEXT: [[EXITCOND_NOT_I:%.*]] = icmp eq i32 [[E_0_I]], 43
-; CHECK-NEXT: br i1 [[EXITCOND_NOT_I]], label %[[EXIT:.*]], label %[[LOOP]], !llvm.loop [[LOOP5:![0-9]+]]
+; CHECK-NEXT: br i1 [[EXITCOND_NOT_I]], label %[[EXIT]], label %[[LOOP]], !llvm.loop [[LOOP8:![0-9]+]]
; CHECK: [[EXIT]]:
; CHECK-NEXT: ret void
;
diff --git a/llvm/test/Transforms/LoopVectorize/X86/pr81872.ll b/llvm/test/Transforms/LoopVectorize/X86/pr81872.ll
index 24d5e36f6b67a..adc65339e7ec6 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/pr81872.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/pr81872.ll
@@ -14,30 +14,21 @@ define void @test(ptr noundef align 8 dereferenceable_or_null(16) %arr) #0 {
; CHECK-LABEL: define void @test(
; CHECK-SAME: ptr noundef align 8 dereferenceable_or_null(16) [[ARR:%.*]]) #[[ATTR0:[0-9]+]] {
; CHECK-NEXT: bb5:
-; CHECK-NEXT: br label [[VECTOR_PH:%.*]]
-; CHECK: vector.ph:
; CHECK-NEXT: br label [[VECTOR_BODY:%.*]]
-; CHECK: vector.body:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, [[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], [[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 99, i64 98, i64 97, i64 96>, [[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], [[VECTOR_BODY]] ]
-; CHECK-NEXT: [[VEC_IND1:%.*]] = phi <4 x i8> [ <i8 0, i8 1, i8 2, i8 3>, [[VECTOR_PH]] ], [ [[VEC_IND_NEXT2:%.*]], [[VECTOR_BODY]] ]
-; CHECK-NEXT: [[TMP0:%.*]] = icmp ule <4 x i8> [[VEC_IND1]], splat (i8 14)
-; CHECK-NEXT: [[OFFSET_IDX:%.*]] = sub i64 99, [[INDEX]]
-; CHECK-NEXT: [[TMP1:%.*]] = and <4 x i64> [[VEC_IND]], splat (i64 1)
-; CHECK-NEXT: [[TMP2:%.*]] = icmp eq <4 x i64> [[TMP1]], zeroinitializer
-; CHECK-NEXT: [[TMP3:%.*]] = select <4 x i1> [[TMP0]], <4 x i1> [[TMP2]], <4 x i1> zeroinitializer
-; CHECK-NEXT: [[TMP4:%.*]] = add i64 [[OFFSET_IDX]], 1
-; CHECK-NEXT: [[TMP5:%.*]] = getelementptr i64, ptr [[ARR]], i64 [[TMP4]]
-; CHECK-NEXT: [[TMP6:%.*]] = getelementptr i64, ptr [[TMP5]], i64 -3
-; CHECK-NEXT: [[REVERSE:%.*]] = shufflevector <4 x i1> [[TMP3]], <4 x i1> poison, <4 x i32> <i32 3, i32 2, i32 1, i32 0>
-; CHECK-NEXT: call void @llvm.masked.store.v4i64.p0(<4 x i64> splat (i64 1), ptr align 8 [[TMP6]], <4 x i1> [[REVERSE]])
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-NEXT: [[VEC_IND_NEXT]] = add nsw <4 x i64> [[VEC_IND]], splat (i64 -4)
-; CHECK-NEXT: [[VEC_IND_NEXT2]] = add nuw <4 x i8> [[VEC_IND1]], splat (i8 4)
-; CHECK-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], 16
-; CHECK-NEXT: br i1 [[TMP7]], label [[MIDDLE_BLOCK:%.*]], label [[VECTOR_BODY]], !prof [[PROF0:![0-9]+]], !llvm.loop [[LOOP1:![0-9]+]]
-; CHECK: middle.block:
-; CHECK-NEXT: br label [[BB6:%.*]]
+; CHECK: loop.header:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 99, [[BB5:%.*]] ], [ [[IV_NEXT:%.*]], [[BB6:%.*]] ]
+; CHECK-NEXT: [[AND:%.*]] = and i64 [[IV]], 1
+; CHECK-NEXT: [[ICMP17:%.*]] = icmp eq i64 [[AND]], 0
+; CHECK-NEXT: br i1 [[ICMP17]], label [[BB18:%.*]], label [[BB6]], !prof [[PROF0:![0-9]+]]
+; CHECK: bb18:
+; CHECK-NEXT: [[OR:%.*]] = or disjoint i64 [[IV]], 1
+; CHECK-NEXT: [[GETELEMENTPTR19:%.*]] = getelementptr inbounds i64, ptr [[ARR]], i64 [[OR]]
+; CHECK-NEXT: store i64 1, ptr [[GETELEMENTPTR19]], align 8
+; CHECK-NEXT: br label [[BB6]]
+; CHECK: loop.latch:
+; CHECK-NEXT: [[IV_NEXT]] = add nsw i64 [[IV]], -1
+; CHECK-NEXT: [[ICMP22:%.*]] = icmp eq i64 [[IV_NEXT]], 84
+; CHECK-NEXT: br i1 [[ICMP22]], label [[BB7:%.*]], label [[VECTOR_BODY]], !prof [[PROF1:![0-9]+]]
; CHECK: bb6:
; CHECK-NEXT: ret void
;
@@ -74,9 +65,6 @@ attributes #0 = {"target-cpu"="haswell" "target-features"="+avx2" }
!21 = !{!"branch_weights", i32 1, i32 1}
!22 = !{!"branch_weights", i32 1, i32 95}
;.
-; CHECK: [[PROF0]] = !{!"branch_weights", i32 1, i32 23}
-; CHECK: [[LOOP1]] = distinct !{[[LOOP1]], [[META2:![0-9]+]], [[META3:![0-9]+]], [[META4:![0-9]+]]}
-; CHECK: [[META2]] = !{!"llvm.loop.isvectorized", i32 1}
-; CHECK: [[META3]] = !{!"llvm.loop.unroll.runtime.disable"}
-; CHECK: [[META4]] = !{!"llvm.loop.estimated_trip_count", i32 24}
+; CHECK: [[PROF0]] = !{!"branch_weights", i32 1, i32 1}
+; CHECK: [[PROF1]] = !{!"branch_weights", i32 1, i32 95}
;.
diff --git a/llvm/test/Transforms/LoopVectorize/X86/reduction-small-size.ll b/llvm/test/Transforms/LoopVectorize/X86/reduction-small-size.ll
index 78c86f241adad..6e6741dee554f 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/reduction-small-size.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/reduction-small-size.ll
@@ -28,8 +28,7 @@ target datalayout = "e-m:e-i64:64-f80:128-n8:16:32:64-S128"
; CHECK: LV: Found an estimated cost of {{[0-9]+}} for VF 1 For instruction: %{{.*}} = trunc
; CHECK: LV: Found an estimated cost of {{[0-9]+}} for VF 1 For instruction: %{{.*}} = icmp
; CHECK: LV: Found an estimated cost of {{[0-9]+}} for VF 1 For instruction: br
-; CHECK: Cost of 1 for VF 2: induction instruction %indvars.iv.next = add nuw nsw i64 %indvars.iv, 1
-; CHECK: Cost of 1 for VF 2: induction instruction %indvars.iv = phi i64 [ %indvars.iv.next, %for.body ], [ 0, %for.body.preheader ]
+; CHECK: Cost of 1 for VF 2: canonical IV increment
; CHECK: Cost of 1 for VF 2: WIDEN-REDUCTION-PHI ir<%sum.013> = phi (add) vp<{{.+}}>, vp<[[EXT:%.+]]>
; CHECK: Cost of 0 for VF 2: vp<[[STEPS:%.+]]> = SCALAR-STEPS vp<[[CAN_IV:%.+]]>, ir<1>
; CHECK: Cost of 0 for VF 2: CLONE ir<%arrayidx> = getelementptr inbounds ir<%a>, vp<[[STEPS]]>
diff --git a/llvm/test/Transforms/LoopVectorize/induction-cost.ll b/llvm/test/Transforms/LoopVectorize/induction-cost.ll
index c7bbd8f56934a..b2a9999ed05ec 100644
--- a/llvm/test/Transforms/LoopVectorize/induction-cost.ll
+++ b/llvm/test/Transforms/LoopVectorize/induction-cost.ll
@@ -1,4 +1,4 @@
-; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --filter "(Cost.*WIDEN-INDUCTION)|(induction instruction)" --filter-out-after "LV: Selecting VF"
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --filter "(Cost.*WIDEN-INDUCTION)|(induction instruction)" --filter-out-after "LV: (Selecting|Using user) VF"
; REQUIRES: asserts
; RUN: opt -passes=loop-vectorize -force-vector-width=4 -force-vector-interleave=1 -debug-only=loop-vectorize -disable-output %s 2>&1 | FileCheck %s
@@ -23,8 +23,6 @@ exit:
define void @fp_induction(ptr noalias %dst, i64 %n) {
; CHECK-LABEL: 'fp_induction'
-; CHECK: Cost of 1 for VF 4: induction instruction %iv.next = add nuw nsw i64 %iv, 1
-; CHECK: Cost of 1 for VF 4: induction instruction %iv = phi i64 [ 0, %entry ], [ %iv.next, %loop ]
; CHECK: Cost of 2 for VF 4: ir<%f.iv> = WIDEN-INDUCTION fast ir<0.000000e+00>, ir<1.000000e+00>, vp<[[VP0:%[0-9]+]]>
;
entry:
More information about the llvm-commits
mailing list