[llvm] [LoopVectorize] Introduce check-first vectorization for multiple early-exit loops (PR #210492)

Arjun H Kumar via llvm-commits llvm-commits at lists.llvm.org
Fri Jul 17 23:58:12 PDT 2026


https://github.com/arjun-harikumar-amd created https://github.com/llvm/llvm-project/pull/210492

### Check-First Early-Exit Loop Vectorization

We introduce check-first strategy for vectorizing loops with multiple early exits. Instead of masking the loop body, the check-first approach evaluates all exit conditions first for a vector chunk and only runs the loop body if no exit fires. This generalizes early-exit vectorization to a much broader set of loops.

Implemented as a VPlan transformation in the loop vectorizer. It restructures the vector loop into three regions:

1. **Check blocks**: One check block per early exit, each holding the minimal condition slice to evaluate that exit. If any exit fires, control transfers to the exit-handling path.
2. **Vector body**: Runs only when no exit fires.
3. **Exit handling**: For the exiting chunk, lanes before the exit still need to execute. We have two strategies:
   - **Scalar replay** (default): fall back to the scalar loop, replaying from the start of the current chunk. 
   - **Masked replay** (experimental): clone the body's stores with lane masks derived from the first active exit lane, avoiding the scalar loop.

### Flags
```
-mllvm -enable-check-first-early-exit-vectorization 
-mllvm -enable-check-first-masked-replay
```

RFC to be posted soon.

Co-authored by @nema-ashutosh 

>From 899f88855c67a2ceb253459e3264b7c928ba0515 Mon Sep 17 00:00:00 2001
From: Arjun H Kumar <Arjun.HKumar at amd.com>
Date: Tue, 30 Jun 2026 23:24:28 +0530
Subject: [PATCH] [LoopVectorize] Introduce check-first vectorization for
 multiple early-exit loops

This patch introduces a new strategy 'Check-first' where we try to generalize
vectorizing loops with 'n' early exits.

The core idea is to check the exiting conditions first and execute the loop body later.
---
 .../Vectorize/LoopVectorizationLegality.h     |  22 +
 .../Vectorize/LoopVectorizationLegality.cpp   | 300 +++++-
 .../Transforms/Vectorize/LoopVectorize.cpp    | 111 ++-
 llvm/lib/Transforms/Vectorize/VPlan.cpp       |  28 +
 llvm/lib/Transforms/Vectorize/VPlan.h         |  91 +-
 .../Vectorize/VPlanConstruction.cpp           |  19 +-
 .../Transforms/Vectorize/VPlanPredicator.cpp  |   2 +
 .../lib/Transforms/Vectorize/VPlanRecipes.cpp |  13 +-
 .../Transforms/Vectorize/VPlanTransforms.cpp  | 935 +++++++++++++++++-
 .../Transforms/Vectorize/VPlanTransforms.h    |  10 +
 .../Transforms/Vectorize/VPlanVerifier.cpp    |   6 +
 .../VPlan/vplan-print-before-after-all.ll     |   1 +
 .../check-first-multi-exit-cascade.ll         | 133 +++
 .../LoopVectorize/check-first-nested-exit.ll  | 205 ++++
 14 files changed, 1840 insertions(+), 36 deletions(-)
 create mode 100644 llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
 create mode 100644 llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll

diff --git a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
index 3e8db73fd79d2..de7eb4466ef6b 100644
--- a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
+++ b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
@@ -440,6 +440,20 @@ class LoopVectorizationLegality {
     return getUncountableExitTrait() == UncountableExitTrait::ReadWrite;
   }
 
+  /// Returns true if every widened exit condition load is
+  /// dereferenceable for the complete trip count.
+  bool exitLoadsAreDereferenceable() const {
+    return AllExitLoadsDereferenceable;
+  }
+
+  /// Returns true if this early exit loop would use the check first strategy if
+  /// enabled.
+  bool wouldUseCheckFirstStyle() const {
+    if (hasUncountableExitWithSideEffects())
+      return true;
+    return hasUncountableEarlyExit() && !AllExitLoadsDereferenceable;
+  }
+
   /// Return true if there is store-load forwarding dependencies.
   bool isSafeForAnyStoreLoadForwardDistances() const {
     return LAI->getDepChecker().isSafeForAnyStoreLoadForwardDistances();
@@ -635,6 +649,10 @@ class LoopVectorizationLegality {
   /// for it.
   bool canUncountableExitConditionLoadBeMoved(BasicBlock *ExitingBlock);
 
+  /// Returns true if the exit conditions can be safely speculated.
+  bool canCheckFirstSpeculateExitConditions(
+      ArrayRef<BasicBlock *> ExitingBlocks);
+
   /// Return true if all of the instructions in the block can be speculatively
   /// executed, and record the loads/stores that require masking.
   /// \p SafePtrs is a list of addresses that are known to be legal and we know
@@ -752,6 +770,10 @@ class LoopVectorizationLegality {
   /// Records whether we have an uncountable early exit in a loop that's
   /// either read-only or read-write.
   UncountableExitTrait UncountableExitType = UncountableExitTrait::None;
+
+  /// Records whether every widened exit condition load is
+  /// dereferenceable for the complete trip count.
+  bool AllExitLoadsDereferenceable = true;
 };
 
 } // namespace llvm
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
index 8d875b2b6e492..c94de8d9951b3 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
@@ -79,6 +79,14 @@ static cl::opt<bool> EnableHistogramVectorization(
     "enable-histogram-loop-vectorization", cl::init(false), cl::Hidden,
     cl::desc("Enables autovectorization of some loops containing histograms"));
 
+static cl::opt<unsigned> MaxUncountableEarlyExits(
+    "max-uncountable-early-exits", cl::init(4), cl::Hidden,
+    cl::desc("Maximum number of uncountable early exits a loop may have to be "
+             "eligible for multi-exit check-first vectorization."));
+
+extern cl::opt<bool> EnableCheckFirstVectorization;
+extern cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects;
+
 /// Maximum vectorization interleave count.
 static const unsigned MaxInterleaveFactor = 16;
 
@@ -1616,6 +1624,98 @@ bool LoopVectorizationLegality::canVectorizeLoopNestCFG(
   return Result;
 }
 
+static unsigned countInLoopPredecessors(const BasicBlock *BB, const Loop *L) {
+  unsigned Count = 0;
+  for (const BasicBlock *Pred : predecessors(BB))
+    if (L->contains(Pred))
+      ++Count;
+  return Count;
+}
+
+/// Walks up the dominator chain from \p Exiting to the loop header, collecting
+/// into \p GuardConds the branch conditions that must hold for control to reach
+/// \p ExitingBB. Returns false if the control flow has a shape that cannot be
+/// represented as such a guard.
+///
+/// At each conditional dominator, one successor leads toward \p ExitingBB and the
+/// other bypasses it.
+/// Records a guard when the bypass leaves the guarded region cleanly:
+/// it targets \p Latch, or rejoins the chain after an if-without-else.
+static bool collectExitGuards(BasicBlock *Exiting, Loop *L, BasicBlock *Latch,
+                              const DominatorTree &DT,
+                              SmallVectorImpl<Value *> &GuardConds,
+                              bool AllowRejoin) {
+  BasicBlock *Header = L->getHeader();
+  BasicBlock *Cur = Exiting;
+  while (Cur != Header) {
+    DomTreeNode *Node = DT.getNode(Cur);
+    if (!Node || !Node->getIDom())
+      return false;
+    BasicBlock *IDom = Node->getIDom()->getBlock();
+    if (!L->contains(IDom))
+      return false;
+    if (auto *Br = dyn_cast<CondBrInst>(IDom->getTerminator())) {
+      BasicBlock *S0 = Br->getSuccessor(0), *S1 = Br->getSuccessor(1);
+      bool S0Dom = DT.dominates(S0, Cur), S1Dom = DT.dominates(S1, Cur);
+      if (S0Dom != S1Dom) {
+        BasicBlock *Interior = S0Dom ? S0 : S1;
+        BasicBlock *Bypass = S0Dom ? S1 : S0;
+        if (L->contains(Bypass)) {
+          if (Bypass == Latch) {
+            GuardConds.push_back(Br->getCondition());
+          } else if (!AllowRejoin) {
+            return false;
+          } else if (countInLoopPredecessors(Interior, L) == 1 &&
+                     countInLoopPredecessors(Bypass, L) > 1) {
+            GuardConds.push_back(Br->getCondition());
+          } else if (countInLoopPredecessors(Interior, L) > 1) {
+          } else {
+            return false;
+          }
+        }
+      }
+    }
+    Cur = IDom;
+  }
+  return true;
+}
+
+/// Collects the loads feeding the exit conditions of early-exits.
+/// condition, which check-first widens speculatively.
+static void collectExitConditionSliceLoads(ArrayRef<BasicBlock *> ExitingBlocks,
+                                           Loop *L, BasicBlock *Latch,
+                                           const DominatorTree &DT,
+                                           SmallVectorImpl<LoadInst *> &Out) {
+  SmallPtrSet<Value *, 16> Visited;
+  SmallVector<Value *, 16> Worklist;
+  for (BasicBlock *BB : ExitingBlocks) {
+    auto *Br = dyn_cast<CondBrInst>(BB->getTerminator());
+    assert(Br && "exiting block must terminate with a conditional branch");
+    Worklist.push_back(Br->getCondition());
+    collectExitGuards(BB, L, Latch, DT, Worklist, /*AllowRejoin=*/true);
+  }
+
+  for (BasicBlock *BB : L->blocks())
+    for (Instruction &I : *BB)
+      if (isa<StoreInst>(&I) && !DT.dominates(BB, Latch))
+        collectExitGuards(BB, L, Latch, DT, Worklist, /*AllowRejoin=*/false);
+
+  while (!Worklist.empty()) {
+    Value *V = Worklist.pop_back_val();
+    if (!Visited.insert(V).second)
+      continue;
+    auto *I = dyn_cast<Instruction>(V);
+    if (!I || !L->contains(I) || isa<PHINode>(I))
+      continue;
+    if (auto *LI = dyn_cast<LoadInst>(I)) {
+      Out.push_back(LI);
+      continue;
+    }
+    for (Value *Op : I->operands())
+      Worklist.push_back(Op);
+  }
+}
+
 bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
   BasicBlock *LatchBB = TheLoop->getLoopLatch();
   if (!LatchBB) {
@@ -1733,10 +1833,15 @@ bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
       return false;
     }
   } else {
-    // Check all uncountable exiting blocks for movable loads.
-    for (BasicBlock *ExitingBB : UncountableExitingBlocks) {
-      if (!canUncountableExitConditionLoadBeMoved(ExitingBB))
+    if (EnableCheckFirstVectorization &&
+        !EnableEarlyExitVectorizationWithSideEffects) {
+      if (!canCheckFirstSpeculateExitConditions(UncountableExitingBlocks))
         return false;
+    } else {
+      for (BasicBlock *ExitingBB : UncountableExitingBlocks) {
+        if (!canUncountableExitConditionLoadBeMoved(ExitingBB))
+          return false;
+      }
     }
   }
 
@@ -1754,6 +1859,57 @@ bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
     }
   }
 
+  // Safe only if every widened condition slice load is dereferenceable.
+  if (HasSideEffects) {
+    SmallVector<LoadInst *, 4> SpeculatedCondLoads;
+    collectExitConditionSliceLoads(UncountableExitingBlocks, TheLoop,
+                                   TheLoop->getLoopLatch(), *DT,
+                                   SpeculatedCondLoads);
+
+    bool AllDeref = true;
+    for (LoadInst *LI : SpeculatedCondLoads) {
+      if (!isDereferenceableAndAlignedInLoop(LI, TheLoop, *PSE.getSE(), *DT,
+                                             AC)) {
+        AllDeref = false;
+        break;
+      }
+    }
+
+    AllExitLoadsDereferenceable = AllDeref;
+
+    LLVM_DEBUG({
+      dbgs() << "LV: check-first early-exit memory-safety strategy: ";
+      if (AllDeref)
+        dbgs() << "all speculated condition-slice loads provably "
+                  "dereferenceable. \n";
+      else
+        dbgs() << "Condition-slice loads not provably dereferenceable. \n";
+    });
+  }
+
+  bool WillUseCheckFirst =
+      HasSideEffects && !EnableEarlyExitVectorizationWithSideEffects;
+  if (WillUseCheckFirst) {
+    const InductionDescriptor *IndDesc = nullptr;
+    if (Inductions.size() == 1) {
+      IndDesc = &Inductions.begin()->second;
+    } else if (PHINode *PrimaryIV = getPrimaryInduction()) {
+      auto It = Inductions.find(PrimaryIV);
+      if (It != Inductions.end())
+        IndDesc = &It->second;
+    }
+    if (!IndDesc ||
+        (IndDesc->getKind() != InductionDescriptor::IK_IntInduction &&
+         IndDesc->getKind() != InductionDescriptor::IK_PtrInduction) ||
+        !IndDesc->getConstIntStepValue()) {
+      reportVectorizationFailure(
+          "Check-first early-exit vectorization requires a single integer or "
+          "pointer induction with a constant step",
+          "UnsupportedCheckFirstInduction", ORE, TheLoop);
+      return false;
+    }
+  }
+
   [[maybe_unused]] const SCEV *SymbolicMaxBTC =
       PSE.getSymbolicMaxBackedgeTakenCount();
   // Since we have an exact exit count for the latch and the early exit
@@ -1854,6 +2010,144 @@ bool LoopVectorizationLegality::canUncountableExitConditionLoadBeMoved(
   return true;
 }
 
+bool LoopVectorizationLegality::canCheckFirstSpeculateExitConditions(
+    ArrayRef<BasicBlock *> ExitingBlocks) {
+  if (ExitingBlocks.size() > MaxUncountableEarlyExits) {
+    reportVectorizationFailure(
+        "Too many uncountable early exits for check-first vectorization",
+        "TooManyEarlyExitsForCheckFirst", ORE, TheLoop);
+    return false;
+  }
+
+  BasicBlock *Latch = TheLoop->getLoopLatch();
+  if (!Latch) {
+    reportVectorizationFailure("Check-first early-exit loop has no latch",
+                               "NoLatchCheckFirstExit", ORE, TheLoop);
+    return false;
+  }
+
+  SmallVector<Value *, 8> GuardConds;
+  for (BasicBlock *BB : ExitingBlocks) {
+    if (!collectExitGuards(BB, TheLoop, Latch, *DT, GuardConds,
+                           /*AllowRejoin=*/true)) {
+      reportVectorizationFailure(
+          "Check-first early-exit vectorization does not support this guarded "
+          "(conditionally-executed) early-exit control-flow shape",
+          "UnsupportedGuardedCheckFirstExit", ORE, TheLoop);
+      return false;
+    }
+  }
+
+  for (BasicBlock *BB : TheLoop->blocks())
+    for (Instruction &I : *BB)
+      if (isa<StoreInst>(&I) && !DT->dominates(BB, Latch)) {
+        SmallVector<Value *, 4> StoreGuards;
+        if (!collectExitGuards(BB, TheLoop, Latch, *DT, StoreGuards,
+                               /*AllowRejoin=*/false)) {
+          reportVectorizationFailure(
+              "Check-first early-exit vectorization does not support this "
+              "conditionally-executed (guarded) store control-flow shape",
+              "GuardedCheckFirstStore", ORE, TheLoop);
+          return false;
+        }
+      }
+
+  SmallPtrSet<LoadInst *, 8> CondLoads;
+  SmallVector<Value *, 16> Worklist;
+  SmallPtrSet<Value *, 16> Visited;
+  for (BasicBlock *BB : ExitingBlocks) {
+    auto *Br = dyn_cast<CondBrInst>(BB->getTerminator());
+    if (!Br) {
+      reportVectorizationFailure(
+          "Exiting block does not terminate with a conditional branch",
+          "UnsupportedCheckFirstExitTerminator", ORE, TheLoop);
+      return false;
+    }
+    Worklist.push_back(Br->getCondition());
+  }
+  append_range(Worklist, GuardConds);
+
+  while (!Worklist.empty()) {
+    Value *V = Worklist.pop_back_val();
+    if (!Visited.insert(V).second)
+      continue;
+    if (TheLoop->isLoopInvariant(V))
+      continue;
+    auto *I = dyn_cast<Instruction>(V);
+    if (!I || !TheLoop->contains(I)) {
+      reportVectorizationFailure(
+          "Early exit condition depends on a value that cannot be "
+          "speculatively evaluated for check-first vectorization",
+          "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+      return false;
+    }
+    if (auto *LI = dyn_cast<LoadInst>(I)) {
+      const auto *AR = dyn_cast<SCEVAddRecExpr>(
+          PSE.getSE()->getSCEV(LI->getPointerOperand()));
+      if (!LI->isSimple() || !AR || AR->getLoop() != TheLoop ||
+          !AR->isAffine()) {
+        reportVectorizationFailure(
+            "Early exit condition depends on a load that is not a simple "
+            "affine (unit-stride) access",
+            "CheckFirstExitLoadInvariantAddress", ORE, TheLoop);
+        return false;
+      }
+      CondLoads.insert(LI);
+      continue;
+    }
+    if (isa<PHINode>(I)) {
+      if (I->getParent() != TheLoop->getHeader()) {
+        reportVectorizationFailure(
+            "Early exit condition depends on a non-header PHI",
+            "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+        return false;
+      }
+      continue;
+    }
+    if (I->mayReadOrWriteMemory() || !isSafeToSpeculativelyExecute(I)) {
+      reportVectorizationFailure(
+          "Early exit condition contains an operation that cannot be "
+          "speculatively executed",
+          "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+      return false;
+    }
+    for (Value *Op : I->operands())
+      Worklist.push_back(Op);
+  }
+
+  SmallPtrSet<const Instruction *, 4> CondLoadSet(CondLoads.begin(),
+                                                  CondLoads.end());
+  ConditionallyExecutedOps.clear();
+  for (auto *BB : TheLoop->blocks()) {
+    for (auto &I : *BB) {
+      if (CondLoadSet.contains(&I) || !I.mayReadOrWriteMemory())
+        continue;
+      ConditionallyExecutedOps.insert(&I);
+      if (isa<LoadInst>(&I))
+        continue;
+      auto *SI = dyn_cast<StoreInst>(&I);
+      if (!SI) {
+        reportVectorizationFailure(
+            "Unsupported memory operation in check-first early-exit loop",
+            "UnsupportedCheckFirstMemOp", ORE, TheLoop);
+        return false;
+      }
+      for (LoadInst *CL : CondLoads) {
+        if (AA->alias(CL->getPointerOperand(), SI->getPointerOperand()) !=
+            AliasResult::NoAlias) {
+          reportVectorizationFailure(
+              "Cannot determine whether an early-exit condition load aliases "
+              "a store (deferred stores must not be observed out of order)",
+              "CheckFirstExitLoadAliasesStore", ORE, TheLoop);
+          return false;
+        }
+      }
+    }
+  }
+
+  return true;
+}
+
 bool LoopVectorizationLegality::canVectorize(bool UseVPlanNativePath) {
   // Store the result and return it at the end instead of exiting early, in case
   // allowExtraAnalysis is used to report multiple reasons for not vectorizing.
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index f1ca4061bfd9e..e7377ff7cea2b 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -411,12 +411,34 @@ static cl::opt<bool> EnableEarlyExitVectorization(
     cl::desc(
         "Enable vectorization of early exit loops with uncountable exits."));
 
-static cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects(
+cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects(
     "enable-early-exit-vectorization-with-side-effects", cl::init(false),
     cl::Hidden,
     cl::desc("Enable vectorization of early exit loops with uncountable exits "
              "and side effects"));
 
+cl::opt<bool> EnableCheckFirstVectorization(
+    "enable-check-first-early-exit-vectorization", cl::init(false), cl::Hidden,
+    cl::desc("Enable check-first vectorization of early exit loops with "
+             "multiple exits."));
+
+cl::opt<bool> EnableCheckFirstMaskedReplay(
+    "enable-check-first-masked-replay", cl::init(false), cl::Hidden,
+    cl::desc("Replace scalar replay in check-first early exit vectorization "
+             "with a masked vector replay."));
+
+/// Returns true if loop uses check first with scalar replay of the
+/// failing chunk.
+static bool usesCheckFirstReplay(const LoopVectorizationLegality *Legal) {
+  if (!Legal->hasUncountableEarlyExit())
+    return false;
+  if (EnableCheckFirstMaskedReplay)
+    return false;
+  if (Legal->hasUncountableExitWithSideEffects())
+    return !EnableEarlyExitVectorizationWithSideEffects;
+  return EnableCheckFirstVectorization && Legal->wouldUseCheckFirstStyle();
+}
+
 // Likelyhood of bypassing the vectorized loop because there are zero trips left
 // after prolog. See `emitIterationCountCheck`.
 static constexpr uint32_t MinItersBypassWeights[] = {1, 127};
@@ -3668,6 +3690,11 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
   if (Plan.hasEarlyExit())
     return 1;
 
+  // Interleaving would break check-first scalar-replay resume wiring. 
+  // So forcing IC=1.
+  if (usesCheckFirstReplay(Legal))
+    return 1;
+
   const bool HasReductions =
       any_of(Plan.getVectorLoopRegion()->getEntryBasicBlock()->phis(),
              IsaPred<VPReductionPHIRecipe>);
@@ -5496,6 +5523,10 @@ void LoopVectorizationPlanner::plan(ElementCount UserVF, unsigned UserIC) {
   if (!MaxFactors) // Cases that should not to be vectorized nor interleaved.
     return;
 
+  // Disable scalable vectorization for check-first early-exit loops for now.
+  if (usesCheckFirstReplay(Legal))
+    MaxFactors.ScalableVF = ElementCount::getScalable(0);
+
   Config.collectInLoopReductions();
   // Cases that may be vectorized may be optimized by unit stride predicates.
   // TODO: Currently unit stride predicates are added unconditionally, even if
@@ -5968,6 +5999,11 @@ DenseMap<const SCEV *, Value *> LoopVectorizationPlanner::executePlan(
   // Regions are dissolved after optimizing for VF and UF, which completely
   // removes unneeded loop regions first.
   RUN_VPLAN_PASS(VPlanTransforms::dissolveLoopRegions, BestVPlan);
+  // Scalar replay routes check.exit to the scalar preheader after region
+  // dissolution.
+  VPlanTransforms::wireCheckFirstExitToScalar(BestVPlan);
+  // Masked replay routes check.exit to the real early exit block instead.
+  VPlanTransforms::wireCheckFirstMaskedReplayToExit(BestVPlan);
   // Expand BranchOnTwoConds after dissolution, when latch has direct access to
   // its successors.
   RUN_VPLAN_PASS(VPlanTransforms::expandBranchOnTwoConds, BestVPlan);
@@ -6178,8 +6214,6 @@ VPRecipeBase *VPRecipeBuilder::tryToWidenMemory(VPInstruction *VPI,
   if (!LoopVectorizationPlanner::getDecisionAndClampRange(WillWiden, Range))
     return nullptr;
 
-  // If a mask is not required, drop it - use unmasked version for safe loads.
-  // TODO: Determine if mask is needed in VPlan.
   VPValue *Mask = CM.isMaskRequired(I) ? VPI->getMask() : nullptr;
 
   // Determine if the pointer operand of the access is either consecutive or
@@ -6580,10 +6614,21 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
   //       the presence of an uncountable exit and the presence of stores in
   //       the loop inside handleEarlyExits itself.
   UncountableExitStyle EEStyle = UncountableExitStyle::NoUncountableExit;
-  if (Legal->hasUncountableEarlyExit())
-    EEStyle = Legal->hasUncountableExitWithSideEffects()
-                  ? UncountableExitStyle::MaskedHandleExitInScalarLoop
-                  : UncountableExitStyle::ReadOnly;
+  if (Legal->hasUncountableEarlyExit()) {
+    if (Legal->hasUncountableExitWithSideEffects()) {
+      if (EnableEarlyExitVectorizationWithSideEffects)
+        EEStyle = UncountableExitStyle::MaskedHandleExitInScalarLoop;
+      else
+        EEStyle = UncountableExitStyle::CheckFirst;
+    } else {
+      EEStyle = UncountableExitStyle::ReadOnly;
+    }
+  }
+
+  assert((EEStyle != UncountableExitStyle::CheckFirst ||
+          Legal->exitLoadsAreDereferenceable()) &&
+         "check-first vectorization reached for a loop whose speculatively "
+         "widened loads could not be made memory-safe");
 
   if (!RUN_VPLAN_PASS(VPlanTransforms::handleEarlyExits, *VPlan0, EEStyle,
                       OrigLoop, PSE, *DT, Legal->getAssumptionCache())) {
@@ -6592,7 +6637,8 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
 
   RUN_VPLAN_PASS(VPlanTransforms::createLoopRegions, *VPlan0,
                  getDebugLocFromInstOrOperands(Legal->getPrimaryInduction()));
-  if (CM.foldTailByMasking())
+  if (CM.foldTailByMasking() &&
+      EEStyle != UncountableExitStyle::CheckFirst)
     RUN_VPLAN_PASS(VPlanTransforms::foldTailByMasking, *VPlan0);
   RUN_VPLAN_PASS(VPlanTransforms::introduceMasksAndLinearize, *VPlan0);
 
@@ -7286,8 +7332,12 @@ static void checkMixedPrecision(Loop *L, OptimizationRemarkEmitter *ORE) {
 /// TODO: This is currently overly pessimistic because the loop may not take
 /// the early exit, but better to keep this conservative for now. In future,
 /// it might be possible to relax this by using branch probabilities.
+///
+/// For check first loops, add the scalar replay cost of the early exit chunk.
 static InstructionCost calculateEarlyExitCost(VPCostContext &CostCtx,
-                                              VPlan &Plan, ElementCount VF) {
+                                              VPlan &Plan, ElementCount VF,
+                                              uint64_t ScalarCostPerIter,
+                                              bool IsCheckFirstReplay) {
   InstructionCost Cost = 0;
   for (auto *ExitVPBB : Plan.getExitBlocks()) {
     for (auto *PredVPBB : ExitVPBB->getPredecessors()) {
@@ -7301,6 +7351,15 @@ static InstructionCost calculateEarlyExitCost(VPCostContext &CostCtx,
       }
     }
   }
+
+  if (IsCheckFirstReplay && VF.isFixed()) {
+    uint64_t ReplayedIters = VF.getFixedValue() / 2;
+    InstructionCost ReplayCost(ScalarCostPerIter * ReplayedIters);
+    LLVM_DEBUG(dbgs() << "LV: Adding check-first scalar-replay cost "
+                      << ReplayCost << " (~" << ReplayedIters
+                      << " scalar iterations) for VF " << VF << ".\n");
+    Cost += ReplayCost;
+  }
   return Cost;
 }
 
@@ -7317,7 +7376,8 @@ static bool isOutsideLoopWorkProfitable(GeneratedRTChecks &Checks,
                                         PredicatedScalarEvolution &PSE,
                                         VPCostContext &CostCtx, VPlan &Plan,
                                         EpilogueLowering SEL,
-                                        std::optional<unsigned> VScale) {
+                                        std::optional<unsigned> VScale,
+                                        bool IsCheckFirstReplay) {
   InstructionCost RtC = Checks.getCost();
   if (!RtC.isValid())
     return false;
@@ -7343,8 +7403,8 @@ static bool isOutsideLoopWorkProfitable(GeneratedRTChecks &Checks,
 
   InstructionCost TotalCost = RtC;
   // Add on the cost of any work required in the vector early exit block, if
-  // one exists.
-  TotalCost += calculateEarlyExitCost(CostCtx, Plan, VF.Width);
+  // one exists plus check first scalar replay cost.
+  TotalCost += calculateEarlyExitCost(CostCtx, Plan, VF.Width, ScalarC, IsCheckFirstReplay);
   TotalCost += Plan.getMiddleBlock()->cost(VF.Width, CostCtx);
 
   // First, compute the minimum iteration count required so that the vector
@@ -7924,6 +7984,19 @@ bool LoopVectorizePass::processLoop(Loop *L) {
     return false;
   }
 
+  // Memory-safety gate: bail to the scalar loop when a speculatively widened
+  // exit-condition load is not provably dereferenceable.
+  if (EnableCheckFirstVectorization &&
+      !EnableEarlyExitVectorizationWithSideEffects &&
+      LVL.wouldUseCheckFirstStyle() && !LVL.exitLoadsAreDereferenceable()) {
+    reportVectorizationFailure(
+        "check-first early-exit memory-safety strategy is Unsafe: a "
+        "speculatively-widened condition load could not be proven "
+        "dereferenceable for the full trip count",
+        "CheckFirstUnsafeMemSafety", ORE, L);
+    return false;
+  }
+
   bool IsInnerLoop = L->isInnermost();
 
   // Outer loops require a computable trip count.
@@ -7940,11 +8013,12 @@ bool LoopVectorizePass::processLoop(Loop *L) {
       return false;
     }
     if (LVL.hasUncountableExitWithSideEffects() &&
-        !EnableEarlyExitVectorizationWithSideEffects) {
-      reportVectorizationFailure("Auto-vectorization of loops with uncountable "
-                                 "early exit and side effects is not enabled",
-                                 "UncountableEarlyExitSideEffectLoopsDisabled",
-                                 ORE, L);
+        !EnableEarlyExitVectorizationWithSideEffects &&
+        !EnableCheckFirstVectorization) {
+      reportVectorizationFailure(
+          "Auto-vectorization of loops with uncountable "
+          "early exit and side effects is not enabled",
+          "UncountableEarlyExitSideEffectLoopsDisabled", ORE, L);
       return false;
     }
   }
@@ -8123,7 +8197,8 @@ bool LoopVectorizePass::processLoop(Loop *L) {
                           CM.PSE, L);
     if (!ForceVectorization &&
         !isOutsideLoopWorkProfitable(Checks, VF, L, PSE, CostCtx, *BestPlanPtr,
-                                     SEL, Config.getVScaleForTuning())) {
+                                     SEL, Config.getVScaleForTuning(),
+                                     usesCheckFirstReplay(&LVL))) {
       ORE->emit([&]() {
         return OptimizationRemarkAnalysisAliasing(
                    DEBUG_TYPE, "CantReorderMemOps", L->getStartLoc(),
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.cpp b/llvm/lib/Transforms/Vectorize/VPlan.cpp
index 9c0bec2350c97..f9260639a5f23 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlan.cpp
@@ -1206,6 +1206,18 @@ static void remapOperands(VPBlockBase *Entry, VPBlockBase *NewEntry,
   }
 }
 
+/// Returns the clone of Old recorded in Old2NewBlocks, or nullptr when
+/// Old is null.
+static VPBasicBlock *
+remapClonedBlock(const DenseMap<VPBlockBase *, VPBlockBase *> &Old2NewBlocks,
+                 VPBasicBlock *Old) {
+  if (!Old)
+    return nullptr;
+  VPBlockBase *New = Old2NewBlocks.lookup(Old);
+  assert(New && "Check-first block not found in cloned plan.");
+  return cast<VPBasicBlock>(New);
+}
+
 VPlan *VPlan::duplicate() {
   unsigned NumBlocksBeforeCloning = CreatedBlocks.size();
   // Clone blocks.
@@ -1291,6 +1303,22 @@ VPlan *VPlan::duplicate() {
       NewPlan->ExitBlocks.push_back(cast<VPIRBasicBlock>(VPB));
   }
 
+  NewPlan->CheckFirst.EarlyExitBlock = CheckFirst.EarlyExitBlock;
+  NewPlan->CheckFirst.InclusiveReplayStores = CheckFirst.InclusiveReplayStores;
+
+  // Map each original block to its clone via a single depth-first traversal.
+  DenseMap<VPBlockBase *, VPBlockBase *> Old2NewBlocks;
+  for (const auto &[OldBB, NewBB] :
+       zip_equal(vp_depth_first_deep(Entry), vp_depth_first_deep(NewEntry)))
+    Old2NewBlocks[OldBB] = NewBB;
+
+  NewPlan->CheckFirst.ExitBlock =
+      remapClonedBlock(Old2NewBlocks, CheckFirst.ExitBlock);
+  NewPlan->CheckFirst.MaskedReplayBlock =
+      remapClonedBlock(Old2NewBlocks, CheckFirst.MaskedReplayBlock);
+  NewPlan->CheckFirst.CheckHeaderBlock =
+      remapClonedBlock(Old2NewBlocks, CheckFirst.CheckHeaderBlock);
+
   return NewPlan;
 }
 
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index eaf9d1433aff7..d6215a0570def 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -87,6 +87,11 @@ enum class UncountableExitStyle {
   /// uncountable exit is taken, then all lanes before the exiting lane will
   /// complete, leaving just the final lane to execute in the scalar tail.
   MaskedHandleExitInScalarLoop,
+  /// Check-first semantics: exit conditions are evaluated at the start of each
+  /// vector iteration before any stores execute. On early exit, the scalar loop
+  /// resumes from the start of the current vector chunk. Stores are placed in a
+  /// separate body block only reachable when no exit fires.
+  CheckFirst,
 };
 
 /// VPBlockBase is the building block of the Hierarchical Control-Flow Graph.
@@ -3746,6 +3751,14 @@ class LLVM_ABI_FOR_TEST VPWidenMemoryRecipe : public VPIRMetadata {
   /// Whether the memory access is masked.
   bool IsMasked = false;
 
+  VPWidenMemoryRecipe(Instruction &I, bool Consecutive,
+                      const VPIRMetadata &Metadata)
+      : VPIRMetadata(Metadata), Ingredient(I),
+        Alignment(getLoadStoreAlignment(&I)), Consecutive(Consecutive) {}
+
+public:
+  virtual ~VPWidenMemoryRecipe() = default;
+
   void setMask(VPValue *Mask) {
     assert(!IsMasked && "cannot re-set mask");
     if (!Mask)
@@ -3756,14 +3769,6 @@ class LLVM_ABI_FOR_TEST VPWidenMemoryRecipe : public VPIRMetadata {
     IsMasked = true;
   }
 
-  VPWidenMemoryRecipe(Instruction &I, bool Consecutive,
-                      const VPIRMetadata &Metadata)
-      : VPIRMetadata(Metadata), Ingredient(I),
-        Alignment(getLoadStoreAlignment(&I)), Consecutive(Consecutive) {}
-
-public:
-  virtual ~VPWidenMemoryRecipe() = default;
-
   /// Return a VPRecipeBase* to the current object.
   virtual VPRecipeBase *getAsRecipe() = 0;
   virtual const VPRecipeBase *getAsRecipe() const = 0;
@@ -4756,6 +4761,28 @@ class VPlan {
   /// VPIRBasicBlock wrapping the header of the original scalar loop.
   VPIRBasicBlock *ScalarHeader;
 
+  /// Used for recording the state for check-first early-exit vectorization.
+  struct CheckFirstEarlyExitState {
+    /// Block for routing check-first early exits to the scalar preheader.
+    VPBasicBlock *ExitBlock = nullptr;
+
+    /// The check.exit block used for masked replay wiring.
+    VPBasicBlock *MaskedReplayBlock = nullptr;
+
+    /// Early-exit IR block for masked replay.
+    BasicBlock *EarlyExitBlock = nullptr;
+
+    /// First block of the check-first cascade.
+    VPBasicBlock *CheckHeaderBlock = nullptr;
+
+    /// Block the masked-replay chain temporarily branches to.
+    VPBlockBase *MaskedReplayTempTarget = nullptr;
+
+    /// Stores that must replay the exiting lane.
+    SmallPtrSet<const Instruction *, 4> InclusiveReplayStores;
+  };
+  CheckFirstEarlyExitState CheckFirst;
+
   /// Immutable list of VPIRBasicBlocks wrapping the exit blocks of the original
   /// scalar loop. Note that some exit blocks may be unreachable at the moment,
   /// e.g. if the scalar epilogue always executes.
@@ -4886,6 +4913,54 @@ class VPlan {
         getScalarHeader()->getSinglePredecessor());
   }
 
+  VPBasicBlock *getCheckFirstExitBlock() const { return CheckFirst.ExitBlock; }
+
+  void setCheckFirstExitBlock(VPBasicBlock *VPBB) {
+    assert((!CheckFirst.ExitBlock || CheckFirst.ExitBlock == VPBB) &&
+           "CheckFirstExitBlock already set");
+    CheckFirst.ExitBlock = VPBB;
+  }
+
+  VPBasicBlock *getCheckFirstMaskedReplayBlock() const {
+    return CheckFirst.MaskedReplayBlock;
+  }
+
+  void setCheckFirstMaskedReplayBlock(VPBasicBlock *VPBB) {
+    CheckFirst.MaskedReplayBlock = VPBB;
+  }
+
+  BasicBlock *getCheckFirstEarlyExitBlock() const {
+    return CheckFirst.EarlyExitBlock;
+  }
+
+  void setCheckFirstEarlyExitBlock(BasicBlock *IRBB) {
+    CheckFirst.EarlyExitBlock = IRBB;
+  }
+
+  void addCheckFirstInclusiveReplayStore(const Instruction *I) {
+    CheckFirst.InclusiveReplayStores.insert(I);
+  }
+
+  bool isCheckFirstInclusiveReplayStore(const Instruction *I) const {
+    return CheckFirst.InclusiveReplayStores.contains(I);
+  }
+
+  VPBasicBlock *getCheckFirstCheckHeaderBlock() const {
+    return CheckFirst.CheckHeaderBlock;
+  }
+
+  void setCheckFirstCheckHeaderBlock(VPBasicBlock *VPBB) {
+    CheckFirst.CheckHeaderBlock = VPBB;
+  }
+
+  VPBlockBase *getCheckFirstMaskedReplayTempTarget() const {
+    return CheckFirst.MaskedReplayTempTarget;
+  }
+
+  void setCheckFirstMaskedReplayTempTarget(VPBlockBase *B) {
+    CheckFirst.MaskedReplayTempTarget = B;
+  }
+
   /// Return the VPIRBasicBlock wrapping the header of the scalar loop.
   VPIRBasicBlock *getScalarHeader() const { return ScalarHeader; }
 
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index 619fea8c10b4d..5581bbe1fe696 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -462,13 +462,25 @@ static void createLoopRegion(VPlan &Plan, VPBlockBase *HeaderVPB, DebugLoc DL) {
   // Transfer latch's successors to the region.
   VPBlockUtils::transferSuccessors(LatchVPBB, R);
 
+  VPBasicBlock *CheckExitVPBB = Plan.getCheckFirstExitBlock();
+  if (CheckExitVPBB) {
+    assert(CheckExitVPBB->empty() &&
+           "check.exit block should be empty before region creation");
+    assert(CheckExitVPBB->getNumSuccessors() == 0 &&
+           "check.exit should have no successors before temporary edge");
+    VPBlockUtils::connectBlocks(CheckExitVPBB, LatchVPBB);
+  }
+
   VPBlockUtils::connectBlocks(PreheaderVPBB, R);
   R->setEntry(HeaderVPB);
   R->setExiting(LatchVPBB);
 
   // All VPBB's reachable shallowly from HeaderVPB belong to the current region.
-  for (VPBlockBase *VPBB : vp_depth_first_shallow(HeaderVPB))
+  for (VPBlockBase *VPBB : vp_depth_first_shallow(HeaderVPB)) {
+    if (VPBB == Plan.getScalarPreheader())
+      continue;
     VPBB->setParent(R);
+  }
 
   if (!IsOutermost)
     return;
@@ -1275,6 +1287,8 @@ bool VPlanTransforms::handleEarlyExits(VPlan &Plan, UncountableExitStyle Style,
     // Dereferenceability is checked separately for uncountable exit loops with
     // stores, as only the loads contributing to the exit condition need to
     // be checked.
+    // ReadOnly needs all loads dereferenceable whereas CheckFirst checks only
+    // condition-slice loads later in handleUncountableEarlyExits.
     if (Style == UncountableExitStyle::ReadOnly &&
         !areAllLoadsDereferenceable(HeaderVPBB, TheLoop, PSE, DT, AC))
       return false;
@@ -1348,7 +1362,8 @@ void VPlanTransforms::createLoopRegions(VPlan &Plan, DebugLoc DL) {
 
   VPRegionBlock *TopRegion = Plan.getVectorLoopRegion();
   TopRegion->setName("vector loop");
-  TopRegion->getEntryBasicBlock()->setName("vector.body");
+  TopRegion->getEntryBasicBlock()->setName(
+      Plan.getCheckFirstCheckHeaderBlock() ? "vector.check" : "vector.body");
 }
 
 void VPlanTransforms::foldTailByMasking(VPlan &Plan) {
diff --git a/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp b/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
index 655ac58e24426..d678b147abb5a 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
@@ -274,6 +274,8 @@ void VPlanTransforms::introduceMasksAndLinearize(VPlan &Plan) {
   // Nested loop regions (outer-loop vectorization) are not supported yet.
   if (Plan.isOuterLoop())
     return;
+  if (Plan.getCheckFirstExitBlock())
+    return;
   VPRegionBlock *LoopRegion = Plan.getVectorLoopRegion();
   // Scan the body of the loop in a topological order to visit each basic block
   // after having visited its predecessor basic blocks.
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index f45b9e4f6c35b..fd0b7974879a1 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -95,6 +95,7 @@ bool VPRecipeBase::mayWriteToMemory() const {
   case VPBlendSC:
   case VPReductionEVLSC:
   case VPReductionSC:
+  case VPVectorEndPointerSC:
   case VPVectorPointerSC:
   case VPWidenCanonicalIVSC:
   case VPWidenCastSC:
@@ -149,6 +150,7 @@ bool VPRecipeBase::mayReadFromMemory() const {
   case VPBlendSC:
   case VPReductionEVLSC:
   case VPReductionSC:
+  case VPVectorEndPointerSC:
   case VPVectorPointerSC:
   case VPWidenCanonicalIVSC:
   case VPWidenCastSC:
@@ -870,9 +872,13 @@ Value *VPInstruction::generate(VPTransformState &State) {
         cast<VPBasicBlock>(getParent()->getSuccessors()[1]);
     BasicBlock *SecondIRSucc = State.CFG.VPBB2IRBB.lookup(SecondVPSucc);
     BasicBlock *IRBB = State.CFG.VPBB2IRBB[getParent()];
-    auto *Br = Builder.CreateCondBr(Cond, IRBB, SecondIRSucc);
+    // Placeholder for a successor. Assigned in connectToPredecessors.
+    auto *Br =
+        Builder.CreateCondBr(Cond, IRBB, SecondIRSucc ? SecondIRSucc : IRBB);
     // First successor is always forward, reset it to nullptr.
     Br->setSuccessor(0, nullptr);
+    if (!SecondIRSucc)
+      Br->setSuccessor(1, nullptr);
     IRBB->getTerminator()->eraseFromParent();
     applyMetadata(*Br);
     return Br;
@@ -1552,6 +1558,11 @@ void VPInstruction::addOperand(VPValue *Op) {
            "matching operand 1's type and i1, respectively");
     break;
   }
+  case Instruction::PHI:
+    assert((getNumOperands() == 0 ||
+            Ty == getOperand(0)->getScalarType()) &&
+           "all incoming values must have the same type");
+    break;
   default:
     llvm_unreachable("opcode does not support growing the operand list "
                      "outside of construction");
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
index ceb7c38ca3e29..2aba5b8363b3d 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
@@ -35,11 +35,13 @@
 #include "llvm/Analysis/MemoryLocation.h"
 #include "llvm/Analysis/ScalarEvolutionPatternMatch.h"
 #include "llvm/Analysis/ScopedNoAliasAA.h"
+#include "llvm/Analysis/ValueTracking.h"
 #include "llvm/Analysis/VectorUtils.h"
 #include "llvm/IR/Intrinsics.h"
 #include "llvm/IR/MDBuilder.h"
 #include "llvm/IR/Metadata.h"
 #include "llvm/Support/Casting.h"
+#include "llvm/Support/ErrorHandling.h"
 #include "llvm/Support/TypeSize.h"
 #include "llvm/Transforms/Utils/LoopUtils.h"
 #include "llvm/Transforms/Utils/ScalarEvolutionExpander.h"
@@ -48,6 +50,8 @@ using namespace llvm;
 using namespace VPlanPatternMatch;
 using namespace SCEVPatternMatch;
 
+extern cl::opt<bool> EnableCheckFirstMaskedReplay;
+
 bool VPlanTransforms::tryToConvertVPInstructionsToVPRecipes(
     VPlan &Plan, const TargetLibraryInfo &TLI) {
 
@@ -78,14 +82,21 @@ bool VPlanTransforms::tryToConvertVPInstructionsToVPRecipes(
         assert(!isa<PHINode>(Inst) && "phis should be handled above");
         // Create VPWidenMemoryRecipe for loads and stores.
         if (LoadInst *Load = dyn_cast<LoadInst>(Inst)) {
+          VPValue *Mask = Ingredient.getNumOperands() > 1
+                              ? Ingredient.getOperand(1)
+                              : nullptr;
           NewRecipe = new VPWidenLoadRecipe(
-              *Load, Ingredient.getOperand(0), nullptr /*Mask*/,
+              *Load, Ingredient.getOperand(0), Mask,
               false /*Consecutive*/, *VPI, Ingredient.getDebugLoc());
         } else if (StoreInst *Store = dyn_cast<StoreInst>(Inst)) {
+          // Stores carry check-first guard predicate
+          // as the third operand. It becomes the store mask.
+          VPValue *Mask = Ingredient.getNumOperands() > 2
+                              ? Ingredient.getOperand(2)
+                              : nullptr;
           NewRecipe = new VPWidenStoreRecipe(
-              *Store, Ingredient.getOperand(1), Ingredient.getOperand(0),
-              nullptr /*Mask*/, false /*Consecutive*/, *VPI,
-              Ingredient.getDebugLoc());
+              *Store, Ingredient.getOperand(1), Ingredient.getOperand(0), Mask,
+              false /*Consecutive*/, *VPI, Ingredient.getDebugLoc());
         } else if (GetElementPtrInst *GEP = dyn_cast<GetElementPtrInst>(Inst)) {
           NewRecipe = new VPWidenGEPRecipe(GEP->getSourceElementType(),
                                            Ingredient.operands(), *VPI,
@@ -2561,6 +2572,15 @@ static bool cannotHoistOrSinkRecipe(VPRecipeBase &R, VPBasicBlock *FirstBB,
       match(&R, m_Intrinsic<Intrinsic::assume>()))
     return vputils::cannotHoistOrSinkRecipe(R, Sinking);
 
+  bool InSingleSuccChain = false;
+  for (VPBlockBase *Succ = FirstBB; Succ; Succ = Succ->getSingleSuccessor())
+    if (Succ == LastBB) {
+      InSingleSuccChain = true;
+      break;
+    }
+  if (!InSingleSuccChain)
+    return true;
+
   // Check that the memory operation doesn't alias between FirstBB and LastBB.
   auto MemLoc = vputils::getMemoryLocation(R);
 
@@ -2864,6 +2884,8 @@ void VPlanTransforms::optimize(VPlan &Plan) {
   RUN_VPLAN_PASS(simplifyReverses, Plan);
   RUN_VPLAN_PASS(removeDeadRecipes, Plan);
 
+  RUN_VPLAN_PASS(maskCheckFirstReplayStores, Plan);
+
   RUN_VPLAN_PASS(createAndOptimizeReplicateRegions, Plan);
   RUN_VPLAN_PASS(mergeBlocksIntoPredecessors, Plan);
   RUN_VPLAN_PASS(licm, Plan);
@@ -4596,6 +4618,29 @@ static bool handleUncountableExitsWithSideEffects(
   return true;
 }
 
+/// Walk backward from ExitCond to collect the recipes needed to evaluate the
+/// exit condition, stopping at PHIs. Returns false if the condition depends on
+/// a memory-writing recipe which cannot be placed in the check block.
+static bool computeConditionSlice(VPValue *ExitCond,
+                                  SmallPtrSetImpl<VPRecipeBase *> &Slice) {
+  SmallVector<VPValue *, 16> Worklist;
+  Worklist.push_back(ExitCond);
+  while (!Worklist.empty()) {
+    VPValue *V = Worklist.pop_back_val();
+    VPRecipeBase *DefR = V->getDefiningRecipe();
+    if (!DefR || DefR->isPhi())
+      continue;
+    if (Slice.contains(DefR))
+      continue;
+    if (DefR->mayWriteToMemory())
+      return false;
+    Slice.insert(DefR);
+    for (VPValue *Op : DefR->operands())
+      Worklist.push_back(Op);
+  }
+  return true;
+}
+
 bool VPlanTransforms::handleUncountableEarlyExits(
     VPlan &Plan, VPBasicBlock *HeaderVPBB, VPBasicBlock *LatchVPBB,
     VPBasicBlock *MiddleVPBB, Loop *TheLoop, PredicatedScalarEvolution &PSE,
@@ -4663,6 +4708,393 @@ bool VPlanTransforms::handleUncountableEarlyExits(
              "RPO sort must place dominating exits before dominated ones");
 #endif
 
+  if (Style == UncountableExitStyle::CheckFirst) {
+
+    if (range_size(TheLoop->getHeader()->phis()) > 1)
+      return false;
+
+    SmallPtrSet<const VPBlockBase *, 4> ExitBlockSet;
+    SmallPtrSet<const VPBlockBase *, 4> EarlyExitingSet;
+    for (const EarlyExitInfo &E : Exits) {
+      ExitBlockSet.insert(E.EarlyExitVPBB);
+      EarlyExitingSet.insert(E.EarlyExitingVPBB);
+    }
+    if (EarlyExitingSet.contains(LatchVPBB))
+      return false;
+    SmallVector<VPBasicBlock *, 4> LoopChain;
+    struct DiamondGuard {
+      VPValue *Cond;
+      bool ChainOnFalseEdge;
+      VPBasicBlock *Rejoin;
+      unsigned RejoinIdx;
+    };
+    DenseMap<VPBasicBlock *, DiamondGuard> ChainDiamond;
+    {
+      SmallPtrSet<VPBasicBlock *, 8> Seen;
+      VPBasicBlock *Cur = HeaderVPBB;
+      while (Cur) {
+        if (!Seen.insert(Cur).second)
+          return false;
+        LoopChain.push_back(Cur);
+        if (Cur != HeaderVPBB && Cur->begin() != Cur->end() &&
+            Cur->begin()->isPhi())
+          return false;
+        if (Cur == LatchVPBB)
+          break;
+        SmallVector<VPBasicBlock *, 2> InLoopSuccs;
+        for (VPBlockBase *S : Cur->getSuccessors()) {
+          if (S == MiddleVPBB || ExitBlockSet.contains(S))
+            continue;
+          auto *SBB = dyn_cast<VPBasicBlock>(S);
+          if (!SBB)
+            return false;
+          InLoopSuccs.push_back(SBB);
+        }
+        VPBasicBlock *NextInLoop = nullptr;
+        if (InLoopSuccs.size() == 1) {
+          NextInLoop = InLoopSuccs[0];
+        } else if (InLoopSuccs.size() == 2) {
+          VPBasicBlock *Cont = nullptr;
+          VPBasicBlock *Bypass = nullptr;
+          if (InLoopSuccs[0] == LatchVPBB) {
+            Bypass = InLoopSuccs[0];
+            Cont = InLoopSuccs[1];
+          } else if (InLoopSuccs[1] == LatchVPBB) {
+            Bypass = InLoopSuccs[1];
+            Cont = InLoopSuccs[0];
+          } else {
+            bool S0Merge = InLoopSuccs[0]->getNumPredecessors() > 1;
+            bool S1Merge = InLoopSuccs[1]->getNumPredecessors() > 1;
+            if (S0Merge && !S1Merge) {
+              Bypass = InLoopSuccs[0];
+              Cont = InLoopSuccs[1];
+            } else if (S1Merge && !S0Merge) {
+              Bypass = InLoopSuccs[1];
+              Cont = InLoopSuccs[0];
+            } else {
+              return false;
+            }
+          }
+          VPValue *GuardCond = nullptr;
+          if (!match(Cur->getTerminator(), m_BranchOnCond(m_VPValue(GuardCond))))
+            return false;
+          bool ChainOnFalseEdge = Cur->getSuccessors()[1] == Cont;
+          ChainDiamond[Cur] = {GuardCond, ChainOnFalseEdge, Bypass, 0};
+          NextInLoop = Cont;
+        } else {
+          return false;
+        }
+        Cur = NextInLoop;
+      }
+      if (LoopChain.empty() || LoopChain.back() != LatchVPBB)
+        return false;
+    }
+
+    DenseMap<const VPBasicBlock *, unsigned> ChainIdx;
+    for (unsigned I = 0, E = LoopChain.size(); I != E; ++I)
+      ChainIdx[LoopChain[I]] = I;
+    for (auto &[DBB, DG] : ChainDiamond) {
+      auto RIt = ChainIdx.find(DG.Rejoin);
+      if (RIt == ChainIdx.end() || RIt->second <= ChainIdx[DBB])
+        return false;
+      DG.RejoinIdx = RIt->second;
+    }
+
+    ArrayRef<VPBasicBlock *> Intermediates =
+        ArrayRef(LoopChain).drop_front().drop_back();
+
+    for (EarlyExitInfo &Exit : Exits) {
+      auto It = llvm::find(LoopChain, Exit.EarlyExitingVPBB);
+      if (It == LoopChain.end())
+        return false;
+      unsigned ExitIdx = std::distance(LoopChain.begin(), It);
+      auto *MC = cast<VPInstruction>(Exit.CondToExit);
+      VPBuilder GuardBuilder(MC);
+      VPValue *CombinedGuard = nullptr;
+      for (unsigned I = 0; I < ExitIdx; ++I) {
+        auto DiamondIt = ChainDiamond.find(LoopChain[I]);
+        if (DiamondIt == ChainDiamond.end() ||
+            ExitIdx >= DiamondIt->second.RejoinIdx)
+          continue;
+        VPValue *Guard = DiamondIt->second.Cond;
+        if (DiamondIt->second.ChainOnFalseEdge)
+          Guard = GuardBuilder.createNot(Guard);
+        CombinedGuard =
+            CombinedGuard ? GuardBuilder.createLogicalAnd(CombinedGuard, Guard)
+                          : Guard;
+      }
+      if (CombinedGuard)
+        MC->setOperand(
+            0, GuardBuilder.createLogicalAnd(CombinedGuard, MC->getOperand(0)));
+    }
+
+    DenseMap<VPRecipeBase *, SmallVector<std::pair<VPValue *, bool>, 2>>
+        StoreGuards;
+    DenseMap<VPRecipeBase *, SmallVector<std::pair<VPValue *, bool>, 2>>
+        BodyLoadGuards;
+    for (unsigned J = 0, E = LoopChain.size(); J != E; ++J) {
+      if (LoopChain[J] == LatchVPBB)
+        continue;
+      SmallVector<std::pair<VPValue *, bool>, 2> Guards;
+      for (unsigned I = 0; I < J; ++I) {
+        auto GIt = ChainDiamond.find(LoopChain[I]);
+        if (GIt != ChainDiamond.end() && J < GIt->second.RejoinIdx)
+          Guards.push_back({GIt->second.Cond, GIt->second.ChainOnFalseEdge});
+      }
+      if (Guards.empty())
+        continue;
+      for (VPRecipeBase &R : *LoopChain[J]) {
+        if (R.mayWriteToMemory()) {
+          StoreGuards[&R] = Guards;
+        } else {
+          auto *VPI = dyn_cast<VPInstruction>(&R);
+          if (VPI && VPI->getUnderlyingValue())
+            if (auto *LI = dyn_cast<LoadInst>(VPI->getUnderlyingValue()))
+              if (!isSafeToSpeculativelyExecute(LI))
+                BodyLoadGuards[&R] = Guards;
+        }
+      }
+    }
+
+    // Compute each exit's condition slice.
+    SmallVector<SmallPtrSet<VPRecipeBase *, 16>, 4> Slices(Exits.size());
+    DenseMap<VPRecipeBase *, unsigned> EarliestCheck;
+    for (unsigned K = 0, E = Exits.size(); K != E; ++K) {
+      if (!computeConditionSlice(Exits[K].CondToExit, Slices[K]))
+        return false;
+      for (VPRecipeBase *R : Slices[K])
+        EarliestCheck.try_emplace(R, K);
+    }
+
+
+    VPIRBasicBlock *MaskedReplayExitBB = nullptr;
+    bool DoMaskedReplay = EnableCheckFirstMaskedReplay && Exits.size() == 1;
+
+    if (DoMaskedReplay) {
+      ScalarEvolution &SE = *PSE.getSE();
+      PHINode *IndVar = TheLoop->getInductionVariable(SE);
+      // Reconstructed as CanonIV + first_active_lane.
+      bool IsUnitFromZero = false;
+      if (IndVar) {
+        if (auto *AR = dyn_cast<SCEVAddRecExpr>(SE.getSCEV(IndVar))) {
+          if (AR->getLoop() == TheLoop && AR->getStart()->isZero()) {
+            const SCEV *Step = AR->getStepRecurrence(SE);
+            IsUnitFromZero = Step->isOne();
+          }
+        }
+      }
+      VPBasicBlock *EEing = Exits[0].EarlyExitingVPBB;
+      for (VPRecipeBase &R : Exits[0].EarlyExitVPBB->phis()) {
+        VPValue *Incoming =
+            cast<VPIRPhi>(&R)->getIncomingValueForBlock(EEing);
+        Value *UV = Incoming->getUnderlyingValue();
+        if (UV)
+          UV = UV->stripPointerCasts();
+        if (!IsUnitFromZero || !UV || UV != IndVar)
+          DoMaskedReplay = false;
+      }
+      if (DoMaskedReplay) {
+        MaskedReplayExitBB = Exits[0].EarlyExitVPBB;
+
+        BasicBlock *EarlyExitIR = MaskedReplayExitBB->getIRBasicBlock();
+        BasicBlock *Latch = TheLoop->getLoopLatch();
+        BasicBlock *ExitingIR = nullptr;
+        for (BasicBlock *Pred : predecessors(EarlyExitIR))
+          if (TheLoop->contains(Pred) && Pred != Latch) {
+            ExitingIR = Pred;
+            break;
+          }
+        if (ExitingIR)
+          for (BasicBlock *BB : TheLoop->blocks())
+            for (Instruction &I : *BB)
+              if (isa<StoreInst>(&I) && DT.dominates(BB, ExitingIR))
+                Plan.addCheckFirstInclusiveReplayStore(&I);
+      }
+    }
+
+    // From here, all modifications are destructive. We cannot bail out.
+
+    for (auto &Exit : Exits) {
+      auto &[EarlyExitingVPBB, EarlyExitVPBB, _] = Exit;
+      for (VPRecipeBase &R : EarlyExitVPBB->phis())
+        cast<VPIRPhi>(&R)->removeIncomingValueFor(EarlyExitingVPBB);
+      EarlyExitingVPBB->getTerminator()->eraseFromParent();
+      VPBlockUtils::disconnectBlocks(EarlyExitingVPBB, EarlyExitVPBB);
+    }
+
+    // Flatten intermediate blocks recipes into the header, then connect the
+    // header straight to the latch. Single-exit chains are left untouched.
+    if (!Intermediates.empty()) {
+      if (!EarlyExitingSet.contains(HeaderVPBB))
+        HeaderVPBB->getTerminator()->eraseFromParent();
+      for (VPBasicBlock *BB : Intermediates)
+        if (!EarlyExitingSet.contains(BB) && BB->getTerminator())
+          BB->getTerminator()->eraseFromParent();
+      for (VPBasicBlock *BB : Intermediates)
+        for (VPRecipeBase &R : make_early_inc_range(*BB))
+          R.moveBefore(*HeaderVPBB, HeaderVPBB->end());
+      for (VPBlockBase *S : to_vector(HeaderVPBB->getSuccessors()))
+        VPBlockUtils::disconnectBlocks(HeaderVPBB, S);
+      for (VPBasicBlock *BB : Intermediates)
+        for (VPBlockBase *S : to_vector(BB->getSuccessors()))
+          VPBlockUtils::disconnectBlocks(BB, S);
+      VPBlockUtils::connectBlocks(HeaderVPBB, LatchVPBB);
+    }
+
+    // Hoist guard producers to the header to avoid use-before-def.
+    if (DoMaskedReplay) {
+      SmallVector<VPRecipeBase *, 8> GuardHoist;
+      auto EnqueueProducers = [&](VPValue *V) {
+        if (VPRecipeBase *Def = V->getDefiningRecipe())
+          if (!EarliestCheck.contains(Def))
+            GuardHoist.push_back(Def);
+      };
+      for (auto &[R, Guards] : StoreGuards)
+        for (auto &[G, ChainOnFalseEdge] : Guards)
+          EnqueueProducers(G);
+      for (auto &[R, Guards] : BodyLoadGuards)
+        for (auto &[G, ChainOnFalseEdge] : Guards)
+          EnqueueProducers(G);
+      while (!GuardHoist.empty()) {
+        VPRecipeBase *R = GuardHoist.pop_back_val();
+        if (EarliestCheck.try_emplace(R, 0).second)
+          for (VPValue *Op : R->operands())
+            EnqueueProducers(Op);
+      }
+    }
+
+    // Create the check cascade. Checks[0] is the header.
+    // Checks[k] is a fresh block holding exit k's condition slice.
+    SmallVector<VPBasicBlock *, 4> Checks;
+    Checks.push_back(HeaderVPBB);
+    for (unsigned K = 1, E = Exits.size(); K != E; ++K)
+      Checks.push_back(Plan.createVPBasicBlock("vector.check"));
+
+    // Body block holds non-slice recipes. Runs when no exit fires.
+    VPBasicBlock *BodyVPBB = Plan.createVPBasicBlock("vector.body");
+
+    // Partition recipes. Slice recipes to their check block, everything else to
+    // the body.
+    auto PartitionBlock = [&](VPBasicBlock *BB) {
+      for (VPRecipeBase &R : make_early_inc_range(*BB)) {
+        if (R.isPhi() || &R == BB->getTerminator())
+          continue;
+        auto It = EarliestCheck.find(&R);
+        VPBasicBlock *Target =
+            It != EarliestCheck.end() ? Checks[It->second] : BodyVPBB;
+        if (Target != BB)
+          R.moveBefore(*Target, Target->end());
+      }
+    };
+    PartitionBlock(HeaderVPBB);
+    if (HeaderVPBB != LatchVPBB)
+      PartitionBlock(LatchVPBB);
+
+    // Attach combined guard mask as a trailing operand to each guarded store.
+    for (auto &[R, Guards] : StoreGuards) {
+      auto *OldStore = cast<VPInstruction>(R);
+      assert(OldStore->getOpcode() == Instruction::Store &&
+             "check-first guarded side effect is not a store");
+      VPBuilder GuardBuilder;
+      if (DoMaskedReplay) {
+        if (VPRecipeBase *Term = HeaderVPBB->getTerminator())
+          GuardBuilder.setInsertPoint(HeaderVPBB, Term->getIterator());
+        else
+          GuardBuilder.setInsertPoint(HeaderVPBB);
+      } else {
+        GuardBuilder.setInsertPoint(OldStore);
+      }
+      VPValue *StGuard = nullptr;
+      for (auto &[G, ChainOnFalseEdge] : Guards) {
+        VPValue *GG = ChainOnFalseEdge ? GuardBuilder.createNot(G) : G;
+        StGuard = StGuard ? GuardBuilder.createLogicalAnd(StGuard, GG) : GG;
+      }
+      assert(StGuard && "guarded store recorded without any guard");
+      SmallVector<VPValue *, 3> Ops(OldStore->operands());
+      Ops.push_back(StGuard);
+      auto *NewStore = new VPInstruction(Instruction::Store, Ops, *OldStore,
+                                         *OldStore, OldStore->getDebugLoc());
+      NewStore->setUnderlyingValue(OldStore->getUnderlyingValue());
+      NewStore->insertBefore(OldStore);
+      OldStore->eraseFromParent();
+    }
+
+    for (auto &[R, Guards] : BodyLoadGuards) {
+      auto *OldLoad = cast<VPInstructionWithType>(R);
+      VPBuilder GuardBuilder(OldLoad);
+      VPValue *LdGuard = nullptr;
+      for (auto &[G, ChainOnFalseEdge] : Guards) {
+        VPValue *GG = ChainOnFalseEdge ? GuardBuilder.createNot(G) : G;
+        LdGuard = LdGuard ? GuardBuilder.createLogicalAnd(LdGuard, GG) : GG;
+      }
+      assert(LdGuard && "masked body load recorded without any guard");
+      SmallVector<VPValue *, 2> Ops(OldLoad->operands());
+      Ops.push_back(LdGuard);
+      auto *NewLoad = new VPInstructionWithType(
+          Instruction::Load, Ops, OldLoad->getResultType(), *OldLoad, *OldLoad,
+          OldLoad->getDebugLoc(), OldLoad->getName(),
+          OldLoad->getUnderlyingValue());
+      NewLoad->insertBefore(OldLoad);
+      OldLoad->replaceAllUsesWith(NewLoad);
+      OldLoad->eraseFromParent();
+    }
+
+    // Routes to the scalar preheader.
+    VPBasicBlock *EarlyExitToScalarVPBB =
+        Plan.createVPBasicBlock("vector.check.exit");
+    Plan.setCheckFirstExitBlock(EarlyExitToScalarVPBB);
+
+    // Masked replay fills this block later else it routes to the scalar PH.
+    if (DoMaskedReplay)
+      Plan.setCheckFirstMaskedReplayBlock(EarlyExitToScalarVPBB);
+
+    // Extract the latch condition before erasing the latch terminator.
+    auto *LatchBranch = cast<VPInstruction>(LatchVPBB->getTerminator());
+    assert(LatchBranch->getOpcode() == VPInstruction::BranchOnCond &&
+           "Unexpected terminator");
+    VPValue *IsLatchExitTaken = LatchBranch->getOperand(0);
+    DebugLoc LatchDL = LatchBranch->getDebugLoc();
+    LatchBranch->eraseFromParent();
+
+    if (HeaderVPBB != LatchVPBB) {
+      for (VPBlockBase *Succ : to_vector(LatchVPBB->getSuccessors()))
+        VPBlockUtils::disconnectBlocks(LatchVPBB, Succ);
+    }
+
+    for (VPBlockBase *Succ : to_vector(HeaderVPBB->getSuccessors()))
+      VPBlockUtils::disconnectBlocks(HeaderVPBB, Succ);
+
+    // Wire each check.k to exit if exit k fires, else fall through to the next
+    // check or body. BranchOnCond takes successor 0 when true: wire exit first.
+    for (unsigned K = 0, E = Exits.size(); K != E; ++K) {
+      VPBasicBlock *CheckBB = Checks[K];
+      VPBuilder CheckBuilder(CheckBB, CheckBB->end());
+      VPValue *IsExitTaken =
+          CheckBuilder.createNaryOp(VPInstruction::AnyOf, {Exits[K].CondToExit});
+      CheckBuilder.createNaryOp(VPInstruction::BranchOnCond, {IsExitTaken});
+      VPBasicBlock *NextBB = (K + 1 != E) ? Checks[K + 1] : BodyVPBB;
+      VPBlockUtils::connectBlocks(CheckBB, EarlyExitToScalarVPBB);
+      VPBlockUtils::connectBlocks(CheckBB, NextBB);
+    }
+
+
+    Plan.setCheckFirstCheckHeaderBlock(HeaderVPBB);
+
+    // Remember the real early-exit block for wireCheckFirstMaskedReplayToExit.
+    if (DoMaskedReplay)
+      Plan.setCheckFirstEarlyExitBlock(MaskedReplayExitBB->getIRBasicBlock());
+
+    VPBuilder BodyBuilder(BodyVPBB, BodyVPBB->end());
+    BodyBuilder.createNaryOp(VPInstruction::BranchOnCond, {IsLatchExitTaken},
+                             LatchDL);
+
+    // Wire: BodyVPBB → {MiddleVPBB , HeaderVPBB}
+    VPBlockUtils::connectBlocks(BodyVPBB, MiddleVPBB);
+    VPBlockUtils::connectBlocks(BodyVPBB, HeaderVPBB);
+
+    return true;
+  }
+
   // Build the AnyOf condition for the latch terminator using logical OR
   // to avoid poison propagation from later exit conditions when an earlier
   // exit is taken.
@@ -4825,6 +5257,498 @@ bool VPlanTransforms::handleUncountableEarlyExits(
   return true;
 }
 
+/// Returns true if Root transitively uses Target through its defining
+/// recipes operands.
+static bool vpValueDependsOn(VPValue *Root, VPValue *Target) {
+  SmallVector<VPValue *, 16> Worklist{Root};
+  SmallPtrSet<VPValue *, 16> Visited;
+  while (!Worklist.empty()) {
+    VPValue *V = Worklist.pop_back_val();
+    if (V == Target)
+      return true;
+    if (!Visited.insert(V).second)
+      continue;
+    if (VPRecipeBase *Def = V->getDefiningRecipe())
+      for (VPValue *Op : Def->operands())
+        Worklist.push_back(Op);
+  }
+  return false;
+}
+
+static VPValue *
+rebuildIVResumeExprImpl(VPBuilder &B, VPValue *V, VPValue *VectorTC,
+                        VPValue *NewIndex,
+                        SmallDenseMap<VPValue *, VPValue *> &Cache) {
+  using namespace VPlanPatternMatch;
+  if (V == VectorTC)
+    return NewIndex;
+  if (auto It = Cache.find(V); It != Cache.end())
+    return It->second;
+  auto Remap = [&](VPValue *Op) {
+    return rebuildIVResumeExprImpl(B, Op, VectorTC, NewIndex, Cache);
+  };
+  auto RemapBinOp = [&](unsigned Opcode, VPValue *LHS, VPValue *RHS,
+                        DebugLoc DL, const Twine &Name) -> VPValue * {
+    VPValue *NL = Remap(LHS), *NR = Remap(RHS);
+    if (NL == LHS && NR == RHS)
+      return V;
+    auto Flags =
+        cast<VPRecipeWithIRFlags>(V->getDefiningRecipe())->getNoWrapFlags();
+    return B.createOverflowingOp(Opcode, {NL, NR}, Flags, DL, Name);
+  };
+  auto RemapCast = [&](Instruction::CastOps Opcode, VPValue *Op) -> VPValue * {
+    VPValue *NO = Remap(Op);
+    if (NO == Op)
+      return V;
+    return B.createScalarCast(Opcode, NO, V->getScalarType(), DebugLoc());
+  };
+
+  VPValue *A, *Bv;
+  VPValue *Result = V;
+  if (match(V, m_VPInstruction<VPInstruction::PtrAdd>(m_VPValue(A),
+                                                      m_VPValue(Bv)))) {
+    VPValue *NB = Remap(Bv);
+    if (NB != Bv)
+      Result = B.createPtrAdd(A, NB, DebugLoc(), "check.exit.iv.resume");
+  } else if (match(V, m_Mul(m_VPValue(A), m_VPValue(Bv)))) {
+    Result = RemapBinOp(Instruction::Mul, A, Bv, DebugLoc::getUnknown(), "");
+  } else if (match(V, m_c_Add(m_VPValue(A), m_VPValue(Bv)))) {
+    Result =
+        RemapBinOp(Instruction::Add, A, Bv, DebugLoc(), "check.exit.iv.resume");
+  } else if (match(V, m_Sub(m_VPValue(A), m_VPValue(Bv)))) {
+    Result = RemapBinOp(Instruction::Sub, A, Bv, DebugLoc::getUnknown(), "");
+  } else if (match(V, m_Trunc(m_VPValue(A)))) {
+    Result = RemapCast(Instruction::Trunc, A);
+  } else if (match(V, m_ZExt(m_VPValue(A)))) {
+    Result = RemapCast(Instruction::ZExt, A);
+  } else if (match(V, m_SExt(m_VPValue(A)))) {
+    Result = RemapCast(Instruction::SExt, A);
+  } else if (vpValueDependsOn(V, VectorTC)) {
+    assert(false && "Unhandled VectorTC-dependent check-first resume "
+                    "value");
+  }
+  Cache[V] = Result;
+  return Result;
+}
+
+/// Returns the induction resume expression Expr rebuilt with VectorTC
+/// replaced by NewIndex.
+static VPValue *rebuildIVResumeExpr(VPBuilder &B, VPValue *Expr,
+                                    VPValue *VectorTC, VPValue *NewIndex) {
+  SmallDenseMap<VPValue *, VPValue *> Cache;
+  return rebuildIVResumeExprImpl(B, Expr, VectorTC, NewIndex, Cache);
+}
+
+void VPlanTransforms::wireCheckFirstExitToScalar(VPlan &Plan) {
+  VPBasicBlock *CheckExitVPBB = Plan.getCheckFirstExitBlock();
+  if (!CheckExitVPBB)
+    return;
+
+  // Masked replay handles check.exit during restructuring.
+  if (Plan.getCheckFirstMaskedReplayBlock())
+    return;
+
+  VPBasicBlock *HeaderVPBB = Plan.getCheckFirstCheckHeaderBlock();
+  assert(HeaderVPBB && "check-first cascade header block not recorded");
+
+  // Replace the temporary check.exit→body edge with check.exit→ScalarPH.
+  assert(CheckExitVPBB->getNumSuccessors() == 1 &&
+         "check.exit should have exactly one successor after dissolution");
+  VPBlockBase *OldSucc = CheckExitVPBB->getSuccessors()[0];
+  VPBlockUtils::disconnectBlocks(CheckExitVPBB, OldSucc);
+
+  VPBasicBlock *ScalarPH = Plan.getScalarPreheader();
+  assert(ScalarPH &&
+         "CheckFirst requires a scalar preheader for early-exit replay. "
+         "Ensure the scalar tail is not removed by earlier passes.");
+  VPBlockUtils::connectBlocks(CheckExitVPBB, ScalarPH);
+
+  assert(HeaderVPBB->getNumPredecessors() == 2 &&
+         "loop header must have exactly two predecessors (preheader, latch)");
+  VPBasicBlock *LatchVPBB = nullptr;
+  for (VPBlockBase *Pred : HeaderVPBB->getPredecessors()) {
+    auto *PredVPBB = cast<VPBasicBlock>(Pred);
+    if (any_of(PredVPBB->getSuccessors(),
+               [&](VPBlockBase *S) { return S != HeaderVPBB; })) {
+      LatchVPBB = PredVPBB;
+      break;
+    }
+  }
+  assert(LatchVPBB &&
+         "could not identify the loop latch (backedge source) among the "
+         "header's predecessors");
+  VPValue *CanonIV = cast<VPPhi>(&*HeaderVPBB->begin());
+
+  auto SPHPhis = ScalarPH->phis();
+  assert(range_size(SPHPhis) == 1 &&
+         "CheckFirst expects exactly one scalar-preheader PHI (the IV). "
+         "Extending to multiple inductions or live-outs requires computing "
+         "proper resume values for each PHI.");
+
+  VPBasicBlock *MiddleVPBB = nullptr;
+  for (VPBlockBase *Succ : LatchVPBB->getSuccessors()) {
+    if (Succ != HeaderVPBB) {
+      MiddleVPBB = cast<VPBasicBlock>(Succ);
+      break;
+    }
+  }
+
+  VPBuilder CheckExitBuilder(CheckExitVPBB, CheckExitVPBB->getFirstNonPhi());
+  VPValue *VectorTC = &Plan.getVectorTripCount();
+
+  using namespace VPlanPatternMatch;
+  for (VPRecipeBase &R : SPHPhis) {
+    auto *Phi = cast<VPPhi>(&R);
+
+    VPValue *MidVal = nullptr;
+    for (unsigned I = 0, E = Phi->getNumIncoming(); I != E; ++I) {
+      if (MiddleVPBB && Phi->getIncomingBlock(I) == MiddleVPBB) {
+        MidVal = Phi->getIncomingValue(I);
+        break;
+      }
+    }
+
+    VPValue *ResumeVal =
+        MidVal ? rebuildIVResumeExpr(CheckExitBuilder, MidVal, VectorTC, CanonIV)
+               : CanonIV;
+
+    assert(!(MidVal && ResumeVal == MidVal &&
+             vpValueDependsOn(MidVal, VectorTC)) &&
+           "check-first early-exit resume value could not be rebuilt from the "
+           "vector trip count");
+
+    Phi->addIncoming(ResumeVal);
+  }
+}
+
+void VPlanTransforms::maskCheckFirstReplayStores(VPlan &Plan) {
+  VPBasicBlock *ReplayBB = Plan.getCheckFirstMaskedReplayBlock();
+  if (!ReplayBB)
+    return;
+
+  // Replay surviving lanes body stores here, masked to [0, first_active_lane).
+  assert(ReplayBB->getNumPredecessors() == 1 &&
+         "masked-replay block must have a single (check) predecessor");
+  auto *CheckBB = cast<VPBasicBlock>(ReplayBB->getPredecessors()[0]);
+
+  // The combined per-lane exit condition is the AnyOf operand of BranchOnCond.
+  auto *CheckTerm = cast<VPInstruction>(CheckBB->getTerminator());
+  assert(CheckTerm->getOpcode() == VPInstruction::BranchOnCond &&
+         "check block terminator must be BranchOnCond");
+  auto *AnyOf = cast<VPInstruction>(CheckTerm->getOperand(0)->getDefiningRecipe());
+  assert(AnyOf->getOpcode() == VPInstruction::AnyOf &&
+         "BranchOnCond operand must be AnyOf");
+  VPValue *Combined = AnyOf->getOperand(0);
+
+  VPBasicBlock *BodyVPBB = nullptr;
+  for (VPBlockBase *Succ : CheckBB->getSuccessors()) {
+    if (Succ != ReplayBB) {
+      BodyVPBB = cast<VPBasicBlock>(Succ);
+      break;
+    }
+  }
+  assert(BodyVPBB && "could not find the body block to replay");
+
+#ifndef NDEBUG
+  // Masked replay requires full-width chunks; tail folding is unsupported.
+  for (VPBlockBase *VPB : vp_depth_first_deep(Plan.getEntry()))
+    if (auto *VPBB = dyn_cast<VPBasicBlock>(VPB))
+      for (VPRecipeBase &R : *VPBB)
+        assert(!isa<VPActiveLaneMaskPHIRecipe>(&R) &&
+               "check-first masked replay is unsound under tail folding");
+#endif
+
+  VPBuilder HeadBuilder(ReplayBB, ReplayBB->begin());
+  VPInstruction *FirstActiveLane = HeadBuilder.createFirstActiveLane(
+      {Combined}, DebugLoc::getUnknown(), "first.active.lane");
+
+  Type *IndexTy = FirstActiveLane->getScalarType();
+  assert(IndexTy->isIntegerTy() &&
+         "FirstActiveLane must produce an integer index for mask bounds");
+  VPValue *Zero = Plan.getZero(IndexTy);
+  VPValue *One = Plan.getConstantInt(IndexTy, 1);
+  VPInstruction *ExclMask = nullptr;
+  VPInstruction *InclMask = nullptr;
+  auto getMaskFor = [&](VPRecipeBase &R) -> VPValue * {
+    const Instruction *SI = nullptr;
+    if (auto *WS = dyn_cast<VPWidenStoreRecipe>(&R))
+      SI = &WS->getIngredient();
+    else if (auto *RR = dyn_cast<VPReplicateRecipe>(&R))
+      SI = RR->getUnderlyingInstr();
+    bool Inclusive = SI && Plan.isCheckFirstInclusiveReplayStore(SI);
+    if (Inclusive) {
+      if (!InclMask) {
+        VPInstruction *FALPlusOne =
+            VPBuilder::getToInsertAfter(FirstActiveLane)
+                .createAdd(FirstActiveLane, One, DebugLoc(),
+                           "first.active.lane.incl",
+                           {/*nuw=*/true, /*nsw=*/false});
+        InclMask = VPBuilder::getToInsertAfter(FALPlusOne)
+                       .createNaryOp(VPInstruction::ActiveLaneMask,
+                                     {Zero, FALPlusOne, One}, DebugLoc(),
+                                     "masked.replay.mask.incl");
+      }
+      return InclMask;
+    }
+    if (!ExclMask)
+      ExclMask = VPBuilder::getToInsertAfter(FirstActiveLane)
+                     .createNaryOp(VPInstruction::ActiveLaneMask,
+                                   {Zero, FirstActiveLane, One}, DebugLoc(),
+                                   "masked.replay.mask");
+    return ExclMask;
+  };
+
+  // Replay only the stores and their backward slice.
+  SmallPtrSet<VPRecipeBase *, 8> Needed;
+  SmallVector<VPRecipeBase *, 8> Work;
+  for (VPRecipeBase &R : *BodyVPBB)
+    if (R.mayWriteToMemory()) {
+      Needed.insert(&R);
+      Work.push_back(&R);
+    }
+  while (!Work.empty()) {
+    VPRecipeBase *R = Work.pop_back_val();
+    for (VPValue *Op : R->operands())
+      if (VPRecipeBase *Def = Op->getDefiningRecipe())
+        if (Def->getParent() == BodyVPBB && Needed.insert(Def).second)
+          Work.push_back(Def);
+  }
+
+  SmallPtrSet<VPRecipeBase *, 4> InclusiveFeeds;
+  {
+    SmallVector<VPRecipeBase *, 4> InclWork;
+    for (VPRecipeBase &R : *BodyVPBB) {
+      if (!Needed.contains(&R))
+        continue;
+      const Instruction *SI = nullptr;
+      if (auto *WS = dyn_cast<VPWidenStoreRecipe>(&R))
+        SI = &WS->getIngredient();
+      else if (auto *RR = dyn_cast<VPReplicateRecipe>(&R))
+        if (RR->getUnderlyingInstr()->mayWriteToMemory())
+          SI = RR->getUnderlyingInstr();
+      if (SI && Plan.isCheckFirstInclusiveReplayStore(SI))
+        InclWork.push_back(&R);
+    }
+    while (!InclWork.empty()) {
+      VPRecipeBase *R = InclWork.pop_back_val();
+      for (VPValue *Op : R->operands()) {
+        VPRecipeBase *Def = Op->getDefiningRecipe();
+        if (Def && Def->getParent() == BodyVPBB && Needed.contains(Def) &&
+            InclusiveFeeds.insert(Def).second)
+          InclWork.push_back(Def);
+      }
+    }
+  }
+
+  VPBuilder Builder(ReplayBB, std::next(FirstActiveLane->getIterator()));
+  DenseMap<VPValue *, VPValue *> OperandMap;
+  for (VPRecipeBase &R : *BodyVPBB) {
+    if (!Needed.contains(&R))
+      continue;
+
+    auto Remap = [&](VPValue *V) -> VPValue * {
+      VPValue *Mapped = OperandMap.lookup(V);
+      return Mapped ? Mapped : V;
+    };
+
+    SmallVector<VPValue *, 4> NewOps;
+    for (VPValue *Op : R.operands())
+      NewOps.push_back(Remap(Op));
+
+    // Reverse the mask for negative-stride memory ops.
+    auto ReverseIfNeeded = [&](VPValue *Mask, VPValue *Addr,
+                               DebugLoc DL) -> VPValue * {
+      if (isa_and_nonnull<VPVectorEndPointerRecipe>(Addr->getDefiningRecipe()))
+        return Builder.createNaryOp(VPInstruction::Reverse, {Mask}, DL,
+                                    "masked.replay.mask.rev");
+      return Mask;
+    };
+
+    VPRecipeBase *Clone = nullptr;
+    if (auto *WStore = dyn_cast<VPWidenStoreRecipe>(&R)) {
+      VPValue *G = WStore->isMasked() ? Remap(WStore->getMask()) : nullptr;
+      VPValue *L =
+          ReverseIfNeeded(getMaskFor(R), NewOps[0], WStore->getDebugLoc());
+      VPValue *M = G ? Builder.createLogicalAnd(G, L) : L;
+      Clone = new VPWidenStoreRecipe(
+          cast<StoreInst>(WStore->getIngredient()), NewOps[0], NewOps[1], M,
+          WStore->isConsecutive(), *WStore, WStore->getDebugLoc());
+      Builder.insert(Clone);
+    } else if (auto *RepR = dyn_cast<VPReplicateRecipe>(&R);
+               RepR && RepR->getUnderlyingInstr()->mayWriteToMemory()) {
+      VPValue *G = RepR->isPredicated() ? Remap(RepR->getMask()) : nullptr;
+      VPValue *L = getMaskFor(R);
+      VPValue *M = G ? Builder.createLogicalAnd(G, L) : L;
+
+      auto *SI = cast<StoreInst>(RepR->getUnderlyingInstr());
+      VPValue *StoredVal = RepR->getOperand(0);
+      VPValue *PtrVal = RepR->getOperand(1);
+
+      // Widen to a masked vector store when the stored value is a vector.
+      VPValue *ReplayStoredVal = Remap(StoredVal);
+      VPValue *ReplayPtrVal = Remap(PtrVal);
+
+      bool CanWiden = ReplayStoredVal->getDefiningRecipe() &&
+                      !isa<VPReplicateRecipe, VPScalarIVStepsRecipe>(
+                          ReplayStoredVal->getDefiningRecipe());
+
+      if (CanWiden) {
+        Type *StoreTy = SI->getValueOperand()->getType();
+        const DataLayout &DL = SI->getDataLayout();
+        auto *StrideTy =
+            DL.getIndexType(SI->getPointerOperand()->getType());
+        VPValue *StrideOne = Plan.getConstantInt(StrideTy, 1);
+        auto *VecPtr = new VPVectorPointerRecipe(
+            ReplayPtrVal, StoreTy, StrideOne, GEPNoWrapFlags::none(),
+            SI->getDebugLoc());
+        Builder.insert(VecPtr);
+
+        auto *WS = new VPWidenStoreRecipe(*SI, VecPtr, ReplayStoredVal, M,
+                                         /*Consecutive=*/true, *RepR,
+                                         RepR->getDebugLoc());
+        Builder.insert(WS);
+        continue;
+      } else {
+        // Cannot widen: recreate as a predicated replicate.
+        SmallVector<VPValue *, 4> ValOps;
+        for (VPValue *Op : RepR->operandsWithoutMask())
+          ValOps.push_back(Remap(Op));
+        auto *PredStore = new VPReplicateRecipe(
+            RepR->getUnderlyingInstr(), ValOps, RepR->isSingleScalar(), M,
+            *RepR, *RepR, RepR->getDebugLoc());
+        Builder.insert(PredStore);
+        Clone = PredStore;
+      }
+    } else if (auto *WLoad = dyn_cast<VPWidenLoadRecipe>(&R)) {
+      // Mask replay load: inclusive/exclusive mask ANDed with remapped guard.
+      VPValue *G = WLoad->isMasked() ? Remap(WLoad->getMask()) : nullptr;
+      VPValue *ReplayMask;
+      if (InclusiveFeeds.contains(&R)) {
+        if (!InclMask) {
+          VPInstruction *FALPlusOne =
+              VPBuilder::getToInsertAfter(FirstActiveLane)
+                  .createAdd(FirstActiveLane, One, DebugLoc(),
+                             "first.active.lane.incl",
+                             {/*nuw=*/true, /*nsw=*/false});
+          InclMask = VPBuilder::getToInsertAfter(FALPlusOne)
+                         .createNaryOp(VPInstruction::ActiveLaneMask,
+                                       {Zero, FALPlusOne, One}, DebugLoc(),
+                                       "masked.replay.mask.incl");
+        }
+        ReplayMask = InclMask;
+      } else {
+        if (!ExclMask)
+          ExclMask = VPBuilder::getToInsertAfter(FirstActiveLane)
+                         .createNaryOp(VPInstruction::ActiveLaneMask,
+                                       {Zero, FirstActiveLane, One}, DebugLoc(),
+                                       "masked.replay.mask");
+        ReplayMask = ExclMask;
+      }
+      ReplayMask = ReverseIfNeeded(ReplayMask, NewOps[0], WLoad->getDebugLoc());
+      VPValue *Mask = G ? Builder.createLogicalAnd(G, ReplayMask) : ReplayMask;
+      Clone = new VPWidenLoadRecipe(
+          cast<LoadInst>(WLoad->getIngredient()), NewOps[0], Mask,
+          WLoad->isConsecutive(), *WLoad, WLoad->getDebugLoc());
+      Builder.insert(Clone);
+    } else if (auto *RepR = dyn_cast<VPReplicateRecipe>(&R);
+               RepR && RepR->isPredicated()) {
+      SmallVector<VPValue *, 4> ValOps;
+      for (VPValue *Op : RepR->operandsWithoutMask())
+        ValOps.push_back(Remap(Op));
+      auto *NewRep = new VPReplicateRecipe(
+          RepR->getUnderlyingInstr(), ValOps, RepR->isSingleScalar(),
+          Remap(RepR->getMask()), *RepR, *RepR, RepR->getDebugLoc());
+      Builder.insert(NewRep);
+      Clone = NewRep;
+    } else {
+      assert(!R.mayWriteToMemory() && !R.mayHaveSideEffects() &&
+             "masked replay reached an unmaskable side-effecting recipe; "
+             "such loops must fall back to scalar replay");
+      Clone = R.clone();
+      for (unsigned I = 0, E = Clone->getNumOperands(); I != E; ++I)
+        Clone->setOperand(I, NewOps[I]);
+      Builder.insert(Clone);
+    }
+
+    for (unsigned I = 0, E = R.getNumDefinedValues(); I != E; ++I)
+      OperandMap[R.getVPValue(I)] = Clone->getVPValue(I);
+  }
+
+  assert(ReplayBB->getNumSuccessors() == 1 &&
+         "masked-replay block must have the single temporary latch edge");
+  Plan.setCheckFirstMaskedReplayTempTarget(ReplayBB->getSuccessors()[0]);
+}
+
+void VPlanTransforms::wireCheckFirstMaskedReplayToExit(VPlan &Plan) {
+  VPBasicBlock *ReplayBB = Plan.getCheckFirstMaskedReplayBlock();
+  if (!ReplayBB)
+    return;
+
+  BasicBlock *EarlyExitIRBB = Plan.getCheckFirstEarlyExitBlock();
+  assert(EarlyExitIRBB && "masked replay requires a captured early-exit block");
+  // Recreate only if it was removed during cloning.
+  VPIRBasicBlock *EarlyExitBB = nullptr;
+  for (VPIRBasicBlock *EB : Plan.getExitBlocks())
+    if (EB->getIRBasicBlock() == EarlyExitIRBB) {
+      EarlyExitBB = EB;
+      break;
+    }
+  if (!EarlyExitBB)
+    EarlyExitBB = Plan.createVPIRBasicBlock(EarlyExitIRBB);
+
+  VPInstruction *FirstActiveLane = nullptr;
+  for (VPRecipeBase &R : *ReplayBB) {
+    if (auto *VPI = dyn_cast<VPInstruction>(&R);
+        VPI && VPI->getOpcode() == VPInstruction::FirstActiveLane) {
+      FirstActiveLane = VPI;
+      break;
+    }
+  }
+  assert(FirstActiveLane && "masked-replay block missing FirstActiveLane");
+
+  assert(ReplayBB->getNumPredecessors() == 1 &&
+         "masked-replay head must have a single (check/header) predecessor");
+  auto *HeaderVPBB = cast<VPBasicBlock>(ReplayBB->getPredecessors()[0]);
+  VPValue *CanonIV = cast<VPPhi>(&*HeaderVPBB->begin());
+  Type *CanonTy = CanonIV->getScalarType();
+
+  // Recover the tail of the chain it is the recorded temp-target's predecessor
+  //  that is not the loop header.
+  VPBlockBase *TempTarget = Plan.getCheckFirstMaskedReplayTempTarget();
+  assert(TempTarget && "masked replay temporary-edge target not recorded");
+  VPBasicBlock *TailVPBB = nullptr;
+  for (VPBlockBase *Pred : TempTarget->getPredecessors()) {
+    if (Pred != HeaderVPBB) {
+      assert(!TailVPBB && "expected a single replay-chain predecessor of the "
+                          "temporary-edge target");
+      TailVPBB = cast<VPBasicBlock>(Pred);
+    }
+  }
+  assert(TailVPBB && "could not locate the masked-replay chain tail");
+
+  // Replace the temporary tail → latch edge with tail → early-exit.
+  VPBlockUtils::disconnectBlocks(TailVPBB, TempTarget);
+  VPBlockUtils::connectBlocks(TailVPBB, EarlyExitBB);
+
+  // Exiting index = chunk_start + first_active_lane.
+  // Build in the tail so it dominates its uses.
+  VPBuilder Builder(TailVPBB, TailVPBB->end());
+  VPValue *FALCast = Builder.createScalarZExtOrTrunc(
+      FirstActiveLane, CanonTy, FirstActiveLane->getScalarType(), DebugLoc());
+  VPValue *ExitIndex =
+      Builder.createAdd(CanonIV, FALCast, DebugLoc(), "masked.replay.exit.idx",
+                        {/*nuw=*/true, /*nsw=*/false});
+
+  // Each live-out is the exiting induction value, cast to the PHI's type.
+  for (VPRecipeBase &R : EarlyExitBB->phis()) {
+    auto *Phi = cast<VPIRPhi>(&R);
+    Type *PhiTy = Phi->getIRPhi().getType();
+    VPValue *LiveOut =
+        Builder.createScalarZExtOrTrunc(ExitIndex, PhiTy, CanonTy, DebugLoc());
+    Phi->addIncoming(LiveOut);
+  }
+}
+
 /// This function tries convert extended in-loop reductions to
 /// VPExpressionRecipe and clamp the \p Range if it is beneficial and
 /// valid. The created recipe must be decomposed to its constituent
@@ -5336,6 +6260,9 @@ void VPlanTransforms::materializeConstantVectorTripCount(
       !isa<VPIRValue>(TC))
     return;
 
+  if (Plan.getCheckFirstExitBlock())
+    return;
+
   // Materialize vector trip counts for constants early if it can simply
   // be computed as (Original TC / VF * UF) * VF * UF.
   // TODO: Compute vector trip counts for loops requiring a scalar epilogue and
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
index 3260526552281..513267c2bb41e 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
@@ -366,6 +366,16 @@ struct VPlanTransforms {
       VPBasicBlock *MiddleVPBB, Loop *TheLoop, PredicatedScalarEvolution &PSE,
       DominatorTree &DT, AssumptionCache *AC, UncountableExitStyle Style);
 
+  /// Connects check-first early-exit blocks to the scalar preheader.
+  static void wireCheckFirstExitToScalar(VPlan &Plan);
+
+  /// Clones the stores inside loop body into the masked-replay block.
+  static void maskCheckFirstReplayStores(VPlan &Plan);
+
+  /// Wires the masked replay block to the early-exit block and rebuilds its
+  /// live-out.
+  static void wireCheckFirstMaskedReplayToExit(VPlan &Plan);
+
   /// Replaces the exit condition from
   ///   (branch-on-cond eq CanonicalIVInc, VectorTripCount)
   /// to
diff --git a/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp b/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
index 362bfe92f573e..2428c718bcaf4 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
@@ -308,6 +308,12 @@ bool VPlanVerifier::verifyVPBasicBlock(const VPBasicBlock *VPBB) {
           continue;
         }
 
+        if (VPBasicBlock *CheckExit =
+                VPBB->getPlan()->getCheckFirstExitBlock()) {
+          if (is_contained(CheckExit->getPredecessors(), VPBB))
+            continue;
+        }
+
         errs() << "Use before def!\n";
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
         VPSlotTracker Tracker(VPBB->getPlan());
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
index 2a250d4c896ff..d91204aba896f 100644
--- a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
@@ -55,6 +55,7 @@
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] removeBranchOnConst
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] simplifyReverses
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] removeDeadRecipes
+; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] maskCheckFirstReplayStores
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] createAndOptimizeReplicateRegions
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] mergeBlocksIntoPredecessors
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] licm
diff --git a/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll b/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
new file mode 100644
index 0000000000000..1760b673f8abf
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
@@ -0,0 +1,133 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -S < %s -p loop-vectorize -force-vector-width=4 -enable-check-first-early-exit-vectorization | FileCheck %s
+
+ at A = global [1024 x i32] zeroinitializer
+ at B = global [1024 x i32] zeroinitializer
+ at D = global [1024 x i32] zeroinitializer
+
+define i32 @multi_exit_cascade() {
+; CHECK-LABEL: define i32 @multi_exit_cascade() {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY4:.*]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT:    [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT:    [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT:    [[TMP6:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-NEXT:    [[TMP7:%.*]] = freeze <4 x i1> [[TMP6]]
+; CHECK-NEXT:    [[TMP8:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP7]])
+; CHECK-NEXT:    br i1 [[TMP8]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_CHECK:.*]]
+; CHECK:       [[VECTOR_CHECK]]:
+; CHECK-NEXT:    [[TMP21:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD4:%.*]] = load <4 x i32>, ptr [[TMP21]], align 4
+; CHECK-NEXT:    [[TMP22:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD4]], zeroinitializer
+; CHECK-NEXT:    [[TMP9:%.*]] = freeze <4 x i1> [[TMP22]]
+; CHECK-NEXT:    [[TMP10:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP9]])
+; CHECK-NEXT:    br i1 [[TMP10]], label %[[VECTOR_CHECK_EXIT]], label %[[VECTOR_BODY4]]
+; CHECK:       [[VECTOR_BODY4]]:
+; CHECK-NEXT:    [[TMP11:%.*]] = add <4 x i32> [[WIDE_LOAD3]], [[WIDE_LOAD4]]
+; CHECK-NEXT:    [[TMP12:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT:    [[TMP14:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT:    [[TMP15:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT:    [[TMP16:%.*]] = extractelement <4 x i32> [[TMP11]], i64 0
+; CHECK-NEXT:    store i32 [[TMP16]], ptr [[TMP12]], align 4
+; CHECK-NEXT:    [[TMP17:%.*]] = extractelement <4 x i32> [[TMP11]], i64 1
+; CHECK-NEXT:    store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT:    [[TMP18:%.*]] = extractelement <4 x i32> [[TMP11]], i64 2
+; CHECK-NEXT:    store i32 [[TMP18]], ptr [[TMP14]], align 4
+; CHECK-NEXT:    [[TMP19:%.*]] = extractelement <4 x i32> [[TMP11]], i64 3
+; CHECK-NEXT:    store i32 [[TMP19]], ptr [[TMP15]], align 4
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP20:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT:    br i1 [[TMP20]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[RET_N:.*]]
+; CHECK:       [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT:    br label %[[SCALAR_PH:.*]]
+; CHECK:       [[SCALAR_PH]]:
+; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
+; CHECK:       [[FOR_BODY]]:
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ [[INDEX]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
+; CHECK-NEXT:    [[GA:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[IV]]
+; CHECK-NEXT:    [[LA:%.*]] = load i32, ptr [[GA]], align 4
+; CHECK-NEXT:    [[CA:%.*]] = icmp eq i32 [[LA]], 0
+; CHECK-NEXT:    br i1 [[CA]], label %[[EXIT0:.*]], label %[[CHECK1:.*]]
+; CHECK:       [[CHECK1]]:
+; CHECK-NEXT:    [[GB:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[IV]]
+; CHECK-NEXT:    [[LB:%.*]] = load i32, ptr [[GB]], align 4
+; CHECK-NEXT:    [[CB:%.*]] = icmp eq i32 [[LB]], 0
+; CHECK-NEXT:    br i1 [[CB]], label %[[EXIT1:.*]], label %[[LATCH]]
+; CHECK:       [[LATCH]]:
+; CHECK-NEXT:    [[SUM:%.*]] = add i32 [[LA]], [[LB]]
+; CHECK-NEXT:    [[GD:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[IV]]
+; CHECK-NEXT:    store i32 [[SUM]], ptr [[GD]], align 4
+; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT:    [[DONE:%.*]] = icmp eq i64 [[IV_NEXT]], 1024
+; CHECK-NEXT:    br i1 [[DONE]], label %[[RET_N]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK:       [[EXIT0]]:
+; CHECK-NEXT:    [[IV_LCSSA:%.*]] = phi i64 [ [[IV]], %[[FOR_BODY]] ]
+; CHECK-NEXT:    [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT:    br label %[[RET:.*]]
+; CHECK:       [[EXIT1]]:
+; CHECK-NEXT:    [[IV_LCSSA1:%.*]] = phi i64 [ [[IV]], %[[CHECK1]] ]
+; CHECK-NEXT:    [[T1:%.*]] = trunc i64 [[IV_LCSSA1]] to i32
+; CHECK-NEXT:    [[NEG:%.*]] = sub i32 0, [[T1]]
+; CHECK-NEXT:    br label %[[RET]]
+; CHECK:       [[RET_N]]:
+; CHECK-NEXT:    br label %[[RET]]
+; CHECK:       [[RET]]:
+; CHECK-NEXT:    [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT0]] ], [ [[NEG]], %[[EXIT1]] ], [ 1024, %[[RET_N]] ]
+; CHECK-NEXT:    ret i32 [[R]]
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+  %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+  %la = load i32, ptr %ga, align 4
+  %ca = icmp eq i32 %la, 0
+  br i1 %ca, label %exit0, label %check1
+
+check1:
+  %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+  %lb = load i32, ptr %gb, align 4
+  %cb = icmp eq i32 %lb, 0
+  br i1 %cb, label %exit1, label %latch
+
+latch:
+  %sum = add i32 %la, %lb
+  %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+  store i32 %sum, ptr %gd, align 4
+  %iv.next = add nuw nsw i64 %iv, 1
+  %done = icmp eq i64 %iv.next, 1024
+  br i1 %done, label %ret.n, label %for.body
+
+exit0:
+  %t0 = trunc i64 %iv to i32
+  br label %ret
+
+exit1:
+  %t1 = trunc i64 %iv to i32
+  %neg = sub i32 0, %t1
+  br label %ret
+
+ret.n:
+  br label %ret
+
+ret:
+  %r = phi i32 [ %t0, %exit0 ], [ %neg, %exit1 ], [ 1024, %ret.n ]
+  ret i32 %r
+}
+;.
+; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
+; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
+; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
+; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META2]], [[META1]]}
+;.
diff --git a/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll b/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll
new file mode 100644
index 0000000000000..428d9cc953425
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll
@@ -0,0 +1,205 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -S < %s -p loop-vectorize -force-vector-width=4 -enable-check-first-early-exit-vectorization -enable-check-first-masked-replay | FileCheck %s
+
+ at A = global [1024 x i32] zeroinitializer
+ at B = global [1024 x i32] zeroinitializer
+ at C = global [1024 x i32] zeroinitializer
+ at D = global [1024 x i32] zeroinitializer
+
+; for (i) { if (A[i] != 0) { if (B[i] == 0) return i; } D[i] = A[i]; }
+define i32 @single_guard() {
+; CHECK-LABEL: define i32 @single_guard() {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY3:.*]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT:    [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT:    [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT:    [[TMP4:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT:    [[TMP5:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP5]], align 4
+; CHECK-NEXT:    [[TMP6:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD2]], zeroinitializer
+; CHECK-NEXT:    [[TMP7:%.*]] = select <4 x i1> [[TMP4]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
+; CHECK-NEXT:    [[TMP8:%.*]] = freeze <4 x i1> [[TMP7]]
+; CHECK-NEXT:    [[TMP9:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP8]])
+; CHECK-NEXT:    br i1 [[TMP9]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_BODY3]]
+; CHECK:       [[VECTOR_BODY3]]:
+; CHECK-NEXT:    [[TMP10:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[TMP11:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT:    [[TMP12:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT:    [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT:    [[TMP14:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 0
+; CHECK-NEXT:    store i32 [[TMP14]], ptr [[TMP10]], align 4
+; CHECK-NEXT:    [[TMP15:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 1
+; CHECK-NEXT:    store i32 [[TMP15]], ptr [[TMP11]], align 4
+; CHECK-NEXT:    [[TMP16:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 2
+; CHECK-NEXT:    store i32 [[TMP16]], ptr [[TMP12]], align 4
+; CHECK-NEXT:    [[TMP17:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 3
+; CHECK-NEXT:    store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP18:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT:    br i1 [[TMP18]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[RET_N:.*]]
+; CHECK:       [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT:    [[FIRST_ACTIVE_LANE:%.*]] = call i64 @llvm.experimental.cttz.elts.i64.v4i1(<4 x i1> [[TMP7]], i1 false)
+; CHECK-NEXT:    [[MASKED_REPLAY_MASK:%.*]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i64(i64 0, i64 [[FIRST_ACTIVE_LANE]])
+; CHECK-NEXT:    [[TMP20:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    call void @llvm.masked.store.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP20]], <4 x i1> [[MASKED_REPLAY_MASK]])
+; CHECK-NEXT:    [[IV_LCSSA:%.*]] = add nuw i64 [[INDEX]], [[FIRST_ACTIVE_LANE]]
+; CHECK-NEXT:    br label %[[EXIT1:.*]]
+; CHECK:       [[EXIT1]]:
+; CHECK-NEXT:    [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT:    br label %[[RET:.*]]
+; CHECK:       [[RET_N]]:
+; CHECK-NEXT:    br label %[[RET]]
+; CHECK:       [[RET]]:
+; CHECK-NEXT:    [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT1]] ], [ 1024, %[[RET_N]] ]
+; CHECK-NEXT:    ret i32 [[R]]
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+  %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+  %la = load i32, ptr %ga, align 4
+  %xg = icmp ne i32 %la, 0
+  br i1 %xg, label %guard, label %latch
+
+guard:
+  %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+  %lb = load i32, ptr %gb, align 4
+  %yb = icmp eq i32 %lb, 0
+  br i1 %yb, label %exit0, label %latch
+
+latch:
+  %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+  store i32 %la, ptr %gd, align 4
+  %iv.next = add nuw nsw i64 %iv, 1
+  %done = icmp eq i64 %iv.next, 1024
+  br i1 %done, label %ret.n, label %for.body
+
+exit0:
+  %t0 = trunc i64 %iv to i32
+  br label %ret
+
+ret.n:
+  br label %ret
+
+ret:
+  %r = phi i32 [ %t0, %exit0 ], [ 1024, %ret.n ]
+  ret i32 %r
+}
+
+; for (i) { if (A[i]!=0) { if (C[i]!=0) { if (B[i]==0) return i; } } D[i]=A[i]; }
+define i32 @two_nested_guards() {
+; CHECK-LABEL: define i32 @two_nested_guards() {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY4:.*]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT:    [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT:    [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT:    [[TMP4:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT:    [[TMP5:%.*]] = getelementptr inbounds [1024 x i32], ptr @C, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP5]], align 4
+; CHECK-NEXT:    [[TMP6:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD2]], zeroinitializer
+; CHECK-NEXT:    [[TMP7:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP7]], align 4
+; CHECK-NEXT:    [[TMP8:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-NEXT:    [[TMP9:%.*]] = select <4 x i1> [[TMP4]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
+; CHECK-NEXT:    [[TMP10:%.*]] = select <4 x i1> [[TMP9]], <4 x i1> [[TMP8]], <4 x i1> zeroinitializer
+; CHECK-NEXT:    [[TMP11:%.*]] = freeze <4 x i1> [[TMP10]]
+; CHECK-NEXT:    [[TMP12:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP11]])
+; CHECK-NEXT:    br i1 [[TMP12]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_BODY4]]
+; CHECK:       [[VECTOR_BODY4]]:
+; CHECK-NEXT:    [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    [[TMP14:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT:    [[TMP15:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT:    [[TMP16:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT:    [[TMP17:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 0
+; CHECK-NEXT:    store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT:    [[TMP18:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 1
+; CHECK-NEXT:    store i32 [[TMP18]], ptr [[TMP14]], align 4
+; CHECK-NEXT:    [[TMP19:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 2
+; CHECK-NEXT:    store i32 [[TMP19]], ptr [[TMP15]], align 4
+; CHECK-NEXT:    [[TMP20:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 3
+; CHECK-NEXT:    store i32 [[TMP20]], ptr [[TMP16]], align 4
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP21:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT:    br i1 [[TMP21]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[SCALAR_PH:.*]]
+; CHECK:       [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT:    [[FIRST_ACTIVE_LANE:%.*]] = call i64 @llvm.experimental.cttz.elts.i64.v4i1(<4 x i1> [[TMP10]], i1 false)
+; CHECK-NEXT:    [[MASKED_REPLAY_MASK:%.*]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i64(i64 0, i64 [[FIRST_ACTIVE_LANE]])
+; CHECK-NEXT:    [[TMP23:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT:    call void @llvm.masked.store.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP23]], <4 x i1> [[MASKED_REPLAY_MASK]])
+; CHECK-NEXT:    [[IV_LCSSA:%.*]] = add nuw i64 [[INDEX]], [[FIRST_ACTIVE_LANE]]
+; CHECK-NEXT:    br label %[[EXIT1:.*]]
+; CHECK:       [[EXIT1]]:
+; CHECK-NEXT:    [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT:    br label %[[RET:.*]]
+; CHECK:       [[SCALAR_PH]]:
+; CHECK-NEXT:    br label %[[RET]]
+; CHECK:       [[RET]]:
+; CHECK-NEXT:    [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT1]] ], [ 1024, %[[SCALAR_PH]] ]
+; CHECK-NEXT:    ret i32 [[R]]
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+  %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+  %la = load i32, ptr %ga, align 4
+  %xg = icmp ne i32 %la, 0
+  br i1 %xg, label %g1, label %latch
+
+g1:
+  %gc = getelementptr inbounds [1024 x i32], ptr @C, i64 0, i64 %iv
+  %lc = load i32, ptr %gc, align 4
+  %zg = icmp ne i32 %lc, 0
+  br i1 %zg, label %g2, label %latch
+
+g2:
+  %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+  %lb = load i32, ptr %gb, align 4
+  %yb = icmp eq i32 %lb, 0
+  br i1 %yb, label %exit0, label %latch
+
+latch:
+  %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+  store i32 %la, ptr %gd, align 4
+  %iv.next = add nuw nsw i64 %iv, 1
+  %done = icmp eq i64 %iv.next, 1024
+  br i1 %done, label %ret.n, label %for.body
+
+exit0:
+  %t0 = trunc i64 %iv to i32
+  br label %ret
+
+ret.n:
+  br label %ret
+
+ret:
+  %r = phi i32 [ %t0, %exit0 ], [ 1024, %ret.n ]
+  ret i32 %r
+}
+;.
+; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
+; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
+; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
+; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META1]], [[META2]]}
+;.



More information about the llvm-commits mailing list