[llvm] [LoopVectorize] Introduce check-first vectorization for multiple early-exit loops (PR #210492)
Arjun H Kumar via llvm-commits
llvm-commits at lists.llvm.org
Fri Jul 17 23:58:12 PDT 2026
https://github.com/arjun-harikumar-amd created https://github.com/llvm/llvm-project/pull/210492
### Check-First Early-Exit Loop Vectorization
We introduce check-first strategy for vectorizing loops with multiple early exits. Instead of masking the loop body, the check-first approach evaluates all exit conditions first for a vector chunk and only runs the loop body if no exit fires. This generalizes early-exit vectorization to a much broader set of loops.
Implemented as a VPlan transformation in the loop vectorizer. It restructures the vector loop into three regions:
1. **Check blocks**: One check block per early exit, each holding the minimal condition slice to evaluate that exit. If any exit fires, control transfers to the exit-handling path.
2. **Vector body**: Runs only when no exit fires.
3. **Exit handling**: For the exiting chunk, lanes before the exit still need to execute. We have two strategies:
- **Scalar replay** (default): fall back to the scalar loop, replaying from the start of the current chunk.
- **Masked replay** (experimental): clone the body's stores with lane masks derived from the first active exit lane, avoiding the scalar loop.
### Flags
```
-mllvm -enable-check-first-early-exit-vectorization
-mllvm -enable-check-first-masked-replay
```
RFC to be posted soon.
Co-authored by @nema-ashutosh
>From 899f88855c67a2ceb253459e3264b7c928ba0515 Mon Sep 17 00:00:00 2001
From: Arjun H Kumar <Arjun.HKumar at amd.com>
Date: Tue, 30 Jun 2026 23:24:28 +0530
Subject: [PATCH] [LoopVectorize] Introduce check-first vectorization for
multiple early-exit loops
This patch introduces a new strategy 'Check-first' where we try to generalize
vectorizing loops with 'n' early exits.
The core idea is to check the exiting conditions first and execute the loop body later.
---
.../Vectorize/LoopVectorizationLegality.h | 22 +
.../Vectorize/LoopVectorizationLegality.cpp | 300 +++++-
.../Transforms/Vectorize/LoopVectorize.cpp | 111 ++-
llvm/lib/Transforms/Vectorize/VPlan.cpp | 28 +
llvm/lib/Transforms/Vectorize/VPlan.h | 91 +-
.../Vectorize/VPlanConstruction.cpp | 19 +-
.../Transforms/Vectorize/VPlanPredicator.cpp | 2 +
.../lib/Transforms/Vectorize/VPlanRecipes.cpp | 13 +-
.../Transforms/Vectorize/VPlanTransforms.cpp | 935 +++++++++++++++++-
.../Transforms/Vectorize/VPlanTransforms.h | 10 +
.../Transforms/Vectorize/VPlanVerifier.cpp | 6 +
.../VPlan/vplan-print-before-after-all.ll | 1 +
.../check-first-multi-exit-cascade.ll | 133 +++
.../LoopVectorize/check-first-nested-exit.ll | 205 ++++
14 files changed, 1840 insertions(+), 36 deletions(-)
create mode 100644 llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll
diff --git a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
index 3e8db73fd79d2..de7eb4466ef6b 100644
--- a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
+++ b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
@@ -440,6 +440,20 @@ class LoopVectorizationLegality {
return getUncountableExitTrait() == UncountableExitTrait::ReadWrite;
}
+ /// Returns true if every widened exit condition load is
+ /// dereferenceable for the complete trip count.
+ bool exitLoadsAreDereferenceable() const {
+ return AllExitLoadsDereferenceable;
+ }
+
+ /// Returns true if this early exit loop would use the check first strategy if
+ /// enabled.
+ bool wouldUseCheckFirstStyle() const {
+ if (hasUncountableExitWithSideEffects())
+ return true;
+ return hasUncountableEarlyExit() && !AllExitLoadsDereferenceable;
+ }
+
/// Return true if there is store-load forwarding dependencies.
bool isSafeForAnyStoreLoadForwardDistances() const {
return LAI->getDepChecker().isSafeForAnyStoreLoadForwardDistances();
@@ -635,6 +649,10 @@ class LoopVectorizationLegality {
/// for it.
bool canUncountableExitConditionLoadBeMoved(BasicBlock *ExitingBlock);
+ /// Returns true if the exit conditions can be safely speculated.
+ bool canCheckFirstSpeculateExitConditions(
+ ArrayRef<BasicBlock *> ExitingBlocks);
+
/// Return true if all of the instructions in the block can be speculatively
/// executed, and record the loads/stores that require masking.
/// \p SafePtrs is a list of addresses that are known to be legal and we know
@@ -752,6 +770,10 @@ class LoopVectorizationLegality {
/// Records whether we have an uncountable early exit in a loop that's
/// either read-only or read-write.
UncountableExitTrait UncountableExitType = UncountableExitTrait::None;
+
+ /// Records whether every widened exit condition load is
+ /// dereferenceable for the complete trip count.
+ bool AllExitLoadsDereferenceable = true;
};
} // namespace llvm
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
index 8d875b2b6e492..c94de8d9951b3 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
@@ -79,6 +79,14 @@ static cl::opt<bool> EnableHistogramVectorization(
"enable-histogram-loop-vectorization", cl::init(false), cl::Hidden,
cl::desc("Enables autovectorization of some loops containing histograms"));
+static cl::opt<unsigned> MaxUncountableEarlyExits(
+ "max-uncountable-early-exits", cl::init(4), cl::Hidden,
+ cl::desc("Maximum number of uncountable early exits a loop may have to be "
+ "eligible for multi-exit check-first vectorization."));
+
+extern cl::opt<bool> EnableCheckFirstVectorization;
+extern cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects;
+
/// Maximum vectorization interleave count.
static const unsigned MaxInterleaveFactor = 16;
@@ -1616,6 +1624,98 @@ bool LoopVectorizationLegality::canVectorizeLoopNestCFG(
return Result;
}
+static unsigned countInLoopPredecessors(const BasicBlock *BB, const Loop *L) {
+ unsigned Count = 0;
+ for (const BasicBlock *Pred : predecessors(BB))
+ if (L->contains(Pred))
+ ++Count;
+ return Count;
+}
+
+/// Walks up the dominator chain from \p Exiting to the loop header, collecting
+/// into \p GuardConds the branch conditions that must hold for control to reach
+/// \p ExitingBB. Returns false if the control flow has a shape that cannot be
+/// represented as such a guard.
+///
+/// At each conditional dominator, one successor leads toward \p ExitingBB and the
+/// other bypasses it.
+/// Records a guard when the bypass leaves the guarded region cleanly:
+/// it targets \p Latch, or rejoins the chain after an if-without-else.
+static bool collectExitGuards(BasicBlock *Exiting, Loop *L, BasicBlock *Latch,
+ const DominatorTree &DT,
+ SmallVectorImpl<Value *> &GuardConds,
+ bool AllowRejoin) {
+ BasicBlock *Header = L->getHeader();
+ BasicBlock *Cur = Exiting;
+ while (Cur != Header) {
+ DomTreeNode *Node = DT.getNode(Cur);
+ if (!Node || !Node->getIDom())
+ return false;
+ BasicBlock *IDom = Node->getIDom()->getBlock();
+ if (!L->contains(IDom))
+ return false;
+ if (auto *Br = dyn_cast<CondBrInst>(IDom->getTerminator())) {
+ BasicBlock *S0 = Br->getSuccessor(0), *S1 = Br->getSuccessor(1);
+ bool S0Dom = DT.dominates(S0, Cur), S1Dom = DT.dominates(S1, Cur);
+ if (S0Dom != S1Dom) {
+ BasicBlock *Interior = S0Dom ? S0 : S1;
+ BasicBlock *Bypass = S0Dom ? S1 : S0;
+ if (L->contains(Bypass)) {
+ if (Bypass == Latch) {
+ GuardConds.push_back(Br->getCondition());
+ } else if (!AllowRejoin) {
+ return false;
+ } else if (countInLoopPredecessors(Interior, L) == 1 &&
+ countInLoopPredecessors(Bypass, L) > 1) {
+ GuardConds.push_back(Br->getCondition());
+ } else if (countInLoopPredecessors(Interior, L) > 1) {
+ } else {
+ return false;
+ }
+ }
+ }
+ }
+ Cur = IDom;
+ }
+ return true;
+}
+
+/// Collects the loads feeding the exit conditions of early-exits.
+/// condition, which check-first widens speculatively.
+static void collectExitConditionSliceLoads(ArrayRef<BasicBlock *> ExitingBlocks,
+ Loop *L, BasicBlock *Latch,
+ const DominatorTree &DT,
+ SmallVectorImpl<LoadInst *> &Out) {
+ SmallPtrSet<Value *, 16> Visited;
+ SmallVector<Value *, 16> Worklist;
+ for (BasicBlock *BB : ExitingBlocks) {
+ auto *Br = dyn_cast<CondBrInst>(BB->getTerminator());
+ assert(Br && "exiting block must terminate with a conditional branch");
+ Worklist.push_back(Br->getCondition());
+ collectExitGuards(BB, L, Latch, DT, Worklist, /*AllowRejoin=*/true);
+ }
+
+ for (BasicBlock *BB : L->blocks())
+ for (Instruction &I : *BB)
+ if (isa<StoreInst>(&I) && !DT.dominates(BB, Latch))
+ collectExitGuards(BB, L, Latch, DT, Worklist, /*AllowRejoin=*/false);
+
+ while (!Worklist.empty()) {
+ Value *V = Worklist.pop_back_val();
+ if (!Visited.insert(V).second)
+ continue;
+ auto *I = dyn_cast<Instruction>(V);
+ if (!I || !L->contains(I) || isa<PHINode>(I))
+ continue;
+ if (auto *LI = dyn_cast<LoadInst>(I)) {
+ Out.push_back(LI);
+ continue;
+ }
+ for (Value *Op : I->operands())
+ Worklist.push_back(Op);
+ }
+}
+
bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
BasicBlock *LatchBB = TheLoop->getLoopLatch();
if (!LatchBB) {
@@ -1733,10 +1833,15 @@ bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
return false;
}
} else {
- // Check all uncountable exiting blocks for movable loads.
- for (BasicBlock *ExitingBB : UncountableExitingBlocks) {
- if (!canUncountableExitConditionLoadBeMoved(ExitingBB))
+ if (EnableCheckFirstVectorization &&
+ !EnableEarlyExitVectorizationWithSideEffects) {
+ if (!canCheckFirstSpeculateExitConditions(UncountableExitingBlocks))
return false;
+ } else {
+ for (BasicBlock *ExitingBB : UncountableExitingBlocks) {
+ if (!canUncountableExitConditionLoadBeMoved(ExitingBB))
+ return false;
+ }
}
}
@@ -1754,6 +1859,57 @@ bool LoopVectorizationLegality::isVectorizableEarlyExitLoop() {
}
}
+ // Safe only if every widened condition slice load is dereferenceable.
+ if (HasSideEffects) {
+ SmallVector<LoadInst *, 4> SpeculatedCondLoads;
+ collectExitConditionSliceLoads(UncountableExitingBlocks, TheLoop,
+ TheLoop->getLoopLatch(), *DT,
+ SpeculatedCondLoads);
+
+ bool AllDeref = true;
+ for (LoadInst *LI : SpeculatedCondLoads) {
+ if (!isDereferenceableAndAlignedInLoop(LI, TheLoop, *PSE.getSE(), *DT,
+ AC)) {
+ AllDeref = false;
+ break;
+ }
+ }
+
+ AllExitLoadsDereferenceable = AllDeref;
+
+ LLVM_DEBUG({
+ dbgs() << "LV: check-first early-exit memory-safety strategy: ";
+ if (AllDeref)
+ dbgs() << "all speculated condition-slice loads provably "
+ "dereferenceable. \n";
+ else
+ dbgs() << "Condition-slice loads not provably dereferenceable. \n";
+ });
+ }
+
+ bool WillUseCheckFirst =
+ HasSideEffects && !EnableEarlyExitVectorizationWithSideEffects;
+ if (WillUseCheckFirst) {
+ const InductionDescriptor *IndDesc = nullptr;
+ if (Inductions.size() == 1) {
+ IndDesc = &Inductions.begin()->second;
+ } else if (PHINode *PrimaryIV = getPrimaryInduction()) {
+ auto It = Inductions.find(PrimaryIV);
+ if (It != Inductions.end())
+ IndDesc = &It->second;
+ }
+ if (!IndDesc ||
+ (IndDesc->getKind() != InductionDescriptor::IK_IntInduction &&
+ IndDesc->getKind() != InductionDescriptor::IK_PtrInduction) ||
+ !IndDesc->getConstIntStepValue()) {
+ reportVectorizationFailure(
+ "Check-first early-exit vectorization requires a single integer or "
+ "pointer induction with a constant step",
+ "UnsupportedCheckFirstInduction", ORE, TheLoop);
+ return false;
+ }
+ }
+
[[maybe_unused]] const SCEV *SymbolicMaxBTC =
PSE.getSymbolicMaxBackedgeTakenCount();
// Since we have an exact exit count for the latch and the early exit
@@ -1854,6 +2010,144 @@ bool LoopVectorizationLegality::canUncountableExitConditionLoadBeMoved(
return true;
}
+bool LoopVectorizationLegality::canCheckFirstSpeculateExitConditions(
+ ArrayRef<BasicBlock *> ExitingBlocks) {
+ if (ExitingBlocks.size() > MaxUncountableEarlyExits) {
+ reportVectorizationFailure(
+ "Too many uncountable early exits for check-first vectorization",
+ "TooManyEarlyExitsForCheckFirst", ORE, TheLoop);
+ return false;
+ }
+
+ BasicBlock *Latch = TheLoop->getLoopLatch();
+ if (!Latch) {
+ reportVectorizationFailure("Check-first early-exit loop has no latch",
+ "NoLatchCheckFirstExit", ORE, TheLoop);
+ return false;
+ }
+
+ SmallVector<Value *, 8> GuardConds;
+ for (BasicBlock *BB : ExitingBlocks) {
+ if (!collectExitGuards(BB, TheLoop, Latch, *DT, GuardConds,
+ /*AllowRejoin=*/true)) {
+ reportVectorizationFailure(
+ "Check-first early-exit vectorization does not support this guarded "
+ "(conditionally-executed) early-exit control-flow shape",
+ "UnsupportedGuardedCheckFirstExit", ORE, TheLoop);
+ return false;
+ }
+ }
+
+ for (BasicBlock *BB : TheLoop->blocks())
+ for (Instruction &I : *BB)
+ if (isa<StoreInst>(&I) && !DT->dominates(BB, Latch)) {
+ SmallVector<Value *, 4> StoreGuards;
+ if (!collectExitGuards(BB, TheLoop, Latch, *DT, StoreGuards,
+ /*AllowRejoin=*/false)) {
+ reportVectorizationFailure(
+ "Check-first early-exit vectorization does not support this "
+ "conditionally-executed (guarded) store control-flow shape",
+ "GuardedCheckFirstStore", ORE, TheLoop);
+ return false;
+ }
+ }
+
+ SmallPtrSet<LoadInst *, 8> CondLoads;
+ SmallVector<Value *, 16> Worklist;
+ SmallPtrSet<Value *, 16> Visited;
+ for (BasicBlock *BB : ExitingBlocks) {
+ auto *Br = dyn_cast<CondBrInst>(BB->getTerminator());
+ if (!Br) {
+ reportVectorizationFailure(
+ "Exiting block does not terminate with a conditional branch",
+ "UnsupportedCheckFirstExitTerminator", ORE, TheLoop);
+ return false;
+ }
+ Worklist.push_back(Br->getCondition());
+ }
+ append_range(Worklist, GuardConds);
+
+ while (!Worklist.empty()) {
+ Value *V = Worklist.pop_back_val();
+ if (!Visited.insert(V).second)
+ continue;
+ if (TheLoop->isLoopInvariant(V))
+ continue;
+ auto *I = dyn_cast<Instruction>(V);
+ if (!I || !TheLoop->contains(I)) {
+ reportVectorizationFailure(
+ "Early exit condition depends on a value that cannot be "
+ "speculatively evaluated for check-first vectorization",
+ "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+ return false;
+ }
+ if (auto *LI = dyn_cast<LoadInst>(I)) {
+ const auto *AR = dyn_cast<SCEVAddRecExpr>(
+ PSE.getSE()->getSCEV(LI->getPointerOperand()));
+ if (!LI->isSimple() || !AR || AR->getLoop() != TheLoop ||
+ !AR->isAffine()) {
+ reportVectorizationFailure(
+ "Early exit condition depends on a load that is not a simple "
+ "affine (unit-stride) access",
+ "CheckFirstExitLoadInvariantAddress", ORE, TheLoop);
+ return false;
+ }
+ CondLoads.insert(LI);
+ continue;
+ }
+ if (isa<PHINode>(I)) {
+ if (I->getParent() != TheLoop->getHeader()) {
+ reportVectorizationFailure(
+ "Early exit condition depends on a non-header PHI",
+ "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+ return false;
+ }
+ continue;
+ }
+ if (I->mayReadOrWriteMemory() || !isSafeToSpeculativelyExecute(I)) {
+ reportVectorizationFailure(
+ "Early exit condition contains an operation that cannot be "
+ "speculatively executed",
+ "UnsupportedCheckFirstExitCondition", ORE, TheLoop);
+ return false;
+ }
+ for (Value *Op : I->operands())
+ Worklist.push_back(Op);
+ }
+
+ SmallPtrSet<const Instruction *, 4> CondLoadSet(CondLoads.begin(),
+ CondLoads.end());
+ ConditionallyExecutedOps.clear();
+ for (auto *BB : TheLoop->blocks()) {
+ for (auto &I : *BB) {
+ if (CondLoadSet.contains(&I) || !I.mayReadOrWriteMemory())
+ continue;
+ ConditionallyExecutedOps.insert(&I);
+ if (isa<LoadInst>(&I))
+ continue;
+ auto *SI = dyn_cast<StoreInst>(&I);
+ if (!SI) {
+ reportVectorizationFailure(
+ "Unsupported memory operation in check-first early-exit loop",
+ "UnsupportedCheckFirstMemOp", ORE, TheLoop);
+ return false;
+ }
+ for (LoadInst *CL : CondLoads) {
+ if (AA->alias(CL->getPointerOperand(), SI->getPointerOperand()) !=
+ AliasResult::NoAlias) {
+ reportVectorizationFailure(
+ "Cannot determine whether an early-exit condition load aliases "
+ "a store (deferred stores must not be observed out of order)",
+ "CheckFirstExitLoadAliasesStore", ORE, TheLoop);
+ return false;
+ }
+ }
+ }
+ }
+
+ return true;
+}
+
bool LoopVectorizationLegality::canVectorize(bool UseVPlanNativePath) {
// Store the result and return it at the end instead of exiting early, in case
// allowExtraAnalysis is used to report multiple reasons for not vectorizing.
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index f1ca4061bfd9e..e7377ff7cea2b 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -411,12 +411,34 @@ static cl::opt<bool> EnableEarlyExitVectorization(
cl::desc(
"Enable vectorization of early exit loops with uncountable exits."));
-static cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects(
+cl::opt<bool> EnableEarlyExitVectorizationWithSideEffects(
"enable-early-exit-vectorization-with-side-effects", cl::init(false),
cl::Hidden,
cl::desc("Enable vectorization of early exit loops with uncountable exits "
"and side effects"));
+cl::opt<bool> EnableCheckFirstVectorization(
+ "enable-check-first-early-exit-vectorization", cl::init(false), cl::Hidden,
+ cl::desc("Enable check-first vectorization of early exit loops with "
+ "multiple exits."));
+
+cl::opt<bool> EnableCheckFirstMaskedReplay(
+ "enable-check-first-masked-replay", cl::init(false), cl::Hidden,
+ cl::desc("Replace scalar replay in check-first early exit vectorization "
+ "with a masked vector replay."));
+
+/// Returns true if loop uses check first with scalar replay of the
+/// failing chunk.
+static bool usesCheckFirstReplay(const LoopVectorizationLegality *Legal) {
+ if (!Legal->hasUncountableEarlyExit())
+ return false;
+ if (EnableCheckFirstMaskedReplay)
+ return false;
+ if (Legal->hasUncountableExitWithSideEffects())
+ return !EnableEarlyExitVectorizationWithSideEffects;
+ return EnableCheckFirstVectorization && Legal->wouldUseCheckFirstStyle();
+}
+
// Likelyhood of bypassing the vectorized loop because there are zero trips left
// after prolog. See `emitIterationCountCheck`.
static constexpr uint32_t MinItersBypassWeights[] = {1, 127};
@@ -3668,6 +3690,11 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
if (Plan.hasEarlyExit())
return 1;
+ // Interleaving would break check-first scalar-replay resume wiring.
+ // So forcing IC=1.
+ if (usesCheckFirstReplay(Legal))
+ return 1;
+
const bool HasReductions =
any_of(Plan.getVectorLoopRegion()->getEntryBasicBlock()->phis(),
IsaPred<VPReductionPHIRecipe>);
@@ -5496,6 +5523,10 @@ void LoopVectorizationPlanner::plan(ElementCount UserVF, unsigned UserIC) {
if (!MaxFactors) // Cases that should not to be vectorized nor interleaved.
return;
+ // Disable scalable vectorization for check-first early-exit loops for now.
+ if (usesCheckFirstReplay(Legal))
+ MaxFactors.ScalableVF = ElementCount::getScalable(0);
+
Config.collectInLoopReductions();
// Cases that may be vectorized may be optimized by unit stride predicates.
// TODO: Currently unit stride predicates are added unconditionally, even if
@@ -5968,6 +5999,11 @@ DenseMap<const SCEV *, Value *> LoopVectorizationPlanner::executePlan(
// Regions are dissolved after optimizing for VF and UF, which completely
// removes unneeded loop regions first.
RUN_VPLAN_PASS(VPlanTransforms::dissolveLoopRegions, BestVPlan);
+ // Scalar replay routes check.exit to the scalar preheader after region
+ // dissolution.
+ VPlanTransforms::wireCheckFirstExitToScalar(BestVPlan);
+ // Masked replay routes check.exit to the real early exit block instead.
+ VPlanTransforms::wireCheckFirstMaskedReplayToExit(BestVPlan);
// Expand BranchOnTwoConds after dissolution, when latch has direct access to
// its successors.
RUN_VPLAN_PASS(VPlanTransforms::expandBranchOnTwoConds, BestVPlan);
@@ -6178,8 +6214,6 @@ VPRecipeBase *VPRecipeBuilder::tryToWidenMemory(VPInstruction *VPI,
if (!LoopVectorizationPlanner::getDecisionAndClampRange(WillWiden, Range))
return nullptr;
- // If a mask is not required, drop it - use unmasked version for safe loads.
- // TODO: Determine if mask is needed in VPlan.
VPValue *Mask = CM.isMaskRequired(I) ? VPI->getMask() : nullptr;
// Determine if the pointer operand of the access is either consecutive or
@@ -6580,10 +6614,21 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
// the presence of an uncountable exit and the presence of stores in
// the loop inside handleEarlyExits itself.
UncountableExitStyle EEStyle = UncountableExitStyle::NoUncountableExit;
- if (Legal->hasUncountableEarlyExit())
- EEStyle = Legal->hasUncountableExitWithSideEffects()
- ? UncountableExitStyle::MaskedHandleExitInScalarLoop
- : UncountableExitStyle::ReadOnly;
+ if (Legal->hasUncountableEarlyExit()) {
+ if (Legal->hasUncountableExitWithSideEffects()) {
+ if (EnableEarlyExitVectorizationWithSideEffects)
+ EEStyle = UncountableExitStyle::MaskedHandleExitInScalarLoop;
+ else
+ EEStyle = UncountableExitStyle::CheckFirst;
+ } else {
+ EEStyle = UncountableExitStyle::ReadOnly;
+ }
+ }
+
+ assert((EEStyle != UncountableExitStyle::CheckFirst ||
+ Legal->exitLoadsAreDereferenceable()) &&
+ "check-first vectorization reached for a loop whose speculatively "
+ "widened loads could not be made memory-safe");
if (!RUN_VPLAN_PASS(VPlanTransforms::handleEarlyExits, *VPlan0, EEStyle,
OrigLoop, PSE, *DT, Legal->getAssumptionCache())) {
@@ -6592,7 +6637,8 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
RUN_VPLAN_PASS(VPlanTransforms::createLoopRegions, *VPlan0,
getDebugLocFromInstOrOperands(Legal->getPrimaryInduction()));
- if (CM.foldTailByMasking())
+ if (CM.foldTailByMasking() &&
+ EEStyle != UncountableExitStyle::CheckFirst)
RUN_VPLAN_PASS(VPlanTransforms::foldTailByMasking, *VPlan0);
RUN_VPLAN_PASS(VPlanTransforms::introduceMasksAndLinearize, *VPlan0);
@@ -7286,8 +7332,12 @@ static void checkMixedPrecision(Loop *L, OptimizationRemarkEmitter *ORE) {
/// TODO: This is currently overly pessimistic because the loop may not take
/// the early exit, but better to keep this conservative for now. In future,
/// it might be possible to relax this by using branch probabilities.
+///
+/// For check first loops, add the scalar replay cost of the early exit chunk.
static InstructionCost calculateEarlyExitCost(VPCostContext &CostCtx,
- VPlan &Plan, ElementCount VF) {
+ VPlan &Plan, ElementCount VF,
+ uint64_t ScalarCostPerIter,
+ bool IsCheckFirstReplay) {
InstructionCost Cost = 0;
for (auto *ExitVPBB : Plan.getExitBlocks()) {
for (auto *PredVPBB : ExitVPBB->getPredecessors()) {
@@ -7301,6 +7351,15 @@ static InstructionCost calculateEarlyExitCost(VPCostContext &CostCtx,
}
}
}
+
+ if (IsCheckFirstReplay && VF.isFixed()) {
+ uint64_t ReplayedIters = VF.getFixedValue() / 2;
+ InstructionCost ReplayCost(ScalarCostPerIter * ReplayedIters);
+ LLVM_DEBUG(dbgs() << "LV: Adding check-first scalar-replay cost "
+ << ReplayCost << " (~" << ReplayedIters
+ << " scalar iterations) for VF " << VF << ".\n");
+ Cost += ReplayCost;
+ }
return Cost;
}
@@ -7317,7 +7376,8 @@ static bool isOutsideLoopWorkProfitable(GeneratedRTChecks &Checks,
PredicatedScalarEvolution &PSE,
VPCostContext &CostCtx, VPlan &Plan,
EpilogueLowering SEL,
- std::optional<unsigned> VScale) {
+ std::optional<unsigned> VScale,
+ bool IsCheckFirstReplay) {
InstructionCost RtC = Checks.getCost();
if (!RtC.isValid())
return false;
@@ -7343,8 +7403,8 @@ static bool isOutsideLoopWorkProfitable(GeneratedRTChecks &Checks,
InstructionCost TotalCost = RtC;
// Add on the cost of any work required in the vector early exit block, if
- // one exists.
- TotalCost += calculateEarlyExitCost(CostCtx, Plan, VF.Width);
+ // one exists plus check first scalar replay cost.
+ TotalCost += calculateEarlyExitCost(CostCtx, Plan, VF.Width, ScalarC, IsCheckFirstReplay);
TotalCost += Plan.getMiddleBlock()->cost(VF.Width, CostCtx);
// First, compute the minimum iteration count required so that the vector
@@ -7924,6 +7984,19 @@ bool LoopVectorizePass::processLoop(Loop *L) {
return false;
}
+ // Memory-safety gate: bail to the scalar loop when a speculatively widened
+ // exit-condition load is not provably dereferenceable.
+ if (EnableCheckFirstVectorization &&
+ !EnableEarlyExitVectorizationWithSideEffects &&
+ LVL.wouldUseCheckFirstStyle() && !LVL.exitLoadsAreDereferenceable()) {
+ reportVectorizationFailure(
+ "check-first early-exit memory-safety strategy is Unsafe: a "
+ "speculatively-widened condition load could not be proven "
+ "dereferenceable for the full trip count",
+ "CheckFirstUnsafeMemSafety", ORE, L);
+ return false;
+ }
+
bool IsInnerLoop = L->isInnermost();
// Outer loops require a computable trip count.
@@ -7940,11 +8013,12 @@ bool LoopVectorizePass::processLoop(Loop *L) {
return false;
}
if (LVL.hasUncountableExitWithSideEffects() &&
- !EnableEarlyExitVectorizationWithSideEffects) {
- reportVectorizationFailure("Auto-vectorization of loops with uncountable "
- "early exit and side effects is not enabled",
- "UncountableEarlyExitSideEffectLoopsDisabled",
- ORE, L);
+ !EnableEarlyExitVectorizationWithSideEffects &&
+ !EnableCheckFirstVectorization) {
+ reportVectorizationFailure(
+ "Auto-vectorization of loops with uncountable "
+ "early exit and side effects is not enabled",
+ "UncountableEarlyExitSideEffectLoopsDisabled", ORE, L);
return false;
}
}
@@ -8123,7 +8197,8 @@ bool LoopVectorizePass::processLoop(Loop *L) {
CM.PSE, L);
if (!ForceVectorization &&
!isOutsideLoopWorkProfitable(Checks, VF, L, PSE, CostCtx, *BestPlanPtr,
- SEL, Config.getVScaleForTuning())) {
+ SEL, Config.getVScaleForTuning(),
+ usesCheckFirstReplay(&LVL))) {
ORE->emit([&]() {
return OptimizationRemarkAnalysisAliasing(
DEBUG_TYPE, "CantReorderMemOps", L->getStartLoc(),
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.cpp b/llvm/lib/Transforms/Vectorize/VPlan.cpp
index 9c0bec2350c97..f9260639a5f23 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlan.cpp
@@ -1206,6 +1206,18 @@ static void remapOperands(VPBlockBase *Entry, VPBlockBase *NewEntry,
}
}
+/// Returns the clone of Old recorded in Old2NewBlocks, or nullptr when
+/// Old is null.
+static VPBasicBlock *
+remapClonedBlock(const DenseMap<VPBlockBase *, VPBlockBase *> &Old2NewBlocks,
+ VPBasicBlock *Old) {
+ if (!Old)
+ return nullptr;
+ VPBlockBase *New = Old2NewBlocks.lookup(Old);
+ assert(New && "Check-first block not found in cloned plan.");
+ return cast<VPBasicBlock>(New);
+}
+
VPlan *VPlan::duplicate() {
unsigned NumBlocksBeforeCloning = CreatedBlocks.size();
// Clone blocks.
@@ -1291,6 +1303,22 @@ VPlan *VPlan::duplicate() {
NewPlan->ExitBlocks.push_back(cast<VPIRBasicBlock>(VPB));
}
+ NewPlan->CheckFirst.EarlyExitBlock = CheckFirst.EarlyExitBlock;
+ NewPlan->CheckFirst.InclusiveReplayStores = CheckFirst.InclusiveReplayStores;
+
+ // Map each original block to its clone via a single depth-first traversal.
+ DenseMap<VPBlockBase *, VPBlockBase *> Old2NewBlocks;
+ for (const auto &[OldBB, NewBB] :
+ zip_equal(vp_depth_first_deep(Entry), vp_depth_first_deep(NewEntry)))
+ Old2NewBlocks[OldBB] = NewBB;
+
+ NewPlan->CheckFirst.ExitBlock =
+ remapClonedBlock(Old2NewBlocks, CheckFirst.ExitBlock);
+ NewPlan->CheckFirst.MaskedReplayBlock =
+ remapClonedBlock(Old2NewBlocks, CheckFirst.MaskedReplayBlock);
+ NewPlan->CheckFirst.CheckHeaderBlock =
+ remapClonedBlock(Old2NewBlocks, CheckFirst.CheckHeaderBlock);
+
return NewPlan;
}
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index eaf9d1433aff7..d6215a0570def 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -87,6 +87,11 @@ enum class UncountableExitStyle {
/// uncountable exit is taken, then all lanes before the exiting lane will
/// complete, leaving just the final lane to execute in the scalar tail.
MaskedHandleExitInScalarLoop,
+ /// Check-first semantics: exit conditions are evaluated at the start of each
+ /// vector iteration before any stores execute. On early exit, the scalar loop
+ /// resumes from the start of the current vector chunk. Stores are placed in a
+ /// separate body block only reachable when no exit fires.
+ CheckFirst,
};
/// VPBlockBase is the building block of the Hierarchical Control-Flow Graph.
@@ -3746,6 +3751,14 @@ class LLVM_ABI_FOR_TEST VPWidenMemoryRecipe : public VPIRMetadata {
/// Whether the memory access is masked.
bool IsMasked = false;
+ VPWidenMemoryRecipe(Instruction &I, bool Consecutive,
+ const VPIRMetadata &Metadata)
+ : VPIRMetadata(Metadata), Ingredient(I),
+ Alignment(getLoadStoreAlignment(&I)), Consecutive(Consecutive) {}
+
+public:
+ virtual ~VPWidenMemoryRecipe() = default;
+
void setMask(VPValue *Mask) {
assert(!IsMasked && "cannot re-set mask");
if (!Mask)
@@ -3756,14 +3769,6 @@ class LLVM_ABI_FOR_TEST VPWidenMemoryRecipe : public VPIRMetadata {
IsMasked = true;
}
- VPWidenMemoryRecipe(Instruction &I, bool Consecutive,
- const VPIRMetadata &Metadata)
- : VPIRMetadata(Metadata), Ingredient(I),
- Alignment(getLoadStoreAlignment(&I)), Consecutive(Consecutive) {}
-
-public:
- virtual ~VPWidenMemoryRecipe() = default;
-
/// Return a VPRecipeBase* to the current object.
virtual VPRecipeBase *getAsRecipe() = 0;
virtual const VPRecipeBase *getAsRecipe() const = 0;
@@ -4756,6 +4761,28 @@ class VPlan {
/// VPIRBasicBlock wrapping the header of the original scalar loop.
VPIRBasicBlock *ScalarHeader;
+ /// Used for recording the state for check-first early-exit vectorization.
+ struct CheckFirstEarlyExitState {
+ /// Block for routing check-first early exits to the scalar preheader.
+ VPBasicBlock *ExitBlock = nullptr;
+
+ /// The check.exit block used for masked replay wiring.
+ VPBasicBlock *MaskedReplayBlock = nullptr;
+
+ /// Early-exit IR block for masked replay.
+ BasicBlock *EarlyExitBlock = nullptr;
+
+ /// First block of the check-first cascade.
+ VPBasicBlock *CheckHeaderBlock = nullptr;
+
+ /// Block the masked-replay chain temporarily branches to.
+ VPBlockBase *MaskedReplayTempTarget = nullptr;
+
+ /// Stores that must replay the exiting lane.
+ SmallPtrSet<const Instruction *, 4> InclusiveReplayStores;
+ };
+ CheckFirstEarlyExitState CheckFirst;
+
/// Immutable list of VPIRBasicBlocks wrapping the exit blocks of the original
/// scalar loop. Note that some exit blocks may be unreachable at the moment,
/// e.g. if the scalar epilogue always executes.
@@ -4886,6 +4913,54 @@ class VPlan {
getScalarHeader()->getSinglePredecessor());
}
+ VPBasicBlock *getCheckFirstExitBlock() const { return CheckFirst.ExitBlock; }
+
+ void setCheckFirstExitBlock(VPBasicBlock *VPBB) {
+ assert((!CheckFirst.ExitBlock || CheckFirst.ExitBlock == VPBB) &&
+ "CheckFirstExitBlock already set");
+ CheckFirst.ExitBlock = VPBB;
+ }
+
+ VPBasicBlock *getCheckFirstMaskedReplayBlock() const {
+ return CheckFirst.MaskedReplayBlock;
+ }
+
+ void setCheckFirstMaskedReplayBlock(VPBasicBlock *VPBB) {
+ CheckFirst.MaskedReplayBlock = VPBB;
+ }
+
+ BasicBlock *getCheckFirstEarlyExitBlock() const {
+ return CheckFirst.EarlyExitBlock;
+ }
+
+ void setCheckFirstEarlyExitBlock(BasicBlock *IRBB) {
+ CheckFirst.EarlyExitBlock = IRBB;
+ }
+
+ void addCheckFirstInclusiveReplayStore(const Instruction *I) {
+ CheckFirst.InclusiveReplayStores.insert(I);
+ }
+
+ bool isCheckFirstInclusiveReplayStore(const Instruction *I) const {
+ return CheckFirst.InclusiveReplayStores.contains(I);
+ }
+
+ VPBasicBlock *getCheckFirstCheckHeaderBlock() const {
+ return CheckFirst.CheckHeaderBlock;
+ }
+
+ void setCheckFirstCheckHeaderBlock(VPBasicBlock *VPBB) {
+ CheckFirst.CheckHeaderBlock = VPBB;
+ }
+
+ VPBlockBase *getCheckFirstMaskedReplayTempTarget() const {
+ return CheckFirst.MaskedReplayTempTarget;
+ }
+
+ void setCheckFirstMaskedReplayTempTarget(VPBlockBase *B) {
+ CheckFirst.MaskedReplayTempTarget = B;
+ }
+
/// Return the VPIRBasicBlock wrapping the header of the scalar loop.
VPIRBasicBlock *getScalarHeader() const { return ScalarHeader; }
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index 619fea8c10b4d..5581bbe1fe696 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -462,13 +462,25 @@ static void createLoopRegion(VPlan &Plan, VPBlockBase *HeaderVPB, DebugLoc DL) {
// Transfer latch's successors to the region.
VPBlockUtils::transferSuccessors(LatchVPBB, R);
+ VPBasicBlock *CheckExitVPBB = Plan.getCheckFirstExitBlock();
+ if (CheckExitVPBB) {
+ assert(CheckExitVPBB->empty() &&
+ "check.exit block should be empty before region creation");
+ assert(CheckExitVPBB->getNumSuccessors() == 0 &&
+ "check.exit should have no successors before temporary edge");
+ VPBlockUtils::connectBlocks(CheckExitVPBB, LatchVPBB);
+ }
+
VPBlockUtils::connectBlocks(PreheaderVPBB, R);
R->setEntry(HeaderVPB);
R->setExiting(LatchVPBB);
// All VPBB's reachable shallowly from HeaderVPB belong to the current region.
- for (VPBlockBase *VPBB : vp_depth_first_shallow(HeaderVPB))
+ for (VPBlockBase *VPBB : vp_depth_first_shallow(HeaderVPB)) {
+ if (VPBB == Plan.getScalarPreheader())
+ continue;
VPBB->setParent(R);
+ }
if (!IsOutermost)
return;
@@ -1275,6 +1287,8 @@ bool VPlanTransforms::handleEarlyExits(VPlan &Plan, UncountableExitStyle Style,
// Dereferenceability is checked separately for uncountable exit loops with
// stores, as only the loads contributing to the exit condition need to
// be checked.
+ // ReadOnly needs all loads dereferenceable whereas CheckFirst checks only
+ // condition-slice loads later in handleUncountableEarlyExits.
if (Style == UncountableExitStyle::ReadOnly &&
!areAllLoadsDereferenceable(HeaderVPBB, TheLoop, PSE, DT, AC))
return false;
@@ -1348,7 +1362,8 @@ void VPlanTransforms::createLoopRegions(VPlan &Plan, DebugLoc DL) {
VPRegionBlock *TopRegion = Plan.getVectorLoopRegion();
TopRegion->setName("vector loop");
- TopRegion->getEntryBasicBlock()->setName("vector.body");
+ TopRegion->getEntryBasicBlock()->setName(
+ Plan.getCheckFirstCheckHeaderBlock() ? "vector.check" : "vector.body");
}
void VPlanTransforms::foldTailByMasking(VPlan &Plan) {
diff --git a/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp b/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
index 655ac58e24426..d678b147abb5a 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanPredicator.cpp
@@ -274,6 +274,8 @@ void VPlanTransforms::introduceMasksAndLinearize(VPlan &Plan) {
// Nested loop regions (outer-loop vectorization) are not supported yet.
if (Plan.isOuterLoop())
return;
+ if (Plan.getCheckFirstExitBlock())
+ return;
VPRegionBlock *LoopRegion = Plan.getVectorLoopRegion();
// Scan the body of the loop in a topological order to visit each basic block
// after having visited its predecessor basic blocks.
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index f45b9e4f6c35b..fd0b7974879a1 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -95,6 +95,7 @@ bool VPRecipeBase::mayWriteToMemory() const {
case VPBlendSC:
case VPReductionEVLSC:
case VPReductionSC:
+ case VPVectorEndPointerSC:
case VPVectorPointerSC:
case VPWidenCanonicalIVSC:
case VPWidenCastSC:
@@ -149,6 +150,7 @@ bool VPRecipeBase::mayReadFromMemory() const {
case VPBlendSC:
case VPReductionEVLSC:
case VPReductionSC:
+ case VPVectorEndPointerSC:
case VPVectorPointerSC:
case VPWidenCanonicalIVSC:
case VPWidenCastSC:
@@ -870,9 +872,13 @@ Value *VPInstruction::generate(VPTransformState &State) {
cast<VPBasicBlock>(getParent()->getSuccessors()[1]);
BasicBlock *SecondIRSucc = State.CFG.VPBB2IRBB.lookup(SecondVPSucc);
BasicBlock *IRBB = State.CFG.VPBB2IRBB[getParent()];
- auto *Br = Builder.CreateCondBr(Cond, IRBB, SecondIRSucc);
+ // Placeholder for a successor. Assigned in connectToPredecessors.
+ auto *Br =
+ Builder.CreateCondBr(Cond, IRBB, SecondIRSucc ? SecondIRSucc : IRBB);
// First successor is always forward, reset it to nullptr.
Br->setSuccessor(0, nullptr);
+ if (!SecondIRSucc)
+ Br->setSuccessor(1, nullptr);
IRBB->getTerminator()->eraseFromParent();
applyMetadata(*Br);
return Br;
@@ -1552,6 +1558,11 @@ void VPInstruction::addOperand(VPValue *Op) {
"matching operand 1's type and i1, respectively");
break;
}
+ case Instruction::PHI:
+ assert((getNumOperands() == 0 ||
+ Ty == getOperand(0)->getScalarType()) &&
+ "all incoming values must have the same type");
+ break;
default:
llvm_unreachable("opcode does not support growing the operand list "
"outside of construction");
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
index ceb7c38ca3e29..2aba5b8363b3d 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
@@ -35,11 +35,13 @@
#include "llvm/Analysis/MemoryLocation.h"
#include "llvm/Analysis/ScalarEvolutionPatternMatch.h"
#include "llvm/Analysis/ScopedNoAliasAA.h"
+#include "llvm/Analysis/ValueTracking.h"
#include "llvm/Analysis/VectorUtils.h"
#include "llvm/IR/Intrinsics.h"
#include "llvm/IR/MDBuilder.h"
#include "llvm/IR/Metadata.h"
#include "llvm/Support/Casting.h"
+#include "llvm/Support/ErrorHandling.h"
#include "llvm/Support/TypeSize.h"
#include "llvm/Transforms/Utils/LoopUtils.h"
#include "llvm/Transforms/Utils/ScalarEvolutionExpander.h"
@@ -48,6 +50,8 @@ using namespace llvm;
using namespace VPlanPatternMatch;
using namespace SCEVPatternMatch;
+extern cl::opt<bool> EnableCheckFirstMaskedReplay;
+
bool VPlanTransforms::tryToConvertVPInstructionsToVPRecipes(
VPlan &Plan, const TargetLibraryInfo &TLI) {
@@ -78,14 +82,21 @@ bool VPlanTransforms::tryToConvertVPInstructionsToVPRecipes(
assert(!isa<PHINode>(Inst) && "phis should be handled above");
// Create VPWidenMemoryRecipe for loads and stores.
if (LoadInst *Load = dyn_cast<LoadInst>(Inst)) {
+ VPValue *Mask = Ingredient.getNumOperands() > 1
+ ? Ingredient.getOperand(1)
+ : nullptr;
NewRecipe = new VPWidenLoadRecipe(
- *Load, Ingredient.getOperand(0), nullptr /*Mask*/,
+ *Load, Ingredient.getOperand(0), Mask,
false /*Consecutive*/, *VPI, Ingredient.getDebugLoc());
} else if (StoreInst *Store = dyn_cast<StoreInst>(Inst)) {
+ // Stores carry check-first guard predicate
+ // as the third operand. It becomes the store mask.
+ VPValue *Mask = Ingredient.getNumOperands() > 2
+ ? Ingredient.getOperand(2)
+ : nullptr;
NewRecipe = new VPWidenStoreRecipe(
- *Store, Ingredient.getOperand(1), Ingredient.getOperand(0),
- nullptr /*Mask*/, false /*Consecutive*/, *VPI,
- Ingredient.getDebugLoc());
+ *Store, Ingredient.getOperand(1), Ingredient.getOperand(0), Mask,
+ false /*Consecutive*/, *VPI, Ingredient.getDebugLoc());
} else if (GetElementPtrInst *GEP = dyn_cast<GetElementPtrInst>(Inst)) {
NewRecipe = new VPWidenGEPRecipe(GEP->getSourceElementType(),
Ingredient.operands(), *VPI,
@@ -2561,6 +2572,15 @@ static bool cannotHoistOrSinkRecipe(VPRecipeBase &R, VPBasicBlock *FirstBB,
match(&R, m_Intrinsic<Intrinsic::assume>()))
return vputils::cannotHoistOrSinkRecipe(R, Sinking);
+ bool InSingleSuccChain = false;
+ for (VPBlockBase *Succ = FirstBB; Succ; Succ = Succ->getSingleSuccessor())
+ if (Succ == LastBB) {
+ InSingleSuccChain = true;
+ break;
+ }
+ if (!InSingleSuccChain)
+ return true;
+
// Check that the memory operation doesn't alias between FirstBB and LastBB.
auto MemLoc = vputils::getMemoryLocation(R);
@@ -2864,6 +2884,8 @@ void VPlanTransforms::optimize(VPlan &Plan) {
RUN_VPLAN_PASS(simplifyReverses, Plan);
RUN_VPLAN_PASS(removeDeadRecipes, Plan);
+ RUN_VPLAN_PASS(maskCheckFirstReplayStores, Plan);
+
RUN_VPLAN_PASS(createAndOptimizeReplicateRegions, Plan);
RUN_VPLAN_PASS(mergeBlocksIntoPredecessors, Plan);
RUN_VPLAN_PASS(licm, Plan);
@@ -4596,6 +4618,29 @@ static bool handleUncountableExitsWithSideEffects(
return true;
}
+/// Walk backward from ExitCond to collect the recipes needed to evaluate the
+/// exit condition, stopping at PHIs. Returns false if the condition depends on
+/// a memory-writing recipe which cannot be placed in the check block.
+static bool computeConditionSlice(VPValue *ExitCond,
+ SmallPtrSetImpl<VPRecipeBase *> &Slice) {
+ SmallVector<VPValue *, 16> Worklist;
+ Worklist.push_back(ExitCond);
+ while (!Worklist.empty()) {
+ VPValue *V = Worklist.pop_back_val();
+ VPRecipeBase *DefR = V->getDefiningRecipe();
+ if (!DefR || DefR->isPhi())
+ continue;
+ if (Slice.contains(DefR))
+ continue;
+ if (DefR->mayWriteToMemory())
+ return false;
+ Slice.insert(DefR);
+ for (VPValue *Op : DefR->operands())
+ Worklist.push_back(Op);
+ }
+ return true;
+}
+
bool VPlanTransforms::handleUncountableEarlyExits(
VPlan &Plan, VPBasicBlock *HeaderVPBB, VPBasicBlock *LatchVPBB,
VPBasicBlock *MiddleVPBB, Loop *TheLoop, PredicatedScalarEvolution &PSE,
@@ -4663,6 +4708,393 @@ bool VPlanTransforms::handleUncountableEarlyExits(
"RPO sort must place dominating exits before dominated ones");
#endif
+ if (Style == UncountableExitStyle::CheckFirst) {
+
+ if (range_size(TheLoop->getHeader()->phis()) > 1)
+ return false;
+
+ SmallPtrSet<const VPBlockBase *, 4> ExitBlockSet;
+ SmallPtrSet<const VPBlockBase *, 4> EarlyExitingSet;
+ for (const EarlyExitInfo &E : Exits) {
+ ExitBlockSet.insert(E.EarlyExitVPBB);
+ EarlyExitingSet.insert(E.EarlyExitingVPBB);
+ }
+ if (EarlyExitingSet.contains(LatchVPBB))
+ return false;
+ SmallVector<VPBasicBlock *, 4> LoopChain;
+ struct DiamondGuard {
+ VPValue *Cond;
+ bool ChainOnFalseEdge;
+ VPBasicBlock *Rejoin;
+ unsigned RejoinIdx;
+ };
+ DenseMap<VPBasicBlock *, DiamondGuard> ChainDiamond;
+ {
+ SmallPtrSet<VPBasicBlock *, 8> Seen;
+ VPBasicBlock *Cur = HeaderVPBB;
+ while (Cur) {
+ if (!Seen.insert(Cur).second)
+ return false;
+ LoopChain.push_back(Cur);
+ if (Cur != HeaderVPBB && Cur->begin() != Cur->end() &&
+ Cur->begin()->isPhi())
+ return false;
+ if (Cur == LatchVPBB)
+ break;
+ SmallVector<VPBasicBlock *, 2> InLoopSuccs;
+ for (VPBlockBase *S : Cur->getSuccessors()) {
+ if (S == MiddleVPBB || ExitBlockSet.contains(S))
+ continue;
+ auto *SBB = dyn_cast<VPBasicBlock>(S);
+ if (!SBB)
+ return false;
+ InLoopSuccs.push_back(SBB);
+ }
+ VPBasicBlock *NextInLoop = nullptr;
+ if (InLoopSuccs.size() == 1) {
+ NextInLoop = InLoopSuccs[0];
+ } else if (InLoopSuccs.size() == 2) {
+ VPBasicBlock *Cont = nullptr;
+ VPBasicBlock *Bypass = nullptr;
+ if (InLoopSuccs[0] == LatchVPBB) {
+ Bypass = InLoopSuccs[0];
+ Cont = InLoopSuccs[1];
+ } else if (InLoopSuccs[1] == LatchVPBB) {
+ Bypass = InLoopSuccs[1];
+ Cont = InLoopSuccs[0];
+ } else {
+ bool S0Merge = InLoopSuccs[0]->getNumPredecessors() > 1;
+ bool S1Merge = InLoopSuccs[1]->getNumPredecessors() > 1;
+ if (S0Merge && !S1Merge) {
+ Bypass = InLoopSuccs[0];
+ Cont = InLoopSuccs[1];
+ } else if (S1Merge && !S0Merge) {
+ Bypass = InLoopSuccs[1];
+ Cont = InLoopSuccs[0];
+ } else {
+ return false;
+ }
+ }
+ VPValue *GuardCond = nullptr;
+ if (!match(Cur->getTerminator(), m_BranchOnCond(m_VPValue(GuardCond))))
+ return false;
+ bool ChainOnFalseEdge = Cur->getSuccessors()[1] == Cont;
+ ChainDiamond[Cur] = {GuardCond, ChainOnFalseEdge, Bypass, 0};
+ NextInLoop = Cont;
+ } else {
+ return false;
+ }
+ Cur = NextInLoop;
+ }
+ if (LoopChain.empty() || LoopChain.back() != LatchVPBB)
+ return false;
+ }
+
+ DenseMap<const VPBasicBlock *, unsigned> ChainIdx;
+ for (unsigned I = 0, E = LoopChain.size(); I != E; ++I)
+ ChainIdx[LoopChain[I]] = I;
+ for (auto &[DBB, DG] : ChainDiamond) {
+ auto RIt = ChainIdx.find(DG.Rejoin);
+ if (RIt == ChainIdx.end() || RIt->second <= ChainIdx[DBB])
+ return false;
+ DG.RejoinIdx = RIt->second;
+ }
+
+ ArrayRef<VPBasicBlock *> Intermediates =
+ ArrayRef(LoopChain).drop_front().drop_back();
+
+ for (EarlyExitInfo &Exit : Exits) {
+ auto It = llvm::find(LoopChain, Exit.EarlyExitingVPBB);
+ if (It == LoopChain.end())
+ return false;
+ unsigned ExitIdx = std::distance(LoopChain.begin(), It);
+ auto *MC = cast<VPInstruction>(Exit.CondToExit);
+ VPBuilder GuardBuilder(MC);
+ VPValue *CombinedGuard = nullptr;
+ for (unsigned I = 0; I < ExitIdx; ++I) {
+ auto DiamondIt = ChainDiamond.find(LoopChain[I]);
+ if (DiamondIt == ChainDiamond.end() ||
+ ExitIdx >= DiamondIt->second.RejoinIdx)
+ continue;
+ VPValue *Guard = DiamondIt->second.Cond;
+ if (DiamondIt->second.ChainOnFalseEdge)
+ Guard = GuardBuilder.createNot(Guard);
+ CombinedGuard =
+ CombinedGuard ? GuardBuilder.createLogicalAnd(CombinedGuard, Guard)
+ : Guard;
+ }
+ if (CombinedGuard)
+ MC->setOperand(
+ 0, GuardBuilder.createLogicalAnd(CombinedGuard, MC->getOperand(0)));
+ }
+
+ DenseMap<VPRecipeBase *, SmallVector<std::pair<VPValue *, bool>, 2>>
+ StoreGuards;
+ DenseMap<VPRecipeBase *, SmallVector<std::pair<VPValue *, bool>, 2>>
+ BodyLoadGuards;
+ for (unsigned J = 0, E = LoopChain.size(); J != E; ++J) {
+ if (LoopChain[J] == LatchVPBB)
+ continue;
+ SmallVector<std::pair<VPValue *, bool>, 2> Guards;
+ for (unsigned I = 0; I < J; ++I) {
+ auto GIt = ChainDiamond.find(LoopChain[I]);
+ if (GIt != ChainDiamond.end() && J < GIt->second.RejoinIdx)
+ Guards.push_back({GIt->second.Cond, GIt->second.ChainOnFalseEdge});
+ }
+ if (Guards.empty())
+ continue;
+ for (VPRecipeBase &R : *LoopChain[J]) {
+ if (R.mayWriteToMemory()) {
+ StoreGuards[&R] = Guards;
+ } else {
+ auto *VPI = dyn_cast<VPInstruction>(&R);
+ if (VPI && VPI->getUnderlyingValue())
+ if (auto *LI = dyn_cast<LoadInst>(VPI->getUnderlyingValue()))
+ if (!isSafeToSpeculativelyExecute(LI))
+ BodyLoadGuards[&R] = Guards;
+ }
+ }
+ }
+
+ // Compute each exit's condition slice.
+ SmallVector<SmallPtrSet<VPRecipeBase *, 16>, 4> Slices(Exits.size());
+ DenseMap<VPRecipeBase *, unsigned> EarliestCheck;
+ for (unsigned K = 0, E = Exits.size(); K != E; ++K) {
+ if (!computeConditionSlice(Exits[K].CondToExit, Slices[K]))
+ return false;
+ for (VPRecipeBase *R : Slices[K])
+ EarliestCheck.try_emplace(R, K);
+ }
+
+
+ VPIRBasicBlock *MaskedReplayExitBB = nullptr;
+ bool DoMaskedReplay = EnableCheckFirstMaskedReplay && Exits.size() == 1;
+
+ if (DoMaskedReplay) {
+ ScalarEvolution &SE = *PSE.getSE();
+ PHINode *IndVar = TheLoop->getInductionVariable(SE);
+ // Reconstructed as CanonIV + first_active_lane.
+ bool IsUnitFromZero = false;
+ if (IndVar) {
+ if (auto *AR = dyn_cast<SCEVAddRecExpr>(SE.getSCEV(IndVar))) {
+ if (AR->getLoop() == TheLoop && AR->getStart()->isZero()) {
+ const SCEV *Step = AR->getStepRecurrence(SE);
+ IsUnitFromZero = Step->isOne();
+ }
+ }
+ }
+ VPBasicBlock *EEing = Exits[0].EarlyExitingVPBB;
+ for (VPRecipeBase &R : Exits[0].EarlyExitVPBB->phis()) {
+ VPValue *Incoming =
+ cast<VPIRPhi>(&R)->getIncomingValueForBlock(EEing);
+ Value *UV = Incoming->getUnderlyingValue();
+ if (UV)
+ UV = UV->stripPointerCasts();
+ if (!IsUnitFromZero || !UV || UV != IndVar)
+ DoMaskedReplay = false;
+ }
+ if (DoMaskedReplay) {
+ MaskedReplayExitBB = Exits[0].EarlyExitVPBB;
+
+ BasicBlock *EarlyExitIR = MaskedReplayExitBB->getIRBasicBlock();
+ BasicBlock *Latch = TheLoop->getLoopLatch();
+ BasicBlock *ExitingIR = nullptr;
+ for (BasicBlock *Pred : predecessors(EarlyExitIR))
+ if (TheLoop->contains(Pred) && Pred != Latch) {
+ ExitingIR = Pred;
+ break;
+ }
+ if (ExitingIR)
+ for (BasicBlock *BB : TheLoop->blocks())
+ for (Instruction &I : *BB)
+ if (isa<StoreInst>(&I) && DT.dominates(BB, ExitingIR))
+ Plan.addCheckFirstInclusiveReplayStore(&I);
+ }
+ }
+
+ // From here, all modifications are destructive. We cannot bail out.
+
+ for (auto &Exit : Exits) {
+ auto &[EarlyExitingVPBB, EarlyExitVPBB, _] = Exit;
+ for (VPRecipeBase &R : EarlyExitVPBB->phis())
+ cast<VPIRPhi>(&R)->removeIncomingValueFor(EarlyExitingVPBB);
+ EarlyExitingVPBB->getTerminator()->eraseFromParent();
+ VPBlockUtils::disconnectBlocks(EarlyExitingVPBB, EarlyExitVPBB);
+ }
+
+ // Flatten intermediate blocks recipes into the header, then connect the
+ // header straight to the latch. Single-exit chains are left untouched.
+ if (!Intermediates.empty()) {
+ if (!EarlyExitingSet.contains(HeaderVPBB))
+ HeaderVPBB->getTerminator()->eraseFromParent();
+ for (VPBasicBlock *BB : Intermediates)
+ if (!EarlyExitingSet.contains(BB) && BB->getTerminator())
+ BB->getTerminator()->eraseFromParent();
+ for (VPBasicBlock *BB : Intermediates)
+ for (VPRecipeBase &R : make_early_inc_range(*BB))
+ R.moveBefore(*HeaderVPBB, HeaderVPBB->end());
+ for (VPBlockBase *S : to_vector(HeaderVPBB->getSuccessors()))
+ VPBlockUtils::disconnectBlocks(HeaderVPBB, S);
+ for (VPBasicBlock *BB : Intermediates)
+ for (VPBlockBase *S : to_vector(BB->getSuccessors()))
+ VPBlockUtils::disconnectBlocks(BB, S);
+ VPBlockUtils::connectBlocks(HeaderVPBB, LatchVPBB);
+ }
+
+ // Hoist guard producers to the header to avoid use-before-def.
+ if (DoMaskedReplay) {
+ SmallVector<VPRecipeBase *, 8> GuardHoist;
+ auto EnqueueProducers = [&](VPValue *V) {
+ if (VPRecipeBase *Def = V->getDefiningRecipe())
+ if (!EarliestCheck.contains(Def))
+ GuardHoist.push_back(Def);
+ };
+ for (auto &[R, Guards] : StoreGuards)
+ for (auto &[G, ChainOnFalseEdge] : Guards)
+ EnqueueProducers(G);
+ for (auto &[R, Guards] : BodyLoadGuards)
+ for (auto &[G, ChainOnFalseEdge] : Guards)
+ EnqueueProducers(G);
+ while (!GuardHoist.empty()) {
+ VPRecipeBase *R = GuardHoist.pop_back_val();
+ if (EarliestCheck.try_emplace(R, 0).second)
+ for (VPValue *Op : R->operands())
+ EnqueueProducers(Op);
+ }
+ }
+
+ // Create the check cascade. Checks[0] is the header.
+ // Checks[k] is a fresh block holding exit k's condition slice.
+ SmallVector<VPBasicBlock *, 4> Checks;
+ Checks.push_back(HeaderVPBB);
+ for (unsigned K = 1, E = Exits.size(); K != E; ++K)
+ Checks.push_back(Plan.createVPBasicBlock("vector.check"));
+
+ // Body block holds non-slice recipes. Runs when no exit fires.
+ VPBasicBlock *BodyVPBB = Plan.createVPBasicBlock("vector.body");
+
+ // Partition recipes. Slice recipes to their check block, everything else to
+ // the body.
+ auto PartitionBlock = [&](VPBasicBlock *BB) {
+ for (VPRecipeBase &R : make_early_inc_range(*BB)) {
+ if (R.isPhi() || &R == BB->getTerminator())
+ continue;
+ auto It = EarliestCheck.find(&R);
+ VPBasicBlock *Target =
+ It != EarliestCheck.end() ? Checks[It->second] : BodyVPBB;
+ if (Target != BB)
+ R.moveBefore(*Target, Target->end());
+ }
+ };
+ PartitionBlock(HeaderVPBB);
+ if (HeaderVPBB != LatchVPBB)
+ PartitionBlock(LatchVPBB);
+
+ // Attach combined guard mask as a trailing operand to each guarded store.
+ for (auto &[R, Guards] : StoreGuards) {
+ auto *OldStore = cast<VPInstruction>(R);
+ assert(OldStore->getOpcode() == Instruction::Store &&
+ "check-first guarded side effect is not a store");
+ VPBuilder GuardBuilder;
+ if (DoMaskedReplay) {
+ if (VPRecipeBase *Term = HeaderVPBB->getTerminator())
+ GuardBuilder.setInsertPoint(HeaderVPBB, Term->getIterator());
+ else
+ GuardBuilder.setInsertPoint(HeaderVPBB);
+ } else {
+ GuardBuilder.setInsertPoint(OldStore);
+ }
+ VPValue *StGuard = nullptr;
+ for (auto &[G, ChainOnFalseEdge] : Guards) {
+ VPValue *GG = ChainOnFalseEdge ? GuardBuilder.createNot(G) : G;
+ StGuard = StGuard ? GuardBuilder.createLogicalAnd(StGuard, GG) : GG;
+ }
+ assert(StGuard && "guarded store recorded without any guard");
+ SmallVector<VPValue *, 3> Ops(OldStore->operands());
+ Ops.push_back(StGuard);
+ auto *NewStore = new VPInstruction(Instruction::Store, Ops, *OldStore,
+ *OldStore, OldStore->getDebugLoc());
+ NewStore->setUnderlyingValue(OldStore->getUnderlyingValue());
+ NewStore->insertBefore(OldStore);
+ OldStore->eraseFromParent();
+ }
+
+ for (auto &[R, Guards] : BodyLoadGuards) {
+ auto *OldLoad = cast<VPInstructionWithType>(R);
+ VPBuilder GuardBuilder(OldLoad);
+ VPValue *LdGuard = nullptr;
+ for (auto &[G, ChainOnFalseEdge] : Guards) {
+ VPValue *GG = ChainOnFalseEdge ? GuardBuilder.createNot(G) : G;
+ LdGuard = LdGuard ? GuardBuilder.createLogicalAnd(LdGuard, GG) : GG;
+ }
+ assert(LdGuard && "masked body load recorded without any guard");
+ SmallVector<VPValue *, 2> Ops(OldLoad->operands());
+ Ops.push_back(LdGuard);
+ auto *NewLoad = new VPInstructionWithType(
+ Instruction::Load, Ops, OldLoad->getResultType(), *OldLoad, *OldLoad,
+ OldLoad->getDebugLoc(), OldLoad->getName(),
+ OldLoad->getUnderlyingValue());
+ NewLoad->insertBefore(OldLoad);
+ OldLoad->replaceAllUsesWith(NewLoad);
+ OldLoad->eraseFromParent();
+ }
+
+ // Routes to the scalar preheader.
+ VPBasicBlock *EarlyExitToScalarVPBB =
+ Plan.createVPBasicBlock("vector.check.exit");
+ Plan.setCheckFirstExitBlock(EarlyExitToScalarVPBB);
+
+ // Masked replay fills this block later else it routes to the scalar PH.
+ if (DoMaskedReplay)
+ Plan.setCheckFirstMaskedReplayBlock(EarlyExitToScalarVPBB);
+
+ // Extract the latch condition before erasing the latch terminator.
+ auto *LatchBranch = cast<VPInstruction>(LatchVPBB->getTerminator());
+ assert(LatchBranch->getOpcode() == VPInstruction::BranchOnCond &&
+ "Unexpected terminator");
+ VPValue *IsLatchExitTaken = LatchBranch->getOperand(0);
+ DebugLoc LatchDL = LatchBranch->getDebugLoc();
+ LatchBranch->eraseFromParent();
+
+ if (HeaderVPBB != LatchVPBB) {
+ for (VPBlockBase *Succ : to_vector(LatchVPBB->getSuccessors()))
+ VPBlockUtils::disconnectBlocks(LatchVPBB, Succ);
+ }
+
+ for (VPBlockBase *Succ : to_vector(HeaderVPBB->getSuccessors()))
+ VPBlockUtils::disconnectBlocks(HeaderVPBB, Succ);
+
+ // Wire each check.k to exit if exit k fires, else fall through to the next
+ // check or body. BranchOnCond takes successor 0 when true: wire exit first.
+ for (unsigned K = 0, E = Exits.size(); K != E; ++K) {
+ VPBasicBlock *CheckBB = Checks[K];
+ VPBuilder CheckBuilder(CheckBB, CheckBB->end());
+ VPValue *IsExitTaken =
+ CheckBuilder.createNaryOp(VPInstruction::AnyOf, {Exits[K].CondToExit});
+ CheckBuilder.createNaryOp(VPInstruction::BranchOnCond, {IsExitTaken});
+ VPBasicBlock *NextBB = (K + 1 != E) ? Checks[K + 1] : BodyVPBB;
+ VPBlockUtils::connectBlocks(CheckBB, EarlyExitToScalarVPBB);
+ VPBlockUtils::connectBlocks(CheckBB, NextBB);
+ }
+
+
+ Plan.setCheckFirstCheckHeaderBlock(HeaderVPBB);
+
+ // Remember the real early-exit block for wireCheckFirstMaskedReplayToExit.
+ if (DoMaskedReplay)
+ Plan.setCheckFirstEarlyExitBlock(MaskedReplayExitBB->getIRBasicBlock());
+
+ VPBuilder BodyBuilder(BodyVPBB, BodyVPBB->end());
+ BodyBuilder.createNaryOp(VPInstruction::BranchOnCond, {IsLatchExitTaken},
+ LatchDL);
+
+ // Wire: BodyVPBB → {MiddleVPBB , HeaderVPBB}
+ VPBlockUtils::connectBlocks(BodyVPBB, MiddleVPBB);
+ VPBlockUtils::connectBlocks(BodyVPBB, HeaderVPBB);
+
+ return true;
+ }
+
// Build the AnyOf condition for the latch terminator using logical OR
// to avoid poison propagation from later exit conditions when an earlier
// exit is taken.
@@ -4825,6 +5257,498 @@ bool VPlanTransforms::handleUncountableEarlyExits(
return true;
}
+/// Returns true if Root transitively uses Target through its defining
+/// recipes operands.
+static bool vpValueDependsOn(VPValue *Root, VPValue *Target) {
+ SmallVector<VPValue *, 16> Worklist{Root};
+ SmallPtrSet<VPValue *, 16> Visited;
+ while (!Worklist.empty()) {
+ VPValue *V = Worklist.pop_back_val();
+ if (V == Target)
+ return true;
+ if (!Visited.insert(V).second)
+ continue;
+ if (VPRecipeBase *Def = V->getDefiningRecipe())
+ for (VPValue *Op : Def->operands())
+ Worklist.push_back(Op);
+ }
+ return false;
+}
+
+static VPValue *
+rebuildIVResumeExprImpl(VPBuilder &B, VPValue *V, VPValue *VectorTC,
+ VPValue *NewIndex,
+ SmallDenseMap<VPValue *, VPValue *> &Cache) {
+ using namespace VPlanPatternMatch;
+ if (V == VectorTC)
+ return NewIndex;
+ if (auto It = Cache.find(V); It != Cache.end())
+ return It->second;
+ auto Remap = [&](VPValue *Op) {
+ return rebuildIVResumeExprImpl(B, Op, VectorTC, NewIndex, Cache);
+ };
+ auto RemapBinOp = [&](unsigned Opcode, VPValue *LHS, VPValue *RHS,
+ DebugLoc DL, const Twine &Name) -> VPValue * {
+ VPValue *NL = Remap(LHS), *NR = Remap(RHS);
+ if (NL == LHS && NR == RHS)
+ return V;
+ auto Flags =
+ cast<VPRecipeWithIRFlags>(V->getDefiningRecipe())->getNoWrapFlags();
+ return B.createOverflowingOp(Opcode, {NL, NR}, Flags, DL, Name);
+ };
+ auto RemapCast = [&](Instruction::CastOps Opcode, VPValue *Op) -> VPValue * {
+ VPValue *NO = Remap(Op);
+ if (NO == Op)
+ return V;
+ return B.createScalarCast(Opcode, NO, V->getScalarType(), DebugLoc());
+ };
+
+ VPValue *A, *Bv;
+ VPValue *Result = V;
+ if (match(V, m_VPInstruction<VPInstruction::PtrAdd>(m_VPValue(A),
+ m_VPValue(Bv)))) {
+ VPValue *NB = Remap(Bv);
+ if (NB != Bv)
+ Result = B.createPtrAdd(A, NB, DebugLoc(), "check.exit.iv.resume");
+ } else if (match(V, m_Mul(m_VPValue(A), m_VPValue(Bv)))) {
+ Result = RemapBinOp(Instruction::Mul, A, Bv, DebugLoc::getUnknown(), "");
+ } else if (match(V, m_c_Add(m_VPValue(A), m_VPValue(Bv)))) {
+ Result =
+ RemapBinOp(Instruction::Add, A, Bv, DebugLoc(), "check.exit.iv.resume");
+ } else if (match(V, m_Sub(m_VPValue(A), m_VPValue(Bv)))) {
+ Result = RemapBinOp(Instruction::Sub, A, Bv, DebugLoc::getUnknown(), "");
+ } else if (match(V, m_Trunc(m_VPValue(A)))) {
+ Result = RemapCast(Instruction::Trunc, A);
+ } else if (match(V, m_ZExt(m_VPValue(A)))) {
+ Result = RemapCast(Instruction::ZExt, A);
+ } else if (match(V, m_SExt(m_VPValue(A)))) {
+ Result = RemapCast(Instruction::SExt, A);
+ } else if (vpValueDependsOn(V, VectorTC)) {
+ assert(false && "Unhandled VectorTC-dependent check-first resume "
+ "value");
+ }
+ Cache[V] = Result;
+ return Result;
+}
+
+/// Returns the induction resume expression Expr rebuilt with VectorTC
+/// replaced by NewIndex.
+static VPValue *rebuildIVResumeExpr(VPBuilder &B, VPValue *Expr,
+ VPValue *VectorTC, VPValue *NewIndex) {
+ SmallDenseMap<VPValue *, VPValue *> Cache;
+ return rebuildIVResumeExprImpl(B, Expr, VectorTC, NewIndex, Cache);
+}
+
+void VPlanTransforms::wireCheckFirstExitToScalar(VPlan &Plan) {
+ VPBasicBlock *CheckExitVPBB = Plan.getCheckFirstExitBlock();
+ if (!CheckExitVPBB)
+ return;
+
+ // Masked replay handles check.exit during restructuring.
+ if (Plan.getCheckFirstMaskedReplayBlock())
+ return;
+
+ VPBasicBlock *HeaderVPBB = Plan.getCheckFirstCheckHeaderBlock();
+ assert(HeaderVPBB && "check-first cascade header block not recorded");
+
+ // Replace the temporary check.exit→body edge with check.exit→ScalarPH.
+ assert(CheckExitVPBB->getNumSuccessors() == 1 &&
+ "check.exit should have exactly one successor after dissolution");
+ VPBlockBase *OldSucc = CheckExitVPBB->getSuccessors()[0];
+ VPBlockUtils::disconnectBlocks(CheckExitVPBB, OldSucc);
+
+ VPBasicBlock *ScalarPH = Plan.getScalarPreheader();
+ assert(ScalarPH &&
+ "CheckFirst requires a scalar preheader for early-exit replay. "
+ "Ensure the scalar tail is not removed by earlier passes.");
+ VPBlockUtils::connectBlocks(CheckExitVPBB, ScalarPH);
+
+ assert(HeaderVPBB->getNumPredecessors() == 2 &&
+ "loop header must have exactly two predecessors (preheader, latch)");
+ VPBasicBlock *LatchVPBB = nullptr;
+ for (VPBlockBase *Pred : HeaderVPBB->getPredecessors()) {
+ auto *PredVPBB = cast<VPBasicBlock>(Pred);
+ if (any_of(PredVPBB->getSuccessors(),
+ [&](VPBlockBase *S) { return S != HeaderVPBB; })) {
+ LatchVPBB = PredVPBB;
+ break;
+ }
+ }
+ assert(LatchVPBB &&
+ "could not identify the loop latch (backedge source) among the "
+ "header's predecessors");
+ VPValue *CanonIV = cast<VPPhi>(&*HeaderVPBB->begin());
+
+ auto SPHPhis = ScalarPH->phis();
+ assert(range_size(SPHPhis) == 1 &&
+ "CheckFirst expects exactly one scalar-preheader PHI (the IV). "
+ "Extending to multiple inductions or live-outs requires computing "
+ "proper resume values for each PHI.");
+
+ VPBasicBlock *MiddleVPBB = nullptr;
+ for (VPBlockBase *Succ : LatchVPBB->getSuccessors()) {
+ if (Succ != HeaderVPBB) {
+ MiddleVPBB = cast<VPBasicBlock>(Succ);
+ break;
+ }
+ }
+
+ VPBuilder CheckExitBuilder(CheckExitVPBB, CheckExitVPBB->getFirstNonPhi());
+ VPValue *VectorTC = &Plan.getVectorTripCount();
+
+ using namespace VPlanPatternMatch;
+ for (VPRecipeBase &R : SPHPhis) {
+ auto *Phi = cast<VPPhi>(&R);
+
+ VPValue *MidVal = nullptr;
+ for (unsigned I = 0, E = Phi->getNumIncoming(); I != E; ++I) {
+ if (MiddleVPBB && Phi->getIncomingBlock(I) == MiddleVPBB) {
+ MidVal = Phi->getIncomingValue(I);
+ break;
+ }
+ }
+
+ VPValue *ResumeVal =
+ MidVal ? rebuildIVResumeExpr(CheckExitBuilder, MidVal, VectorTC, CanonIV)
+ : CanonIV;
+
+ assert(!(MidVal && ResumeVal == MidVal &&
+ vpValueDependsOn(MidVal, VectorTC)) &&
+ "check-first early-exit resume value could not be rebuilt from the "
+ "vector trip count");
+
+ Phi->addIncoming(ResumeVal);
+ }
+}
+
+void VPlanTransforms::maskCheckFirstReplayStores(VPlan &Plan) {
+ VPBasicBlock *ReplayBB = Plan.getCheckFirstMaskedReplayBlock();
+ if (!ReplayBB)
+ return;
+
+ // Replay surviving lanes body stores here, masked to [0, first_active_lane).
+ assert(ReplayBB->getNumPredecessors() == 1 &&
+ "masked-replay block must have a single (check) predecessor");
+ auto *CheckBB = cast<VPBasicBlock>(ReplayBB->getPredecessors()[0]);
+
+ // The combined per-lane exit condition is the AnyOf operand of BranchOnCond.
+ auto *CheckTerm = cast<VPInstruction>(CheckBB->getTerminator());
+ assert(CheckTerm->getOpcode() == VPInstruction::BranchOnCond &&
+ "check block terminator must be BranchOnCond");
+ auto *AnyOf = cast<VPInstruction>(CheckTerm->getOperand(0)->getDefiningRecipe());
+ assert(AnyOf->getOpcode() == VPInstruction::AnyOf &&
+ "BranchOnCond operand must be AnyOf");
+ VPValue *Combined = AnyOf->getOperand(0);
+
+ VPBasicBlock *BodyVPBB = nullptr;
+ for (VPBlockBase *Succ : CheckBB->getSuccessors()) {
+ if (Succ != ReplayBB) {
+ BodyVPBB = cast<VPBasicBlock>(Succ);
+ break;
+ }
+ }
+ assert(BodyVPBB && "could not find the body block to replay");
+
+#ifndef NDEBUG
+ // Masked replay requires full-width chunks; tail folding is unsupported.
+ for (VPBlockBase *VPB : vp_depth_first_deep(Plan.getEntry()))
+ if (auto *VPBB = dyn_cast<VPBasicBlock>(VPB))
+ for (VPRecipeBase &R : *VPBB)
+ assert(!isa<VPActiveLaneMaskPHIRecipe>(&R) &&
+ "check-first masked replay is unsound under tail folding");
+#endif
+
+ VPBuilder HeadBuilder(ReplayBB, ReplayBB->begin());
+ VPInstruction *FirstActiveLane = HeadBuilder.createFirstActiveLane(
+ {Combined}, DebugLoc::getUnknown(), "first.active.lane");
+
+ Type *IndexTy = FirstActiveLane->getScalarType();
+ assert(IndexTy->isIntegerTy() &&
+ "FirstActiveLane must produce an integer index for mask bounds");
+ VPValue *Zero = Plan.getZero(IndexTy);
+ VPValue *One = Plan.getConstantInt(IndexTy, 1);
+ VPInstruction *ExclMask = nullptr;
+ VPInstruction *InclMask = nullptr;
+ auto getMaskFor = [&](VPRecipeBase &R) -> VPValue * {
+ const Instruction *SI = nullptr;
+ if (auto *WS = dyn_cast<VPWidenStoreRecipe>(&R))
+ SI = &WS->getIngredient();
+ else if (auto *RR = dyn_cast<VPReplicateRecipe>(&R))
+ SI = RR->getUnderlyingInstr();
+ bool Inclusive = SI && Plan.isCheckFirstInclusiveReplayStore(SI);
+ if (Inclusive) {
+ if (!InclMask) {
+ VPInstruction *FALPlusOne =
+ VPBuilder::getToInsertAfter(FirstActiveLane)
+ .createAdd(FirstActiveLane, One, DebugLoc(),
+ "first.active.lane.incl",
+ {/*nuw=*/true, /*nsw=*/false});
+ InclMask = VPBuilder::getToInsertAfter(FALPlusOne)
+ .createNaryOp(VPInstruction::ActiveLaneMask,
+ {Zero, FALPlusOne, One}, DebugLoc(),
+ "masked.replay.mask.incl");
+ }
+ return InclMask;
+ }
+ if (!ExclMask)
+ ExclMask = VPBuilder::getToInsertAfter(FirstActiveLane)
+ .createNaryOp(VPInstruction::ActiveLaneMask,
+ {Zero, FirstActiveLane, One}, DebugLoc(),
+ "masked.replay.mask");
+ return ExclMask;
+ };
+
+ // Replay only the stores and their backward slice.
+ SmallPtrSet<VPRecipeBase *, 8> Needed;
+ SmallVector<VPRecipeBase *, 8> Work;
+ for (VPRecipeBase &R : *BodyVPBB)
+ if (R.mayWriteToMemory()) {
+ Needed.insert(&R);
+ Work.push_back(&R);
+ }
+ while (!Work.empty()) {
+ VPRecipeBase *R = Work.pop_back_val();
+ for (VPValue *Op : R->operands())
+ if (VPRecipeBase *Def = Op->getDefiningRecipe())
+ if (Def->getParent() == BodyVPBB && Needed.insert(Def).second)
+ Work.push_back(Def);
+ }
+
+ SmallPtrSet<VPRecipeBase *, 4> InclusiveFeeds;
+ {
+ SmallVector<VPRecipeBase *, 4> InclWork;
+ for (VPRecipeBase &R : *BodyVPBB) {
+ if (!Needed.contains(&R))
+ continue;
+ const Instruction *SI = nullptr;
+ if (auto *WS = dyn_cast<VPWidenStoreRecipe>(&R))
+ SI = &WS->getIngredient();
+ else if (auto *RR = dyn_cast<VPReplicateRecipe>(&R))
+ if (RR->getUnderlyingInstr()->mayWriteToMemory())
+ SI = RR->getUnderlyingInstr();
+ if (SI && Plan.isCheckFirstInclusiveReplayStore(SI))
+ InclWork.push_back(&R);
+ }
+ while (!InclWork.empty()) {
+ VPRecipeBase *R = InclWork.pop_back_val();
+ for (VPValue *Op : R->operands()) {
+ VPRecipeBase *Def = Op->getDefiningRecipe();
+ if (Def && Def->getParent() == BodyVPBB && Needed.contains(Def) &&
+ InclusiveFeeds.insert(Def).second)
+ InclWork.push_back(Def);
+ }
+ }
+ }
+
+ VPBuilder Builder(ReplayBB, std::next(FirstActiveLane->getIterator()));
+ DenseMap<VPValue *, VPValue *> OperandMap;
+ for (VPRecipeBase &R : *BodyVPBB) {
+ if (!Needed.contains(&R))
+ continue;
+
+ auto Remap = [&](VPValue *V) -> VPValue * {
+ VPValue *Mapped = OperandMap.lookup(V);
+ return Mapped ? Mapped : V;
+ };
+
+ SmallVector<VPValue *, 4> NewOps;
+ for (VPValue *Op : R.operands())
+ NewOps.push_back(Remap(Op));
+
+ // Reverse the mask for negative-stride memory ops.
+ auto ReverseIfNeeded = [&](VPValue *Mask, VPValue *Addr,
+ DebugLoc DL) -> VPValue * {
+ if (isa_and_nonnull<VPVectorEndPointerRecipe>(Addr->getDefiningRecipe()))
+ return Builder.createNaryOp(VPInstruction::Reverse, {Mask}, DL,
+ "masked.replay.mask.rev");
+ return Mask;
+ };
+
+ VPRecipeBase *Clone = nullptr;
+ if (auto *WStore = dyn_cast<VPWidenStoreRecipe>(&R)) {
+ VPValue *G = WStore->isMasked() ? Remap(WStore->getMask()) : nullptr;
+ VPValue *L =
+ ReverseIfNeeded(getMaskFor(R), NewOps[0], WStore->getDebugLoc());
+ VPValue *M = G ? Builder.createLogicalAnd(G, L) : L;
+ Clone = new VPWidenStoreRecipe(
+ cast<StoreInst>(WStore->getIngredient()), NewOps[0], NewOps[1], M,
+ WStore->isConsecutive(), *WStore, WStore->getDebugLoc());
+ Builder.insert(Clone);
+ } else if (auto *RepR = dyn_cast<VPReplicateRecipe>(&R);
+ RepR && RepR->getUnderlyingInstr()->mayWriteToMemory()) {
+ VPValue *G = RepR->isPredicated() ? Remap(RepR->getMask()) : nullptr;
+ VPValue *L = getMaskFor(R);
+ VPValue *M = G ? Builder.createLogicalAnd(G, L) : L;
+
+ auto *SI = cast<StoreInst>(RepR->getUnderlyingInstr());
+ VPValue *StoredVal = RepR->getOperand(0);
+ VPValue *PtrVal = RepR->getOperand(1);
+
+ // Widen to a masked vector store when the stored value is a vector.
+ VPValue *ReplayStoredVal = Remap(StoredVal);
+ VPValue *ReplayPtrVal = Remap(PtrVal);
+
+ bool CanWiden = ReplayStoredVal->getDefiningRecipe() &&
+ !isa<VPReplicateRecipe, VPScalarIVStepsRecipe>(
+ ReplayStoredVal->getDefiningRecipe());
+
+ if (CanWiden) {
+ Type *StoreTy = SI->getValueOperand()->getType();
+ const DataLayout &DL = SI->getDataLayout();
+ auto *StrideTy =
+ DL.getIndexType(SI->getPointerOperand()->getType());
+ VPValue *StrideOne = Plan.getConstantInt(StrideTy, 1);
+ auto *VecPtr = new VPVectorPointerRecipe(
+ ReplayPtrVal, StoreTy, StrideOne, GEPNoWrapFlags::none(),
+ SI->getDebugLoc());
+ Builder.insert(VecPtr);
+
+ auto *WS = new VPWidenStoreRecipe(*SI, VecPtr, ReplayStoredVal, M,
+ /*Consecutive=*/true, *RepR,
+ RepR->getDebugLoc());
+ Builder.insert(WS);
+ continue;
+ } else {
+ // Cannot widen: recreate as a predicated replicate.
+ SmallVector<VPValue *, 4> ValOps;
+ for (VPValue *Op : RepR->operandsWithoutMask())
+ ValOps.push_back(Remap(Op));
+ auto *PredStore = new VPReplicateRecipe(
+ RepR->getUnderlyingInstr(), ValOps, RepR->isSingleScalar(), M,
+ *RepR, *RepR, RepR->getDebugLoc());
+ Builder.insert(PredStore);
+ Clone = PredStore;
+ }
+ } else if (auto *WLoad = dyn_cast<VPWidenLoadRecipe>(&R)) {
+ // Mask replay load: inclusive/exclusive mask ANDed with remapped guard.
+ VPValue *G = WLoad->isMasked() ? Remap(WLoad->getMask()) : nullptr;
+ VPValue *ReplayMask;
+ if (InclusiveFeeds.contains(&R)) {
+ if (!InclMask) {
+ VPInstruction *FALPlusOne =
+ VPBuilder::getToInsertAfter(FirstActiveLane)
+ .createAdd(FirstActiveLane, One, DebugLoc(),
+ "first.active.lane.incl",
+ {/*nuw=*/true, /*nsw=*/false});
+ InclMask = VPBuilder::getToInsertAfter(FALPlusOne)
+ .createNaryOp(VPInstruction::ActiveLaneMask,
+ {Zero, FALPlusOne, One}, DebugLoc(),
+ "masked.replay.mask.incl");
+ }
+ ReplayMask = InclMask;
+ } else {
+ if (!ExclMask)
+ ExclMask = VPBuilder::getToInsertAfter(FirstActiveLane)
+ .createNaryOp(VPInstruction::ActiveLaneMask,
+ {Zero, FirstActiveLane, One}, DebugLoc(),
+ "masked.replay.mask");
+ ReplayMask = ExclMask;
+ }
+ ReplayMask = ReverseIfNeeded(ReplayMask, NewOps[0], WLoad->getDebugLoc());
+ VPValue *Mask = G ? Builder.createLogicalAnd(G, ReplayMask) : ReplayMask;
+ Clone = new VPWidenLoadRecipe(
+ cast<LoadInst>(WLoad->getIngredient()), NewOps[0], Mask,
+ WLoad->isConsecutive(), *WLoad, WLoad->getDebugLoc());
+ Builder.insert(Clone);
+ } else if (auto *RepR = dyn_cast<VPReplicateRecipe>(&R);
+ RepR && RepR->isPredicated()) {
+ SmallVector<VPValue *, 4> ValOps;
+ for (VPValue *Op : RepR->operandsWithoutMask())
+ ValOps.push_back(Remap(Op));
+ auto *NewRep = new VPReplicateRecipe(
+ RepR->getUnderlyingInstr(), ValOps, RepR->isSingleScalar(),
+ Remap(RepR->getMask()), *RepR, *RepR, RepR->getDebugLoc());
+ Builder.insert(NewRep);
+ Clone = NewRep;
+ } else {
+ assert(!R.mayWriteToMemory() && !R.mayHaveSideEffects() &&
+ "masked replay reached an unmaskable side-effecting recipe; "
+ "such loops must fall back to scalar replay");
+ Clone = R.clone();
+ for (unsigned I = 0, E = Clone->getNumOperands(); I != E; ++I)
+ Clone->setOperand(I, NewOps[I]);
+ Builder.insert(Clone);
+ }
+
+ for (unsigned I = 0, E = R.getNumDefinedValues(); I != E; ++I)
+ OperandMap[R.getVPValue(I)] = Clone->getVPValue(I);
+ }
+
+ assert(ReplayBB->getNumSuccessors() == 1 &&
+ "masked-replay block must have the single temporary latch edge");
+ Plan.setCheckFirstMaskedReplayTempTarget(ReplayBB->getSuccessors()[0]);
+}
+
+void VPlanTransforms::wireCheckFirstMaskedReplayToExit(VPlan &Plan) {
+ VPBasicBlock *ReplayBB = Plan.getCheckFirstMaskedReplayBlock();
+ if (!ReplayBB)
+ return;
+
+ BasicBlock *EarlyExitIRBB = Plan.getCheckFirstEarlyExitBlock();
+ assert(EarlyExitIRBB && "masked replay requires a captured early-exit block");
+ // Recreate only if it was removed during cloning.
+ VPIRBasicBlock *EarlyExitBB = nullptr;
+ for (VPIRBasicBlock *EB : Plan.getExitBlocks())
+ if (EB->getIRBasicBlock() == EarlyExitIRBB) {
+ EarlyExitBB = EB;
+ break;
+ }
+ if (!EarlyExitBB)
+ EarlyExitBB = Plan.createVPIRBasicBlock(EarlyExitIRBB);
+
+ VPInstruction *FirstActiveLane = nullptr;
+ for (VPRecipeBase &R : *ReplayBB) {
+ if (auto *VPI = dyn_cast<VPInstruction>(&R);
+ VPI && VPI->getOpcode() == VPInstruction::FirstActiveLane) {
+ FirstActiveLane = VPI;
+ break;
+ }
+ }
+ assert(FirstActiveLane && "masked-replay block missing FirstActiveLane");
+
+ assert(ReplayBB->getNumPredecessors() == 1 &&
+ "masked-replay head must have a single (check/header) predecessor");
+ auto *HeaderVPBB = cast<VPBasicBlock>(ReplayBB->getPredecessors()[0]);
+ VPValue *CanonIV = cast<VPPhi>(&*HeaderVPBB->begin());
+ Type *CanonTy = CanonIV->getScalarType();
+
+ // Recover the tail of the chain it is the recorded temp-target's predecessor
+ // that is not the loop header.
+ VPBlockBase *TempTarget = Plan.getCheckFirstMaskedReplayTempTarget();
+ assert(TempTarget && "masked replay temporary-edge target not recorded");
+ VPBasicBlock *TailVPBB = nullptr;
+ for (VPBlockBase *Pred : TempTarget->getPredecessors()) {
+ if (Pred != HeaderVPBB) {
+ assert(!TailVPBB && "expected a single replay-chain predecessor of the "
+ "temporary-edge target");
+ TailVPBB = cast<VPBasicBlock>(Pred);
+ }
+ }
+ assert(TailVPBB && "could not locate the masked-replay chain tail");
+
+ // Replace the temporary tail → latch edge with tail → early-exit.
+ VPBlockUtils::disconnectBlocks(TailVPBB, TempTarget);
+ VPBlockUtils::connectBlocks(TailVPBB, EarlyExitBB);
+
+ // Exiting index = chunk_start + first_active_lane.
+ // Build in the tail so it dominates its uses.
+ VPBuilder Builder(TailVPBB, TailVPBB->end());
+ VPValue *FALCast = Builder.createScalarZExtOrTrunc(
+ FirstActiveLane, CanonTy, FirstActiveLane->getScalarType(), DebugLoc());
+ VPValue *ExitIndex =
+ Builder.createAdd(CanonIV, FALCast, DebugLoc(), "masked.replay.exit.idx",
+ {/*nuw=*/true, /*nsw=*/false});
+
+ // Each live-out is the exiting induction value, cast to the PHI's type.
+ for (VPRecipeBase &R : EarlyExitBB->phis()) {
+ auto *Phi = cast<VPIRPhi>(&R);
+ Type *PhiTy = Phi->getIRPhi().getType();
+ VPValue *LiveOut =
+ Builder.createScalarZExtOrTrunc(ExitIndex, PhiTy, CanonTy, DebugLoc());
+ Phi->addIncoming(LiveOut);
+ }
+}
+
/// This function tries convert extended in-loop reductions to
/// VPExpressionRecipe and clamp the \p Range if it is beneficial and
/// valid. The created recipe must be decomposed to its constituent
@@ -5336,6 +6260,9 @@ void VPlanTransforms::materializeConstantVectorTripCount(
!isa<VPIRValue>(TC))
return;
+ if (Plan.getCheckFirstExitBlock())
+ return;
+
// Materialize vector trip counts for constants early if it can simply
// be computed as (Original TC / VF * UF) * VF * UF.
// TODO: Compute vector trip counts for loops requiring a scalar epilogue and
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
index 3260526552281..513267c2bb41e 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
@@ -366,6 +366,16 @@ struct VPlanTransforms {
VPBasicBlock *MiddleVPBB, Loop *TheLoop, PredicatedScalarEvolution &PSE,
DominatorTree &DT, AssumptionCache *AC, UncountableExitStyle Style);
+ /// Connects check-first early-exit blocks to the scalar preheader.
+ static void wireCheckFirstExitToScalar(VPlan &Plan);
+
+ /// Clones the stores inside loop body into the masked-replay block.
+ static void maskCheckFirstReplayStores(VPlan &Plan);
+
+ /// Wires the masked replay block to the early-exit block and rebuilds its
+ /// live-out.
+ static void wireCheckFirstMaskedReplayToExit(VPlan &Plan);
+
/// Replaces the exit condition from
/// (branch-on-cond eq CanonicalIVInc, VectorTripCount)
/// to
diff --git a/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp b/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
index 362bfe92f573e..2428c718bcaf4 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanVerifier.cpp
@@ -308,6 +308,12 @@ bool VPlanVerifier::verifyVPBasicBlock(const VPBasicBlock *VPBB) {
continue;
}
+ if (VPBasicBlock *CheckExit =
+ VPBB->getPlan()->getCheckFirstExitBlock()) {
+ if (is_contained(CheckExit->getPredecessors(), VPBB))
+ continue;
+ }
+
errs() << "Use before def!\n";
#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
VPSlotTracker Tracker(VPBB->getPlan());
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
index 2a250d4c896ff..d91204aba896f 100644
--- a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
@@ -55,6 +55,7 @@
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] removeBranchOnConst
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] simplifyReverses
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] removeDeadRecipes
+; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] maskCheckFirstReplayStores
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] createAndOptimizeReplicateRegions
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] mergeBlocksIntoPredecessors
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] licm
diff --git a/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll b/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
new file mode 100644
index 0000000000000..1760b673f8abf
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/check-first-multi-exit-cascade.ll
@@ -0,0 +1,133 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -S < %s -p loop-vectorize -force-vector-width=4 -enable-check-first-early-exit-vectorization | FileCheck %s
+
+ at A = global [1024 x i32] zeroinitializer
+ at B = global [1024 x i32] zeroinitializer
+ at D = global [1024 x i32] zeroinitializer
+
+define i32 @multi_exit_cascade() {
+; CHECK-LABEL: define i32 @multi_exit_cascade() {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY4:.*]] ]
+; CHECK-NEXT: [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT: [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT: [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP6:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-NEXT: [[TMP7:%.*]] = freeze <4 x i1> [[TMP6]]
+; CHECK-NEXT: [[TMP8:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP7]])
+; CHECK-NEXT: br i1 [[TMP8]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_CHECK:.*]]
+; CHECK: [[VECTOR_CHECK]]:
+; CHECK-NEXT: [[TMP21:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD4:%.*]] = load <4 x i32>, ptr [[TMP21]], align 4
+; CHECK-NEXT: [[TMP22:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD4]], zeroinitializer
+; CHECK-NEXT: [[TMP9:%.*]] = freeze <4 x i1> [[TMP22]]
+; CHECK-NEXT: [[TMP10:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP9]])
+; CHECK-NEXT: br i1 [[TMP10]], label %[[VECTOR_CHECK_EXIT]], label %[[VECTOR_BODY4]]
+; CHECK: [[VECTOR_BODY4]]:
+; CHECK-NEXT: [[TMP11:%.*]] = add <4 x i32> [[WIDE_LOAD3]], [[WIDE_LOAD4]]
+; CHECK-NEXT: [[TMP12:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT: [[TMP14:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT: [[TMP15:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT: [[TMP16:%.*]] = extractelement <4 x i32> [[TMP11]], i64 0
+; CHECK-NEXT: store i32 [[TMP16]], ptr [[TMP12]], align 4
+; CHECK-NEXT: [[TMP17:%.*]] = extractelement <4 x i32> [[TMP11]], i64 1
+; CHECK-NEXT: store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT: [[TMP18:%.*]] = extractelement <4 x i32> [[TMP11]], i64 2
+; CHECK-NEXT: store i32 [[TMP18]], ptr [[TMP14]], align 4
+; CHECK-NEXT: [[TMP19:%.*]] = extractelement <4 x i32> [[TMP11]], i64 3
+; CHECK-NEXT: store i32 [[TMP19]], ptr [[TMP15]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[TMP20:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT: br i1 [[TMP20]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: br label %[[RET_N:.*]]
+; CHECK: [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
+; CHECK: [[SCALAR_PH]]:
+; CHECK-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK: [[FOR_BODY]]:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ [[INDEX]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[LATCH:.*]] ]
+; CHECK-NEXT: [[GA:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[IV]]
+; CHECK-NEXT: [[LA:%.*]] = load i32, ptr [[GA]], align 4
+; CHECK-NEXT: [[CA:%.*]] = icmp eq i32 [[LA]], 0
+; CHECK-NEXT: br i1 [[CA]], label %[[EXIT0:.*]], label %[[CHECK1:.*]]
+; CHECK: [[CHECK1]]:
+; CHECK-NEXT: [[GB:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[IV]]
+; CHECK-NEXT: [[LB:%.*]] = load i32, ptr [[GB]], align 4
+; CHECK-NEXT: [[CB:%.*]] = icmp eq i32 [[LB]], 0
+; CHECK-NEXT: br i1 [[CB]], label %[[EXIT1:.*]], label %[[LATCH]]
+; CHECK: [[LATCH]]:
+; CHECK-NEXT: [[SUM:%.*]] = add i32 [[LA]], [[LB]]
+; CHECK-NEXT: [[GD:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[IV]]
+; CHECK-NEXT: store i32 [[SUM]], ptr [[GD]], align 4
+; CHECK-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-NEXT: [[DONE:%.*]] = icmp eq i64 [[IV_NEXT]], 1024
+; CHECK-NEXT: br i1 [[DONE]], label %[[RET_N]], label %[[FOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK: [[EXIT0]]:
+; CHECK-NEXT: [[IV_LCSSA:%.*]] = phi i64 [ [[IV]], %[[FOR_BODY]] ]
+; CHECK-NEXT: [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT: br label %[[RET:.*]]
+; CHECK: [[EXIT1]]:
+; CHECK-NEXT: [[IV_LCSSA1:%.*]] = phi i64 [ [[IV]], %[[CHECK1]] ]
+; CHECK-NEXT: [[T1:%.*]] = trunc i64 [[IV_LCSSA1]] to i32
+; CHECK-NEXT: [[NEG:%.*]] = sub i32 0, [[T1]]
+; CHECK-NEXT: br label %[[RET]]
+; CHECK: [[RET_N]]:
+; CHECK-NEXT: br label %[[RET]]
+; CHECK: [[RET]]:
+; CHECK-NEXT: [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT0]] ], [ [[NEG]], %[[EXIT1]] ], [ 1024, %[[RET_N]] ]
+; CHECK-NEXT: ret i32 [[R]]
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+ %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+ %la = load i32, ptr %ga, align 4
+ %ca = icmp eq i32 %la, 0
+ br i1 %ca, label %exit0, label %check1
+
+check1:
+ %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+ %lb = load i32, ptr %gb, align 4
+ %cb = icmp eq i32 %lb, 0
+ br i1 %cb, label %exit1, label %latch
+
+latch:
+ %sum = add i32 %la, %lb
+ %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+ store i32 %sum, ptr %gd, align 4
+ %iv.next = add nuw nsw i64 %iv, 1
+ %done = icmp eq i64 %iv.next, 1024
+ br i1 %done, label %ret.n, label %for.body
+
+exit0:
+ %t0 = trunc i64 %iv to i32
+ br label %ret
+
+exit1:
+ %t1 = trunc i64 %iv to i32
+ %neg = sub i32 0, %t1
+ br label %ret
+
+ret.n:
+ br label %ret
+
+ret:
+ %r = phi i32 [ %t0, %exit0 ], [ %neg, %exit1 ], [ 1024, %ret.n ]
+ ret i32 %r
+}
+;.
+; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
+; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
+; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
+; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META2]], [[META1]]}
+;.
diff --git a/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll b/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll
new file mode 100644
index 0000000000000..428d9cc953425
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/check-first-nested-exit.ll
@@ -0,0 +1,205 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --version 6
+; RUN: opt -S < %s -p loop-vectorize -force-vector-width=4 -enable-check-first-early-exit-vectorization -enable-check-first-masked-replay | FileCheck %s
+
+ at A = global [1024 x i32] zeroinitializer
+ at B = global [1024 x i32] zeroinitializer
+ at C = global [1024 x i32] zeroinitializer
+ at D = global [1024 x i32] zeroinitializer
+
+; for (i) { if (A[i] != 0) { if (B[i] == 0) return i; } D[i] = A[i]; }
+define i32 @single_guard() {
+; CHECK-LABEL: define i32 @single_guard() {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY3:.*]] ]
+; CHECK-NEXT: [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT: [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT: [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP4:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT: [[TMP5:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP5]], align 4
+; CHECK-NEXT: [[TMP6:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD2]], zeroinitializer
+; CHECK-NEXT: [[TMP7:%.*]] = select <4 x i1> [[TMP4]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
+; CHECK-NEXT: [[TMP8:%.*]] = freeze <4 x i1> [[TMP7]]
+; CHECK-NEXT: [[TMP9:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP8]])
+; CHECK-NEXT: br i1 [[TMP9]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_BODY3]]
+; CHECK: [[VECTOR_BODY3]]:
+; CHECK-NEXT: [[TMP10:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[TMP11:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT: [[TMP12:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT: [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT: [[TMP14:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 0
+; CHECK-NEXT: store i32 [[TMP14]], ptr [[TMP10]], align 4
+; CHECK-NEXT: [[TMP15:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 1
+; CHECK-NEXT: store i32 [[TMP15]], ptr [[TMP11]], align 4
+; CHECK-NEXT: [[TMP16:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 2
+; CHECK-NEXT: store i32 [[TMP16]], ptr [[TMP12]], align 4
+; CHECK-NEXT: [[TMP17:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 3
+; CHECK-NEXT: store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[TMP18:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT: br i1 [[TMP18]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: br label %[[RET_N:.*]]
+; CHECK: [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT: [[FIRST_ACTIVE_LANE:%.*]] = call i64 @llvm.experimental.cttz.elts.i64.v4i1(<4 x i1> [[TMP7]], i1 false)
+; CHECK-NEXT: [[MASKED_REPLAY_MASK:%.*]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i64(i64 0, i64 [[FIRST_ACTIVE_LANE]])
+; CHECK-NEXT: [[TMP20:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP20]], <4 x i1> [[MASKED_REPLAY_MASK]])
+; CHECK-NEXT: [[IV_LCSSA:%.*]] = add nuw i64 [[INDEX]], [[FIRST_ACTIVE_LANE]]
+; CHECK-NEXT: br label %[[EXIT1:.*]]
+; CHECK: [[EXIT1]]:
+; CHECK-NEXT: [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT: br label %[[RET:.*]]
+; CHECK: [[RET_N]]:
+; CHECK-NEXT: br label %[[RET]]
+; CHECK: [[RET]]:
+; CHECK-NEXT: [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT1]] ], [ 1024, %[[RET_N]] ]
+; CHECK-NEXT: ret i32 [[R]]
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+ %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+ %la = load i32, ptr %ga, align 4
+ %xg = icmp ne i32 %la, 0
+ br i1 %xg, label %guard, label %latch
+
+guard:
+ %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+ %lb = load i32, ptr %gb, align 4
+ %yb = icmp eq i32 %lb, 0
+ br i1 %yb, label %exit0, label %latch
+
+latch:
+ %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+ store i32 %la, ptr %gd, align 4
+ %iv.next = add nuw nsw i64 %iv, 1
+ %done = icmp eq i64 %iv.next, 1024
+ br i1 %done, label %ret.n, label %for.body
+
+exit0:
+ %t0 = trunc i64 %iv to i32
+ br label %ret
+
+ret.n:
+ br label %ret
+
+ret:
+ %r = phi i32 [ %t0, %exit0 ], [ 1024, %ret.n ]
+ ret i32 %r
+}
+
+; for (i) { if (A[i]!=0) { if (C[i]!=0) { if (B[i]==0) return i; } } D[i]=A[i]; }
+define i32 @two_nested_guards() {
+; CHECK-LABEL: define i32 @two_nested_guards() {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY4:.*]] ]
+; CHECK-NEXT: [[TMP0:%.*]] = add i64 [[INDEX]], 1
+; CHECK-NEXT: [[TMP1:%.*]] = add i64 [[INDEX]], 2
+; CHECK-NEXT: [[TMP2:%.*]] = add i64 [[INDEX]], 3
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP4:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT: [[TMP5:%.*]] = getelementptr inbounds [1024 x i32], ptr @C, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP5]], align 4
+; CHECK-NEXT: [[TMP6:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD2]], zeroinitializer
+; CHECK-NEXT: [[TMP7:%.*]] = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP7]], align 4
+; CHECK-NEXT: [[TMP8:%.*]] = icmp eq <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-NEXT: [[TMP9:%.*]] = select <4 x i1> [[TMP4]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
+; CHECK-NEXT: [[TMP10:%.*]] = select <4 x i1> [[TMP9]], <4 x i1> [[TMP8]], <4 x i1> zeroinitializer
+; CHECK-NEXT: [[TMP11:%.*]] = freeze <4 x i1> [[TMP10]]
+; CHECK-NEXT: [[TMP12:%.*]] = call i1 @llvm.vector.reduce.or.v4i1(<4 x i1> [[TMP11]])
+; CHECK-NEXT: br i1 [[TMP12]], label %[[VECTOR_CHECK_EXIT:.*]], label %[[VECTOR_BODY4]]
+; CHECK: [[VECTOR_BODY4]]:
+; CHECK-NEXT: [[TMP13:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: [[TMP14:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP0]]
+; CHECK-NEXT: [[TMP15:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP1]]
+; CHECK-NEXT: [[TMP16:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[TMP2]]
+; CHECK-NEXT: [[TMP17:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 0
+; CHECK-NEXT: store i32 [[TMP17]], ptr [[TMP13]], align 4
+; CHECK-NEXT: [[TMP18:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 1
+; CHECK-NEXT: store i32 [[TMP18]], ptr [[TMP14]], align 4
+; CHECK-NEXT: [[TMP19:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 2
+; CHECK-NEXT: store i32 [[TMP19]], ptr [[TMP15]], align 4
+; CHECK-NEXT: [[TMP20:%.*]] = extractelement <4 x i32> [[WIDE_LOAD]], i64 3
+; CHECK-NEXT: store i32 [[TMP20]], ptr [[TMP16]], align 4
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[TMP21:%.*]] = icmp eq i64 [[INDEX_NEXT]], 1024
+; CHECK-NEXT: br i1 [[TMP21]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: br label %[[SCALAR_PH:.*]]
+; CHECK: [[VECTOR_CHECK_EXIT]]:
+; CHECK-NEXT: [[FIRST_ACTIVE_LANE:%.*]] = call i64 @llvm.experimental.cttz.elts.i64.v4i1(<4 x i1> [[TMP10]], i1 false)
+; CHECK-NEXT: [[MASKED_REPLAY_MASK:%.*]] = call <4 x i1> @llvm.get.active.lane.mask.v4i1.i64(i64 0, i64 [[FIRST_ACTIVE_LANE]])
+; CHECK-NEXT: [[TMP23:%.*]] = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 [[INDEX]]
+; CHECK-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP23]], <4 x i1> [[MASKED_REPLAY_MASK]])
+; CHECK-NEXT: [[IV_LCSSA:%.*]] = add nuw i64 [[INDEX]], [[FIRST_ACTIVE_LANE]]
+; CHECK-NEXT: br label %[[EXIT1:.*]]
+; CHECK: [[EXIT1]]:
+; CHECK-NEXT: [[T0:%.*]] = trunc i64 [[IV_LCSSA]] to i32
+; CHECK-NEXT: br label %[[RET:.*]]
+; CHECK: [[SCALAR_PH]]:
+; CHECK-NEXT: br label %[[RET]]
+; CHECK: [[RET]]:
+; CHECK-NEXT: [[R:%.*]] = phi i32 [ [[T0]], %[[EXIT1]] ], [ 1024, %[[SCALAR_PH]] ]
+; CHECK-NEXT: ret i32 [[R]]
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+ %ga = getelementptr inbounds [1024 x i32], ptr @A, i64 0, i64 %iv
+ %la = load i32, ptr %ga, align 4
+ %xg = icmp ne i32 %la, 0
+ br i1 %xg, label %g1, label %latch
+
+g1:
+ %gc = getelementptr inbounds [1024 x i32], ptr @C, i64 0, i64 %iv
+ %lc = load i32, ptr %gc, align 4
+ %zg = icmp ne i32 %lc, 0
+ br i1 %zg, label %g2, label %latch
+
+g2:
+ %gb = getelementptr inbounds [1024 x i32], ptr @B, i64 0, i64 %iv
+ %lb = load i32, ptr %gb, align 4
+ %yb = icmp eq i32 %lb, 0
+ br i1 %yb, label %exit0, label %latch
+
+latch:
+ %gd = getelementptr inbounds [1024 x i32], ptr @D, i64 0, i64 %iv
+ store i32 %la, ptr %gd, align 4
+ %iv.next = add nuw nsw i64 %iv, 1
+ %done = icmp eq i64 %iv.next, 1024
+ br i1 %done, label %ret.n, label %for.body
+
+exit0:
+ %t0 = trunc i64 %iv to i32
+ br label %ret
+
+ret.n:
+ br label %ret
+
+ret:
+ %r = phi i32 [ %t0, %exit0 ], [ 1024, %ret.n ]
+ ret i32 %r
+}
+;.
+; CHECK: [[LOOP0]] = distinct !{[[LOOP0]], [[META1:![0-9]+]], [[META2:![0-9]+]]}
+; CHECK: [[META1]] = !{!"llvm.loop.isvectorized", i32 1}
+; CHECK: [[META2]] = !{!"llvm.loop.unroll.runtime.disable"}
+; CHECK: [[LOOP3]] = distinct !{[[LOOP3]], [[META1]], [[META2]]}
+;.
More information about the llvm-commits
mailing list