[llvm] [LoopVectorize] Support vectorization of compressing patterns (PR #214491)
Benjamin Maxwell via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 05:25:21 PDT 2026
https://github.com/MacDue updated https://github.com/llvm/llvm-project/pull/214491
>From 36b2c7948db9cea9f8c67202f7d5c9606043513b Mon Sep 17 00:00:00 2001
From: Sergey Kachkov <sergey.kachkov at syntacore.com>
Date: Wed, 5 Aug 2026 18:04:38 +0000
Subject: [PATCH 1/3] [LoopVectorize] Support vectorization of compressing
patterns in VPlan
RFC link: https://discourse.llvm.org/t/rfc-loop-vectorization-of-compress-store-expand-load-patterns/86442
This adds loop vectorizer support for "compressing" patterns,
for example:
```
int dst_idx = 0;
for (int i = 0; i < n; i++) {
if (cond[i])
dst[dst_idx++] = src[i];
}
```
Can be vectorized with a `llvm.masked.compressstore` as:
```
int dst_idx = 0;
for (int i = 0; i < n; i++) {
%cond = load(%cond) != 0
%src = masked.load(%src[i], %cond)
masked.compressstore(%src, %dst[dst_idx], %cond)
dst_idx += num.active.lanes(%cond)
}
```
and:
```
int src_idx = 0;
for (int i = 0; i < n; i++) {
if (cond[i])
dst[i] = src[src_idx++];
}
```
Can be vectorized with a `llvm.masked.expandload` as:
```
int src_idx = 0;
for (int i = 0; i < n; i++) {
%cond = load(%cond) != 0
%src = masked.expandload(%src[src_idx], %cond)
masked.store(%src, %dst[i], %cond)
src_idx += num.active.lanes(%cond)
}
```
This uses the new `MonotonicDescriptor` to recognize
monotonic/compressing patterns. The phis are mapped to a new
`VPMonotonicPHIRecipe`, this will map to a scalar phi. We only
allow uniform uses of monotonic phis in the loop (e.g., as the pointer
to a compressed load/store).
Compressed loads/stores are recognized with
`LoopVectorizationLegality::isCompressedPtr`. Currently, we only allow
cases where:
- The (monotonic) pointer has a stride equal to the access size
- The memory operation is predicated with the same condition as the increment
This is a continuation Sergey Kachkov's patch (#140723).
There are a number of changes from the initial patch:
- Expandloads/compresstores directly use `VPWidenMemIntrinsic`
- ComputeMonotonicResult is replaced with existing VP instructions
- AArch64, VPlan, and target-agnostic tests have been added
- This style of vectorization if off by default
- The switch can be flipped soon after this patch lands
Tests, fixes, and design rework
Add comment
Add out-of-loop use check
Fixups
Fixups
Fixups
Add test
Remove AArch64 flags
---
llvm/include/llvm/Analysis/VectorUtils.h | 15 +
.../include/llvm/Transforms/Utils/LoopUtils.h | 11 +
.../Vectorize/LoopVectorizationLegality.h | 48 ++
llvm/lib/Analysis/VectorUtils.cpp | 45 +
llvm/lib/IR/IntrinsicInst.cpp | 2 +
llvm/lib/Transforms/Utils/LoopUtils.cpp | 74 ++
.../Vectorize/LoopVectorizationLegality.cpp | 45 +
.../Vectorize/LoopVectorizationPlanner.cpp | 7 +
.../Vectorize/LoopVectorizationPlanner.h | 6 +
.../Transforms/Vectorize/LoopVectorize.cpp | 101 ++-
.../Transforms/Vectorize/VPRecipeBuilder.h | 6 +
llvm/lib/Transforms/Vectorize/VPlan.cpp | 6 +-
llvm/lib/Transforms/Vectorize/VPlan.h | 71 +-
.../Vectorize/VPlanConstruction.cpp | 9 +
.../lib/Transforms/Vectorize/VPlanRecipes.cpp | 37 +-
.../Transforms/Vectorize/VPlanTransforms.cpp | 87 ++
.../Transforms/Vectorize/VPlanTransforms.h | 10 +
llvm/lib/Transforms/Vectorize/VPlanUtils.cpp | 4 +-
.../LoopVectorize/AArch64/compress-idioms.ll | 126 +++
.../LoopVectorize/VPlan/compress-idioms.ll | 153 ++++
.../VPlan/vplan-print-before-after-all.ll | 1 +
.../compress-idioms-negative-tests.ll | 268 ++++++
.../LoopVectorize/compress-idioms.ll | 797 ++++++++++++++++++
.../compress-store-vec-epilogue.ll | 95 +++
.../Transforms/Vectorize/VPlanTestBase.h | 1 +
25 files changed, 1995 insertions(+), 30 deletions(-)
create mode 100644 llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/compress-idioms.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
diff --git a/llvm/include/llvm/Analysis/VectorUtils.h b/llvm/include/llvm/Analysis/VectorUtils.h
index b177d9eec2189..42058ac97e3fa 100644
--- a/llvm/include/llvm/Analysis/VectorUtils.h
+++ b/llvm/include/llvm/Analysis/VectorUtils.h
@@ -165,6 +165,21 @@ LLVM_ABI bool
isVectorIntrinsicWithOverloadTypeAtArg(Intrinsic::ID ID, int OpdIdx,
const TargetTransformInfo *TTI);
+/// Returns the argument index of the pointer parameter for the vector memory
+/// intrinsic \p ID, or `std::nullopt` if the intrinsic does not have a pointer
+/// operand.
+LLVM_ABI std::optional<unsigned>
+getVectorMemoryIntrinsicPointerArgIdx(Intrinsic::ID ID);
+
+/// Returns the argument index of the data value of the vector store intrinsic
+/// \p ID, or `std::nullopt` if the intrinsic does not have a data operand.
+LLVM_ABI std::optional<unsigned>
+getVectorStoreIntrinsicDataArgIdx(Intrinsic::ID ID);
+
+/// Returns the argument index of the mask for the vector intrinsic \p ID, or
+/// `std::nullopt` if the intrinsic does not have a mask operand.
+LLVM_ABI std::optional<unsigned> getVectorIntrinsicMaskArgIdx(Intrinsic::ID ID);
+
/// Identifies if the vector form of the intrinsic that returns a struct is
/// overloaded at the struct element index \p RetIdx. /// \p TTI is used to
/// consider target specific intrinsics, if no target specific intrinsics
diff --git a/llvm/include/llvm/Transforms/Utils/LoopUtils.h b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
index 74c549be35ddf..cab79ddd72685 100644
--- a/llvm/include/llvm/Transforms/Utils/LoopUtils.h
+++ b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
@@ -44,6 +44,8 @@ class TargetLibraryInfo;
class LPPassManager;
class Instruction;
struct RuntimeCheckingPtrGroup;
+class MonotonicDescriptor;
+
typedef std::pair<const RuntimeCheckingPtrGroup *,
const RuntimeCheckingPtrGroup *>
RuntimePointerCheck;
@@ -707,6 +709,15 @@ LLVM_ABI std::optional<IVConditionInfo>
hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
const MemorySSA &MSSA, AAResults &AA);
+/// Collects pointer values (used by loads/stores) whose addresses are derived
+/// from the monotonic PHI described by \p MD. The pointer operands and
+/// approximate SCEV expressions (assuming the monotonic PHI always increments)
+/// for the pointers are placed in \p CompressedPtrs. Returns true if all
+/// in-loop users of the monotonic PHI are loads/stores.
+bool collectCompressedPtrs(DenseMap<Value *, const SCEV *> &CompressedPtrs,
+ const Loop &L, const MonotonicDescriptor &MD,
+ ScalarEvolution &SE);
+
} // end namespace llvm
#endif // LLVM_TRANSFORMS_UTILS_LOOPUTILS_H
diff --git a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
index 7b8b27c6541e1..cd062eef9535c 100644
--- a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
+++ b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
@@ -247,6 +247,13 @@ struct HistogramInfo {
: Load(Load), Update(Update), Store(Store) {}
};
+/// Holds details about a "compressed" pointer: the monotonic PHI used to
+/// derive the pointer and the SCEV expression for the pointer.
+struct CompressedPtrInfo {
+ PHINode *MonotonicPHI;
+ const SCEVAddRecExpr *PtrSCEV;
+};
+
/// Indicates the characteristics of a loop with an uncountable exit.
/// * None -- No uncountable exit present.
/// * ReadOnly -- At least one uncountable exit in a readonly loop.
@@ -287,6 +294,10 @@ class LoopVectorizationLegality {
/// induction descriptor.
using InductionList = MapVector<PHINode *, InductionDescriptor>;
+ /// MonotonicPHIList saves monotonic phi variables and maps them to the
+ /// monotonic phi descriptor.
+ using MonotonicPHIList = MapVector<PHINode *, MonotonicDescriptor>;
+
/// RecurrenceSet contains the phi nodes that are recurrences other than
/// inductions and reductions.
using RecurrenceSet = SmallPtrSet<const PHINode *, 8>;
@@ -330,6 +341,11 @@ class LoopVectorizationLegality {
/// Returns the induction variables found in the loop.
const InductionList &getInductionVars() const { return Inductions; }
+ /// Returns the monotonic phi variables found in the loop.
+ const MonotonicPHIList &getMonotonicPHIs() const { return MonotonicPHIs; }
+
+ bool hasMonotonicPHIs() const { return !MonotonicPHIs.empty(); }
+
/// Return the fixed-order recurrences found in the loop.
RecurrenceSet &getFixedOrderRecurrences() { return FixedOrderRecurrences; }
@@ -475,6 +491,26 @@ class LoopVectorizationLegality {
/// Returns a list of all known histogram operations in the loop.
bool hasHistograms() const { return !Histograms.empty(); }
+ /// Returns the CompressedPtrInfo for \p Ptr if the pointer is defined via
+ /// a monotonic PHI, otherwise std::nullptr.
+ std::optional<CompressedPtrInfo>
+ getCompressedPtrInfo(const Value *Ptr) const {
+ auto It = CompressedPtrs.find(Ptr);
+ if (It != CompressedPtrs.end())
+ return It->second;
+ return std::nullopt;
+ }
+
+ /// Returns the CompressedPtrInfo for \p I if it corresponds to a compressed
+ /// load or store (which can map to an llvm.masked.expandload or
+ /// llvm.masked.compressstore), otherwise std::nullopt.
+ std::optional<CompressedPtrInfo>
+ isCompressedLoadOrStore(const Instruction *I) {
+ if (isa<LoadInst, StoreInst>(I))
+ return getCompressedPtrInfo(getLoadStorePointerOperand(I));
+ return std::nullopt;
+ }
+
PredicatedScalarEvolution *getPredicatedScalarEvolution() const {
return &PSE;
}
@@ -645,6 +681,10 @@ class LoopVectorizationLegality {
/// better choice for the main induction than the existing one.
void addInductionPhi(PHINode *Phi, const InductionDescriptor &ID);
+ /// Adds \p Phi to the monotonic PHI list and collects load/store users of
+ /// the phi. Returns true if all users of \p Phi are legal for vectorization.
+ bool addMonotonicPHI(PHINode *Phi, const MonotonicDescriptor &MD);
+
/// The loop that we evaluate.
Loop *TheLoop;
@@ -689,6 +729,9 @@ class LoopVectorizationLegality {
/// variables can be pointers.
InductionList Inductions;
+ /// Holds all of the monotonic phi variables that we found in the loop.
+ MonotonicPHIList MonotonicPHIs;
+
/// Holds all the casts that participate in the update chain of the induction
/// variables, and that have been proven to be redundant (possibly under a
/// runtime guard). These casts can be ignored when creating the vectorized
@@ -726,6 +769,11 @@ class LoopVectorizationLegality {
/// may work on the same memory location.
SmallVector<HistogramInfo, 1> Histograms;
+ /// Contains all pointers used in the loop that are defined using an index
+ /// derived from a monotonic PHI. Loads/stores to these pointers map to
+ /// expandloads or compressstores.
+ SmallDenseMap<const Value *, CompressedPtrInfo> CompressedPtrs;
+
/// Whether or not creating SCEV predicates is allowed.
bool AllowRuntimeSCEVChecks;
diff --git a/llvm/lib/Analysis/VectorUtils.cpp b/llvm/lib/Analysis/VectorUtils.cpp
index f28fc4afc70ef..f5bdd8654d781 100644
--- a/llvm/lib/Analysis/VectorUtils.cpp
+++ b/llvm/lib/Analysis/VectorUtils.cpp
@@ -155,6 +155,7 @@ bool llvm::isVectorIntrinsicWithScalarOpAtArg(Intrinsic::ID ID,
case Intrinsic::is_fpclass:
case Intrinsic::powi:
case Intrinsic::vector_extract:
+ case Intrinsic::masked_compressstore:
return (ScalarOpdIdx == 1);
case Intrinsic::smul_fix:
case Intrinsic::smul_fix_sat:
@@ -169,6 +170,8 @@ bool llvm::isVectorIntrinsicWithScalarOpAtArg(Intrinsic::ID ID,
return ScalarOpdIdx == 0 || ScalarOpdIdx == 1;
case Intrinsic::experimental_vp_strided_store:
return ScalarOpdIdx == 1 || ScalarOpdIdx == 2;
+ case Intrinsic::masked_expandload:
+ return ScalarOpdIdx == 0;
case Intrinsic::loop_dependence_war_mask:
return true;
default:
@@ -194,6 +197,7 @@ bool llvm::isVectorIntrinsicWithOverloadTypeAtArg(
case Intrinsic::scmp:
case Intrinsic::vector_extract:
case Intrinsic::loop_dependence_war_mask:
+ case Intrinsic::masked_expandload:
return OpdIdx == -1 || OpdIdx == 0;
case Intrinsic::modf:
case Intrinsic::sincos:
@@ -207,11 +211,52 @@ bool llvm::isVectorIntrinsicWithOverloadTypeAtArg(
return OpdIdx == -1 || OpdIdx == 0 || OpdIdx == 1;
case Intrinsic::experimental_vp_strided_store:
return OpdIdx == 0 || OpdIdx == 1 || OpdIdx == 2;
+ case Intrinsic::masked_compressstore:
+ return OpdIdx == 0 || OpdIdx == 1;
default:
return OpdIdx == -1;
}
}
+std::optional<unsigned>
+llvm::getVectorMemoryIntrinsicPointerArgIdx(Intrinsic::ID ID) {
+ if (auto PtrPos = VPIntrinsic::getMemoryPointerParamPos(ID))
+ return PtrPos;
+ switch (ID) {
+ case Intrinsic::masked_compressstore:
+ return 1;
+ case Intrinsic::masked_expandload:
+ return 0;
+ default:
+ return std::nullopt;
+ }
+}
+
+std::optional<unsigned>
+llvm::getVectorStoreIntrinsicDataArgIdx(Intrinsic::ID ID) {
+ if (auto DataPos = VPIntrinsic::getMemoryDataParamPos(ID))
+ return DataPos;
+ switch (ID) {
+ case Intrinsic::masked_expandload:
+ return 2;
+ default:
+ return std::nullopt;
+ }
+}
+
+std::optional<unsigned> llvm::getVectorIntrinsicMaskArgIdx(Intrinsic::ID ID) {
+ if (auto MaskPos = VPIntrinsic::getMaskParamPos(ID))
+ return MaskPos;
+ switch (ID) {
+ case Intrinsic::masked_compressstore:
+ return 2;
+ case Intrinsic::masked_expandload:
+ return 1;
+ default:
+ return std::nullopt;
+ }
+}
+
bool llvm::isVectorIntrinsicWithStructReturnOverloadAtField(
Intrinsic::ID ID, int RetIdx, const TargetTransformInfo *TTI) {
diff --git a/llvm/lib/IR/IntrinsicInst.cpp b/llvm/lib/IR/IntrinsicInst.cpp
index 684aaf1a8f2d3..1997aa3f27b9e 100644
--- a/llvm/lib/IR/IntrinsicInst.cpp
+++ b/llvm/lib/IR/IntrinsicInst.cpp
@@ -439,10 +439,12 @@ VPIntrinsic::getMemoryPointerParamPos(Intrinsic::ID VPID) {
switch (VPID) {
default:
return std::nullopt;
+ case Intrinsic::masked_compressstore:
case Intrinsic::vp_store:
case Intrinsic::vp_scatter:
case Intrinsic::experimental_vp_strided_store:
return 1;
+ case Intrinsic::masked_expandload:
case Intrinsic::vp_load:
case Intrinsic::vp_load_ff:
case Intrinsic::vp_gather:
diff --git a/llvm/lib/Transforms/Utils/LoopUtils.cpp b/llvm/lib/Transforms/Utils/LoopUtils.cpp
index a2e544801b9c4..4715cb14d941d 100644
--- a/llvm/lib/Transforms/Utils/LoopUtils.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUtils.cpp
@@ -21,6 +21,7 @@
#include "llvm/Analysis/BasicAliasAnalysis.h"
#include "llvm/Analysis/DomTreeUpdater.h"
#include "llvm/Analysis/GlobalsModRef.h"
+#include "llvm/Analysis/IVDescriptors.h"
#include "llvm/Analysis/InstSimplifyFolder.h"
#include "llvm/Analysis/LoopAccessAnalysis.h"
#include "llvm/Analysis/LoopInfo.h"
@@ -2537,3 +2538,76 @@ llvm::hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
return {};
}
+
+bool llvm::collectCompressedPtrs(
+ DenseMap<Value *, const SCEV *> &CompressedPtrs, const Loop &L,
+ const MonotonicDescriptor &MD, ScalarEvolution &SE) {
+ // Over-approximates the monotonic PHI as a SCEVAddRec assuming the condition
+ // is always true.
+ const SCEV *ApproximatePhiSCEV = SE.getAddRecExpr(
+ MD.getStartSCEV(), MD.getStepSCEV(), &L, SCEV::FlagAnyWrap);
+
+ // TODO: Take into account the non-wrap flags of the MD when rewriting the
+ // SCEV expressions for pointers. This should allow folding away zext/sext
+ // operations.
+ ValueToSCEVMapTy PhiMap{{MD.getHeaderPHI(), ApproximatePhiSCEV}};
+
+ auto GetCompressedPtrSCEV = [&](Value *Ptr, Type *AccessTy) -> const SCEV * {
+ const SCEV *PtrSCEV =
+ SCEVParameterRewriter::rewrite(SE.getSCEV(Ptr), SE, PhiMap);
+ auto *AddRec = dyn_cast<SCEVAddRecExpr>(PtrSCEV);
+ if (!AddRec || !AddRec->isAffine())
+ return nullptr;
+
+ // Check if pointer step equals access size.
+ SCEVUse Step = AddRec->getStepRecurrence(SE);
+ if (Step != SE.getSizeOfExpr(Step->getType(), AccessTy))
+ return nullptr;
+
+ return PtrSCEV;
+ };
+
+ SmallPtrSet<Use *, 16> Seen;
+ SmallVector<Use *> Worklist(make_pointer_range(MD.getHeaderPHI()->uses()));
+ while (!Worklist.empty()) {
+ Use *U = Worklist.pop_back_val();
+ if (!Seen.insert(U).second)
+ continue;
+
+ // Always allow uses outside the loop or by the backedge update.
+ auto *I = cast<Instruction>(U->getUser());
+ if (I == MD.getBackedgePHI() || !L.contains(I))
+ continue;
+
+ Value *CurrentVal = U->get();
+ if (isa<LoadInst, StoreInst>(I)) {
+ // Disallow any store that uses the monotonic value as the stored value.
+ auto *SI = dyn_cast<StoreInst>(I);
+ if (SI && SI->getValueOperand() == CurrentVal)
+ return false;
+
+ Value *Ptr = getLoadStorePointerOperand(I);
+ const SCEV *PrtSCEV = GetCompressedPtrSCEV(Ptr, getLoadStoreType(I));
+ if (!PrtSCEV)
+ return false;
+ CompressedPtrs.insert({Ptr, PrtSCEV});
+ continue;
+ }
+
+ auto LoopVariantOp = [&](Value *V, bool /*AllowRepeats*/) -> Value * {
+ return L.isLoopInvariant(V) ? nullptr : V;
+ };
+
+ // Non-memory users may use any opcode (select/and/or/etc.), but they must
+ // only have CurrentVal as their only loop-varying input. That prevents
+ // mixing in a second loop-varying term. GetCompressedPtrSCEV rewrites the
+ // full leaf pointer SCEV and rejects it unless the entire address still
+ // simplifies to the required affine AddRec.
+ if (I->use_empty() ||
+ find_singleton<Value>(I->operands(), LoopVariantOp) != CurrentVal)
+ return false;
+ append_range(Worklist, make_pointer_range(I->uses()));
+ }
+
+ return true;
+}
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
index 0c080d9434ea8..8db26d29ed6dc 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
@@ -48,6 +48,10 @@ AllowStridedPointerIVs("lv-strided-pointer-ivs", cl::init(false), cl::Hidden,
cl::desc("Enable recognition of non-constant strided "
"pointer induction variables."));
+static cl::opt<bool> EnableMonotonicPatterns(
+ "lv-monotonic-patterns", cl::init(false), cl::Hidden,
+ cl::desc("Enable recognition of monotonic patterns."));
+
static cl::opt<bool>
HintsAllowReordering("hints-allow-reordering", cl::init(true), cl::Hidden,
cl::desc("Allow enabling loop hints to reorder "
@@ -463,6 +467,13 @@ int LoopVectorizationLegality::isConsecutivePtr(Type *AccessTy,
const auto &Strides = LAI && AllowRuntimeSCEVChecks
? LAI->getSymbolicStrides()
: SymbolicStrideMap();
+
+ // Check if the pointer is derived from a a monotonic PHI. If so, return a
+ // conservative stride (assuming the PHI is always updated).
+ if (std::optional<CompressedPtrInfo> PtrInfo = getCompressedPtrInfo(Ptr))
+ return getStrideFromAddRec(PtrInfo->PtrSCEV, TheLoop, AccessTy, Ptr, PSE)
+ .value_or(0);
+
SmallVector<const SCEVPredicate *> Predicates;
int Stride = getPtrStride(PSE, AccessTy, Ptr, TheLoop, *DT, Strides, false,
AllowRuntimeSCEVChecks ? &Predicates : nullptr)
@@ -746,6 +757,34 @@ void LoopVectorizationLegality::addInductionPhi(PHINode *Phi,
LLVM_DEBUG(dbgs() << "LV: Found an induction variable.\n");
}
+bool LoopVectorizationLegality::addMonotonicPHI(PHINode *Phi,
+ const MonotonicDescriptor &MD) {
+ for (User *U : Phi->users()) {
+ if (!TheLoop->contains(cast<Instruction>(U))) {
+ reportVectorizationFailure(
+ "Unsupported out-of-loop user of monotonic phi",
+ "UnsupportedMonotonicUse", ORE, TheLoop);
+ return false;
+ }
+ }
+
+ MonotonicPHIs[Phi] = MD;
+ DenseMap<Value *, const SCEV *> CompressedPtrsForMD;
+ if (!collectCompressedPtrs(CompressedPtrsForMD, *TheLoop, MD, *PSE.getSE())) {
+ reportVectorizationFailure("Unsupported user of monotonic phi in loop",
+ "UnsupportedMonotonicUse", ORE, TheLoop);
+ return false;
+ }
+
+ for (auto [Ptr, PtrSCEV] : CompressedPtrsForMD) {
+ auto *PtrAddRec = cast<SCEVAddRecExpr>(PtrSCEV);
+ assert(PtrAddRec->isAffine() && "Expected affine SCEVAddRecExpr");
+ CompressedPtrs[Ptr] = CompressedPtrInfo{Phi, PtrAddRec};
+ }
+
+ return true;
+}
+
bool LoopVectorizationLegality::setupOuterLoopInductions() {
BasicBlock *Header = TheLoop->getHeader();
@@ -905,6 +944,12 @@ bool LoopVectorizationLegality::canVectorizeInstr(Instruction &I) {
return true;
}
+ MonotonicDescriptor MD;
+ if (EnableMonotonicPatterns &&
+ MonotonicDescriptor::isMonotonicPHI(Phi, TheLoop, MD, *PSE.getSE())) {
+ return addMonotonicPHI(Phi, MD);
+ }
+
if (RecurrenceDescriptor::isFixedOrderRecurrence(Phi, TheLoop, DT)) {
FixedOrderRecurrences.insert(Phi);
return true;
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
index 17ea2932ddd50..196649cf8198e 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
@@ -159,6 +159,13 @@ bool VFSelectionContext::isLegalGatherOrScatter(bool IsLoad, Type *ScalarTy,
: TTI.isLegalMaskedScatter(VectorTy, Alignment));
}
+bool VFSelectionContext::isLegalExpandLoadOrCompressStore(
+ bool IsLoad, Type *ScalarTy, Align Alignment) const {
+ return ForceTargetSupportsMaskedMemoryOps ||
+ (IsLoad ? TTI.isLegalMaskedExpandLoad(ScalarTy, Alignment)
+ : TTI.isLegalMaskedCompressStore(ScalarTy, Alignment));
+}
+
bool VFSelectionContext::supportsScalableVectors() const {
return TTI.supportsScalableVectors() || ForceTargetSupportsScalableVectors ||
VectorizerParams::VectorizationFactor.isScalable();
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
index 3d886e723ee90..06ac16c3d01a2 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
@@ -810,6 +810,12 @@ class VFSelectionContext {
bool isLegalGatherOrScatter(bool IsLoad, Type *ScalarTy, Align Alignment,
ElementCount VF) const;
+ /// Returns true if the target machine supports a masked expand load (if \p
+ /// IsLoad) or masked compress store of scalar type \p ScalarTy with \p
+ /// Alignment.
+ bool isLegalExpandLoadOrCompressStore(bool IsLoad, Type *ScalarTy,
+ Align Alignment) const;
+
/// Split reductions into those that happen in the loop, and those that
/// happen outside. In-loop reductions are collected into InLoopReductions.
/// InLoopReductionImmediateChains is filled with each in-loop reduction
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 5c4e20ee66b65..9d3614d3b36c8 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -1064,6 +1064,10 @@ class LoopVectorizationCostModel {
/// data type and alignment.
bool isLegalGatherOrScatter(Instruction *I, ElementCount VF) const;
+ /// Returns true if the target machine supports a masked expand load or masked
+ /// compress store for \p I's data type and alignment.
+ bool isLegalExpandLoadOrCompressStore(Instruction *I) const;
+
/// Check if \p Instr belongs to any interleaved access group.
bool isAccessInterleaved(Instruction *Instr) const {
return InterleaveInfo.isInterleaved(Instr);
@@ -2383,6 +2387,13 @@ bool LoopVectorizationCostModel::isLegalGatherOrScatter(Instruction *I,
getLoadStoreAlignment(I), VF);
}
+bool LoopVectorizationCostModel::isLegalExpandLoadOrCompressStore(
+ Instruction *I) const {
+ assert(isa<LoadInst>(I) || isa<StoreInst>(I));
+ return Config.isLegalExpandLoadOrCompressStore(
+ isa<LoadInst>(I), getLoadStoreType(I), getLoadStoreAlignment(I));
+}
+
bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
ElementCount VF) {
if (!isPredicatedInst(I))
@@ -2403,9 +2414,13 @@ bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
}
case Instruction::Load:
case Instruction::Store: {
- bool IsConsecutive = Legal->isConsecutivePtr(getLoadStoreType(I),
- getLoadStorePointerOperand(I));
- return !(IsConsecutive && isLegalMaskedLoadOrStore(I, VF)) &&
+ Type *ScalarTy = getLoadStoreType(I);
+ Value *Ptr = getLoadStorePointerOperand(I);
+ bool IsCompressed = Legal->getCompressedPtrInfo(Ptr).has_value();
+ bool IsConsecutive = Legal->isConsecutivePtr(ScalarTy, Ptr);
+ return !(IsConsecutive && !IsCompressed &&
+ isLegalMaskedLoadOrStore(I, VF)) &&
+ !(IsCompressed && isLegalExpandLoadOrCompressStore(I)) &&
!isLegalGatherOrScatter(I, VF);
}
case Instruction::UDiv:
@@ -2644,8 +2659,11 @@ LoopVectorizationCostModel::memoryInstructionCanBeWidened(Instruction *I,
auto *Ptr = getLoadStorePointerOperand(I);
auto *ScalarTy = getLoadStoreType(I);
- // In order to be widened, the pointer should be consecutive, first of all.
+ // In order to be widened, the pointer should be consecutive or compressed.
int Stride = Legal->isConsecutivePtr(ScalarTy, Ptr);
+ assert((Stride == 1 || !Legal->isCompressedLoadOrStore(I)) &&
+ "Compressed memory ops must be consecutive");
+
if (!Stride)
return std::nullopt;
@@ -3273,6 +3291,7 @@ static bool willGenerateVectors(VPlan &Plan, ElementCount VF,
case VPRecipeBase::VPExpandSCEVSC:
case VPRecipeBase::VPPredInstPHISC:
case VPRecipeBase::VPBranchOnMaskSC:
+ case VPRecipeBase::VPMonotonicPHISC:
continue;
case VPRecipeBase::VPReductionSC:
case VPRecipeBase::VPActiveLaneMaskPHISC:
@@ -3670,6 +3689,10 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
if (Plan.hasEarlyExit())
return 1;
+ // Monotonic vars don't support interleaving.
+ if (Legal->hasMonotonicPHIs())
+ return 1;
+
const bool HasReductions =
any_of(Plan.getVectorLoopRegion()->getEntryBasicBlock()->phis(),
IsaPred<VPReductionPHIRecipe>);
@@ -4029,8 +4052,10 @@ void LoopVectorizationCostModel::collectInstsToScalarize(ElementCount VF) {
// of the instruction.
// 2. Scalable VF, as that would lead to invalid scalarization costs.
// 3. Emulated masked memrefs, if a hacked cost is needed.
+ // 4. Compressed loads/stores (which do not support scalarization)
if (!isScalarAfterVectorization(&I, VF) && !VF.isScalable() &&
!useEmulatedMaskMemRefHack(&I, VF) &&
+ !Legal->isCompressedLoadOrStore(&I) &&
computePredInstDiscount(&I, ScalarCosts, VF) >= 0) {
for (const auto &[I, IC] : ScalarCosts)
ScalarCostsVF.insert({I, IC});
@@ -4289,9 +4314,14 @@ InstructionCost LoopVectorizationCostModel::getConsecutiveMemOpCost(
const Align Alignment = getLoadStoreAlignment(I);
InstructionCost Cost = 0;
if (isMaskRequired(I)) {
- unsigned IID = I->getOpcode() == Instruction::Load
- ? Intrinsic::masked_load
- : Intrinsic::masked_store;
+ Intrinsic::ID LoadIID = Intrinsic::masked_load;
+ Intrinsic::ID StoreIID = Intrinsic::masked_store;
+ if (Legal->isCompressedLoadOrStore(I)) {
+ LoadIID = Intrinsic::masked_expandload;
+ StoreIID = Intrinsic::masked_compressstore;
+ }
+
+ unsigned IID = I->getOpcode() == Instruction::Load ? LoadIID : StoreIID;
Cost += TTI.getMemIntrinsicInstrCost(
MemIntrinsicCostAttributes(IID, VectorTy, Alignment, AS),
Config.CostKind);
@@ -5071,6 +5101,9 @@ LoopVectorizationCostModel::getInstructionCost(Instruction *I,
return TTI::CastContextHint::Interleave;
case LoopVectorizationCostModel::CM_Scalarize:
case LoopVectorizationCostModel::CM_Widen:
+ // TODO: Add 'Compressed' hint (not needed for any targets yet).
+ if (Legal->isCompressedLoadOrStore(I))
+ return TTI::CastContextHint::None;
return isPredicatedInst(I) ? TTI::CastContextHint::Masked
: TTI::CastContextHint::Normal;
case LoopVectorizationCostModel::CM_Widen_Reverse:
@@ -5986,6 +6019,7 @@ VPRecipeBase *VPRecipeBuilder::tryToWidenMemory(VPInstruction *VPI,
// reverse consecutive.
LoopVectorizationCostModel::InstWidening Decision =
CM.getWideningDecision(I, Range.Start);
+
bool Reverse = Decision == LoopVectorizationCostModel::CM_Widen_Reverse;
bool Consecutive =
Reverse || Decision == LoopVectorizationCostModel::CM_Widen;
@@ -6143,6 +6177,39 @@ VPHistogramRecipe *VPRecipeBuilder::widenIfHistogram(VPInstruction *VPI) {
VPI->getDebugLoc());
}
+VPWidenMemIntrinsicRecipe *
+VPRecipeBuilder::widenIfCompressedLoadOrStore(VPInstruction *VPI,
+ VPMonotonicPHIRecipe *PhiR) {
+ Instruction *I = VPI->getUnderlyingInstr();
+
+ std::optional<CompressedPtrInfo> Info = Legal->isCompressedLoadOrStore(I);
+ if (!Info || Info->MonotonicPHI != PhiR->getPHINode())
+ return nullptr;
+
+ VPBuilder::InsertPointGuard Guard(Builder);
+ Builder.setInsertPoint(VPI);
+
+ VPValue *Mask = VPI->getMask();
+ Type *AccessTy = getLoadStoreType(I);
+ Align Alignment = getLoadStoreAlignment(I);
+
+ VPValue *Ptr = VPI->getOpcode() == Instruction::Load ? VPI->getOperand(0)
+ : VPI->getOperand(1);
+ Ptr = Builder.createConsecutiveVectorPointer(Ptr, AccessTy,
+ /*Reverse=*/false,
+ VPI->getDebugLoc());
+
+ if (VPI->getOpcode() == Instruction::Load)
+ return new VPWidenMemIntrinsicRecipe(
+ Intrinsic::masked_expandload, {Ptr, Mask, Plan.getPoison(AccessTy)},
+ AccessTy, Alignment, *VPI, I->getDebugLoc());
+
+ VPValue *StoredValue = VPI->getOperand(0);
+ return new VPWidenMemIntrinsicRecipe(Intrinsic::masked_compressstore,
+ {StoredValue, Ptr, Mask}, AccessTy,
+ Alignment, *VPI, I->getDebugLoc());
+}
+
bool VPRecipeBuilder::replaceWithFinalIfReductionStore(
VPInstruction *VPI, VPBuilder &FinalRedStoresBuilder) {
StoreInst *SI;
@@ -6402,8 +6469,8 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
if (!RUN_VPLAN_PASS(
VPlanTransforms::createHeaderPhiRecipes, *VPlan0, PSE, *OrigLoop,
VPDT, Legal->getInductionVars(), Legal->getReductionVars(),
- Legal->getFixedOrderRecurrences(), Config.getInLoopReductions(),
- Config.getHints().allowReordering())) {
+ Legal->getMonotonicPHIs(), Legal->getFixedOrderRecurrences(),
+ Config.getInLoopReductions(), Config.getHints().allowReordering())) {
return nullptr;
}
@@ -6603,6 +6670,10 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan(VPlanPtr Plan,
ReversePostOrderTraversal<VPBlockShallowTraversalWrapper<VPBlockBase *>> RPOT(
HeaderVPBB);
+ if (!RUN_VPLAN_PASS(VPlanTransforms::handleCompressingPatterns, *Plan,
+ HeaderVPBB, RecipeBuilder))
+ return nullptr;
+
RUN_VPLAN_PASS(VPlanTransforms::createInLoopReductionRecipes, *Plan,
Range.Start);
@@ -7501,7 +7572,8 @@ static SmallVector<Instruction *> preparePlanForEpilogueVectorLoop(
}
} else {
// Retrieve the induction resume value via ResumeForEpilogue.
- PHINode *IndPhi = cast<VPWidenInductionRecipe>(&R)->getPHINode();
+ assert(isa<VPWidenInductionRecipe>(&R) || isa<VPMonotonicPHIRecipe>(&R));
+ PHINode *IndPhi = cast<VPHeaderPHIRecipe>(&R)->getPHINode();
ResumeV = IRPhiToResumeForEpi.at(IndPhi)->getUnderlyingValue();
}
assert(ResumeV && "Must have a resume value");
@@ -7910,6 +7982,15 @@ bool LoopVectorizePass::processLoop(Loop *L) {
IC = LVP.selectInterleaveCount(*BestPlanPtr, VF.Width, VF.Cost);
unsigned SelectedIC = std::max(IC, UserIC);
+
+ if (LVL.hasMonotonicPHIs() && SelectedIC > 1) {
+ reportVectorizationFailure(
+ "Interleaving of loop with monotonic vars",
+ "Interleaving of loops with monotonic vars is not supported",
+ "CantInterleaveWithMonotonicVars", ORE, L);
+ return false;
+ }
+
// Optimistically generate runtime checks if they are needed. Drop them if
// they turn out to not be profitable.
if (VF.Width.isVector() || SelectedIC > 1) {
diff --git a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
index 1303b62d5faf4..e4cef6f372732 100644
--- a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
+++ b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
@@ -78,6 +78,12 @@ class VPRecipeBuilder {
/// scalar loop.
VPHistogramRecipe *widenIfHistogram(VPInstruction *VPI);
+ /// If \p VPI represents a compressed load or store (as determined by
+ /// LoopVectorizationLegality) whose pointer is derived from \p PhiR, lower it
+ /// to a llvm.masked.expandload or llvm.masked.compressstore intrinsic.
+ VPWidenMemIntrinsicRecipe *
+ widenIfCompressedLoadOrStore(VPInstruction *VPI, VPMonotonicPHIRecipe *PhiR);
+
/// If \p VPI is a store of a reduction into an invariant address, delete it.
/// If it is the final store of a reduction result, a uniform store recipe
/// will be created for it in the middle block. Returns `true` if replacement
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.cpp b/llvm/lib/Transforms/Vectorize/VPlan.cpp
index f3bdf910a6e7b..0cb7535187d18 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlan.cpp
@@ -370,9 +370,9 @@ void VPTransformState::fixupHeaderPhis() {
for (VPRecipeBase &R : Header->phis()) {
auto *PhiR = cast<VPSingleDefRecipe>(&R);
- bool NeedsScalar =
- isa<VPPhi>(PhiR) || (isa<VPReductionPHIRecipe>(PhiR) &&
- cast<VPReductionPHIRecipe>(PhiR)->isInLoop());
+ bool NeedsScalar = isa<VPPhi>(PhiR) || isa<VPMonotonicPHIRecipe>(PhiR) ||
+ (isa<VPReductionPHIRecipe>(PhiR) &&
+ cast<VPReductionPHIRecipe>(PhiR)->isInLoop());
Value *Phi = get(PhiR, NeedsScalar);
Value *Val = get(PhiR->getOperand(1), NeedsScalar);
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index daabca902724b..016606a2481f2 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -462,12 +462,13 @@ class LLVM_ABI_FOR_TEST VPRecipeBase
VPWidenIntOrFpInductionSC,
VPWidenPointerInductionSC,
VPReductionPHISC,
+ VPMonotonicPHISC,
// END: SubclassID for recipes that inherit VPHeaderPHIRecipe
// END: Phi-like recipes
VPFirstPHISC = VPWidenPHISC,
VPFirstHeaderPHISC = VPCurrentIterationPHISC,
- VPLastHeaderPHISC = VPReductionPHISC,
- VPLastPHISC = VPReductionPHISC,
+ VPLastHeaderPHISC = VPMonotonicPHISC,
+ VPLastPHISC = VPMonotonicPHISC,
};
VPRecipeBase(VPRecipeTy SC, ArrayRef<VPValue *> Operands,
@@ -660,6 +661,7 @@ class LLVM_ABI_FOR_TEST VPSingleDefRecipe : public VPRecipeBase,
case VPRecipeBase::VPReductionPHISC:
case VPRecipeBase::VPWidenLoadEVLSC:
case VPRecipeBase::VPWidenLoadSC:
+ case VPRecipeBase::VPMonotonicPHISC:
return true;
case VPRecipeBase::VPBranchOnMaskSC:
case VPRecipeBase::VPInterleaveEVLSC:
@@ -2075,7 +2077,9 @@ class VPWidenMemIntrinsicRecipe final : public VPWidenIntrinsicRecipe {
DL),
Alignment(Alignment) {
assert((VectorIntrinsicID == Intrinsic::experimental_vp_strided_load ||
- VectorIntrinsicID == Intrinsic::experimental_vp_strided_store) &&
+ VectorIntrinsicID == Intrinsic::experimental_vp_strided_store ||
+ VectorIntrinsicID == Intrinsic::masked_compressstore ||
+ VectorIntrinsicID == Intrinsic::masked_expandload) &&
"Unexpected intrinsic");
}
@@ -2092,6 +2096,9 @@ class VPWidenMemIntrinsicRecipe final : public VPWidenIntrinsicRecipe {
/// Produce a widened version of the vector memory intrinsic.
void execute(VPTransformState &State) override;
+ /// Returns the mask of a predicated VPWidenMemIntrinsicRecipe.
+ VPValue *getMask() const;
+
/// Helper function for computing the cost of vector memory intrinsic.
static InstructionCost computeMemIntrinsicCost(Intrinsic::ID IID, Type *Ty,
bool IsMasked, Align Alignment,
@@ -2504,6 +2511,11 @@ class LLVM_ABI_FOR_TEST VPHeaderPHIRecipe : public VPSingleDefRecipe,
VPUser::addOperand(V);
}
+ /// Returns the underlying PHINode if one exists, or null otherwise.
+ PHINode *getPHINode() const {
+ return cast_if_present<PHINode>(getUnderlyingValue());
+ }
+
protected:
#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
/// Print the recipe.
@@ -2577,11 +2589,6 @@ class VPWidenInductionRecipe : public VPHeaderPHIRecipe {
/// incoming value, its start value.
unsigned getNumIncoming() const override { return 1; }
- /// Returns the underlying PHINode if one exists, or null otherwise.
- PHINode *getPHINode() const {
- return cast_if_present<PHINode>(getUnderlyingValue());
- }
-
/// Returns the induction descriptor for the recipe.
const InductionDescriptor &getInductionDescriptor() const { return IndDesc; }
@@ -2952,6 +2959,51 @@ class VPReductionPHIRecipe : public VPHeaderPHIRecipe, public VPIRFlags {
#endif
};
+/// A recipe for handling monotonic phis. The start value is the first operand
+/// of the recipe, the incoming value from the backedge is the second
+/// operand, and the third operand is the step.
+class VPMonotonicPHIRecipe : public VPHeaderPHIRecipe {
+public:
+ VPMonotonicPHIRecipe(PHINode &Phi, VPValue &Start, VPValue &BackedgeValue,
+ VPValue &Step)
+ : VPHeaderPHIRecipe(VPRecipeBase::VPMonotonicPHISC, &Phi, &Start) {
+ addOperand(&BackedgeValue);
+ addOperand(&Step);
+ }
+
+ VPValue *getStep() const { return getOperand(2); }
+
+ unsigned getNumIncoming() const override { return 2; }
+
+ ~VPMonotonicPHIRecipe() override = default;
+
+ VPMonotonicPHIRecipe *clone() override {
+ return new VPMonotonicPHIRecipe(*getPHINode(), *getStartValue(),
+ *getBackedgeValue(), *getStep());
+ }
+
+ VP_CLASSOF_IMPL(VPRecipeBase::VPMonotonicPHISC)
+
+ static inline bool classof(const VPHeaderPHIRecipe *R) {
+ return R->getVPRecipeID() == VPRecipeBase::VPMonotonicPHISC;
+ }
+
+ void execute(VPTransformState &State) override;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+ /// Print the recipe.
+ void printRecipe(raw_ostream &O, const Twine &Indent,
+ VPSlotTracker &SlotTracker) const override;
+#endif
+
+ /// Returns true if the recipe only uses the first lane of operand \p Op.
+ bool usesFirstLaneOnly(const VPValue *Op) const override {
+ assert(is_contained(operands(), Op) &&
+ "Op must be an operand of the recipe");
+ return true;
+ }
+};
+
/// A recipe for vectorizing a phi-node as a sequence of mask-based select
/// instructions.
class LLVM_ABI_FOR_TEST VPBlendRecipe : public VPRecipeWithIRFlags {
@@ -4357,7 +4409,8 @@ struct CastInfoMixinImpl
template <>
struct CastInfo<VPPhiAccessors, VPRecipeBase *>
: vpdetail::CastInfoMixinImpl<VPPhiAccessors, VPPhi, VPIRPhi,
- VPWidenPHIRecipe, VPHeaderPHIRecipe> {};
+ VPWidenPHIRecipe, VPHeaderPHIRecipe,
+ VPMonotonicPHIRecipe> {};
template <>
struct CastInfo<VPPhiAccessors, const VPRecipeBase *>
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index 4cd455adb7a90..e5c932f0ba716 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -928,6 +928,7 @@ bool VPlanTransforms::createHeaderPhiRecipes(
const VPDominatorTree &VPDT,
const MapVector<PHINode *, InductionDescriptor> &Inductions,
const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
+ const MapVector<PHINode *, MonotonicDescriptor> &MonotonicPHIs,
const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering) {
// Retrieve the header manually from the intial plain-CFG VPlan.
@@ -960,6 +961,14 @@ bool VPlanTransforms::createHeaderPhiRecipes(
Plan, PSE, OrigLoop,
PhiR->getDebugLoc());
+ auto MonotonicIt = MonotonicPHIs.find(Phi);
+ if (MonotonicIt != MonotonicPHIs.end()) {
+ const MonotonicDescriptor &MD = MonotonicIt->second;
+ VPValue *Step =
+ vputils::getOrCreateVPValueForSCEVExpr(Plan, MD.getStepSCEV());
+ return new VPMonotonicPHIRecipe(*Phi, *Start, *BackedgeValue, *Step);
+ }
+
assert(Reductions.contains(Phi) && "only reductions are expected now");
const RecurrenceDescriptor &RdxDesc = Reductions.lookup(Phi);
assert(RdxDesc.getRecurrenceStartValue() ==
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index 5997b6e0fef76..a1ed8ac5791cf 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -151,6 +151,7 @@ bool VPRecipeBase::mayReadFromMemory() const {
case VPWidenStoreEVLSC:
case VPWidenStoreSC:
case VPExpandSCEVSC:
+ case VPMonotonicPHISC:
return false;
case VPBlendSC:
case VPReductionEVLSC:
@@ -1494,6 +1495,12 @@ InstructionCost VPInstruction::computeCost(ElementCount VF,
VectorTy, Ctx.CostKind, /*Mask=*/{},
/*Index=*/0);
}
+ case VPInstruction::NumActiveLanes: {
+ Type *ElementTy = getOperand(0)->getScalarType();
+ auto *VectorTy = cast<VectorType>(toVectorTy(ElementTy, VF));
+ return Ctx.TTI.getArithmeticReductionCost(Instruction::Add, VectorTy,
+ std::nullopt, Ctx.CostKind);
+ }
case VPInstruction::ExtractLastLane: {
// Add on the cost of extracting the element.
auto *VecTy = toVectorTy(getOperand(0)->getScalarType(), VF);
@@ -2426,7 +2433,7 @@ void VPWidenIntrinsicRecipe::printRecipe(raw_ostream &O, const Twine &Indent,
void VPWidenMemIntrinsicRecipe::execute(VPTransformState &State) {
CallInst *MemI = createVectorCall(State);
- auto PtrPos = VPIntrinsic::getMemoryPointerParamPos(getVectorIntrinsicID());
+ auto PtrPos = getVectorMemoryIntrinsicPointerArgIdx(getVectorIntrinsicID());
assert(PtrPos && "Expected a memory intrinsic with a valid pointer position");
MemI->addParamAttr(
*PtrPos, Attribute::getWithAlignment(MemI->getContext(), Alignment));
@@ -2434,6 +2441,12 @@ void VPWidenMemIntrinsicRecipe::execute(VPTransformState &State) {
State.set(this, MemI);
}
+VPValue *VPWidenMemIntrinsicRecipe::getMask() const {
+ auto MaskPos = getVectorIntrinsicMaskArgIdx(getVectorIntrinsicID());
+ assert(MaskPos && "Expected a memory intrinsic with a valid mask position");
+ return getOperand(*MaskPos);
+}
+
InstructionCost VPWidenMemIntrinsicRecipe::computeMemIntrinsicCost(
Intrinsic::ID IID, Type *Ty, bool IsMasked, Align Alignment,
VPCostContext &Ctx) {
@@ -2446,17 +2459,14 @@ InstructionCost
VPWidenMemIntrinsicRecipe::computeCost(ElementCount VF,
VPCostContext &Ctx) const {
Type *DataTy;
- if (auto DataPos = VPIntrinsic::getMemoryDataParamPos(getVectorIntrinsicID()))
+ if (auto DataPos = getVectorStoreIntrinsicDataArgIdx(getVectorIntrinsicID()))
DataTy = getOperand(*DataPos)->getScalarType();
else
DataTy = getScalarType();
assert(!DataTy->isVoidTy() && "Expected a non-void data type");
Type *Ty = toVectorTy(DataTy, VF);
- auto MaskPos = VPIntrinsic::getMaskParamPos(getVectorIntrinsicID());
- assert(MaskPos && "Expected a memory intrinsic with a valid mask position");
return computeMemIntrinsicCost(getVectorIntrinsicID(), Ty,
- !match(getOperand(*MaskPos), m_True()),
- Alignment, Ctx);
+ !match(getMask(), m_True()), Alignment, Ctx);
}
void VPHistogramRecipe::execute(VPTransformState &State) {
@@ -5137,6 +5147,21 @@ bool VPBlendRecipe::usesFirstLaneOnly(const VPValue *Op) const {
return vputils::onlyFirstLaneUsed(this);
}
+void VPMonotonicPHIRecipe::execute(VPTransformState &State) {
+ executePhiRecipe(this, *this, State, /*IsScalar=*/true, "monotonic.iv");
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+void VPMonotonicPHIRecipe::printRecipe(raw_ostream &O, const Twine &Indent,
+ VPSlotTracker &SlotTracker) const {
+ O << Indent << "MONOTONIC-PHI ";
+
+ printAsOperand(O, SlotTracker);
+ O << " = phi ";
+ printOperands(O, SlotTracker);
+}
+#endif
+
void VPWidenPHIRecipe::execute(VPTransformState &State) {
executePhiRecipe(this, *this, State, /*IsScalar=*/false, Name);
}
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
index d78a3d3b1e344..56a6144c577cb 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
@@ -6028,3 +6028,90 @@ void VPlanTransforms::convertToStridedAccesses(VPlan &Plan,
}
}
}
+
+bool VPlanTransforms::handleCompressingPatterns(
+ VPlan &Plan, VPBasicBlock *HeaderVPBB, VPRecipeBuilder &RecipeBuilder) {
+ SmallVector<VPInstruction *> MemOps;
+ for (VPBasicBlock *VPBB :
+ VPBlockUtils::blocksOnly<VPBasicBlock>(vp_depth_first_shallow(
+ Plan.getVectorLoopRegion()->getEntryBasicBlock()))) {
+ for (VPRecipeBase &R : *VPBB) {
+ auto *VPI = dyn_cast<VPInstruction>(&R);
+ if (VPI && VPI->getUnderlyingValue() &&
+ is_contained({Instruction::Load, Instruction::Store},
+ VPI->getOpcode()))
+ MemOps.push_back(VPI);
+ }
+ }
+
+ VPBuilder Builder;
+ for (VPRecipeBase &R : HeaderVPBB->phis()) {
+ auto *MonotonicPhi = dyn_cast<VPMonotonicPHIRecipe>(&R);
+ if (!MonotonicPhi)
+ continue;
+
+ // Obtain the mask for the monotonic phi update from the VPBlendRecipe.
+ auto *BlendR = cast<VPBlendRecipe>(MonotonicPhi->getBackedgeValue());
+ VPValue *Mask = nullptr;
+ for (unsigned I = 0, E = BlendR->getNumIncomingValues(); I != E; ++I)
+ if (auto *IncomingVal = BlendR->getIncomingValue(I);
+ IncomingVal != MonotonicPhi) {
+ Mask = BlendR->getMask(I);
+ break;
+ }
+ assert(Mask);
+
+ // Replace all "compressed" loads and stores with expandload and
+ // compressstore respectively.
+ for (VPInstruction *&VPI : MemOps) {
+ auto *CompressedMemOp =
+ RecipeBuilder.widenIfCompressedLoadOrStore(VPI, MonotonicPhi);
+ if (!CompressedMemOp)
+ continue;
+
+ Builder.setInsertPoint(VPI);
+ Builder.insert(CompressedMemOp);
+
+ // Bail out if the mask for the memory op does not match the condition
+ // used to update the montontic phi.
+ VPValue *MemOpMask = CompressedMemOp->getMask();
+ if (MemOpMask != Mask)
+ return false;
+
+ if (VPI->getOpcode() == Instruction::Load)
+ VPI->replaceAllUsesWith(CompressedMemOp->getVPSingleValue());
+ VPI->eraseFromParent();
+ VPI = nullptr; // Mark handled instructions with a nullptr.
+ }
+
+ // Remove all memory operations we've handled.
+ MemOps.erase(
+ remove_if(MemOps, [](VPInstruction *VPI) { return VPI == nullptr; }),
+ MemOps.end());
+
+ // Update the monotonic PHI to increment by the number of active lanes in
+ // the mask.
+ auto *BackedgeVal = MonotonicPhi->getBackedgeValue();
+ auto *InsertBlock = BackedgeVal->getDefiningRecipe()->getParent();
+ Builder.setInsertPoint(InsertBlock, InsertBlock->getFirstNonPhi());
+
+ Type *UpdateType = MonotonicPhi->getScalarType();
+ if (UpdateType->isPointerTy())
+ UpdateType = Plan.getDataLayout().getIndexType(UpdateType);
+
+ auto *HandledLanes = Builder.createNaryOp(
+ VPInstruction::NumActiveLanes, {Mask}, nullptr, {}, {},
+ DebugLoc::getUnknown(), "handled.lanes", UpdateType);
+ VPValue *Offset = Builder.createOverflowingOp(
+ Instruction::Mul, {MonotonicPhi->getStep(), HandledLanes});
+ VPValue *Update;
+ if (MonotonicPhi->getScalarType()->isPointerTy())
+ Update = Builder.createPtrAdd(MonotonicPhi, Offset);
+ else
+ Update = Builder.createAdd(MonotonicPhi, Offset, {}, "monotonic.add");
+
+ BackedgeVal->replaceAllUsesWith(Update);
+ }
+
+ return true;
+}
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
index 1c7b17942b795..1c4bbb0f18c32 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
@@ -172,6 +172,7 @@ struct VPlanTransforms {
const VPDominatorTree &VPDT,
const MapVector<PHINode *, InductionDescriptor> &Inductions,
const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
+ const MapVector<PHINode *, MonotonicDescriptor> &MonotonicPHIs,
const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering);
@@ -261,6 +262,15 @@ struct VPlanTransforms {
/// was unsuccessful.
static bool handleFindLastReductions(VPlan &Plan);
+ /// Handles compressing memory loads/stores. Loads/stores where the pointer
+ /// is derived from a monotonic PHI are replaced with expandloads or
+ /// compressstores respectively. The backedge value of the monotonic PHI is
+ /// updated to increment by the number of active lanes of the block mask.
+ /// Returns false if any memory operation could not be updated (e.g., due to
+ /// having a mask that does not match the PHI).
+ static bool handleCompressingPatterns(VPlan &Plan, VPBasicBlock *HeaderVPBB,
+ VPRecipeBuilder &RecipeBuilder);
+
/// Clear NSW/NUW flags from reduction instructions if necessary.
static void clearReductionWrapFlags(VPlan &Plan);
diff --git a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
index c440d37bef517..908d6d040023f 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
@@ -461,8 +461,8 @@ bool vputils::isSingleScalar(const VPValue *VPV) {
all_of(VPI->operands(), isSingleScalar));
if (auto *RR = dyn_cast<VPReductionRecipe>(VPV))
return !RR->isPartialReduction();
- if (isa<VPVectorPointerRecipe, VPVectorEndPointerRecipe, VPDerivedIVRecipe>(
- VPV))
+ if (isa<VPVectorPointerRecipe, VPVectorEndPointerRecipe, VPDerivedIVRecipe,
+ VPMonotonicPHIRecipe>(VPV))
return true;
if (auto *Expr = dyn_cast<VPExpressionRecipe>(VPV))
return Expr->isVectorToScalar();
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
new file mode 100644
index 0000000000000..93bc4113fac9b
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
@@ -0,0 +1,126 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^scalar.ph:" --version 5
+; RUN: opt < %s -lv-monotonic-patterns=true -mtriple=aarch64 -mattr=+sve2p2 -passes=loop-vectorize -S 2>&1 | FileCheck %s
+
+; SVE compresstore/expandload vectorization (requires +sve2p2 for expandload and +sve for compresstore).
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: define void @compress_store(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT: [[VECTOR_PH:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT: [[TMP2:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP2]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH1:.*]]
+; CHECK: [[VECTOR_PH1]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP2]]
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP4:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP5]]
+; CHECK-NEXT: call void @llvm.masked.compressstore.nxv4i32.p0(<vscale x 4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP6]], <vscale x 4 x i1> [[TMP4]])
+; CHECK-NEXT: [[TMP7:%.*]] = zext <vscale x 4 x i1> [[TMP4]] to <vscale x 4 x i64>
+; CHECK-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP7]])
+; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP8]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP2]]
+; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP9]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[IF_THEN]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: define void @expand_load(
+; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
+; CHECK-NEXT: [[VECTOR_PH:.*:]]
+; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT: [[TMP2:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP2]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH1:.*]]
+; CHECK: [[VECTOR_PH1]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP2]]
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP3]], align 4
+; CHECK-NEXT: [[TMP4:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[TMP5]]
+; CHECK-NEXT: [[TMP7:%.*]] = call <vscale x 4 x i32> @llvm.masked.expandload.nxv4i32.p0(ptr align 4 [[TMP6]], <vscale x 4 x i1> [[TMP4]], <vscale x 4 x i32> poison)
+; CHECK-NEXT: call void @llvm.masked.store.nxv4i32.p0(<vscale x 4 x i32> [[TMP7]], ptr align 4 [[TMP3]], <vscale x 4 x i1> [[TMP4]])
+; CHECK-NEXT: [[TMP8:%.*]] = zext <vscale x 4 x i1> [[TMP4]] to <vscale x 4 x i64>
+; CHECK-NEXT: [[TMP9:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP8]])
+; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP9]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP2]]
+; CHECK-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP10]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK: [[IF_THEN]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK: [[SCALAR_PH]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ %cmp = icmp slt i32 %load.dst, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
new file mode 100644
index 0000000000000..af74e9ec2a936
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
@@ -0,0 +1,153 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --filter-out-after "^scalar.ph:" --version 6
+; RUN: opt -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -vplan-print-after=printOptimizedVPlan -disable-output %s -S 2>&1 | FileCheck %s
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: VPlan for loop in 'compress_store'
+; CHECK: VPlan 'Initial VPlan for VF={4},UF>=1' {
+; CHECK-NEXT: Live-in vp<[[VP0:%[0-9]+]]> = VF
+; CHECK-NEXT: Live-in vp<[[VP1:%[0-9]+]]> = VF * UF
+; CHECK-NEXT: Live-in vp<[[VP2:%[0-9]+]]> = vector-trip-count
+; CHECK-NEXT: Live-in ir<%n> = original trip-count
+; CHECK-EMPTY:
+; CHECK-NEXT: ir-bb<entry>:
+; CHECK-NEXT: Successor(s): scalar.ph, vector.ph
+; CHECK-EMPTY:
+; CHECK-NEXT: vector.ph:
+; CHECK-NEXT: Successor(s): vector loop
+; CHECK-EMPTY:
+; CHECK-NEXT: <x1> vector loop: {
+; CHECK-NEXT: vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
+; CHECK-EMPTY:
+; CHECK-NEXT: vector.body:
+; CHECK-NEXT: MONOTONIC-PHI ir<%idx> = phi ir<0>, vp<%monotonic.add>, ir<1>
+; CHECK-NEXT: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
+; CHECK-NEXT: CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
+; CHECK-NEXT: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
+; CHECK-NEXT: WIDEN ir<%load.src> = load vp<[[VP5]]>
+; CHECK-NEXT: WIDEN ir<%cmp> = icmp slt ir<%load.src>, ir<%c>
+; CHECK-NEXT: CLONE ir<%dst.ptr> = getelementptr inbounds ir<%dst>, ir<%idx>
+; CHECK-NEXT: vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.compressstore(ir<%load.src>, vp<[[VP6]]>, ir<%cmp>)
+; CHECK-NEXT: EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
+; CHECK-NEXT: EMIT vp<%monotonic.add> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT: EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
+; CHECK-NEXT: EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
+; CHECK-NEXT: No successors
+; CHECK-NEXT: }
+; CHECK-NEXT: Successor(s): middle.block
+; CHECK-EMPTY:
+; CHECK-NEXT: middle.block:
+; CHECK-NEXT: EMIT vp<[[VP9:%[0-9]+]]> = extract-last-part vp<%monotonic.add>
+; CHECK-NEXT: EMIT vp<[[VP10:%[0-9]+]]> = extract-last-lane vp<[[VP9]]>
+; CHECK-NEXT: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
+; CHECK-NEXT: EMIT branch-on-cond vp<%cmp.n>
+; CHECK-NEXT: Successor(s): ir-bb<exit>, scalar.ph
+; CHECK-EMPTY:
+; CHECK-NEXT: ir-bb<exit>:
+; CHECK-NEXT: No successors
+; CHECK-EMPTY:
+; CHECK-NEXT: scalar.ph:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: VPlan for loop in 'expand_load'
+; CHECK: VPlan 'Initial VPlan for VF={4},UF>=1' {
+; CHECK-NEXT: Live-in vp<[[VP0:%[0-9]+]]> = VF
+; CHECK-NEXT: Live-in vp<[[VP1:%[0-9]+]]> = VF * UF
+; CHECK-NEXT: Live-in vp<[[VP2:%[0-9]+]]> = vector-trip-count
+; CHECK-NEXT: Live-in ir<%n> = original trip-count
+; CHECK-EMPTY:
+; CHECK-NEXT: ir-bb<entry>:
+; CHECK-NEXT: Successor(s): scalar.ph, vector.ph
+; CHECK-EMPTY:
+; CHECK-NEXT: vector.ph:
+; CHECK-NEXT: Successor(s): vector loop
+; CHECK-EMPTY:
+; CHECK-NEXT: <x1> vector loop: {
+; CHECK-NEXT: vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
+; CHECK-EMPTY:
+; CHECK-NEXT: vector.body:
+; CHECK-NEXT: MONOTONIC-PHI ir<%idx> = phi ir<0>, vp<%monotonic.add>, ir<1>
+; CHECK-NEXT: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
+; CHECK-NEXT: CLONE ir<%dst.ptr> = getelementptr ir<%dst>, vp<[[VP4]]>
+; CHECK-NEXT: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT: WIDEN ir<%load.dst> = load vp<[[VP5]]>
+; CHECK-NEXT: WIDEN ir<%cmp> = icmp slt ir<%load.dst>, ir<%c>
+; CHECK-NEXT: CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, ir<%idx>
+; CHECK-NEXT: vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
+; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.expandload(vp<[[VP6]]>, ir<%cmp>, ir<poison>)
+; CHECK-NEXT: vp<[[VP8:%[0-9]+]]> = vector-pointer i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT: WIDEN store vp<[[VP8]]>, vp<[[VP7]]>, ir<%cmp>
+; CHECK-NEXT: EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
+; CHECK-NEXT: EMIT vp<%monotonic.add> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT: EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
+; CHECK-NEXT: EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
+; CHECK-NEXT: No successors
+; CHECK-NEXT: }
+; CHECK-NEXT: Successor(s): middle.block
+; CHECK-EMPTY:
+; CHECK-NEXT: middle.block:
+; CHECK-NEXT: EMIT vp<[[VP10:%[0-9]+]]> = extract-last-part vp<%monotonic.add>
+; CHECK-NEXT: EMIT vp<[[VP11:%[0-9]+]]> = extract-last-lane vp<[[VP10]]>
+; CHECK-NEXT: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
+; CHECK-NEXT: EMIT branch-on-cond vp<%cmp.n>
+; CHECK-NEXT: Successor(s): ir-bb<exit>, scalar.ph
+; CHECK-EMPTY:
+; CHECK-NEXT: ir-bb<exit>:
+; CHECK-NEXT: No successors
+; CHECK-EMPTY:
+; CHECK-NEXT: scalar.ph:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ %cmp = icmp slt i32 %load.dst, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
index b5ee488ae5bd5..59bc89a291888 100644
--- a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
@@ -21,6 +21,7 @@
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::handleCountableEarlyExits
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::createLoopRegions
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::introduceMasksAndLinearize
+; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::handleCompressingPatterns
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::createInLoopReductionRecipes
; CHECK-BEFORE: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::makeMemOpWideningDecisions
; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] lowerMemoryIdioms
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
new file mode 100644
index 0000000000000..ef3c8d79d6d59
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
@@ -0,0 +1,268 @@
+; RUN: opt < %s -lv-monotonic-patterns=true -enable-early-exit-vectorization-with-side-effects -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -disable-output -pass-remarks-missed=".*" 2>&1 | FileCheck %s
+
+; CHECK: loop not vectorized
+
+; Negative test: Conditional pointer (rather than index) increments are not supported yet (needs LAA support).
+define void @test_compress_store_with_pointer(ptr writeonly noalias %init.dst, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %dst = phi ptr [ %init.dst, %entry ], [ %dst.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.inc = getelementptr inbounds i8, ptr %dst, i64 4
+ store i32 %load.src, ptr %dst, align 4
+ br label %for.inc
+
+for.inc:
+ %dst.1 = phi ptr [ %dst.inc, %if.then ], [ %dst, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; CHECK: loop not vectorized
+
+; Negative test: Storing the conditionally incremented phi is invalid.
+
+define void @test_store_conditionally_incremented_value(ptr writeonly noalias %dst, ptr writeonly noalias %dst2, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i32 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ store i32 %idx, ptr %dst.ptr, align 4
+ %idx.next = add nsw i32 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i32 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; CHECK: loop not vectorized
+
+; Pre-increment is currently not matched as we require one use of the step instruction.
+define void @test_pre_increment_compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %idx.next = add nsw i64 %idx, 1
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx.next
+ store i32 %load.src, ptr %dst.ptr, align 4
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; CHECK: the cost-model indicates that vectorization is not beneficial
+
+; Negative test: In this case the %idx is incremented when %cond.val != 0,
+; but the store occurs when %cond.val > 100. The store mask does not match the
+; PHI mask, so the loop is not vectorized.
+define void @compress_mismatched_mask(ptr noalias %dst, ptr noalias %src, ptr noalias %cond, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.next, %for.inc ]
+ %cond.ptr = getelementptr inbounds nuw [4 x i8], ptr %cond, i64 %iv
+ %cond.val = load i32, ptr %cond.ptr, align 4
+ %cond.bool = icmp eq i32 %cond.val, 0
+ br i1 %cond.bool, label %for.inc, label %if.then
+
+if.then:
+ %cmp.cond = icmp sgt i32 %cond.val, 100
+ br i1 %cmp.cond, label %if.then1, label %if.end
+
+if.then1:
+ %src.ptr = getelementptr inbounds nuw [4 x i8], ptr %src, i64 %iv
+ %src.val = load i32, ptr %src.ptr, align 4
+ %dst.ptr = getelementptr inbounds [4 x i8], ptr %dst, i64 %idx
+ store i32 %src.val, ptr %dst.ptr, align 4
+ br label %if.end
+
+if.end:
+ %inc = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.next = phi i64 [ %inc, %if.end ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; CHECK: the cost-model indicates that vectorization is not beneficial
+
+; Negative test: Simple early exit loop with a compressstore. This fails in VPlan handling for early exits.
+define i32 @compress_store_with_early_exit(ptr dereferenceable(1024) %dst, ptr noalias dereferenceable(1024) %src, ptr noalias dereferenceable(1024) %cond, ptr noalias dereferenceable(1024) %exit_cond) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.2, %for.inc ]
+ %cond.ptr = getelementptr inbounds nuw i32, ptr %cond, i64 %iv
+ %cond.val = load i32, ptr %cond.ptr, align 4
+ %compress.cond = icmp eq i32 %cond.val, 0
+ %exit.ptr = getelementptr inbounds nuw i32, ptr %exit_cond, i64 %iv
+ %exit.val = load i32, ptr %exit.ptr, align 4
+ br i1 %compress.cond, label %for.inc, label %if.then
+
+if.then:
+ %src.ptr = getelementptr inbounds nuw i32, ptr %src, i64 %iv
+ %src.val = load i32, ptr %src.ptr, align 4
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %src.val, ptr %dst.ptr, align 4
+ %not.exit.cond = icmp eq i32 %exit.val, 0
+ %inc = add nsw i64 %idx, 1
+ br i1 %not.exit.cond, label %for.inc, label %early.exit
+
+for.inc:
+ %idx.2 = phi i64 [ %idx, %for.body ], [ %inc, %if.then ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, 128
+ br i1 %exitcond.not, label %early.exit, label %for.body
+
+early.exit:
+ %ret = phi i32 [ 1, %if.then ], [ 0, %for.inc ]
+ ret i32 %ret
+}
+
+; CHECK: loop not vectorized
+
+; Negative test: Using the monotonic phi outside the loop is not supported.
+define i64 @out_of_loop_use_of_monotonic_phi(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret i64 %idx
+}
+
+; CHECK: loop not vectorized
+
+; Negative test: Matching an extended monotonic phi index is not supported yet.
+; Note: We should be able to support this case by using the no-wrap flags on %idx.next.
+define void @test_compress_store_with_extended_index_with_nsw(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i32 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.idx = sext i32 %idx to i64
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %dst.idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i32 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i32 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; CHECK: loop not vectorized
+
+; Negative test: We can't vectorize a extended monotonic phi use without no-wrap flags on the step.
+define void @test_compress_store_with_extended_index(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i8 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr i8, ptr %dst, i8 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add i8 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i8 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
new file mode 100644
index 0000000000000..19010788ceac5
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
@@ -0,0 +1,797 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^for.body:" --version 5
+; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-IC1
+; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -tail-folding-policy=must-fold-tail -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-TF
+; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -force-vector-interleave=2 -passes=loop-vectorize -disable-output -pass-remarks-analysis=loop-vectorize 2>&1 | FileCheck %s --check-prefix=IC2
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+define void @test_compress_store_with_index(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_compress_store_with_index(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
+; CHECK-IC1: [[SCALAR_PH1]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
+; CHECK-IC1: [[FOR_BODY1]]:
+;
+; CHECK-TF-LABEL: define void @test_compress_store_with_index(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK1]]:
+; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
+; CHECK-TF: [[EXIT1]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+define void @test_expand_load_with_index(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_expand_load_with_index(
+; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP3]], <4 x i1> [[TMP2]], <4 x i32> poison)
+; CHECK-IC1-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP4]], ptr align 4 [[TMP1]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP6]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
+; CHECK-IC1: [[SCALAR_PH1]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
+; CHECK-IC1: [[FOR_BODY1]]:
+;
+; CHECK-TF-LABEL: define void @test_expand_load_with_index(
+; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP6:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP5]], <4 x i1> [[TMP4]], <4 x i32> poison)
+; CHECK-TF-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP6]], ptr align 4 [[TMP2]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP8]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK1]]:
+; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
+; CHECK-TF: [[EXIT1]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ %cmp = icmp slt i32 %load.dst, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+define i64 @test_conditionally_incremented_phi_liveout(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
+; CHECK-IC1: [[SCALAR_PH1]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
+; CHECK-IC1: [[FOR_BODY1]]:
+;
+; CHECK-TF-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK1]]:
+; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
+; CHECK-TF: [[EXIT1]]:
+; CHECK-TF-NEXT: ret i64 [[MONOTONIC_ADD]]
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret i64 %idx.1
+}
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+define void @test_compress_store_with_scaled_pointer(ptr writeonly noalias %dst.bytes, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_compress_store_with_scaled_pointer(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST_BYTES:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = shl nsw i64 [[TMP3]], 2
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP4]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP3]], [[TMP6]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+;
+; CHECK-TF-LABEL: define void @test_compress_store_with_scaled_pointer(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST_BYTES:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP6:%.*]] = shl nsw i64 [[TMP5]], 2
+; CHECK-TF-NEXT: [[TMP7:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP6]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP7]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT: [[TMP9:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP9]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP8]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %idx.bytes = shl nsw i64 %idx, 2
+ %dst.ptr = getelementptr inbounds i8, ptr %dst.bytes, i64 %idx.bytes
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+; Test a nested conditional compress store, where the phi is only updated on iterations where the store takes place.
+define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_nested_conditional_compress_store(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = icmp sgt <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = select <4 x i1> [[TMP2]], <4 x i1> [[TMP4]], <4 x i1> zeroinitializer
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP5]])
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP5]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+;
+; CHECK-TF-LABEL: define void @test_nested_conditional_compress_store(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = icmp sgt <4 x i32> [[WIDE_MASKED_LOAD]], zeroinitializer
+; CHECK-TF-NEXT: [[TMP6:%.*]] = select <4 x i1> [[TMP3]], <4 x i1> [[TMP5]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP7:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP4]], <4 x i1> [[TMP7]])
+; CHECK-TF-NEXT: [[TMP8:%.*]] = zext <4 x i1> [[TMP7]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP9:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP8]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP9]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %latch ]
+ %idx = phi i64 [ 0, %entry ], [ %phi, %latch ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %update.block, label %latch
+
+update.block:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ %update = add nsw i64 %idx, 1
+ %cmp2 = icmp sgt i32 %load.src, 0
+ br i1 %cmp2, label %store.block, label %latch
+
+store.block:
+ store i32 %load.src, ptr %dst.ptr, align 4
+ br label %latch
+
+latch:
+ %phi = phi i64 [ %idx, %for.body ], [ %idx, %update.block ], [ %update, %store.block ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; An unconditional increment should lower as a simple induction (not a monotonic PHI).
+define void @test_unconditional_increment(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_unconditional_increment(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = add i64 15, [[N_VEC]]
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = add i64 15, [[INDEX]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP2]]
+; CHECK-IC1-NEXT: store <4 x i32> [[WIDE_LOAD]], ptr [[TMP4]], align 4
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP12:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL1:%.*]] = phi i64 [ [[TMP1]], %[[MIDDLE_BLOCK]] ], [ 15, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+;
+; CHECK-TF-LABEL: define void @test_unconditional_increment(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = add i64 15, [[INDEX]]
+; CHECK-TF-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP3]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP2]]
+; CHECK-TF-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP4]], <4 x i1> [[TMP1]])
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP7:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 15, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br label %inc.step
+
+inc.step:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %inc.step ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+
+define void @test_multiple_monotonic_phis(ptr %dst, ptr noalias %dst2, ptr noalias %src, ptr noalias %cond, ptr noalias %cond2, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_multiple_monotonic_phis(
+; CHECK-IC1-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
+; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1: [[VECTOR_PH]]:
+; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[MONOTONIC_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD4:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP14:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i8>
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP14]], ptr align 1 [[TMP4]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP6]]
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP7]], align 4
+; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-IC1-NEXT: [[TMP9:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i16>
+; CHECK-IC1-NEXT: [[TMP10:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[MONOTONIC_IV1]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP9]], ptr align 2 [[TMP10]], <4 x i1> [[TMP8]])
+; CHECK-IC1-NEXT: [[TMP11:%.*]] = zext <4 x i1> [[TMP8]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP12:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP11]])
+; CHECK-IC1-NEXT: [[MONOTONIC_ADD4]] = add i64 [[MONOTONIC_IV1]], [[TMP12]]
+; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP13]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP14:![0-9]+]]
+; CHECK-IC1: [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX5:%.*]] = phi i64 [ [[MONOTONIC_ADD4]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+;
+; CHECK-TF-LABEL: define void @test_multiple_monotonic_phis(
+; CHECK-TF-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
+; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT: [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-TF: [[VECTOR_BODY]]:
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[MONOTONIC_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD4:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp ne <4 x i32> [[WIDE_MASKED_LOAD]], zeroinitializer
+; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD2:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP4]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP5:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP17:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i8>
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP17]], ptr align 1 [[TMP6]], <4 x i1> [[TMP5]])
+; CHECK-TF-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP5]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP8]]
+; CHECK-TF-NEXT: [[TMP9:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD3:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP9]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp ne <4 x i32> [[WIDE_MASKED_LOAD3]], zeroinitializer
+; CHECK-TF-NEXT: [[TMP11:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP10]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP12:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i16>
+; CHECK-TF-NEXT: [[TMP13:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[MONOTONIC_IV1]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP12]], ptr align 2 [[TMP13]], <4 x i1> [[TMP11]])
+; CHECK-TF-NEXT: [[TMP14:%.*]] = zext <4 x i1> [[TMP11]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP15:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP14]])
+; CHECK-TF-NEXT: [[MONOTONIC_ADD4]] = add i64 [[MONOTONIC_IV1]], [[TMP15]]
+; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT: [[TMP16:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %dst.idx = phi i64 [ 0, %entry ], [ %dst.inc, %for.inc ]
+ %dst2.idx = phi i64 [ 0, %entry ], [ %dst2.inc, %for.inc ]
+ %cond.gep = getelementptr inbounds i32, ptr %cond, i64 %iv
+ %cond.val = load i32, ptr %cond.gep, align 4
+ %cond.is.zero = icmp eq i32 %cond.val, 0
+ %src.gep = getelementptr inbounds i32, ptr %src, i64 %iv
+ %src.val = load i32, ptr %src.gep, align 4
+ br i1 %cond.is.zero, label %if.end, label %if.then0
+
+if.then0:
+ %dst.idx.next = add nsw i64 %dst.idx, 1
+ %dst.gep = getelementptr inbounds i8, ptr %dst, i64 %dst.idx
+ %dst.val.trunc = trunc i32 %src.val to i8
+ store i8 %dst.val.trunc, ptr %dst.gep, align 1
+ br label %if.end
+
+if.end:
+ %dst.inc = phi i64 [ %dst.idx.next, %if.then0 ], [ %dst.idx, %for.body ]
+ %cond2.gep = getelementptr inbounds i32, ptr %cond2, i64 %iv
+ %cond2.val = load i32, ptr %cond2.gep, align 4
+ %cond2.is.zero = icmp eq i32 %cond2.val, 0
+ br i1 %cond2.is.zero, label %for.inc, label %if.then1
+
+if.then1:
+ %dst2.val.trunc = trunc i32 %src.val to i16
+ %dst2.idx.next = add nsw i64 %dst2.idx, 1
+ %dst2.gep = getelementptr inbounds i16, ptr %dst2, i64 %dst2.idx
+ store i16 %dst2.val.trunc, ptr %dst2.gep, align 2
+ br label %for.inc
+
+for.inc:
+ %dst2.inc = phi i64 [ %dst2.idx.next, %if.then1 ], [ %dst2.idx, %if.end ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+; Test expand load with always false update (this probably can be simplified).
+define void @test_expand_load_always_false_cond(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_expand_load_always_false_cond(
+; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-IC1-NEXT: [[SCALAR_PH:.*:]]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
+;
+; CHECK-TF-LABEL: define void @test_expand_load_always_false_cond(
+; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i32 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ br i1 0, label %if.then, label %for.inc
+
+if.then:
+ %src.idx = sext i32 %idx to i64
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %src.idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i32 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i32 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
new file mode 100644
index 0000000000000..346f7e5c2f6ec
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
@@ -0,0 +1,95 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^for.body:" --version 5
+; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=16 -epilogue-vectorization-force-VF=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefixes=CHECK
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: define void @compress_store(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-NEXT: [[ITER_CHECK:.*]]:
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-NEXT: [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[N]], 16
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK1]], label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[TMP0:%.*]] = and i64 [[N]], 15
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK: [[VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 42, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-NEXT: [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-NEXT: call void @llvm.masked.compressstore.v16i32.p0(<16 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <16 x i1> [[TMP2]])
+; CHECK-NEXT: [[TMP4:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
+; CHECK-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP4]])
+; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK: [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-NEXT: [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP0]], 4
+; CHECK-NEXT: br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3:![0-9]+]]
+; CHECK: [[VEC_EPILOG_PH]]:
+; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[TMP7:%.*]] = and i64 [[N]], 3
+; CHECK-NEXT: [[N_VEC2:%.*]] = sub i64 [[N]], [[TMP7]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT4:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT3]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
+; CHECK: [[VEC_EPILOG_VECTOR_BODY]]:
+; CHECK-NEXT: [[INDEX5:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT9:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT: [[MONOTONIC_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[MONOTONIC_ADD8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP8:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX5]]
+; CHECK-NEXT: [[WIDE_LOAD7:%.*]] = load <4 x i32>, ptr [[TMP8]], align 4
+; CHECK-NEXT: [[TMP9:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
+; CHECK-NEXT: [[TMP10:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV6]]
+; CHECK-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD7]], ptr align 4 [[TMP10]], <4 x i1> [[TMP9]])
+; CHECK-NEXT: [[TMP11:%.*]] = zext <4 x i1> [[TMP9]] to <4 x i64>
+; CHECK-NEXT: [[TMP12:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP11]])
+; CHECK-NEXT: [[MONOTONIC_ADD8]] = add i64 [[MONOTONIC_IV6]], [[TMP12]]
+; CHECK-NEXT: [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 4
+; CHECK-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT9]], [[N_VEC2]]
+; CHECK-NEXT: br i1 [[TMP13]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK: [[VEC_EPILOG_MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[CMP_N10:%.*]] = icmp eq i64 [[N]], [[N_VEC2]]
+; CHECK-NEXT: br i1 [[CMP_N10]], [[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
+; CHECK: [[VEC_EPILOG_SCALAR_PH]]:
+; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC2]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX11:%.*]] = phi i64 [ [[MONOTONIC_ADD8]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[MONOTONIC_ADD]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[ITER_CHECK]] ]
+; CHECK-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK: [[FOR_BODY]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 42, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
index 81ac481fbabb9..bd51bc1fb007f 100644
--- a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
+++ b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
@@ -100,6 +100,7 @@ class VPlanTestIRBase : public testing::Test {
VPlanTransforms::createHeaderPhiRecipes(
*Plan, PSE, *L, VPDT, Inductions,
MapVector<PHINode *, RecurrenceDescriptor>(),
+ MapVector<PHINode *, MonotonicDescriptor>(),
SmallPtrSet<const PHINode *, 1>(), SmallPtrSet<PHINode *, 1>(),
/*AllowReordering=*/false);
}
>From f7c2ac8e9b46ebc6a6a8bb66a6f87ce6d78b1803 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Wed, 16 Sep 2026 10:40:46 +0000
Subject: [PATCH 2/3] Rename MonotonicPHI (and related) to
ConditionalInduction*
---
.../include/llvm/Transforms/Utils/LoopUtils.h | 7 +-
.../Vectorize/LoopVectorizationLegality.h | 41 ++-
llvm/lib/Transforms/Utils/LoopUtils.cpp | 15 +-
.../Vectorize/LoopVectorizationLegality.cpp | 41 ++-
.../Transforms/Vectorize/LoopVectorize.cpp | 26 +-
.../Transforms/Vectorize/VPRecipeBuilder.h | 3 +-
llvm/lib/Transforms/Vectorize/VPlan.cpp | 3 +-
llvm/lib/Transforms/Vectorize/VPlan.h | 37 +-
.../Vectorize/VPlanConstruction.cpp | 15 +-
.../lib/Transforms/Vectorize/VPlanRecipes.cpp | 12 +-
.../Transforms/Vectorize/VPlanTransforms.cpp | 36 +-
.../Transforms/Vectorize/VPlanTransforms.h | 10 +-
llvm/lib/Transforms/Vectorize/VPlanUtils.cpp | 2 +-
.../LoopVectorize/AArch64/compress-idioms.ll | 84 ++---
.../LoopVectorize/VPlan/compress-idioms.ll | 20 +-
.../compress-idioms-negative-tests.ll | 10 +-
.../LoopVectorize/compress-idioms.ll | 334 +++++++++---------
.../compress-store-vec-epilogue.ll | 18 +-
.../Transforms/Vectorize/VPlanTestBase.h | 2 +-
19 files changed, 371 insertions(+), 345 deletions(-)
diff --git a/llvm/include/llvm/Transforms/Utils/LoopUtils.h b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
index cab79ddd72685..1245c122e3906 100644
--- a/llvm/include/llvm/Transforms/Utils/LoopUtils.h
+++ b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
@@ -44,7 +44,7 @@ class TargetLibraryInfo;
class LPPassManager;
class Instruction;
struct RuntimeCheckingPtrGroup;
-class MonotonicDescriptor;
+class ConditionalInductionDescriptor;
typedef std::pair<const RuntimeCheckingPtrGroup *,
const RuntimeCheckingPtrGroup *>
@@ -713,9 +713,10 @@ hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
/// from the monotonic PHI described by \p MD. The pointer operands and
/// approximate SCEV expressions (assuming the monotonic PHI always increments)
/// for the pointers are placed in \p CompressedPtrs. Returns true if all
-/// in-loop users of the monotonic PHI are loads/stores.
+/// in-loop users of the conditional induction are loads/stores.
bool collectCompressedPtrs(DenseMap<Value *, const SCEV *> &CompressedPtrs,
- const Loop &L, const MonotonicDescriptor &MD,
+ const Loop &L,
+ const ConditionalInductionDescriptor &CondID,
ScalarEvolution &SE);
} // end namespace llvm
diff --git a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
index cd062eef9535c..29a86fe304dbd 100644
--- a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
+++ b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
@@ -247,10 +247,10 @@ struct HistogramInfo {
: Load(Load), Update(Update), Store(Store) {}
};
-/// Holds details about a "compressed" pointer: the monotonic PHI used to
-/// derive the pointer and the SCEV expression for the pointer.
+/// Holds details about a "compressed" pointer: the conditional induction PHI
+/// used to derive the pointer and the SCEV expression for the pointer.
struct CompressedPtrInfo {
- PHINode *MonotonicPHI;
+ PHINode *ConditionalInductionPHI;
const SCEVAddRecExpr *PtrSCEV;
};
@@ -294,9 +294,10 @@ class LoopVectorizationLegality {
/// induction descriptor.
using InductionList = MapVector<PHINode *, InductionDescriptor>;
- /// MonotonicPHIList saves monotonic phi variables and maps them to the
- /// monotonic phi descriptor.
- using MonotonicPHIList = MapVector<PHINode *, MonotonicDescriptor>;
+ /// ConditionalInductionList saves conditional inductions and maps them to
+ /// their descriptors.
+ using ConditionalInductionList =
+ MapVector<PHINode *, ConditionalInductionDescriptor>;
/// RecurrenceSet contains the phi nodes that are recurrences other than
/// inductions and reductions.
@@ -341,10 +342,14 @@ class LoopVectorizationLegality {
/// Returns the induction variables found in the loop.
const InductionList &getInductionVars() const { return Inductions; }
- /// Returns the monotonic phi variables found in the loop.
- const MonotonicPHIList &getMonotonicPHIs() const { return MonotonicPHIs; }
+ /// Returns the conditional inductions found in the loop.
+ const ConditionalInductionList &getConditionalInductions() const {
+ return ConditionalInductions;
+ }
- bool hasMonotonicPHIs() const { return !MonotonicPHIs.empty(); }
+ bool hasConditionalInductions() const {
+ return !ConditionalInductions.empty();
+ }
/// Return the fixed-order recurrences found in the loop.
RecurrenceSet &getFixedOrderRecurrences() { return FixedOrderRecurrences; }
@@ -492,7 +497,7 @@ class LoopVectorizationLegality {
bool hasHistograms() const { return !Histograms.empty(); }
/// Returns the CompressedPtrInfo for \p Ptr if the pointer is defined via
- /// a monotonic PHI, otherwise std::nullptr.
+ /// a conditional induction PHI, otherwise std::nullopt.
std::optional<CompressedPtrInfo>
getCompressedPtrInfo(const Value *Ptr) const {
auto It = CompressedPtrs.find(Ptr);
@@ -681,9 +686,11 @@ class LoopVectorizationLegality {
/// better choice for the main induction than the existing one.
void addInductionPhi(PHINode *Phi, const InductionDescriptor &ID);
- /// Adds \p Phi to the monotonic PHI list and collects load/store users of
- /// the phi. Returns true if all users of \p Phi are legal for vectorization.
- bool addMonotonicPHI(PHINode *Phi, const MonotonicDescriptor &MD);
+ /// Adds \p Phi to the conditional induction list and collects load/store
+ /// users of the PHI. Returns true if all users of \p Phi are legal for
+ /// vectorization.
+ bool addConditionalInduction(PHINode *Phi,
+ const ConditionalInductionDescriptor &CondID);
/// The loop that we evaluate.
Loop *TheLoop;
@@ -729,8 +736,8 @@ class LoopVectorizationLegality {
/// variables can be pointers.
InductionList Inductions;
- /// Holds all of the monotonic phi variables that we found in the loop.
- MonotonicPHIList MonotonicPHIs;
+ /// Holds all of the conditional inductions found in the loop.
+ ConditionalInductionList ConditionalInductions;
/// Holds all the casts that participate in the update chain of the induction
/// variables, and that have been proven to be redundant (possibly under a
@@ -770,8 +777,8 @@ class LoopVectorizationLegality {
SmallVector<HistogramInfo, 1> Histograms;
/// Contains all pointers used in the loop that are defined using an index
- /// derived from a monotonic PHI. Loads/stores to these pointers map to
- /// expandloads or compressstores.
+ /// derived from a conditional induction PHI. Loads/stores to these pointers
+ /// map to expandloads or compressstores.
SmallDenseMap<const Value *, CompressedPtrInfo> CompressedPtrs;
/// Whether or not creating SCEV predicates is allowed.
diff --git a/llvm/lib/Transforms/Utils/LoopUtils.cpp b/llvm/lib/Transforms/Utils/LoopUtils.cpp
index 4715cb14d941d..02a3ca95d8cba 100644
--- a/llvm/lib/Transforms/Utils/LoopUtils.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUtils.cpp
@@ -2541,16 +2541,16 @@ llvm::hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
bool llvm::collectCompressedPtrs(
DenseMap<Value *, const SCEV *> &CompressedPtrs, const Loop &L,
- const MonotonicDescriptor &MD, ScalarEvolution &SE) {
- // Over-approximates the monotonic PHI as a SCEVAddRec assuming the condition
- // is always true.
+ const ConditionalInductionDescriptor &CondID, ScalarEvolution &SE) {
+ // Over-approximates the conditional induction as a SCEVAddRec assuming the
+ // condition is always true.
const SCEV *ApproximatePhiSCEV = SE.getAddRecExpr(
- MD.getStartSCEV(), MD.getStepSCEV(), &L, SCEV::FlagAnyWrap);
+ CondID.getStartSCEV(), CondID.getStepSCEV(), &L, SCEV::FlagAnyWrap);
// TODO: Take into account the non-wrap flags of the MD when rewriting the
// SCEV expressions for pointers. This should allow folding away zext/sext
// operations.
- ValueToSCEVMapTy PhiMap{{MD.getHeaderPHI(), ApproximatePhiSCEV}};
+ ValueToSCEVMapTy PhiMap{{CondID.getHeaderPHI(), ApproximatePhiSCEV}};
auto GetCompressedPtrSCEV = [&](Value *Ptr, Type *AccessTy) -> const SCEV * {
const SCEV *PtrSCEV =
@@ -2568,7 +2568,8 @@ bool llvm::collectCompressedPtrs(
};
SmallPtrSet<Use *, 16> Seen;
- SmallVector<Use *> Worklist(make_pointer_range(MD.getHeaderPHI()->uses()));
+ SmallVector<Use *> Worklist(
+ make_pointer_range(CondID.getHeaderPHI()->uses()));
while (!Worklist.empty()) {
Use *U = Worklist.pop_back_val();
if (!Seen.insert(U).second)
@@ -2576,7 +2577,7 @@ bool llvm::collectCompressedPtrs(
// Always allow uses outside the loop or by the backedge update.
auto *I = cast<Instruction>(U->getUser());
- if (I == MD.getBackedgePHI() || !L.contains(I))
+ if (I == CondID.getBackedgePHI() || !L.contains(I))
continue;
Value *CurrentVal = U->get();
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
index 8db26d29ed6dc..e38f86bec3a96 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
@@ -48,9 +48,9 @@ AllowStridedPointerIVs("lv-strided-pointer-ivs", cl::init(false), cl::Hidden,
cl::desc("Enable recognition of non-constant strided "
"pointer induction variables."));
-static cl::opt<bool> EnableMonotonicPatterns(
- "lv-monotonic-patterns", cl::init(false), cl::Hidden,
- cl::desc("Enable recognition of monotonic patterns."));
+static cl::opt<bool> EnableCompressingPatterns(
+ "lv-compressing-patterns", cl::init(false), cl::Hidden,
+ cl::desc("Enable recognition of compressing patterns."));
static cl::opt<bool>
HintsAllowReordering("hints-allow-reordering", cl::init(true), cl::Hidden,
@@ -468,8 +468,8 @@ int LoopVectorizationLegality::isConsecutivePtr(Type *AccessTy,
? LAI->getSymbolicStrides()
: SymbolicStrideMap();
- // Check if the pointer is derived from a a monotonic PHI. If so, return a
- // conservative stride (assuming the PHI is always updated).
+ // Check if the pointer is derived from a conditional induction PHI. If so,
+ // return a conservative stride (assuming the PHI is always updated).
if (std::optional<CompressedPtrInfo> PtrInfo = getCompressedPtrInfo(Ptr))
return getStrideFromAddRec(PtrInfo->PtrSCEV, TheLoop, AccessTy, Ptr, PSE)
.value_or(0);
@@ -757,26 +757,28 @@ void LoopVectorizationLegality::addInductionPhi(PHINode *Phi,
LLVM_DEBUG(dbgs() << "LV: Found an induction variable.\n");
}
-bool LoopVectorizationLegality::addMonotonicPHI(PHINode *Phi,
- const MonotonicDescriptor &MD) {
+bool LoopVectorizationLegality::addConditionalInduction(
+ PHINode *Phi, const ConditionalInductionDescriptor &CondID) {
for (User *U : Phi->users()) {
if (!TheLoop->contains(cast<Instruction>(U))) {
reportVectorizationFailure(
- "Unsupported out-of-loop user of monotonic phi",
- "UnsupportedMonotonicUse", ORE, TheLoop);
+ "Unsupported out-of-loop user of conditional induction phi",
+ "UnsupportedConditionalInductionUse", ORE, TheLoop);
return false;
}
}
- MonotonicPHIs[Phi] = MD;
- DenseMap<Value *, const SCEV *> CompressedPtrsForMD;
- if (!collectCompressedPtrs(CompressedPtrsForMD, *TheLoop, MD, *PSE.getSE())) {
- reportVectorizationFailure("Unsupported user of monotonic phi in loop",
- "UnsupportedMonotonicUse", ORE, TheLoop);
+ ConditionalInductions[Phi] = CondID;
+ DenseMap<Value *, const SCEV *> CompressedPtrsForCondID;
+ if (!collectCompressedPtrs(CompressedPtrsForCondID, *TheLoop, CondID,
+ *PSE.getSE())) {
+ reportVectorizationFailure(
+ "Unsupported user of conditional induction phi in loop",
+ "UnsupportedConditionalInductionUse", ORE, TheLoop);
return false;
}
- for (auto [Ptr, PtrSCEV] : CompressedPtrsForMD) {
+ for (auto [Ptr, PtrSCEV] : CompressedPtrsForCondID) {
auto *PtrAddRec = cast<SCEVAddRecExpr>(PtrSCEV);
assert(PtrAddRec->isAffine() && "Expected affine SCEVAddRecExpr");
CompressedPtrs[Ptr] = CompressedPtrInfo{Phi, PtrAddRec};
@@ -944,10 +946,11 @@ bool LoopVectorizationLegality::canVectorizeInstr(Instruction &I) {
return true;
}
- MonotonicDescriptor MD;
- if (EnableMonotonicPatterns &&
- MonotonicDescriptor::isMonotonicPHI(Phi, TheLoop, MD, *PSE.getSE())) {
- return addMonotonicPHI(Phi, MD);
+ ConditionalInductionDescriptor CondID;
+ if (EnableCompressingPatterns &&
+ ConditionalInductionDescriptor::isConditionalInductionPHI(
+ Phi, TheLoop, CondID, *PSE.getSE())) {
+ return addConditionalInduction(Phi, CondID);
}
if (RecurrenceDescriptor::isFixedOrderRecurrence(Phi, TheLoop, DT)) {
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 9d3614d3b36c8..195dc2099ecd0 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -3291,7 +3291,7 @@ static bool willGenerateVectors(VPlan &Plan, ElementCount VF,
case VPRecipeBase::VPExpandSCEVSC:
case VPRecipeBase::VPPredInstPHISC:
case VPRecipeBase::VPBranchOnMaskSC:
- case VPRecipeBase::VPMonotonicPHISC:
+ case VPRecipeBase::VPConditionalInductionPHISC:
continue;
case VPRecipeBase::VPReductionSC:
case VPRecipeBase::VPActiveLaneMaskPHISC:
@@ -3689,8 +3689,8 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
if (Plan.hasEarlyExit())
return 1;
- // Monotonic vars don't support interleaving.
- if (Legal->hasMonotonicPHIs())
+ // Conditional inductions don't support interleaving.
+ if (Legal->hasConditionalInductions())
return 1;
const bool HasReductions =
@@ -6177,13 +6177,12 @@ VPHistogramRecipe *VPRecipeBuilder::widenIfHistogram(VPInstruction *VPI) {
VPI->getDebugLoc());
}
-VPWidenMemIntrinsicRecipe *
-VPRecipeBuilder::widenIfCompressedLoadOrStore(VPInstruction *VPI,
- VPMonotonicPHIRecipe *PhiR) {
+VPWidenMemIntrinsicRecipe *VPRecipeBuilder::widenIfCompressedLoadOrStore(
+ VPInstruction *VPI, VPConditionalInductionPHIRecipe *PhiR) {
Instruction *I = VPI->getUnderlyingInstr();
std::optional<CompressedPtrInfo> Info = Legal->isCompressedLoadOrStore(I);
- if (!Info || Info->MonotonicPHI != PhiR->getPHINode())
+ if (!Info || Info->ConditionalInductionPHI != PhiR->getPHINode())
return nullptr;
VPBuilder::InsertPointGuard Guard(Builder);
@@ -6469,7 +6468,7 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
if (!RUN_VPLAN_PASS(
VPlanTransforms::createHeaderPhiRecipes, *VPlan0, PSE, *OrigLoop,
VPDT, Legal->getInductionVars(), Legal->getReductionVars(),
- Legal->getMonotonicPHIs(), Legal->getFixedOrderRecurrences(),
+ Legal->getConditionalInductions(), Legal->getFixedOrderRecurrences(),
Config.getInLoopReductions(), Config.getHints().allowReordering())) {
return nullptr;
}
@@ -7572,7 +7571,8 @@ static SmallVector<Instruction *> preparePlanForEpilogueVectorLoop(
}
} else {
// Retrieve the induction resume value via ResumeForEpilogue.
- assert(isa<VPWidenInductionRecipe>(&R) || isa<VPMonotonicPHIRecipe>(&R));
+ assert(isa<VPWidenInductionRecipe>(&R) ||
+ isa<VPConditionalInductionPHIRecipe>(&R));
PHINode *IndPhi = cast<VPHeaderPHIRecipe>(&R)->getPHINode();
ResumeV = IRPhiToResumeForEpi.at(IndPhi)->getUnderlyingValue();
}
@@ -7983,11 +7983,11 @@ bool LoopVectorizePass::processLoop(Loop *L) {
unsigned SelectedIC = std::max(IC, UserIC);
- if (LVL.hasMonotonicPHIs() && SelectedIC > 1) {
+ if (LVL.hasConditionalInductions() && SelectedIC > 1) {
reportVectorizationFailure(
- "Interleaving of loop with monotonic vars",
- "Interleaving of loops with monotonic vars is not supported",
- "CantInterleaveWithMonotonicVars", ORE, L);
+ "Interleaving of loop with conditional inductions",
+ "Interleaving of loops with conditional inductions is not supported",
+ "CantInterleaveWithConditionalInductions", ORE, L);
return false;
}
diff --git a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
index e4cef6f372732..79bb6f2fbaa77 100644
--- a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
+++ b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
@@ -82,7 +82,8 @@ class VPRecipeBuilder {
/// LoopVectorizationLegality) whose pointer is derived from \p PhiR, lower it
/// to a llvm.masked.expandload or llvm.masked.compressstore intrinsic.
VPWidenMemIntrinsicRecipe *
- widenIfCompressedLoadOrStore(VPInstruction *VPI, VPMonotonicPHIRecipe *PhiR);
+ widenIfCompressedLoadOrStore(VPInstruction *VPI,
+ VPConditionalInductionPHIRecipe *PhiR);
/// If \p VPI is a store of a reduction into an invariant address, delete it.
/// If it is the final store of a reduction result, a uniform store recipe
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.cpp b/llvm/lib/Transforms/Vectorize/VPlan.cpp
index 0cb7535187d18..649775ffb9981 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlan.cpp
@@ -370,7 +370,8 @@ void VPTransformState::fixupHeaderPhis() {
for (VPRecipeBase &R : Header->phis()) {
auto *PhiR = cast<VPSingleDefRecipe>(&R);
- bool NeedsScalar = isa<VPPhi>(PhiR) || isa<VPMonotonicPHIRecipe>(PhiR) ||
+ bool NeedsScalar = isa<VPPhi>(PhiR) ||
+ isa<VPConditionalInductionPHIRecipe>(PhiR) ||
(isa<VPReductionPHIRecipe>(PhiR) &&
cast<VPReductionPHIRecipe>(PhiR)->isInLoop());
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index 016606a2481f2..75f9d30057135 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -462,13 +462,13 @@ class LLVM_ABI_FOR_TEST VPRecipeBase
VPWidenIntOrFpInductionSC,
VPWidenPointerInductionSC,
VPReductionPHISC,
- VPMonotonicPHISC,
+ VPConditionalInductionPHISC,
// END: SubclassID for recipes that inherit VPHeaderPHIRecipe
// END: Phi-like recipes
VPFirstPHISC = VPWidenPHISC,
VPFirstHeaderPHISC = VPCurrentIterationPHISC,
- VPLastHeaderPHISC = VPMonotonicPHISC,
- VPLastPHISC = VPMonotonicPHISC,
+ VPLastHeaderPHISC = VPConditionalInductionPHISC,
+ VPLastPHISC = VPConditionalInductionPHISC,
};
VPRecipeBase(VPRecipeTy SC, ArrayRef<VPValue *> Operands,
@@ -661,7 +661,7 @@ class LLVM_ABI_FOR_TEST VPSingleDefRecipe : public VPRecipeBase,
case VPRecipeBase::VPReductionPHISC:
case VPRecipeBase::VPWidenLoadEVLSC:
case VPRecipeBase::VPWidenLoadSC:
- case VPRecipeBase::VPMonotonicPHISC:
+ case VPRecipeBase::VPConditionalInductionPHISC:
return true;
case VPRecipeBase::VPBranchOnMaskSC:
case VPRecipeBase::VPInterleaveEVLSC:
@@ -2959,14 +2959,15 @@ class VPReductionPHIRecipe : public VPHeaderPHIRecipe, public VPIRFlags {
#endif
};
-/// A recipe for handling monotonic phis. The start value is the first operand
-/// of the recipe, the incoming value from the backedge is the second
-/// operand, and the third operand is the step.
-class VPMonotonicPHIRecipe : public VPHeaderPHIRecipe {
+/// A recipe for handling conditional induction PHIs. The start value is the
+/// first operand of the recipe, the incoming value from the backedge is the
+/// second operand, and the third operand is the step.
+class VPConditionalInductionPHIRecipe : public VPHeaderPHIRecipe {
public:
- VPMonotonicPHIRecipe(PHINode &Phi, VPValue &Start, VPValue &BackedgeValue,
- VPValue &Step)
- : VPHeaderPHIRecipe(VPRecipeBase::VPMonotonicPHISC, &Phi, &Start) {
+ VPConditionalInductionPHIRecipe(PHINode &Phi, VPValue &Start,
+ VPValue &BackedgeValue, VPValue &Step)
+ : VPHeaderPHIRecipe(VPRecipeBase::VPConditionalInductionPHISC, &Phi,
+ &Start) {
addOperand(&BackedgeValue);
addOperand(&Step);
}
@@ -2975,17 +2976,17 @@ class VPMonotonicPHIRecipe : public VPHeaderPHIRecipe {
unsigned getNumIncoming() const override { return 2; }
- ~VPMonotonicPHIRecipe() override = default;
+ ~VPConditionalInductionPHIRecipe() override = default;
- VPMonotonicPHIRecipe *clone() override {
- return new VPMonotonicPHIRecipe(*getPHINode(), *getStartValue(),
- *getBackedgeValue(), *getStep());
+ VPConditionalInductionPHIRecipe *clone() override {
+ return new VPConditionalInductionPHIRecipe(*getPHINode(), *getStartValue(),
+ *getBackedgeValue(), *getStep());
}
- VP_CLASSOF_IMPL(VPRecipeBase::VPMonotonicPHISC)
+ VP_CLASSOF_IMPL(VPRecipeBase::VPConditionalInductionPHISC)
static inline bool classof(const VPHeaderPHIRecipe *R) {
- return R->getVPRecipeID() == VPRecipeBase::VPMonotonicPHISC;
+ return R->getVPRecipeID() == VPRecipeBase::VPConditionalInductionPHISC;
}
void execute(VPTransformState &State) override;
@@ -4410,7 +4411,7 @@ template <>
struct CastInfo<VPPhiAccessors, VPRecipeBase *>
: vpdetail::CastInfoMixinImpl<VPPhiAccessors, VPPhi, VPIRPhi,
VPWidenPHIRecipe, VPHeaderPHIRecipe,
- VPMonotonicPHIRecipe> {};
+ VPConditionalInductionPHIRecipe> {};
template <>
struct CastInfo<VPPhiAccessors, const VPRecipeBase *>
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index e5c932f0ba716..be4896e1e4625 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -928,7 +928,8 @@ bool VPlanTransforms::createHeaderPhiRecipes(
const VPDominatorTree &VPDT,
const MapVector<PHINode *, InductionDescriptor> &Inductions,
const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
- const MapVector<PHINode *, MonotonicDescriptor> &MonotonicPHIs,
+ const MapVector<PHINode *, ConditionalInductionDescriptor>
+ &ConditionalInductions,
const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering) {
// Retrieve the header manually from the intial plain-CFG VPlan.
@@ -961,12 +962,14 @@ bool VPlanTransforms::createHeaderPhiRecipes(
Plan, PSE, OrigLoop,
PhiR->getDebugLoc());
- auto MonotonicIt = MonotonicPHIs.find(Phi);
- if (MonotonicIt != MonotonicPHIs.end()) {
- const MonotonicDescriptor &MD = MonotonicIt->second;
+ auto ConditionalInductionIt = ConditionalInductions.find(Phi);
+ if (ConditionalInductionIt != ConditionalInductions.end()) {
+ const ConditionalInductionDescriptor &CondID =
+ ConditionalInductionIt->second;
VPValue *Step =
- vputils::getOrCreateVPValueForSCEVExpr(Plan, MD.getStepSCEV());
- return new VPMonotonicPHIRecipe(*Phi, *Start, *BackedgeValue, *Step);
+ vputils::getOrCreateVPValueForSCEVExpr(Plan, CondID.getStepSCEV());
+ return new VPConditionalInductionPHIRecipe(*Phi, *Start, *BackedgeValue,
+ *Step);
}
assert(Reductions.contains(Phi) && "only reductions are expected now");
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index a1ed8ac5791cf..a260a1a3f4b98 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -151,7 +151,7 @@ bool VPRecipeBase::mayReadFromMemory() const {
case VPWidenStoreEVLSC:
case VPWidenStoreSC:
case VPExpandSCEVSC:
- case VPMonotonicPHISC:
+ case VPConditionalInductionPHISC:
return false;
case VPBlendSC:
case VPReductionEVLSC:
@@ -5147,14 +5147,14 @@ bool VPBlendRecipe::usesFirstLaneOnly(const VPValue *Op) const {
return vputils::onlyFirstLaneUsed(this);
}
-void VPMonotonicPHIRecipe::execute(VPTransformState &State) {
- executePhiRecipe(this, *this, State, /*IsScalar=*/true, "monotonic.iv");
+void VPConditionalInductionPHIRecipe::execute(VPTransformState &State) {
+ executePhiRecipe(this, *this, State, /*IsScalar=*/true, "conditional.iv");
}
#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
-void VPMonotonicPHIRecipe::printRecipe(raw_ostream &O, const Twine &Indent,
- VPSlotTracker &SlotTracker) const {
- O << Indent << "MONOTONIC-PHI ";
+void VPConditionalInductionPHIRecipe::printRecipe(
+ raw_ostream &O, const Twine &Indent, VPSlotTracker &SlotTracker) const {
+ O << Indent << "CONDITIONAL-INDUCTION-PHI ";
printAsOperand(O, SlotTracker);
O << " = phi ";
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
index 56a6144c577cb..e08c9663c6e1a 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
@@ -6046,16 +6046,19 @@ bool VPlanTransforms::handleCompressingPatterns(
VPBuilder Builder;
for (VPRecipeBase &R : HeaderVPBB->phis()) {
- auto *MonotonicPhi = dyn_cast<VPMonotonicPHIRecipe>(&R);
- if (!MonotonicPhi)
+ auto *ConditionalInductionPhi =
+ dyn_cast<VPConditionalInductionPHIRecipe>(&R);
+ if (!ConditionalInductionPhi)
continue;
- // Obtain the mask for the monotonic phi update from the VPBlendRecipe.
- auto *BlendR = cast<VPBlendRecipe>(MonotonicPhi->getBackedgeValue());
+ // Obtain the mask for the conditional induction update from the
+ // VPBlendRecipe.
+ auto *BlendR =
+ cast<VPBlendRecipe>(ConditionalInductionPhi->getBackedgeValue());
VPValue *Mask = nullptr;
for (unsigned I = 0, E = BlendR->getNumIncomingValues(); I != E; ++I)
if (auto *IncomingVal = BlendR->getIncomingValue(I);
- IncomingVal != MonotonicPhi) {
+ IncomingVal != ConditionalInductionPhi) {
Mask = BlendR->getMask(I);
break;
}
@@ -6064,8 +6067,8 @@ bool VPlanTransforms::handleCompressingPatterns(
// Replace all "compressed" loads and stores with expandload and
// compressstore respectively.
for (VPInstruction *&VPI : MemOps) {
- auto *CompressedMemOp =
- RecipeBuilder.widenIfCompressedLoadOrStore(VPI, MonotonicPhi);
+ auto *CompressedMemOp = RecipeBuilder.widenIfCompressedLoadOrStore(
+ VPI, ConditionalInductionPhi);
if (!CompressedMemOp)
continue;
@@ -6073,7 +6076,7 @@ bool VPlanTransforms::handleCompressingPatterns(
Builder.insert(CompressedMemOp);
// Bail out if the mask for the memory op does not match the condition
- // used to update the montontic phi.
+ // used to update the conditional induction.
VPValue *MemOpMask = CompressedMemOp->getMask();
if (MemOpMask != Mask)
return false;
@@ -6089,13 +6092,13 @@ bool VPlanTransforms::handleCompressingPatterns(
remove_if(MemOps, [](VPInstruction *VPI) { return VPI == nullptr; }),
MemOps.end());
- // Update the monotonic PHI to increment by the number of active lanes in
- // the mask.
- auto *BackedgeVal = MonotonicPhi->getBackedgeValue();
+ // Update the conditional induction to increment by the number of active
+ // lanes in the mask.
+ auto *BackedgeVal = ConditionalInductionPhi->getBackedgeValue();
auto *InsertBlock = BackedgeVal->getDefiningRecipe()->getParent();
Builder.setInsertPoint(InsertBlock, InsertBlock->getFirstNonPhi());
- Type *UpdateType = MonotonicPhi->getScalarType();
+ Type *UpdateType = ConditionalInductionPhi->getScalarType();
if (UpdateType->isPointerTy())
UpdateType = Plan.getDataLayout().getIndexType(UpdateType);
@@ -6103,12 +6106,13 @@ bool VPlanTransforms::handleCompressingPatterns(
VPInstruction::NumActiveLanes, {Mask}, nullptr, {}, {},
DebugLoc::getUnknown(), "handled.lanes", UpdateType);
VPValue *Offset = Builder.createOverflowingOp(
- Instruction::Mul, {MonotonicPhi->getStep(), HandledLanes});
+ Instruction::Mul, {ConditionalInductionPhi->getStep(), HandledLanes});
VPValue *Update;
- if (MonotonicPhi->getScalarType()->isPointerTy())
- Update = Builder.createPtrAdd(MonotonicPhi, Offset);
+ if (ConditionalInductionPhi->getScalarType()->isPointerTy())
+ Update = Builder.createPtrAdd(ConditionalInductionPhi, Offset);
else
- Update = Builder.createAdd(MonotonicPhi, Offset, {}, "monotonic.add");
+ Update = Builder.createAdd(ConditionalInductionPhi, Offset, {},
+ "conditional.step");
BackedgeVal->replaceAllUsesWith(Update);
}
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
index 1c4bbb0f18c32..ff88de8394c1d 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
@@ -172,7 +172,8 @@ struct VPlanTransforms {
const VPDominatorTree &VPDT,
const MapVector<PHINode *, InductionDescriptor> &Inductions,
const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
- const MapVector<PHINode *, MonotonicDescriptor> &MonotonicPHIs,
+ const MapVector<PHINode *, ConditionalInductionDescriptor>
+ &ConditionalInductions,
const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering);
@@ -263,9 +264,10 @@ struct VPlanTransforms {
static bool handleFindLastReductions(VPlan &Plan);
/// Handles compressing memory loads/stores. Loads/stores where the pointer
- /// is derived from a monotonic PHI are replaced with expandloads or
- /// compressstores respectively. The backedge value of the monotonic PHI is
- /// updated to increment by the number of active lanes of the block mask.
+ /// is derived from a conditional induction PHI are replaced with expandloads
+ /// or compressstores respectively. The backedge value of the conditional
+ /// induction PHI is updated to increment by the number of active lanes of
+ /// the block mask.
/// Returns false if any memory operation could not be updated (e.g., due to
/// having a mask that does not match the PHI).
static bool handleCompressingPatterns(VPlan &Plan, VPBasicBlock *HeaderVPBB,
diff --git a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
index 908d6d040023f..ec9e0f930f7c1 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
@@ -462,7 +462,7 @@ bool vputils::isSingleScalar(const VPValue *VPV) {
if (auto *RR = dyn_cast<VPReductionRecipe>(VPV))
return !RR->isPartialReduction();
if (isa<VPVectorPointerRecipe, VPVectorEndPointerRecipe, VPDerivedIVRecipe,
- VPMonotonicPHIRecipe>(VPV))
+ VPConditionalInductionPHIRecipe>(VPV))
return true;
if (auto *Expr = dyn_cast<VPExpressionRecipe>(VPV))
return Expr->isVectorToScalar();
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
index 93bc4113fac9b..332365fe62150 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
@@ -1,37 +1,37 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^scalar.ph:" --version 5
-; RUN: opt < %s -lv-monotonic-patterns=true -mtriple=aarch64 -mattr=+sve2p2 -passes=loop-vectorize -S 2>&1 | FileCheck %s
+; RUN: opt < %s -lv-compressing-patterns=true -mtriple=aarch64 -mattr=+sve2p2 -passes=loop-vectorize -S 2>&1 | FileCheck %s
; SVE compresstore/expandload vectorization (requires +sve2p2 for expandload and +sve for compresstore).
define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-LABEL: define void @compress_store(
; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
-; CHECK-NEXT: [[VECTOR_PH:.*:]]
+; CHECK-NEXT: [[ENTRY:.*:]]
; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
-; CHECK-NEXT: [[TMP2:%.*]] = shl nuw i64 [[TMP0]], 2
-; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP2]]
-; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH1:.*]]
-; CHECK: [[VECTOR_PH1]]:
-; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP2]]
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP3]], align 4
-; CHECK-NEXT: [[TMP4:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP5]]
-; CHECK-NEXT: call void @llvm.masked.compressstore.nxv4i32.p0(<vscale x 4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP6]], <vscale x 4 x i1> [[TMP4]])
-; CHECK-NEXT: [[TMP7:%.*]] = zext <vscale x 4 x i1> [[TMP4]] to <vscale x 4 x i64>
-; CHECK-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP7]])
-; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP8]]
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP2]]
-; CHECK-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-NEXT: br i1 [[TMP9]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
-; CHECK: [[IF_THEN]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP2]], align 4
+; CHECK-NEXT: [[TMP3:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT: call void @llvm.masked.compressstore.nxv4i32.p0(<vscale x 4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP4]], <vscale x 4 x i1> [[TMP3]])
+; CHECK-NEXT: [[TMP5:%.*]] = zext <vscale x 4 x i1> [[TMP3]] to <vscale x 4 x i64>
+; CHECK-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP5]])
+; CHECK-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
@@ -66,33 +66,33 @@ exit:
define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-LABEL: define void @expand_load(
; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-NEXT: [[VECTOR_PH:.*:]]
+; CHECK-NEXT: [[ENTRY:.*:]]
; CHECK-NEXT: [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
-; CHECK-NEXT: [[TMP2:%.*]] = shl nuw i64 [[TMP0]], 2
-; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP2]]
-; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH1:.*]]
-; CHECK: [[VECTOR_PH1]]:
-; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP2]]
+; CHECK-NEXT: [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
-; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH1]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[TMP3:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
-; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP3]], align 4
-; CHECK-NEXT: [[TMP4:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-NEXT: [[TMP6:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[TMP5]]
-; CHECK-NEXT: [[TMP7:%.*]] = call <vscale x 4 x i32> @llvm.masked.expandload.nxv4i32.p0(ptr align 4 [[TMP6]], <vscale x 4 x i1> [[TMP4]], <vscale x 4 x i32> poison)
-; CHECK-NEXT: call void @llvm.masked.store.nxv4i32.p0(<vscale x 4 x i32> [[TMP7]], ptr align 4 [[TMP3]], <vscale x 4 x i1> [[TMP4]])
-; CHECK-NEXT: [[TMP8:%.*]] = zext <vscale x 4 x i1> [[TMP4]] to <vscale x 4 x i64>
-; CHECK-NEXT: [[TMP9:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP8]])
-; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP9]]
-; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP2]]
-; CHECK-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-NEXT: br i1 [[TMP10]], label %[[IF_THEN:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
-; CHECK: [[IF_THEN]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[TMP2:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP2]], align 4
+; CHECK-NEXT: [[TMP3:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT: [[TMP5:%.*]] = call <vscale x 4 x i32> @llvm.masked.expandload.nxv4i32.p0(ptr align 4 [[TMP4]], <vscale x 4 x i1> [[TMP3]], <vscale x 4 x i32> poison)
+; CHECK-NEXT: call void @llvm.masked.store.nxv4i32.p0(<vscale x 4 x i32> [[TMP5]], ptr align 4 [[TMP2]], <vscale x 4 x i1> [[TMP3]])
+; CHECK-NEXT: [[TMP6:%.*]] = zext <vscale x 4 x i1> [[TMP3]] to <vscale x 4 x i64>
+; CHECK-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP6]])
+; CHECK-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
; CHECK: [[SCALAR_PH]]:
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
index af74e9ec2a936..5187429c8f702 100644
--- a/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
@@ -1,5 +1,5 @@
; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --filter-out-after "^scalar.ph:" --version 6
-; RUN: opt -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -vplan-print-after=printOptimizedVPlan -disable-output %s -S 2>&1 | FileCheck %s
+; RUN: opt -lv-compressing-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -vplan-print-after=printOptimizedVPlan -disable-output %s -S 2>&1 | FileCheck %s
define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-LABEL: VPlan for loop in 'compress_store'
@@ -19,7 +19,7 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
; CHECK-EMPTY:
; CHECK-NEXT: vector.body:
-; CHECK-NEXT: MONOTONIC-PHI ir<%idx> = phi ir<0>, vp<%monotonic.add>, ir<1>
+; CHECK-NEXT: CONDITIONAL-INDUCTION-PHI ir<%idx> = phi ir<0>, vp<%conditional.step>, ir<1>
; CHECK-NEXT: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK-NEXT: CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
; CHECK-NEXT: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
@@ -27,9 +27,9 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: WIDEN ir<%cmp> = icmp slt ir<%load.src>, ir<%c>
; CHECK-NEXT: CLONE ir<%dst.ptr> = getelementptr inbounds ir<%dst>, ir<%idx>
; CHECK-NEXT: vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
-; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.compressstore(ir<%load.src>, vp<[[VP6]]>, ir<%cmp>)
+; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.compressstore(ir<%load.src>, vp<[[VP6]]>, ir<%cmp>) (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
; CHECK-NEXT: EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
-; CHECK-NEXT: EMIT vp<%monotonic.add> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT: EMIT vp<%conditional.step> = add ir<%idx>, vp<%handled.lanes>
; CHECK-NEXT: EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
; CHECK-NEXT: EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
; CHECK-NEXT: No successors
@@ -37,7 +37,7 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: Successor(s): middle.block
; CHECK-EMPTY:
; CHECK-NEXT: middle.block:
-; CHECK-NEXT: EMIT vp<[[VP9:%[0-9]+]]> = extract-last-part vp<%monotonic.add>
+; CHECK-NEXT: EMIT vp<[[VP9:%[0-9]+]]> = extract-last-part vp<%conditional.step>
; CHECK-NEXT: EMIT vp<[[VP10:%[0-9]+]]> = extract-last-lane vp<[[VP9]]>
; CHECK-NEXT: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK-NEXT: EMIT branch-on-cond vp<%cmp.n>
@@ -93,7 +93,7 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-NEXT: vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
; CHECK-EMPTY:
; CHECK-NEXT: vector.body:
-; CHECK-NEXT: MONOTONIC-PHI ir<%idx> = phi ir<0>, vp<%monotonic.add>, ir<1>
+; CHECK-NEXT: CONDITIONAL-INDUCTION-PHI ir<%idx> = phi ir<0>, vp<%conditional.step>, ir<1>
; CHECK-NEXT: vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
; CHECK-NEXT: CLONE ir<%dst.ptr> = getelementptr ir<%dst>, vp<[[VP4]]>
; CHECK-NEXT: vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
@@ -101,11 +101,11 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-NEXT: WIDEN ir<%cmp> = icmp slt ir<%load.dst>, ir<%c>
; CHECK-NEXT: CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, ir<%idx>
; CHECK-NEXT: vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
-; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.expandload(vp<[[VP6]]>, ir<%cmp>, ir<poison>)
+; CHECK-NEXT: WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.expandload(vp<[[VP6]]>, ir<%cmp>, ir<poison>) (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
; CHECK-NEXT: vp<[[VP8:%[0-9]+]]> = vector-pointer i32, ir<%dst.ptr>, ir<1>
-; CHECK-NEXT: WIDEN store vp<[[VP8]]>, vp<[[VP7]]>, ir<%cmp>
+; CHECK-NEXT: WIDEN store vp<[[VP8]]>, vp<[[VP7]]>, ir<%cmp> (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
; CHECK-NEXT: EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
-; CHECK-NEXT: EMIT vp<%monotonic.add> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT: EMIT vp<%conditional.step> = add ir<%idx>, vp<%handled.lanes>
; CHECK-NEXT: EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
; CHECK-NEXT: EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
; CHECK-NEXT: No successors
@@ -113,7 +113,7 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-NEXT: Successor(s): middle.block
; CHECK-EMPTY:
; CHECK-NEXT: middle.block:
-; CHECK-NEXT: EMIT vp<[[VP10:%[0-9]+]]> = extract-last-part vp<%monotonic.add>
+; CHECK-NEXT: EMIT vp<[[VP10:%[0-9]+]]> = extract-last-part vp<%conditional.step>
; CHECK-NEXT: EMIT vp<[[VP11:%[0-9]+]]> = extract-last-lane vp<[[VP10]]>
; CHECK-NEXT: EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
; CHECK-NEXT: EMIT branch-on-cond vp<%cmp.n>
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
index ef3c8d79d6d59..3d0ffa2b912db 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
@@ -1,4 +1,4 @@
-; RUN: opt < %s -lv-monotonic-patterns=true -enable-early-exit-vectorization-with-side-effects -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -disable-output -pass-remarks-missed=".*" 2>&1 | FileCheck %s
+; RUN: opt < %s -lv-compressing-patterns=true -enable-early-exit-vectorization-with-side-effects -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -disable-output -pass-remarks-missed=".*" 2>&1 | FileCheck %s
; CHECK: loop not vectorized
@@ -174,8 +174,8 @@ early.exit:
; CHECK: loop not vectorized
-; Negative test: Using the monotonic phi outside the loop is not supported.
-define i64 @out_of_loop_use_of_monotonic_phi(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; Negative test: Using the conditional induction outside the loop is not supported.
+define i64 @out_of_loop_use_of_conditional_induction(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
entry:
br label %for.body
@@ -205,7 +205,7 @@ exit:
; CHECK: loop not vectorized
-; Negative test: Matching an extended monotonic phi index is not supported yet.
+; Negative test: Matching an extended conditional induction index is not supported yet.
; Note: We should be able to support this case by using the no-wrap flags on %idx.next.
define void @test_compress_store_with_extended_index_with_nsw(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
entry:
@@ -238,7 +238,7 @@ exit:
; CHECK: loop not vectorized
-; Negative test: We can't vectorize a extended monotonic phi use without no-wrap flags on the step.
+; Negative test: We can't vectorize an extended conditional induction use without no-wrap flags on the step.
define void @test_compress_store_with_extended_index(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
entry:
br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
index 19010788ceac5..8d1581a124914 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
@@ -1,50 +1,50 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^for.body:" --version 5
-; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-IC1
-; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -tail-folding-policy=must-fold-tail -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-TF
-; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -force-vector-interleave=2 -passes=loop-vectorize -disable-output -pass-remarks-analysis=loop-vectorize 2>&1 | FileCheck %s --check-prefix=IC2
+; RUN: opt < %s -lv-compressing-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-IC1
+; RUN: opt < %s -lv-compressing-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -tail-folding-policy=must-fold-tail -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefix=CHECK-TF
+; RUN: opt < %s -lv-compressing-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=4 -force-vector-interleave=2 -passes=loop-vectorize -disable-output -pass-remarks-analysis=loop-vectorize 2>&1 | FileCheck %s --check-prefix=IC2
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
define void @test_compress_store_with_index(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define void @test_compress_store_with_index(
; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
-; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; CHECK-IC1: [[VECTOR_PH]]:
; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
-; CHECK-IC1: [[FOR_BODY]]:
-; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP2]])
; CHECK-IC1-NEXT: [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
; CHECK-IC1-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-IC1-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
; CHECK-IC1: [[MIDDLE_BLOCK]]:
; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
-; CHECK-IC1: [[SCALAR_PH1]]:
-; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
-; CHECK-IC1: [[FOR_BODY1]]:
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
;
; CHECK-TF-LABEL: define void @test_compress_store_with_index(
; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
-; CHECK-TF-NEXT: br label %[[EXIT:.*]]
-; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
@@ -55,26 +55,26 @@ define void @test_compress_store_with_index(ptr writeonly noalias %dst, ptr read
; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
-; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP4]])
; CHECK-TF-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
; CHECK-TF-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
; CHECK-TF-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
-; CHECK-TF: [[MIDDLE_BLOCK1]]:
-; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
-; CHECK-TF: [[EXIT1]]:
+; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
; CHECK-TF-NEXT: ret void
;
entry:
@@ -104,49 +104,49 @@ exit:
ret void
}
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
define void @test_expand_load_with_index(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define void @test_expand_load_with_index(
; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
-; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; CHECK-IC1: [[VECTOR_PH]]:
; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
-; CHECK-IC1: [[FOR_BODY]]:
-; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
; CHECK-IC1-NEXT: [[TMP4:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP3]], <4 x i1> [[TMP2]], <4 x i32> poison)
; CHECK-IC1-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP4]], ptr align 4 [[TMP1]], <4 x i1> [[TMP2]])
; CHECK-IC1-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP6]]
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-IC1-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
; CHECK-IC1: [[MIDDLE_BLOCK]]:
; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
-; CHECK-IC1: [[SCALAR_PH1]]:
-; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
-; CHECK-IC1: [[FOR_BODY1]]:
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
;
; CHECK-TF-LABEL: define void @test_expand_load_with_index(
; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
-; CHECK-TF-NEXT: br label %[[EXIT:.*]]
-; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
@@ -157,27 +157,27 @@ define void @test_expand_load_with_index(ptr noalias %dst, ptr readonly %src, i3
; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
-; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
; CHECK-TF-NEXT: [[TMP6:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP5]], <4 x i1> [[TMP4]], <4 x i32> poison)
; CHECK-TF-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP6]], ptr align 4 [[TMP2]], <4 x i1> [[TMP4]])
; CHECK-TF-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP8]]
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP8]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
; CHECK-TF-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-TF-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
-; CHECK-TF: [[MIDDLE_BLOCK1]]:
-; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
-; CHECK-TF: [[EXIT1]]:
+; CHECK-TF-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
; CHECK-TF-NEXT: ret void
;
entry:
@@ -208,48 +208,48 @@ exit:
ret void
}
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
define i64 @test_conditionally_incremented_phi_liveout(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-IC1-NEXT: [[SCALAR_PH:.*]]:
+; CHECK-IC1-NEXT: [[ENTRY:.*]]:
; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
-; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH1:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-IC1-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
; CHECK-IC1: [[VECTOR_PH]]:
; CHECK-IC1-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
; CHECK-IC1-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
; CHECK-IC1-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
; CHECK-IC1-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
-; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
-; CHECK-IC1: [[FOR_BODY]]:
-; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[FOR_BODY]] ]
+; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-IC1: [[VECTOR_BODY]]:
+; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP2]])
; CHECK-IC1-NEXT: [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
; CHECK-IC1-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-IC1-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK-IC1-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
; CHECK-IC1: [[MIDDLE_BLOCK]]:
; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH1]]
-; CHECK-IC1: [[SCALAR_PH1]]:
-; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[SCALAR_PH]] ]
-; CHECK-IC1-NEXT: br label %[[FOR_BODY1:.*]]
-; CHECK-IC1: [[FOR_BODY1]]:
+; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-IC1: [[SCALAR_PH]]:
+; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-IC1: [[FOR_BODY]]:
;
; CHECK-TF-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
-; CHECK-TF-NEXT: br label %[[EXIT:.*]]
-; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
+; CHECK-TF: [[VECTOR_PH]]:
; CHECK-TF-NEXT: [[N_RND_UP:%.*]] = add i64 [[N]], 3
; CHECK-TF-NEXT: [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
; CHECK-TF-NEXT: [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
@@ -260,27 +260,27 @@ define i64 @test_conditionally_incremented_phi_liveout(ptr writeonly noalias %ds
; CHECK-TF-NEXT: [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
-; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[EXIT]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[EXIT]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP4]])
; CHECK-TF-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
; CHECK-TF-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
; CHECK-TF-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK1:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
-; CHECK-TF: [[MIDDLE_BLOCK1]]:
-; CHECK-TF-NEXT: br label %[[EXIT1:.*]]
-; CHECK-TF: [[EXIT1]]:
-; CHECK-TF-NEXT: ret i64 [[MONOTONIC_ADD]]
+; CHECK-TF-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-TF: [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT: br label %[[EXIT:.*]]
+; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: ret i64 [[CONDITIONAL_STEP]]
;
entry:
br label %for.body
@@ -309,7 +309,7 @@ exit:
ret i64 %idx.1
}
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
define void @test_compress_store_with_scaled_pointer(ptr writeonly noalias %dst.bytes, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define void @test_compress_store_with_scaled_pointer(
@@ -325,25 +325,25 @@ define void @test_compress_store_with_scaled_pointer(ptr writeonly noalias %dst.
; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-IC1: [[VECTOR_BODY]]:
; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT: [[TMP3:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-IC1-NEXT: [[TMP4:%.*]] = shl nsw i64 [[TMP3]], 2
-; CHECK-IC1-NEXT: [[TMP5:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP4]]
-; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP2]])
-; CHECK-IC1-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
-; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP3]], [[TMP6]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = shl nsw i64 [[CONDITIONAL_IV]], 2
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP3]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP4]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
; CHECK-IC1: [[MIDDLE_BLOCK]]:
; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
; CHECK-IC1: [[SCALAR_PH]]:
; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-IC1: [[FOR_BODY]]:
;
@@ -363,23 +363,23 @@ define void @test_compress_store_with_scaled_pointer(ptr writeonly noalias %dst.
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[TMP5:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
; CHECK-TF-NEXT: [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP6:%.*]] = shl nsw i64 [[TMP5]], 2
-; CHECK-TF-NEXT: [[TMP7:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP6]]
-; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP7]], <4 x i1> [[TMP4]])
-; CHECK-TF-NEXT: [[TMP9:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
-; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP9]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[TMP5]], [[TMP8]]
+; CHECK-TF-NEXT: [[TMP5:%.*]] = shl nsw i64 [[CONDITIONAL_IV]], 2
+; CHECK-TF-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP5]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP6]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP8]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
-; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-TF-NEXT: br i1 [[TMP10]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
+; CHECK-TF-NEXT: [[TMP9:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
; CHECK-TF: [[MIDDLE_BLOCK]]:
; CHECK-TF-NEXT: br label %[[EXIT:.*]]
; CHECK-TF: [[EXIT]]:
@@ -413,7 +413,7 @@ exit:
ret void
}
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
; Test a nested conditional compress store, where the phi is only updated on iterations where the store takes place.
define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
@@ -430,17 +430,17 @@ define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst,
; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-IC1: [[VECTOR_BODY]]:
; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-IC1-NEXT: [[TMP4:%.*]] = icmp sgt <4 x i32> [[WIDE_LOAD]], zeroinitializer
; CHECK-IC1-NEXT: [[TMP5:%.*]] = select <4 x i1> [[TMP2]], <4 x i1> [[TMP4]], <4 x i1> zeroinitializer
; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP5]])
; CHECK-IC1-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP5]] to <4 x i64>
; CHECK-IC1-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP7]]
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-IC1-NEXT: br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
@@ -449,7 +449,7 @@ define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst,
; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
; CHECK-IC1: [[SCALAR_PH]]:
; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-IC1: [[FOR_BODY]]:
;
@@ -469,20 +469,20 @@ define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst,
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
-; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-TF-NEXT: [[TMP5:%.*]] = icmp sgt <4 x i32> [[WIDE_MASKED_LOAD]], zeroinitializer
; CHECK-TF-NEXT: [[TMP6:%.*]] = select <4 x i1> [[TMP3]], <4 x i1> [[TMP5]], <4 x i1> zeroinitializer
; CHECK-TF-NEXT: [[TMP7:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP6]], <4 x i1> zeroinitializer
; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP4]], <4 x i1> [[TMP7]])
; CHECK-TF-NEXT: [[TMP8:%.*]] = zext <4 x i1> [[TMP7]] to <4 x i64>
; CHECK-TF-NEXT: [[TMP9:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP8]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP9]]
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP9]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
@@ -523,7 +523,7 @@ exit:
ret void
}
-; An unconditional increment should lower as a simple induction (not a monotonic PHI).
+; An unconditional increment should lower as a simple induction (not a conditional induction).
define void @test_unconditional_increment(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define void @test_unconditional_increment(
; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
@@ -611,10 +611,10 @@ exit:
ret void
}
-; IC2: loop not vectorized: Interleaving of loops with monotonic vars is not supported
+; IC2: loop not vectorized: Interleaving of loops with conditional inductions is not supported
-define void @test_multiple_monotonic_phis(ptr %dst, ptr noalias %dst2, ptr noalias %src, ptr noalias %cond, ptr noalias %cond2, i64 %n) {
-; CHECK-IC1-LABEL: define void @test_multiple_monotonic_phis(
+define void @test_multiple_conditional_inductions(ptr %dst, ptr noalias %dst2, ptr noalias %src, ptr noalias %cond, ptr noalias %cond2, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_multiple_conditional_inductions(
; CHECK-IC1-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]], i64 [[N:%.*]]) {
; CHECK-IC1-NEXT: [[ENTRY:.*]]:
; CHECK-IC1-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
@@ -625,42 +625,42 @@ define void @test_multiple_monotonic_phis(ptr %dst, ptr noalias %dst2, ptr noali
; CHECK-IC1-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-IC1: [[VECTOR_BODY]]:
; CHECK-IC1-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT: [[MONOTONIC_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD4:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT: [[CONDITIONAL_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP4:%.*]], %[[VECTOR_BODY]] ]
; CHECK-IC1-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
; CHECK-IC1-NEXT: [[TMP2:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
; CHECK-IC1-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-IC1-NEXT: [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP3]], align 4
-; CHECK-IC1-NEXT: [[TMP4:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[MONOTONIC_IV]]
-; CHECK-IC1-NEXT: [[TMP14:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i8>
-; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP14]], ptr align 1 [[TMP4]], <4 x i1> [[TMP2]])
-; CHECK-IC1-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
-; CHECK-IC1-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP6]]
-; CHECK-IC1-NEXT: [[TMP7:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
-; CHECK-IC1-NEXT: [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP7]], align 4
-; CHECK-IC1-NEXT: [[TMP8:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD3]], zeroinitializer
-; CHECK-IC1-NEXT: [[TMP9:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i16>
-; CHECK-IC1-NEXT: [[TMP10:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[MONOTONIC_IV1]]
-; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP9]], ptr align 2 [[TMP10]], <4 x i1> [[TMP8]])
-; CHECK-IC1-NEXT: [[TMP11:%.*]] = zext <4 x i1> [[TMP8]] to <4 x i64>
-; CHECK-IC1-NEXT: [[TMP12:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP11]])
-; CHECK-IC1-NEXT: [[MONOTONIC_ADD4]] = add i64 [[MONOTONIC_IV1]], [[TMP12]]
+; CHECK-IC1-NEXT: [[TMP4:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-IC1-NEXT: [[TMP5:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i8>
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP5]], ptr align 1 [[TMP4]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT: [[TMP6:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
+; CHECK-IC1-NEXT: [[TMP8:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
+; CHECK-IC1-NEXT: [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP8]], align 4
+; CHECK-IC1-NEXT: [[TMP9:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-IC1-NEXT: [[TMP10:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i16>
+; CHECK-IC1-NEXT: [[TMP11:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[CONDITIONAL_IV1]]
+; CHECK-IC1-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP10]], ptr align 2 [[TMP11]], <4 x i1> [[TMP9]])
+; CHECK-IC1-NEXT: [[TMP12:%.*]] = zext <4 x i1> [[TMP9]] to <4 x i64>
+; CHECK-IC1-NEXT: [[TMP13:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP12]])
+; CHECK-IC1-NEXT: [[CONDITIONAL_STEP4]] = add i64 [[CONDITIONAL_IV1]], [[TMP13]]
; CHECK-IC1-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-IC1-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-IC1-NEXT: br i1 [[TMP13]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP14:![0-9]+]]
+; CHECK-IC1-NEXT: [[TMP14:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT: br i1 [[TMP14]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP14:![0-9]+]]
; CHECK-IC1: [[MIDDLE_BLOCK]]:
; CHECK-IC1-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
; CHECK-IC1-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
; CHECK-IC1: [[SCALAR_PH]]:
; CHECK-IC1-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
-; CHECK-IC1-NEXT: [[BC_MERGE_RDX5:%.*]] = phi i64 [ [[MONOTONIC_ADD4]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT: [[BC_MERGE_RDX5:%.*]] = phi i64 [ [[CONDITIONAL_STEP4]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-IC1: [[FOR_BODY]]:
;
-; CHECK-TF-LABEL: define void @test_multiple_monotonic_phis(
+; CHECK-TF-LABEL: define void @test_multiple_conditional_inductions(
; CHECK-TF-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]], i64 [[N:%.*]]) {
; CHECK-TF-NEXT: [[ENTRY:.*:]]
; CHECK-TF-NEXT: br label %[[VECTOR_PH:.*]]
@@ -674,8 +674,8 @@ define void @test_multiple_monotonic_phis(ptr %dst, ptr noalias %dst2, ptr noali
; CHECK-TF-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK-TF: [[VECTOR_BODY]]:
; CHECK-TF-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT: [[MONOTONIC_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD4:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT: [[CONDITIONAL_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP4:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
; CHECK-TF-NEXT: [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
; CHECK-TF-NEXT: [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
@@ -684,26 +684,26 @@ define void @test_multiple_monotonic_phis(ptr %dst, ptr noalias %dst2, ptr noali
; CHECK-TF-NEXT: [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD2:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP4]], <4 x i1> [[TMP1]], <4 x i32> poison)
; CHECK-TF-NEXT: [[TMP5:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[MONOTONIC_IV]]
-; CHECK-TF-NEXT: [[TMP17:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i8>
-; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP17]], ptr align 1 [[TMP6]], <4 x i1> [[TMP5]])
-; CHECK-TF-NEXT: [[TMP7:%.*]] = zext <4 x i1> [[TMP5]] to <4 x i64>
-; CHECK-TF-NEXT: [[TMP8:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP7]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP8]]
-; CHECK-TF-NEXT: [[TMP9:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
-; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD3:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP9]], <4 x i1> [[TMP1]], <4 x i32> poison)
-; CHECK-TF-NEXT: [[TMP10:%.*]] = icmp ne <4 x i32> [[WIDE_MASKED_LOAD3]], zeroinitializer
-; CHECK-TF-NEXT: [[TMP11:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP10]], <4 x i1> zeroinitializer
-; CHECK-TF-NEXT: [[TMP12:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i16>
-; CHECK-TF-NEXT: [[TMP13:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[MONOTONIC_IV1]]
-; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP12]], ptr align 2 [[TMP13]], <4 x i1> [[TMP11]])
-; CHECK-TF-NEXT: [[TMP14:%.*]] = zext <4 x i1> [[TMP11]] to <4 x i64>
-; CHECK-TF-NEXT: [[TMP15:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP14]])
-; CHECK-TF-NEXT: [[MONOTONIC_ADD4]] = add i64 [[MONOTONIC_IV1]], [[TMP15]]
+; CHECK-TF-NEXT: [[TMP6:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-TF-NEXT: [[TMP7:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i8>
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP7]], ptr align 1 [[TMP6]], <4 x i1> [[TMP5]])
+; CHECK-TF-NEXT: [[TMP8:%.*]] = zext <4 x i1> [[TMP5]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP9:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP8]])
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP9]]
+; CHECK-TF-NEXT: [[TMP10:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
+; CHECK-TF-NEXT: [[WIDE_MASKED_LOAD3:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP10]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT: [[TMP11:%.*]] = icmp ne <4 x i32> [[WIDE_MASKED_LOAD3]], zeroinitializer
+; CHECK-TF-NEXT: [[TMP12:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP11]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT: [[TMP13:%.*]] = trunc <4 x i32> [[WIDE_MASKED_LOAD2]] to <4 x i16>
+; CHECK-TF-NEXT: [[TMP14:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[CONDITIONAL_IV1]]
+; CHECK-TF-NEXT: call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP13]], ptr align 2 [[TMP14]], <4 x i1> [[TMP12]])
+; CHECK-TF-NEXT: [[TMP15:%.*]] = zext <4 x i1> [[TMP12]] to <4 x i64>
+; CHECK-TF-NEXT: [[TMP16:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP15]])
+; CHECK-TF-NEXT: [[CONDITIONAL_STEP4]] = add i64 [[CONDITIONAL_IV1]], [[TMP16]]
; CHECK-TF-NEXT: [[INDEX_NEXT]] = add i64 [[INDEX]], 4
; CHECK-TF-NEXT: [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
-; CHECK-TF-NEXT: [[TMP16:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
-; CHECK-TF-NEXT: br i1 [[TMP16]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK-TF-NEXT: [[TMP17:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT: br i1 [[TMP17]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
; CHECK-TF: [[MIDDLE_BLOCK]]:
; CHECK-TF-NEXT: br label %[[EXIT:.*]]
; CHECK-TF: [[EXIT]]:
@@ -758,15 +758,15 @@ exit:
define void @test_expand_load_always_false_cond(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-IC1-LABEL: define void @test_expand_load_always_false_cond(
; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-IC1-NEXT: [[SCALAR_PH:.*:]]
+; CHECK-IC1-NEXT: [[ENTRY:.*:]]
; CHECK-IC1-NEXT: br label %[[FOR_BODY:.*]]
; CHECK-IC1: [[FOR_BODY]]:
;
; CHECK-TF-LABEL: define void @test_expand_load_always_false_cond(
; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
-; CHECK-TF-NEXT: [[MIDDLE_BLOCK:.*:]]
-; CHECK-TF-NEXT: br label %[[EXIT:.*]]
-; CHECK-TF: [[EXIT]]:
+; CHECK-TF-NEXT: [[ENTRY:.*:]]
+; CHECK-TF-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-TF: [[FOR_BODY]]:
;
entry:
br label %for.body
@@ -795,3 +795,5 @@ for.inc:
exit:
ret void
}
+;; NOTE: These prefixes are unused and the list is autogenerated. Do not add tests below this line:
+; IC2: {{.*}}
diff --git a/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
index 346f7e5c2f6ec..1c8fc92d6035d 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
@@ -1,5 +1,5 @@
; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^for.body:" --version 5
-; RUN: opt < %s -lv-monotonic-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=16 -epilogue-vectorization-force-VF=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefixes=CHECK
+; RUN: opt < %s -lv-compressing-patterns=true -force-target-supports-masked-memory-ops -force-vector-width=16 -epilogue-vectorization-force-VF=4 -passes=loop-vectorize -S 2>&1 | FileCheck %s -check-prefixes=CHECK
define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
; CHECK-LABEL: define void @compress_store(
@@ -18,15 +18,15 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: br label %[[VECTOR_BODY:.*]]
; CHECK: [[VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-NEXT: [[MONOTONIC_IV:%.*]] = phi i64 [ 42, %[[VECTOR_PH]] ], [ [[MONOTONIC_ADD:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 42, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
; CHECK-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
; CHECK-NEXT: [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
-; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV]]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
; CHECK-NEXT: call void @llvm.masked.compressstore.v16i32.p0(<16 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <16 x i1> [[TMP2]])
; CHECK-NEXT: [[TMP4:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
; CHECK-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP4]])
-; CHECK-NEXT: [[MONOTONIC_ADD]] = add i64 [[MONOTONIC_IV]], [[TMP5]]
+; CHECK-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
; CHECK-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
; CHECK-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
@@ -38,7 +38,7 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3:![0-9]+]]
; CHECK: [[VEC_EPILOG_PH]]:
; CHECK-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
-; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[MONOTONIC_ADD]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
; CHECK-NEXT: [[TMP7:%.*]] = and i64 [[N]], 3
; CHECK-NEXT: [[N_VEC2:%.*]] = sub i64 [[N]], [[TMP7]]
; CHECK-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
@@ -46,15 +46,15 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
; CHECK: [[VEC_EPILOG_VECTOR_BODY]]:
; CHECK-NEXT: [[INDEX5:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT9:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
-; CHECK-NEXT: [[MONOTONIC_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[MONOTONIC_ADD8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT: [[CONDITIONAL_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[CONDITIONAL_STEP8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
; CHECK-NEXT: [[TMP8:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX5]]
; CHECK-NEXT: [[WIDE_LOAD7:%.*]] = load <4 x i32>, ptr [[TMP8]], align 4
; CHECK-NEXT: [[TMP9:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
-; CHECK-NEXT: [[TMP10:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[MONOTONIC_IV6]]
+; CHECK-NEXT: [[TMP10:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV6]]
; CHECK-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD7]], ptr align 4 [[TMP10]], <4 x i1> [[TMP9]])
; CHECK-NEXT: [[TMP11:%.*]] = zext <4 x i1> [[TMP9]] to <4 x i64>
; CHECK-NEXT: [[TMP12:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP11]])
-; CHECK-NEXT: [[MONOTONIC_ADD8]] = add i64 [[MONOTONIC_IV6]], [[TMP12]]
+; CHECK-NEXT: [[CONDITIONAL_STEP8]] = add i64 [[CONDITIONAL_IV6]], [[TMP12]]
; CHECK-NEXT: [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 4
; CHECK-NEXT: [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT9]], [[N_VEC2]]
; CHECK-NEXT: br i1 [[TMP13]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
@@ -63,7 +63,7 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
; CHECK-NEXT: br i1 [[CMP_N10]], [[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
; CHECK: [[VEC_EPILOG_SCALAR_PH]]:
; CHECK-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC2]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
-; CHECK-NEXT: [[BC_MERGE_RDX11:%.*]] = phi i64 [ [[MONOTONIC_ADD8]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[MONOTONIC_ADD]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[ITER_CHECK]] ]
+; CHECK-NEXT: [[BC_MERGE_RDX11:%.*]] = phi i64 [ [[CONDITIONAL_STEP8]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[ITER_CHECK]] ]
; CHECK-NEXT: br label %[[FOR_BODY:.*]]
; CHECK: [[FOR_BODY]]:
;
diff --git a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
index bd51bc1fb007f..6568a4fe307ec 100644
--- a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
+++ b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
@@ -100,7 +100,7 @@ class VPlanTestIRBase : public testing::Test {
VPlanTransforms::createHeaderPhiRecipes(
*Plan, PSE, *L, VPDT, Inductions,
MapVector<PHINode *, RecurrenceDescriptor>(),
- MapVector<PHINode *, MonotonicDescriptor>(),
+ MapVector<PHINode *, ConditionalInductionDescriptor>(),
SmallPtrSet<const PHINode *, 1>(), SmallPtrSet<PHINode *, 1>(),
/*AllowReordering=*/false);
}
>From c89c91ac74184d1cecc3534bf99b26ad14bc3f73 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Mon, 21 Sep 2026 12:24:56 +0000
Subject: [PATCH 3/3] Add other target tests
---
.../LoopVectorize/RISCV/compress-idioms.ll | 120 ++++++++
.../LoopVectorize/X86/compress-idioms.ll | 287 ++++++++++++++++++
2 files changed, 407 insertions(+)
create mode 100644 llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
create mode 100644 llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
new file mode 100644
index 0000000000000..759c7e8ef00a7
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
@@ -0,0 +1,120 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^scalar.ph:" --version 5
+; RUN: opt < %s -lv-compressing-patterns=true -mtriple=riscv64 -mattr=+v -passes=loop-vectorize -force-vector-width=4 -S 2>&1 | FileCheck %s
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: define void @compress_store(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[EXIT:.*]], label %[[VECTOR_PH:.*]]
+; CHECK: [[VECTOR_PH]]:
+; CHECK-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK: [[FOR_BODY]]:
+; CHECK-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_BODY]] ]
+; CHECK-NEXT: [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[FOR_BODY]] ]
+; CHECK-NEXT: [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[SRC_PTR]], align 4
+; CHECK-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
+; CHECK-NEXT: call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[DST_PTR]], <4 x i1> [[TMP2]])
+; CHECK-NEXT: [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-NEXT: [[CONDITIONAL_STEP]] = add i64 [[IDX]], [[TMP5]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[IV]], 4
+; CHECK-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP6]], label %[[FOR_INC:.*]], label %[[FOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK: [[FOR_INC]]:
+; CHECK-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[CMP_N]], [[EXIT1:label %.*]], label %[[EXIT]]
+; CHECK: [[EXIT]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: define void @expand_load(
+; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
+; CHECK-NEXT: [[ENTRY:.*:]]
+; CHECK-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[FOR_BODY:.*]], label %[[IF_THEN:.*]]
+; CHECK: [[IF_THEN]]:
+; CHECK-NEXT: [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT: br label %[[FOR_INC:.*]]
+; CHECK: [[FOR_INC]]:
+; CHECK-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[IF_THEN]] ], [ [[INDEX_NEXT:%.*]], %[[FOR_INC]] ]
+; CHECK-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[IF_THEN]] ], [ [[CONDITIONAL_STEP:%.*]], %[[FOR_INC]] ]
+; CHECK-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT: [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-NEXT: [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT: [[TMP4:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP3]], <4 x i1> [[TMP2]], <4 x i32> poison)
+; CHECK-NEXT: call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP4]], ptr align 4 [[TMP1]], <4 x i1> [[TMP2]])
+; CHECK-NEXT: [[TMP5:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[FOR_INC]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK: [[MIDDLE_BLOCK]]:
+; CHECK-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT: br i1 [[EXITCOND_NOT]], [[EXIT:label %.*]], label %[[FOR_BODY]]
+; CHECK: [[FOR_BODY]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ %cmp = icmp slt i32 %load.dst, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
new file mode 100644
index 0000000000000..4e25987a3f967
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
@@ -0,0 +1,287 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --filter-out-after "^scalar.ph:" --version 5
+; RUN: opt < %s -lv-compressing-patterns=true -mtriple=x86_64-- -mcpu=x86-64-v4 -passes=loop-vectorize -S 2>&1 | FileCheck %s --check-prefix=CHECK-V4
+; RUN: opt < %s -lv-compressing-patterns=true -mtriple=x86_64-- -mcpu=icelake-server -passes=loop-vectorize -S 2>&1 | FileCheck %s --check-prefix=CHECK-ICELAKE
+; RUN: opt < %s -lv-compressing-patterns=true -mtriple=x86_64-- -mcpu=znver4 -passes=loop-vectorize -S 2>&1 | FileCheck %s --check-prefix=CHECK-ZNVER4
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-V4-LABEL: define void @compress_store(
+; CHECK-V4-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-V4-NEXT: [[ENTRY:.*]]:
+; CHECK-V4-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-V4: [[FOR_BODY]]:
+; CHECK-V4-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-V4-NEXT: [[IDX:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-V4-NEXT: [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
+; CHECK-V4-NEXT: [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
+; CHECK-V4-NEXT: [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
+; CHECK-V4-NEXT: br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
+; CHECK-V4: [[IF_THEN]]:
+; CHECK-V4-NEXT: [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
+; CHECK-V4-NEXT: store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
+; CHECK-V4-NEXT: [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
+; CHECK-V4-NEXT: br label %[[FOR_INC]]
+; CHECK-V4: [[FOR_INC]]:
+; CHECK-V4-NEXT: [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
+; CHECK-V4-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-V4-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; CHECK-V4-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
+; CHECK-V4: [[EXIT]]:
+; CHECK-V4-NEXT: ret void
+;
+; CHECK-ICELAKE-LABEL: define void @compress_store(
+; CHECK-ICELAKE-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-ICELAKE-NEXT: [[ENTRY:.*]]:
+; CHECK-ICELAKE-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-ICELAKE: [[FOR_BODY]]:
+; CHECK-ICELAKE-NEXT: [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-ICELAKE-NEXT: [[IDX:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-ICELAKE-NEXT: [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
+; CHECK-ICELAKE-NEXT: [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
+; CHECK-ICELAKE-NEXT: [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
+; CHECK-ICELAKE-NEXT: br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
+; CHECK-ICELAKE: [[IF_THEN]]:
+; CHECK-ICELAKE-NEXT: [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
+; CHECK-ICELAKE-NEXT: store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
+; CHECK-ICELAKE-NEXT: [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
+; CHECK-ICELAKE-NEXT: br label %[[FOR_INC]]
+; CHECK-ICELAKE: [[FOR_INC]]:
+; CHECK-ICELAKE-NEXT: [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
+; CHECK-ICELAKE-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-ICELAKE-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; CHECK-ICELAKE-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
+; CHECK-ICELAKE: [[EXIT]]:
+; CHECK-ICELAKE-NEXT: ret void
+;
+; CHECK-ZNVER4-LABEL: define void @compress_store(
+; CHECK-ZNVER4-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
+; CHECK-ZNVER4-NEXT: [[ENTRY:.*:]]
+; CHECK-ZNVER4-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 48
+; CHECK-ZNVER4-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ZNVER4: [[VECTOR_PH]]:
+; CHECK-ZNVER4-NEXT: [[TMP0:%.*]] = and i64 [[N]], 15
+; CHECK-ZNVER4-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-ZNVER4: [[VECTOR_BODY]]:
+; CHECK-ZNVER4-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-ZNVER4-NEXT: [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-ZNVER4-NEXT: [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-ZNVER4-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-ZNVER4-NEXT: call void @llvm.masked.compressstore.v16i32.p0(<16 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <16 x i1> [[TMP2]])
+; CHECK-ZNVER4-NEXT: [[TMP4:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
+; CHECK-ZNVER4-NEXT: [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP4]])
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
+; CHECK-ZNVER4-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-ZNVER4-NEXT: [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT: br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-ZNVER4: [[MIDDLE_BLOCK]]:
+; CHECK-ZNVER4-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-ZNVER4: [[SCALAR_PH]]:
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+ %load.src = load i32, ptr %src.ptr, align 4
+ %cmp = icmp slt i32 %load.src, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
+
+define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-V4-LABEL: define void @expand_load(
+; CHECK-V4-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
+; CHECK-V4-NEXT: [[ENTRY:.*:]]
+; CHECK-V4-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-V4-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-V4: [[VECTOR_PH]]:
+; CHECK-V4-NEXT: [[TMP0:%.*]] = and i64 [[N]], 7
+; CHECK-V4-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-V4-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-V4-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT]], <8 x i32> poison, <8 x i32> zeroinitializer
+; CHECK-V4-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-V4: [[VECTOR_BODY]]:
+; CHECK-V4-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-V4-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-V4-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-V4-NEXT: [[WIDE_LOAD:%.*]] = load <8 x i32>, ptr [[TMP1]], align 4
+; CHECK-V4-NEXT: [[TMP2:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-V4-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-V4-NEXT: [[TMP4:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP3]], <8 x i1> [[TMP2]], <8 x i32> poison)
+; CHECK-V4-NEXT: call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP4]], ptr align 4 [[TMP1]], <8 x i1> [[TMP2]])
+; CHECK-V4-NEXT: [[TMP5:%.*]] = zext <8 x i1> [[TMP2]] to <8 x i64>
+; CHECK-V4-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP5]])
+; CHECK-V4-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-V4-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; CHECK-V4-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-V4-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-V4: [[MIDDLE_BLOCK]]:
+; CHECK-V4-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-V4-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-V4: [[SCALAR_PH]]:
+;
+; CHECK-ICELAKE-LABEL: define void @expand_load(
+; CHECK-ICELAKE-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
+; CHECK-ICELAKE-NEXT: [[ENTRY:.*:]]
+; CHECK-ICELAKE-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-ICELAKE-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ICELAKE: [[VECTOR_PH]]:
+; CHECK-ICELAKE-NEXT: [[TMP0:%.*]] = and i64 [[N]], 7
+; CHECK-ICELAKE-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ICELAKE-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-ICELAKE-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT]], <8 x i32> poison, <8 x i32> zeroinitializer
+; CHECK-ICELAKE-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-ICELAKE: [[VECTOR_BODY]]:
+; CHECK-ICELAKE-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ICELAKE-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ICELAKE-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-ICELAKE-NEXT: [[WIDE_LOAD:%.*]] = load <8 x i32>, ptr [[TMP1]], align 4
+; CHECK-ICELAKE-NEXT: [[TMP2:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-ICELAKE-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-ICELAKE-NEXT: [[TMP4:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP3]], <8 x i1> [[TMP2]], <8 x i32> poison)
+; CHECK-ICELAKE-NEXT: call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP4]], ptr align 4 [[TMP1]], <8 x i1> [[TMP2]])
+; CHECK-ICELAKE-NEXT: [[TMP5:%.*]] = zext <8 x i1> [[TMP2]] to <8 x i64>
+; CHECK-ICELAKE-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP5]])
+; CHECK-ICELAKE-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-ICELAKE-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; CHECK-ICELAKE-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ICELAKE-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK-ICELAKE: [[MIDDLE_BLOCK]]:
+; CHECK-ICELAKE-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ICELAKE-NEXT: br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK-ICELAKE: [[SCALAR_PH]]:
+;
+; CHECK-ZNVER4-LABEL: define void @expand_load(
+; CHECK-ZNVER4-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
+; CHECK-ZNVER4-NEXT: [[ITER_CHECK:.*]]:
+; CHECK-ZNVER4-NEXT: [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-ZNVER4-NEXT: br i1 [[MIN_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK-ZNVER4: [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-ZNVER4-NEXT: [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[N]], 16
+; CHECK-ZNVER4-NEXT: br i1 [[MIN_ITERS_CHECK1]], label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ZNVER4: [[VECTOR_PH]]:
+; CHECK-ZNVER4-NEXT: [[TMP0:%.*]] = and i64 [[N]], 15
+; CHECK-ZNVER4-NEXT: [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT: br label %[[VECTOR_BODY:.*]]
+; CHECK-ZNVER4: [[VECTOR_BODY]]:
+; CHECK-ZNVER4-NEXT: [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-ZNVER4-NEXT: [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-ZNVER4-NEXT: [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-ZNVER4-NEXT: [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-ZNVER4-NEXT: [[TMP4:%.*]] = call <16 x i32> @llvm.masked.expandload.v16i32.p0(ptr align 4 [[TMP3]], <16 x i1> [[TMP2]], <16 x i32> poison)
+; CHECK-ZNVER4-NEXT: call void @llvm.masked.store.v16i32.p0(<16 x i32> [[TMP4]], ptr align 4 [[TMP1]], <16 x i1> [[TMP2]])
+; CHECK-ZNVER4-NEXT: [[TMP5:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
+; CHECK-ZNVER4-NEXT: [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP5]])
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-ZNVER4-NEXT: [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-ZNVER4-NEXT: [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT: br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-ZNVER4: [[MIDDLE_BLOCK]]:
+; CHECK-ZNVER4-NEXT: [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT: br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK-ZNVER4: [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-ZNVER4-NEXT: [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP0]], 8
+; CHECK-ZNVER4-NEXT: br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF5:![0-9]+]]
+; CHECK-ZNVER4: [[VEC_EPILOG_PH]]:
+; CHECK-ZNVER4-NEXT: [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT: [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT: [[TMP8:%.*]] = and i64 [[N]], 7
+; CHECK-ZNVER4-NEXT: [[N_VEC2:%.*]] = sub i64 [[N]], [[TMP8]]
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT: [[BROADCAST_SPLAT4:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT3]], <8 x i32> poison, <8 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT: br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
+; CHECK-ZNVER4: [[VEC_EPILOG_VECTOR_BODY]]:
+; CHECK-ZNVER4-NEXT: [[INDEX5:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT9:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[CONDITIONAL_STEP8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[TMP9:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX5]]
+; CHECK-ZNVER4-NEXT: [[WIDE_LOAD7:%.*]] = load <8 x i32>, ptr [[TMP9]], align 4
+; CHECK-ZNVER4-NEXT: [[TMP10:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
+; CHECK-ZNVER4-NEXT: [[TMP11:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV6]]
+; CHECK-ZNVER4-NEXT: [[TMP12:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP11]], <8 x i1> [[TMP10]], <8 x i32> poison)
+; CHECK-ZNVER4-NEXT: call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP12]], ptr align 4 [[TMP9]], <8 x i1> [[TMP10]])
+; CHECK-ZNVER4-NEXT: [[TMP13:%.*]] = zext <8 x i1> [[TMP10]] to <8 x i64>
+; CHECK-ZNVER4-NEXT: [[TMP14:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP13]])
+; CHECK-ZNVER4-NEXT: [[CONDITIONAL_STEP8]] = add i64 [[CONDITIONAL_IV6]], [[TMP14]]
+; CHECK-ZNVER4-NEXT: [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 8
+; CHECK-ZNVER4-NEXT: [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT9]], [[N_VEC2]]
+; CHECK-ZNVER4-NEXT: br i1 [[TMP15]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK-ZNVER4: [[VEC_EPILOG_MIDDLE_BLOCK]]:
+; CHECK-ZNVER4-NEXT: [[CMP_N10:%.*]] = icmp eq i64 [[N]], [[N_VEC2]]
+; CHECK-ZNVER4-NEXT: br i1 [[CMP_N10]], label %[[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
+; CHECK-ZNVER4: [[VEC_EPILOG_SCALAR_PH]]:
+; CHECK-ZNVER4-NEXT: [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC2]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT: [[BC_MERGE_RDX11:%.*]] = phi i64 [ [[CONDITIONAL_STEP8]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT: br label %[[FOR_BODY:.*]]
+; CHECK-ZNVER4: [[FOR_BODY]]:
+; CHECK-ZNVER4-NEXT: [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-ZNVER4-NEXT: [[IDX:%.*]] = phi i64 [ [[BC_MERGE_RDX11]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-ZNVER4-NEXT: [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
+; CHECK-ZNVER4-NEXT: [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
+; CHECK-ZNVER4-NEXT: [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
+; CHECK-ZNVER4-NEXT: br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
+; CHECK-ZNVER4: [[IF_THEN]]:
+; CHECK-ZNVER4-NEXT: [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IDX]]
+; CHECK-ZNVER4-NEXT: [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
+; CHECK-ZNVER4-NEXT: store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
+; CHECK-ZNVER4-NEXT: [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
+; CHECK-ZNVER4-NEXT: br label %[[FOR_INC]]
+; CHECK-ZNVER4: [[FOR_INC]]:
+; CHECK-ZNVER4-NEXT: [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
+; CHECK-ZNVER4-NEXT: [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
+; CHECK-ZNVER4-NEXT: [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; CHECK-ZNVER4-NEXT: br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP7:![0-9]+]]
+; CHECK-ZNVER4: [[EXIT]]:
+; CHECK-ZNVER4-NEXT: ret void
+;
+entry:
+ br label %for.body
+
+for.body:
+ %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+ %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+ %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+ %load.dst = load i32, ptr %dst.ptr, align 4
+ %cmp = icmp slt i32 %load.dst, %c
+ br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+ %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+ %load.src = load i32, ptr %src.ptr, align 4
+ store i32 %load.src, ptr %dst.ptr, align 4
+ %idx.next = add nsw i64 %idx, 1
+ br label %for.inc
+
+for.inc:
+ %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+ %iv.next = add nuw nsw i64 %iv, 1
+ %exitcond.not = icmp eq i64 %iv.next, %n
+ br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+ ret void
+}
More information about the llvm-commits
mailing list