[llvm] [LoopVectorize] Support vectorization of compressing patterns (PR #214491)

Benjamin Maxwell via llvm-commits llvm-commits at lists.llvm.org
Fri Sep 25 08:13:12 PDT 2026


https://github.com/MacDue updated https://github.com/llvm/llvm-project/pull/214491

>From 65ecccbc5569b43ef5cba656e6d858f51754aaf1 Mon Sep 17 00:00:00 2001
From: Sergey Kachkov <sergey.kachkov at syntacore.com>
Date: Wed, 5 Aug 2026 18:04:38 +0000
Subject: [PATCH 1/6] [LoopVectorize] Support vectorization of compressing
 patterns in VPlan

RFC link: https://discourse.llvm.org/t/rfc-loop-vectorization-of-compress-store-expand-load-patterns/86442

This adds loop vectorizer support for "compressing" patterns,
for example:

```
int dst_idx = 0;
for (int i = 0; i < n; i++) {
  if (cond[i])
    dst[dst_idx++] = src[i];
}
```

Can be vectorized with a `llvm.masked.compressstore` as:

```
int dst_idx = 0;
for (int i = 0; i < n; i++) {
  %cond = load(%cond) != 0
  %src = masked.load(%src[i], %cond)
  masked.compressstore(%src, %dst[dst_idx], %cond)
  dst_idx += num.active.lanes(%cond)
}
```

and:

```
int src_idx = 0;
for (int i = 0; i < n; i++) {
  if (cond[i])
    dst[i] = src[src_idx++];
}
```

Can be vectorized with a `llvm.masked.expandload` as:

```
int src_idx = 0;
for (int i = 0; i < n; i++) {
  %cond = load(%cond) != 0
  %src = masked.expandload(%src[src_idx], %cond)
  masked.store(%src, %dst[i], %cond)
  src_idx += num.active.lanes(%cond)
}
```

This uses the new `MonotonicDescriptor` to recognize
monotonic/compressing patterns. The phis are mapped to a new
`VPMonotonicPHIRecipe`, this will map to a scalar phi. We only
allow uniform uses of monotonic phis in the loop (e.g., as the pointer
to a compressed load/store).

Compressed loads/stores are recognized with
`LoopVectorizationLegality::isCompressedPtr`. Currently, we only allow
cases where:

- The (monotonic) pointer has a stride equal to the access size
- The memory operation is predicated with the same condition as the increment

This is a continuation Sergey Kachkov's patch (#140723).

There are a number of changes from the initial patch:

- Expandloads/compresstores directly use `VPWidenMemIntrinsic`
- ComputeMonotonicResult is replaced with existing VP instructions
- AArch64, VPlan, and target-agnostic tests have been added
- This style of vectorization if off by default
  - The switch can be flipped soon after this patch lands

Tests, fixes, and design rework

Add comment

Add out-of-loop use check

Fixups

Fixups

Fixups

Add test

Remove AArch64 flags

Rename MonotonicPHI (and related) to ConditionalInduction*

Add other target tests

Flip flag
---
 llvm/include/llvm/Analysis/VectorUtils.h      |  15 ++
 .../include/llvm/Transforms/Utils/LoopUtils.h |  12 ++
 .../Vectorize/LoopVectorizationLegality.h     |  56 ++++++-
 llvm/lib/Analysis/VectorUtils.cpp             |  45 ++++++
 llvm/lib/IR/IntrinsicInst.cpp                 |   2 +
 llvm/lib/Transforms/Utils/LoopUtils.cpp       |  75 +++++++++
 .../Vectorize/LoopVectorizationLegality.cpp   |  48 ++++++
 .../Vectorize/LoopVectorizationPlanner.cpp    |   7 +
 .../Vectorize/LoopVectorizationPlanner.h      |   6 +
 .../Transforms/Vectorize/LoopVectorize.cpp    | 103 ++++++++++--
 .../Transforms/Vectorize/VPRecipeBuilder.h    |   7 +
 llvm/lib/Transforms/Vectorize/VPlan.cpp       |   7 +-
 llvm/lib/Transforms/Vectorize/VPlan.h         |  72 +++++++--
 .../Vectorize/VPlanConstruction.cpp           |  12 ++
 .../lib/Transforms/Vectorize/VPlanRecipes.cpp |  37 ++++-
 .../Transforms/Vectorize/VPlanTransforms.cpp  |  91 +++++++++++
 .../Transforms/Vectorize/VPlanTransforms.h    |  12 ++
 llvm/lib/Transforms/Vectorize/VPlanUtils.cpp  |   4 +-
 .../LoopVectorize/VPlan/compress-idioms.ll    | 153 ++++++++++++++++++
 .../VPlan/vplan-print-before-after-all.ll     |   1 +
 .../compress-idioms-negative-tests.ll         |   4 +-
 .../Transforms/Vectorize/VPlanTestBase.h      |   1 +
 22 files changed, 736 insertions(+), 34 deletions(-)
 create mode 100644 llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll

diff --git a/llvm/include/llvm/Analysis/VectorUtils.h b/llvm/include/llvm/Analysis/VectorUtils.h
index b177d9eec2189..42058ac97e3fa 100644
--- a/llvm/include/llvm/Analysis/VectorUtils.h
+++ b/llvm/include/llvm/Analysis/VectorUtils.h
@@ -165,6 +165,21 @@ LLVM_ABI bool
 isVectorIntrinsicWithOverloadTypeAtArg(Intrinsic::ID ID, int OpdIdx,
                                        const TargetTransformInfo *TTI);
 
+/// Returns the argument index of the pointer parameter for the vector memory
+/// intrinsic \p ID, or `std::nullopt` if the intrinsic does not have a pointer
+/// operand.
+LLVM_ABI std::optional<unsigned>
+getVectorMemoryIntrinsicPointerArgIdx(Intrinsic::ID ID);
+
+/// Returns the argument index of the data value of the vector store intrinsic
+/// \p ID, or `std::nullopt` if the intrinsic does not have a data operand.
+LLVM_ABI std::optional<unsigned>
+getVectorStoreIntrinsicDataArgIdx(Intrinsic::ID ID);
+
+/// Returns the argument index of the mask for the vector intrinsic \p ID, or
+/// `std::nullopt` if the intrinsic does not have a mask operand.
+LLVM_ABI std::optional<unsigned> getVectorIntrinsicMaskArgIdx(Intrinsic::ID ID);
+
 /// Identifies if the vector form of the intrinsic that returns a struct is
 /// overloaded at the struct element index \p RetIdx. /// \p TTI is used to
 /// consider target specific intrinsics, if no target specific intrinsics
diff --git a/llvm/include/llvm/Transforms/Utils/LoopUtils.h b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
index 74c549be35ddf..1245c122e3906 100644
--- a/llvm/include/llvm/Transforms/Utils/LoopUtils.h
+++ b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
@@ -44,6 +44,8 @@ class TargetLibraryInfo;
 class LPPassManager;
 class Instruction;
 struct RuntimeCheckingPtrGroup;
+class ConditionalInductionDescriptor;
+
 typedef std::pair<const RuntimeCheckingPtrGroup *,
                   const RuntimeCheckingPtrGroup *>
     RuntimePointerCheck;
@@ -707,6 +709,16 @@ LLVM_ABI std::optional<IVConditionInfo>
 hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
                       const MemorySSA &MSSA, AAResults &AA);
 
+/// Collects pointer values (used by loads/stores) whose addresses are derived
+/// from the monotonic PHI described by \p MD. The pointer operands and
+/// approximate SCEV expressions (assuming the monotonic PHI always increments)
+/// for the pointers are placed in \p CompressedPtrs. Returns true if all
+/// in-loop users of the conditional induction are loads/stores.
+bool collectCompressedPtrs(DenseMap<Value *, const SCEV *> &CompressedPtrs,
+                           const Loop &L,
+                           const ConditionalInductionDescriptor &CondID,
+                           ScalarEvolution &SE);
+
 } // end namespace llvm
 
 #endif // LLVM_TRANSFORMS_UTILS_LOOPUTILS_H
diff --git a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
index 1d35938a8e13f..588c62bdf692c 100644
--- a/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
+++ b/llvm/include/llvm/Transforms/Vectorize/LoopVectorizationLegality.h
@@ -247,6 +247,13 @@ struct HistogramInfo {
       : Load(Load), Update(Update), Store(Store) {}
 };
 
+/// Holds details about a "compressed" pointer: the conditional induction PHI
+/// used to derive the pointer and the SCEV expression for the pointer.
+struct CompressedPtrInfo {
+  PHINode *ConditionalInductionPHI;
+  const SCEVAddRecExpr *PtrSCEV;
+};
+
 /// Indicates the characteristics of a loop with an uncountable exit.
 /// * None      -- No uncountable exit present.
 /// * ReadOnly  -- At least one uncountable exit in a readonly loop.
@@ -287,6 +294,11 @@ class LoopVectorizationLegality {
   /// induction descriptor.
   using InductionList = MapVector<PHINode *, InductionDescriptor>;
 
+  /// ConditionalInductionList saves conditional inductions and maps them to
+  /// their descriptors.
+  using ConditionalInductionList =
+      MapVector<PHINode *, ConditionalInductionDescriptor>;
+
   /// RecurrenceSet contains the phi nodes that are recurrences other than
   /// inductions and reductions.
   using RecurrenceSet = SmallPtrSet<const PHINode *, 8>;
@@ -330,6 +342,15 @@ class LoopVectorizationLegality {
   /// Returns the induction variables found in the loop.
   const InductionList &getInductionVars() const { return Inductions; }
 
+  /// Returns the conditional inductions found in the loop.
+  const ConditionalInductionList &getConditionalInductions() const {
+    return ConditionalInductions;
+  }
+
+  bool hasConditionalInductions() const {
+    return !ConditionalInductions.empty();
+  }
+
   /// Return the fixed-order recurrences found in the loop.
   RecurrenceSet &getFixedOrderRecurrences() { return FixedOrderRecurrences; }
 
@@ -462,6 +483,26 @@ class LoopVectorizationLegality {
   /// Returns a list of all known histogram operations in the loop.
   bool hasHistograms() const { return !Histograms.empty(); }
 
+  /// Returns the CompressedPtrInfo for \p Ptr if the pointer is defined via
+  /// a conditional induction PHI, otherwise std::nullopt.
+  std::optional<CompressedPtrInfo>
+  getCompressedPtrInfo(const Value *Ptr) const {
+    auto It = CompressedPtrs.find(Ptr);
+    if (It != CompressedPtrs.end())
+      return It->second;
+    return std::nullopt;
+  }
+
+  /// Returns the CompressedPtrInfo for \p I if it corresponds to a compressed
+  /// load or store (which can map to an llvm.masked.expandload or
+  /// llvm.masked.compressstore), otherwise std::nullopt.
+  std::optional<CompressedPtrInfo>
+  isCompressedLoadOrStore(const Instruction *I) {
+    if (isa<LoadInst, StoreInst>(I))
+      return getCompressedPtrInfo(getLoadStorePointerOperand(I));
+    return std::nullopt;
+  }
+
   PredicatedScalarEvolution *getPredicatedScalarEvolution() const {
     return &PSE;
   }
@@ -632,6 +673,12 @@ class LoopVectorizationLegality {
   /// better choice for the main induction than the existing one.
   void addInductionPhi(PHINode *Phi, const InductionDescriptor &ID);
 
+  /// Adds \p Phi to the conditional induction list and collects load/store
+  /// users of the PHI. Returns true if all users of \p Phi are legal for
+  /// vectorization.
+  bool addConditionalInduction(PHINode *Phi,
+                               const ConditionalInductionDescriptor &CondID);
+
   /// The loop that we evaluate.
   Loop *TheLoop;
 
@@ -676,12 +723,14 @@ class LoopVectorizationLegality {
   /// variables can be pointers.
   InductionList Inductions;
 
+  /// Holds all of the conditional inductions found in the loop.
+  ConditionalInductionList ConditionalInductions;
+
   /// Holds the phi nodes that are fixed-order recurrences.
   RecurrenceSet FixedOrderRecurrences;
 
   /// Holds the widest induction type encountered.
   IntegerType *WidestIndTy = nullptr;
-
   /// Vectorization requirements that will go through late-evaluation.
   LoopVectorizationRequirements *Requirements;
 
@@ -707,6 +756,11 @@ class LoopVectorizationLegality {
   /// may work on the same memory location.
   SmallVector<HistogramInfo, 1> Histograms;
 
+  /// Contains all pointers used in the loop that are defined using an index
+  /// derived from a conditional induction PHI. Loads/stores to these pointers
+  /// map to expandloads or compressstores.
+  SmallDenseMap<const Value *, CompressedPtrInfo> CompressedPtrs;
+
   /// Whether or not creating SCEV predicates is allowed.
   bool AllowRuntimeSCEVChecks;
 
diff --git a/llvm/lib/Analysis/VectorUtils.cpp b/llvm/lib/Analysis/VectorUtils.cpp
index 1c105ebb772b3..1270dbc2b413f 100644
--- a/llvm/lib/Analysis/VectorUtils.cpp
+++ b/llvm/lib/Analysis/VectorUtils.cpp
@@ -157,6 +157,7 @@ bool llvm::isVectorIntrinsicWithScalarOpAtArg(Intrinsic::ID ID,
   case Intrinsic::is_fpclass:
   case Intrinsic::powi:
   case Intrinsic::vector_extract:
+  case Intrinsic::masked_compressstore:
     return (ScalarOpdIdx == 1);
   case Intrinsic::smul_fix:
   case Intrinsic::smul_fix_sat:
@@ -171,6 +172,8 @@ bool llvm::isVectorIntrinsicWithScalarOpAtArg(Intrinsic::ID ID,
     return ScalarOpdIdx == 0 || ScalarOpdIdx == 1;
   case Intrinsic::experimental_vp_strided_store:
     return ScalarOpdIdx == 1 || ScalarOpdIdx == 2;
+  case Intrinsic::masked_expandload:
+    return ScalarOpdIdx == 0;
   case Intrinsic::loop_dependence_war_mask:
     return true;
   default:
@@ -196,6 +199,7 @@ bool llvm::isVectorIntrinsicWithOverloadTypeAtArg(
   case Intrinsic::scmp:
   case Intrinsic::vector_extract:
   case Intrinsic::loop_dependence_war_mask:
+  case Intrinsic::masked_expandload:
     return OpdIdx == -1 || OpdIdx == 0;
   case Intrinsic::modf:
   case Intrinsic::sincos:
@@ -209,11 +213,52 @@ bool llvm::isVectorIntrinsicWithOverloadTypeAtArg(
     return OpdIdx == -1 || OpdIdx == 0 || OpdIdx == 1;
   case Intrinsic::experimental_vp_strided_store:
     return OpdIdx == 0 || OpdIdx == 1 || OpdIdx == 2;
+  case Intrinsic::masked_compressstore:
+    return OpdIdx == 0 || OpdIdx == 1;
   default:
     return OpdIdx == -1;
   }
 }
 
+std::optional<unsigned>
+llvm::getVectorMemoryIntrinsicPointerArgIdx(Intrinsic::ID ID) {
+  if (auto PtrPos = VPIntrinsic::getMemoryPointerParamPos(ID))
+    return PtrPos;
+  switch (ID) {
+  case Intrinsic::masked_compressstore:
+    return 1;
+  case Intrinsic::masked_expandload:
+    return 0;
+  default:
+    return std::nullopt;
+  }
+}
+
+std::optional<unsigned>
+llvm::getVectorStoreIntrinsicDataArgIdx(Intrinsic::ID ID) {
+  if (auto DataPos = VPIntrinsic::getMemoryDataParamPos(ID))
+    return DataPos;
+  switch (ID) {
+  case Intrinsic::masked_expandload:
+    return 2;
+  default:
+    return std::nullopt;
+  }
+}
+
+std::optional<unsigned> llvm::getVectorIntrinsicMaskArgIdx(Intrinsic::ID ID) {
+  if (auto MaskPos = VPIntrinsic::getMaskParamPos(ID))
+    return MaskPos;
+  switch (ID) {
+  case Intrinsic::masked_compressstore:
+    return 2;
+  case Intrinsic::masked_expandload:
+    return 1;
+  default:
+    return std::nullopt;
+  }
+}
+
 bool llvm::isVectorIntrinsicWithStructReturnOverloadAtField(
     Intrinsic::ID ID, int RetIdx, const TargetTransformInfo *TTI) {
 
diff --git a/llvm/lib/IR/IntrinsicInst.cpp b/llvm/lib/IR/IntrinsicInst.cpp
index 684aaf1a8f2d3..1997aa3f27b9e 100644
--- a/llvm/lib/IR/IntrinsicInst.cpp
+++ b/llvm/lib/IR/IntrinsicInst.cpp
@@ -439,10 +439,12 @@ VPIntrinsic::getMemoryPointerParamPos(Intrinsic::ID VPID) {
   switch (VPID) {
   default:
     return std::nullopt;
+  case Intrinsic::masked_compressstore:
   case Intrinsic::vp_store:
   case Intrinsic::vp_scatter:
   case Intrinsic::experimental_vp_strided_store:
     return 1;
+  case Intrinsic::masked_expandload:
   case Intrinsic::vp_load:
   case Intrinsic::vp_load_ff:
   case Intrinsic::vp_gather:
diff --git a/llvm/lib/Transforms/Utils/LoopUtils.cpp b/llvm/lib/Transforms/Utils/LoopUtils.cpp
index 784c833152611..a981dcf6e4b40 100644
--- a/llvm/lib/Transforms/Utils/LoopUtils.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUtils.cpp
@@ -21,6 +21,7 @@
 #include "llvm/Analysis/BasicAliasAnalysis.h"
 #include "llvm/Analysis/DomTreeUpdater.h"
 #include "llvm/Analysis/GlobalsModRef.h"
+#include "llvm/Analysis/IVDescriptors.h"
 #include "llvm/Analysis/InstSimplifyFolder.h"
 #include "llvm/Analysis/LoopAccessAnalysis.h"
 #include "llvm/Analysis/LoopInfo.h"
@@ -2522,3 +2523,77 @@ llvm::hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
 
   return {};
 }
+
+bool llvm::collectCompressedPtrs(
+    DenseMap<Value *, const SCEV *> &CompressedPtrs, const Loop &L,
+    const ConditionalInductionDescriptor &CondID, ScalarEvolution &SE) {
+  // Over-approximates the conditional induction as a SCEVAddRec assuming the
+  // condition is always true.
+  const SCEV *ApproximatePhiSCEV = SE.getAddRecExpr(
+      CondID.getStartSCEV(), CondID.getStepSCEV(), &L, SCEV::FlagAnyWrap);
+
+  // TODO: Take into account the non-wrap flags of the MD when rewriting the
+  // SCEV expressions for pointers. This should allow folding away zext/sext
+  // operations.
+  ValueToSCEVMapTy PhiMap{{CondID.getHeaderPHI(), ApproximatePhiSCEV}};
+
+  auto GetCompressedPtrSCEV = [&](Value *Ptr, Type *AccessTy) -> const SCEV * {
+    const SCEV *PtrSCEV =
+        SCEVParameterRewriter::rewrite(SE.getSCEV(Ptr), SE, PhiMap);
+    auto *AddRec = dyn_cast<SCEVAddRecExpr>(PtrSCEV);
+    if (!AddRec || !AddRec->isAffine())
+      return nullptr;
+
+    // Check if pointer step equals access size.
+    SCEVUse Step = AddRec->getStepRecurrence(SE);
+    if (Step != SE.getSizeOfExpr(Step->getType(), AccessTy))
+      return nullptr;
+
+    return PtrSCEV;
+  };
+
+  SmallPtrSet<Use *, 16> Seen;
+  SmallVector<Use *> Worklist(
+      make_pointer_range(CondID.getHeaderPHI()->uses()));
+  while (!Worklist.empty()) {
+    Use *U = Worklist.pop_back_val();
+    if (!Seen.insert(U).second)
+      continue;
+
+    // Always allow uses outside the loop or by the backedge update.
+    auto *I = cast<Instruction>(U->getUser());
+    if (I == CondID.getBackedgePHI() || !L.contains(I))
+      continue;
+
+    Value *CurrentVal = U->get();
+    if (isa<LoadInst, StoreInst>(I)) {
+      // Disallow any store that uses the monotonic value as the stored value.
+      auto *SI = dyn_cast<StoreInst>(I);
+      if (SI && SI->getValueOperand() == CurrentVal)
+        return false;
+
+      Value *Ptr = getLoadStorePointerOperand(I);
+      const SCEV *PrtSCEV = GetCompressedPtrSCEV(Ptr, getLoadStoreType(I));
+      if (!PrtSCEV)
+        return false;
+      CompressedPtrs.insert({Ptr, PrtSCEV});
+      continue;
+    }
+
+    auto LoopVariantOp = [&](Value *V, bool /*AllowRepeats*/) -> Value * {
+      return L.isLoopInvariant(V) ? nullptr : V;
+    };
+
+    // Non-memory users may use any opcode (select/and/or/etc.), but they must
+    // only have CurrentVal as their only loop-varying input. That prevents
+    // mixing in a second loop-varying term. GetCompressedPtrSCEV rewrites the
+    // full leaf pointer SCEV and rejects it unless the entire address still
+    // simplifies to the required affine AddRec.
+    if (I->use_empty() ||
+        find_singleton<Value>(I->operands(), LoopVariantOp) != CurrentVal)
+      return false;
+    append_range(Worklist, make_pointer_range(I->uses()));
+  }
+
+  return true;
+}
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
index 862470ce0de6d..e5a0d91e7be34 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationLegality.cpp
@@ -48,6 +48,10 @@ AllowStridedPointerIVs("lv-strided-pointer-ivs", cl::init(false), cl::Hidden,
                        cl::desc("Enable recognition of non-constant strided "
                                 "pointer induction variables."));
 
+static cl::opt<bool> EnableCompressingPatterns(
+    "lv-compressing-patterns", cl::init(true), cl::Hidden,
+    cl::desc("Enable recognition of compressing patterns."));
+
 static cl::opt<bool>
     HintsAllowReordering("hints-allow-reordering", cl::init(true), cl::Hidden,
                          cl::desc("Allow enabling loop hints to reorder "
@@ -463,6 +467,13 @@ int LoopVectorizationLegality::isConsecutivePtr(Type *AccessTy,
   const auto &Strides = LAI && AllowRuntimeSCEVChecks
                             ? LAI->getSymbolicStrides()
                             : SymbolicStrideMap();
+
+  // Check if the pointer is derived from a conditional induction PHI. If so,
+  // return a conservative stride (assuming the PHI is always updated).
+  if (std::optional<CompressedPtrInfo> PtrInfo = getCompressedPtrInfo(Ptr))
+    return getStrideFromAddRec(PtrInfo->PtrSCEV, TheLoop, AccessTy, Ptr, PSE)
+        .value_or(0);
+
   SmallVector<const SCEVPredicate *> Predicates;
   int Stride = getPtrStride(PSE, AccessTy, Ptr, TheLoop, *DT, Strides, false,
                             AllowRuntimeSCEVChecks ? &Predicates : nullptr)
@@ -738,6 +749,36 @@ void LoopVectorizationLegality::addInductionPhi(PHINode *Phi,
   LLVM_DEBUG(dbgs() << "LV: Found an induction variable.\n");
 }
 
+bool LoopVectorizationLegality::addConditionalInduction(
+    PHINode *Phi, const ConditionalInductionDescriptor &CondID) {
+  for (User *U : Phi->users()) {
+    if (!TheLoop->contains(cast<Instruction>(U))) {
+      reportVectorizationFailure(
+          "Unsupported out-of-loop user of conditional induction phi",
+          "UnsupportedConditionalInductionUse", ORE, TheLoop);
+      return false;
+    }
+  }
+
+  ConditionalInductions[Phi] = CondID;
+  DenseMap<Value *, const SCEV *> CompressedPtrsForCondID;
+  if (!collectCompressedPtrs(CompressedPtrsForCondID, *TheLoop, CondID,
+                             *PSE.getSE())) {
+    reportVectorizationFailure(
+        "Unsupported user of conditional induction phi in loop",
+        "UnsupportedConditionalInductionUse", ORE, TheLoop);
+    return false;
+  }
+
+  for (auto [Ptr, PtrSCEV] : CompressedPtrsForCondID) {
+    auto *PtrAddRec = cast<SCEVAddRecExpr>(PtrSCEV);
+    assert(PtrAddRec->isAffine() && "Expected affine SCEVAddRecExpr");
+    CompressedPtrs[Ptr] = CompressedPtrInfo{Phi, PtrAddRec};
+  }
+
+  return true;
+}
+
 bool LoopVectorizationLegality::setupOuterLoopInductions() {
   BasicBlock *Header = TheLoop->getHeader();
 
@@ -897,6 +938,13 @@ bool LoopVectorizationLegality::canVectorizeInstr(Instruction &I) {
       return true;
     }
 
+    ConditionalInductionDescriptor CondID;
+    if (EnableCompressingPatterns &&
+        ConditionalInductionDescriptor::isConditionalInductionPHI(
+            Phi, TheLoop, CondID, *PSE.getSE())) {
+      return addConditionalInduction(Phi, CondID);
+    }
+
     if (RecurrenceDescriptor::isFixedOrderRecurrence(Phi, TheLoop, DT)) {
       FixedOrderRecurrences.insert(Phi);
       return true;
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
index 4b38c5dad753c..26d92284d31b0 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
@@ -159,6 +159,13 @@ bool VFSelectionContext::isLegalGatherOrScatter(bool IsLoad, Type *ScalarTy,
                  : TTI.isLegalMaskedScatter(VectorTy, Alignment));
 }
 
+bool VFSelectionContext::isLegalExpandLoadOrCompressStore(
+    bool IsLoad, Type *ScalarTy, Align Alignment) const {
+  return ForceTargetSupportsMaskedMemoryOps ||
+         (IsLoad ? TTI.isLegalMaskedExpandLoad(ScalarTy, Alignment)
+                 : TTI.isLegalMaskedCompressStore(ScalarTy, Alignment));
+}
+
 bool VFSelectionContext::supportsScalableVectors() const {
   return TTI.supportsScalableVectors() || ForceTargetSupportsScalableVectors ||
          VectorizerParams::VectorizationFactor.isScalable();
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
index 465f9331ba0f3..4fe32f4467104 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
@@ -839,6 +839,12 @@ class VFSelectionContext {
   bool isLegalGatherOrScatter(bool IsLoad, Type *ScalarTy, Align Alignment,
                               ElementCount VF) const;
 
+  /// Returns true if the target machine supports a masked expand load (if \p
+  /// IsLoad) or masked compress store of scalar type \p ScalarTy with \p
+  /// Alignment.
+  bool isLegalExpandLoadOrCompressStore(bool IsLoad, Type *ScalarTy,
+                                        Align Alignment) const;
+
   /// Split reductions into those that happen in the loop, and those that
   /// happen outside. In-loop reductions are collected into InLoopReductions.
   void collectInLoopReductions();
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 2f1fc4398654a..2078473933465 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -1073,6 +1073,10 @@ class LoopVectorizationCostModel {
   /// data type and alignment.
   bool isLegalGatherOrScatter(Instruction *I, ElementCount VF) const;
 
+  /// Returns true if the target machine supports a masked expand load or masked
+  /// compress store for \p I's data type and alignment.
+  bool isLegalExpandLoadOrCompressStore(Instruction *I) const;
+
   /// Check if \p Instr belongs to any interleaved access group.
   bool isAccessInterleaved(Instruction *Instr) const {
     return InterleaveInfo.isInterleaved(Instr);
@@ -2386,6 +2390,13 @@ bool LoopVectorizationCostModel::isLegalGatherOrScatter(Instruction *I,
                                        getLoadStoreAlignment(I), VF);
 }
 
+bool LoopVectorizationCostModel::isLegalExpandLoadOrCompressStore(
+    Instruction *I) const {
+  assert(isa<LoadInst>(I) || isa<StoreInst>(I));
+  return Config.isLegalExpandLoadOrCompressStore(
+      isa<LoadInst>(I), getLoadStoreType(I), getLoadStoreAlignment(I));
+}
+
 bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
                                                          ElementCount VF) {
   if (!isPredicatedInst(I))
@@ -2406,9 +2417,13 @@ bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
   }
   case Instruction::Load:
   case Instruction::Store: {
-    bool IsConsecutive = Legal->isConsecutivePtr(getLoadStoreType(I),
-                                                 getLoadStorePointerOperand(I));
-    return !(IsConsecutive && isLegalMaskedLoadOrStore(I, VF)) &&
+    Type *ScalarTy = getLoadStoreType(I);
+    Value *Ptr = getLoadStorePointerOperand(I);
+    bool IsCompressed = Legal->getCompressedPtrInfo(Ptr).has_value();
+    bool IsConsecutive = Legal->isConsecutivePtr(ScalarTy, Ptr);
+    return !(IsConsecutive && !IsCompressed &&
+             isLegalMaskedLoadOrStore(I, VF)) &&
+           !(IsCompressed && isLegalExpandLoadOrCompressStore(I)) &&
            !isLegalGatherOrScatter(I, VF);
   }
   case Instruction::UDiv:
@@ -2647,8 +2662,11 @@ LoopVectorizationCostModel::memoryInstructionCanBeWidened(Instruction *I,
   auto *Ptr = getLoadStorePointerOperand(I);
   auto *ScalarTy = getLoadStoreType(I);
 
-  // In order to be widened, the pointer should be consecutive, first of all.
+  // In order to be widened, the pointer should be consecutive or compressed.
   int Stride = Legal->isConsecutivePtr(ScalarTy, Ptr);
+  assert((Stride == 1 || !Legal->isCompressedLoadOrStore(I)) &&
+         "Compressed memory ops must be consecutive");
+
   if (!Stride)
     return std::nullopt;
 
@@ -3306,6 +3324,7 @@ static bool willGenerateVectors(VPlan &Plan, ElementCount VF,
       case VPRecipeBase::VPExpandSCEVSC:
       case VPRecipeBase::VPPredInstPHISC:
       case VPRecipeBase::VPBranchOnMaskSC:
+      case VPRecipeBase::VPConditionalInductionPHISC:
         continue;
       case VPRecipeBase::VPReductionSC:
       case VPRecipeBase::VPActiveLaneMaskPHISC:
@@ -3735,6 +3754,10 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
   if (Plan.hasEarlyExit())
     return 1;
 
+  // Conditional inductions don't support interleaving.
+  if (Legal->hasConditionalInductions())
+    return 1;
+
   const bool HasReductions =
       any_of(Plan.getVectorLoopRegion()->getEntryBasicBlock()->phis(),
              IsaPred<VPReductionPHIRecipe>);
@@ -4089,8 +4112,10 @@ void LoopVectorizationCostModel::collectInstsToScalarize(ElementCount VF) {
         // of the instruction.
         // 2. Scalable VF, as that would lead to invalid scalarization costs.
         // 3. Emulated masked memrefs, if a hacked cost is needed.
+        // 4. Compressed loads/stores (which do not support scalarization)
         if (!isScalarAfterVectorization(&I, VF) && !VF.isScalable() &&
             !useEmulatedMaskMemRefHack(&I, VF) &&
+            !Legal->isCompressedLoadOrStore(&I) &&
             computePredInstDiscount(&I, ScalarCosts, VF) >= 0) {
           for (const auto &[I, IC] : ScalarCosts)
             ScalarCostsVF.insert({I, IC});
@@ -4349,9 +4374,14 @@ InstructionCost LoopVectorizationCostModel::getConsecutiveMemOpCost(
   const Align Alignment = getLoadStoreAlignment(I);
   InstructionCost Cost = 0;
   if (isMaskRequired(I)) {
-    unsigned IID = I->getOpcode() == Instruction::Load
-                       ? Intrinsic::masked_load
-                       : Intrinsic::masked_store;
+    Intrinsic::ID LoadIID = Intrinsic::masked_load;
+    Intrinsic::ID StoreIID = Intrinsic::masked_store;
+    if (Legal->isCompressedLoadOrStore(I)) {
+      LoadIID = Intrinsic::masked_expandload;
+      StoreIID = Intrinsic::masked_compressstore;
+    }
+
+    unsigned IID = I->getOpcode() == Instruction::Load ? LoadIID : StoreIID;
     Cost += TTI.getMemIntrinsicInstrCost(
         MemIntrinsicCostAttributes(IID, VectorTy, Alignment, AS),
         Config.CostKind);
@@ -5131,6 +5161,9 @@ LoopVectorizationCostModel::getInstructionCost(Instruction *I,
         return TTI::CastContextHint::Interleave;
       case LoopVectorizationCostModel::CM_Scalarize:
       case LoopVectorizationCostModel::CM_Widen:
+        // TODO: Add 'Compressed' hint (not needed for any targets yet).
+        if (Legal->isCompressedLoadOrStore(I))
+          return TTI::CastContextHint::None;
         return isPredicatedInst(I) ? TTI::CastContextHint::Masked
                                    : TTI::CastContextHint::Normal;
       case LoopVectorizationCostModel::CM_Widen_Reverse:
@@ -6080,6 +6113,7 @@ VPRecipeBase *VPRecipeBuilder::tryToWidenMemory(VPInstruction *VPI,
   // reverse consecutive.
   LoopVectorizationCostModel::InstWidening Decision =
       CM.getWideningDecision(I, Range.Start);
+
   bool Reverse = Decision == LoopVectorizationCostModel::CM_Widen_Reverse;
   bool Consecutive =
       Reverse || Decision == LoopVectorizationCostModel::CM_Widen;
@@ -6205,6 +6239,38 @@ VPHistogramRecipe *VPRecipeBuilder::widenIfHistogram(VPInstruction *VPI) {
                                VPI->getDebugLoc());
 }
 
+VPWidenMemIntrinsicRecipe *VPRecipeBuilder::widenIfCompressedLoadOrStore(
+    VPInstruction *VPI, VPConditionalInductionPHIRecipe *PhiR) {
+  Instruction *I = VPI->getUnderlyingInstr();
+
+  std::optional<CompressedPtrInfo> Info = Legal->isCompressedLoadOrStore(I);
+  if (!Info || Info->ConditionalInductionPHI != PhiR->getPHINode())
+    return nullptr;
+
+  VPBuilder::InsertPointGuard Guard(Builder);
+  Builder.setInsertPoint(VPI);
+
+  VPValue *Mask = VPI->getMask();
+  Type *AccessTy = getLoadStoreType(I);
+  Align Alignment = getLoadStoreAlignment(I);
+
+  VPValue *Ptr = VPI->getOpcode() == Instruction::Load ? VPI->getOperand(0)
+                                                       : VPI->getOperand(1);
+  Ptr = Builder.createConsecutiveVectorPointer(Ptr, AccessTy,
+                                               /*Reverse=*/false,
+                                               VPI->getDebugLoc());
+
+  if (VPI->getOpcode() == Instruction::Load)
+    return new VPWidenMemIntrinsicRecipe(
+        Intrinsic::masked_expandload, {Ptr, Mask, Plan.getPoison(AccessTy)},
+        AccessTy, Alignment, *VPI, I->getDebugLoc());
+
+  VPValue *StoredValue = VPI->getOperand(0);
+  return new VPWidenMemIntrinsicRecipe(Intrinsic::masked_compressstore,
+                                       {StoredValue, Ptr, Mask}, AccessTy,
+                                       Alignment, *VPI, I->getDebugLoc());
+}
+
 bool VPRecipeBuilder::replaceWithFinalIfReductionStore(
     VPInstruction *VPI, VPBuilder &FinalRedStoresBuilder) {
   StoreInst *SI;
@@ -6456,8 +6522,8 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan1() {
   if (!RUN_VPLAN_PASS(
           VPlanTransforms::createHeaderPhiRecipes, *VPlan0, PSE, *OrigLoop,
           VPDT, Legal->getInductionVars(), Legal->getReductionVars(),
-          Legal->getFixedOrderRecurrences(), Config.getInLoopReductions(),
-          Config.getHints().allowReordering())) {
+          Legal->getConditionalInductions(), Legal->getFixedOrderRecurrences(),
+          Config.getInLoopReductions(), Config.getHints().allowReordering())) {
     return nullptr;
   }
 
@@ -6650,6 +6716,11 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan(VPlanPtr Plan,
   // ---------------------------------------------------------------------------
   VPRecipeBuilder RecipeBuilder(*Plan, Legal, *CM, Builder);
 
+  VPBasicBlock *HeaderVPBB = LoopRegion->getEntryBasicBlock();
+  if (!RUN_VPLAN_PASS(VPlanTransforms::handleCompressingPatterns, *Plan,
+                      HeaderVPBB, RecipeBuilder))
+    return nullptr;
+
   RUN_VPLAN_PASS(VPlanTransforms::createInLoopReductionRecipes, *Plan,
                  Range.Start);
 
@@ -6668,7 +6739,6 @@ VPlanPtr LoopVectorizationPlanner::tryToBuildVPlan(VPlanPtr Plan,
 
   // Convert remaining VPInstructions to widen or replicate recipes.
   // TODO: This legacy code should eventually be migrated to VPlan.
-  VPBasicBlock *HeaderVPBB = LoopRegion->getEntryBasicBlock();
   for (VPBasicBlock *VPBB : VPBlockUtils::blocksOnly<VPBasicBlock>(
            vp_depth_first_shallow(HeaderVPBB))) {
     // All types but VPInstructions are already widened and don't need extra
@@ -7554,7 +7624,9 @@ static SmallVector<Instruction *> preparePlanForEpilogueVectorLoop(
       }
     } else {
       // Retrieve the induction resume value via ResumeForEpilogue.
-      PHINode *IndPhi = cast<VPWidenInductionRecipe>(&R)->getPHINode();
+      assert(isa<VPWidenInductionRecipe>(&R) ||
+             isa<VPConditionalInductionPHIRecipe>(&R));
+      PHINode *IndPhi = cast<VPHeaderPHIRecipe>(&R)->getPHINode();
       ResumeV = IRPhiToResumeForEpi.at(IndPhi)->getUnderlyingValue();
     }
     assert(ResumeV && "Must have a resume value");
@@ -7919,6 +7991,15 @@ bool LoopVectorizePass::processLoop(Loop *L) {
     IC = LVP.selectInterleaveCount(*BestPlanPtr, VF.Width, VF.Cost);
 
     unsigned SelectedIC = std::max(IC, UserIC);
+
+    if (LVL.hasConditionalInductions() && SelectedIC > 1) {
+      reportVectorizationFailure(
+          "Interleaving of loop with conditional inductions",
+          "Interleaving of loops with conditional inductions is not supported",
+          "CantInterleaveWithConditionalInductions", ORE, L);
+      return false;
+    }
+
     //  Optimistically generate runtime checks if they are needed. Drop them if
     //  they turn out to not be profitable.
     if (VF.Width.isVector() || SelectedIC > 1) {
diff --git a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
index 3af91e68ef427..caceae2f39647 100644
--- a/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
+++ b/llvm/lib/Transforms/Vectorize/VPRecipeBuilder.h
@@ -73,6 +73,13 @@ class VPRecipeBuilder {
   /// scalar loop.
   VPHistogramRecipe *widenIfHistogram(VPInstruction *VPI);
 
+  /// If \p VPI represents a compressed load or store (as determined by
+  /// LoopVectorizationLegality) whose pointer is derived from \p PhiR, lower it
+  /// to a llvm.masked.expandload or llvm.masked.compressstore intrinsic.
+  VPWidenMemIntrinsicRecipe *
+  widenIfCompressedLoadOrStore(VPInstruction *VPI,
+                               VPConditionalInductionPHIRecipe *PhiR);
+
   /// If \p VPI is a store of a reduction into an invariant address, delete it.
   /// If it is the final store of a reduction result, a uniform store recipe
   /// will be created for it in the middle block. Returns `true` if replacement
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.cpp b/llvm/lib/Transforms/Vectorize/VPlan.cpp
index 80d98af8354d2..669d2865bf969 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlan.cpp
@@ -350,9 +350,10 @@ void VPTransformState::fixupHeaderPhis() {
 
     for (VPRecipeBase &R : Header->phis()) {
       auto *PhiR = cast<VPSingleDefRecipe>(&R);
-      bool NeedsScalar =
-          isa<VPPhi>(PhiR) || (isa<VPReductionPHIRecipe>(PhiR) &&
-                               cast<VPReductionPHIRecipe>(PhiR)->isInLoop());
+      bool NeedsScalar = isa<VPPhi>(PhiR) ||
+                         isa<VPConditionalInductionPHIRecipe>(PhiR) ||
+                         (isa<VPReductionPHIRecipe>(PhiR) &&
+                          cast<VPReductionPHIRecipe>(PhiR)->isInLoop());
 
       Value *Phi = get(PhiR, NeedsScalar);
       Value *Val = get(PhiR->getOperand(1), NeedsScalar);
diff --git a/llvm/lib/Transforms/Vectorize/VPlan.h b/llvm/lib/Transforms/Vectorize/VPlan.h
index e4778489a4060..02959b27c16c3 100644
--- a/llvm/lib/Transforms/Vectorize/VPlan.h
+++ b/llvm/lib/Transforms/Vectorize/VPlan.h
@@ -464,12 +464,13 @@ class LLVM_ABI_FOR_TEST VPRecipeBase
     VPWidenIntOrFpInductionSC,
     VPWidenPointerInductionSC,
     VPReductionPHISC,
+    VPConditionalInductionPHISC,
     // END: SubclassID for recipes that inherit VPHeaderPHIRecipe
     // END: Phi-like recipes
     VPFirstPHISC = VPWidenPHISC,
     VPFirstHeaderPHISC = VPCurrentIterationPHISC,
-    VPLastHeaderPHISC = VPReductionPHISC,
-    VPLastPHISC = VPReductionPHISC,
+    VPLastHeaderPHISC = VPConditionalInductionPHISC,
+    VPLastPHISC = VPConditionalInductionPHISC,
   };
 
   VPRecipeBase(VPRecipeTy SC, ArrayRef<VPValue *> Operands,
@@ -662,6 +663,7 @@ class LLVM_ABI_FOR_TEST VPSingleDefRecipe : public VPRecipeBase,
     case VPRecipeBase::VPReductionPHISC:
     case VPRecipeBase::VPWidenLoadEVLSC:
     case VPRecipeBase::VPWidenLoadSC:
+    case VPRecipeBase::VPConditionalInductionPHISC:
       return true;
     case VPRecipeBase::VPBranchOnMaskSC:
     case VPRecipeBase::VPInterleaveEVLSC:
@@ -2079,7 +2081,9 @@ class VPWidenMemIntrinsicRecipe final : public VPWidenIntrinsicRecipe {
                                DL),
         Alignment(Alignment) {
     assert((VectorIntrinsicID == Intrinsic::experimental_vp_strided_load ||
-            VectorIntrinsicID == Intrinsic::experimental_vp_strided_store) &&
+            VectorIntrinsicID == Intrinsic::experimental_vp_strided_store ||
+            VectorIntrinsicID == Intrinsic::masked_compressstore ||
+            VectorIntrinsicID == Intrinsic::masked_expandload) &&
            "Unexpected intrinsic");
   }
 
@@ -2096,6 +2100,9 @@ class VPWidenMemIntrinsicRecipe final : public VPWidenIntrinsicRecipe {
   /// Produce a widened version of the vector memory intrinsic.
   void execute(VPTransformState &State) override;
 
+  /// Returns the mask of a predicated VPWidenMemIntrinsicRecipe.
+  VPValue *getMask() const;
+
   /// Helper function for computing the cost of vector memory intrinsic.
   static InstructionCost computeMemIntrinsicCost(Intrinsic::ID IID, Type *Ty,
                                                  bool IsMasked, Align Alignment,
@@ -2515,6 +2522,11 @@ class LLVM_ABI_FOR_TEST VPHeaderPHIRecipe : public VPSingleDefRecipe,
     VPUser::addOperand(V);
   }
 
+  /// Returns the underlying PHINode if one exists, or null otherwise.
+  PHINode *getPHINode() const {
+    return cast_if_present<PHINode>(getUnderlyingValue());
+  }
+
 protected:
 #if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
   /// Print the recipe.
@@ -2588,11 +2600,6 @@ class VPWidenInductionRecipe : public VPHeaderPHIRecipe {
   /// incoming value, its start value.
   unsigned getNumIncoming() const override { return 1; }
 
-  /// Returns the underlying PHINode if one exists, or null otherwise.
-  PHINode *getPHINode() const {
-    return cast_if_present<PHINode>(getUnderlyingValue());
-  }
-
   /// Returns the induction descriptor for the recipe.
   const InductionDescriptor &getInductionDescriptor() const { return IndDesc; }
 
@@ -2963,6 +2970,52 @@ class VPReductionPHIRecipe : public VPHeaderPHIRecipe, public VPIRFlags {
 #endif
 };
 
+/// A recipe for handling conditional induction PHIs. The start value is the
+/// first operand of the recipe, the incoming value from the backedge is the
+/// second operand, and the third operand is the step.
+class VPConditionalInductionPHIRecipe : public VPHeaderPHIRecipe {
+public:
+  VPConditionalInductionPHIRecipe(PHINode &Phi, VPValue &Start,
+                                  VPValue &BackedgeValue, VPValue &Step)
+      : VPHeaderPHIRecipe(VPRecipeBase::VPConditionalInductionPHISC, &Phi,
+                          &Start) {
+    addOperand(&BackedgeValue);
+    addOperand(&Step);
+  }
+
+  VPValue *getStep() const { return getOperand(2); }
+
+  unsigned getNumIncoming() const override { return 2; }
+
+  ~VPConditionalInductionPHIRecipe() override = default;
+
+  VPConditionalInductionPHIRecipe *clone() override {
+    return new VPConditionalInductionPHIRecipe(*getPHINode(), *getStartValue(),
+                                               *getBackedgeValue(), *getStep());
+  }
+
+  VP_CLASSOF_IMPL(VPRecipeBase::VPConditionalInductionPHISC)
+
+  static inline bool classof(const VPHeaderPHIRecipe *R) {
+    return R->getVPRecipeID() == VPRecipeBase::VPConditionalInductionPHISC;
+  }
+
+  void execute(VPTransformState &State) override;
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+  /// Print the recipe.
+  void printRecipe(raw_ostream &O, const Twine &Indent,
+                   VPSlotTracker &SlotTracker) const override;
+#endif
+
+  /// Returns true if the recipe only uses the first lane of operand \p Op.
+  bool usesFirstLaneOnly(const VPValue *Op) const override {
+    assert(is_contained(operands(), Op) &&
+           "Op must be an operand of the recipe");
+    return true;
+  }
+};
+
 /// A recipe for vectorizing a phi-node as a sequence of mask-based select
 /// instructions.
 class LLVM_ABI_FOR_TEST VPBlendRecipe : public VPRecipeWithIRFlags {
@@ -4368,7 +4421,8 @@ struct CastInfoMixinImpl
 template <>
 struct CastInfo<VPPhiAccessors, VPRecipeBase *>
     : vpdetail::CastInfoMixinImpl<VPPhiAccessors, VPPhi, VPIRPhi,
-                                  VPWidenPHIRecipe, VPHeaderPHIRecipe> {};
+                                  VPWidenPHIRecipe, VPHeaderPHIRecipe,
+                                  VPConditionalInductionPHIRecipe> {};
 
 template <>
 struct CastInfo<VPPhiAccessors, const VPRecipeBase *>
diff --git a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
index 96653c9a1c3ad..8a3f1cb54df54 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanConstruction.cpp
@@ -940,6 +940,8 @@ bool VPlanTransforms::createHeaderPhiRecipes(
     const VPDominatorTree &VPDT,
     const MapVector<PHINode *, InductionDescriptor> &Inductions,
     const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
+    const MapVector<PHINode *, ConditionalInductionDescriptor>
+        &ConditionalInductions,
     const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
     const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering) {
   // Retrieve the header manually from the intial plain-CFG VPlan.
@@ -972,6 +974,16 @@ bool VPlanTransforms::createHeaderPhiRecipes(
                                         Plan, PSE, OrigLoop,
                                         PhiR->getDebugLoc());
 
+    auto ConditionalInductionIt = ConditionalInductions.find(Phi);
+    if (ConditionalInductionIt != ConditionalInductions.end()) {
+      const ConditionalInductionDescriptor &CondID =
+          ConditionalInductionIt->second;
+      VPValue *Step =
+          vputils::getOrCreateVPValueForSCEVExpr(Plan, CondID.getStepSCEV());
+      return new VPConditionalInductionPHIRecipe(*Phi, *Start, *BackedgeValue,
+                                                 *Step);
+    }
+
     assert(Reductions.contains(Phi) && "only reductions are expected now");
     const RecurrenceDescriptor &RdxDesc = Reductions.lookup(Phi);
     assert(RdxDesc.getRecurrenceStartValue() ==
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index 4faf7c3db1e92..7300ee30ff162 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -151,6 +151,7 @@ bool VPRecipeBase::mayReadFromMemory() const {
   case VPWidenStoreEVLSC:
   case VPWidenStoreSC:
   case VPExpandSCEVSC:
+  case VPConditionalInductionPHISC:
     return false;
   case VPBlendSC:
   case VPReductionEVLSC:
@@ -1500,6 +1501,12 @@ InstructionCost VPInstruction::computeCost(ElementCount VF,
                                   VectorTy, Ctx.CostKind, /*Mask=*/{},
                                   /*Index=*/0);
   }
+  case VPInstruction::NumActiveLanes: {
+    Type *ElementTy = getOperand(0)->getScalarType();
+    auto *VectorTy = cast<VectorType>(toVectorTy(ElementTy, VF));
+    return Ctx.TTI.getArithmeticReductionCost(Instruction::Add, VectorTy,
+                                              std::nullopt, Ctx.CostKind);
+  }
   case VPInstruction::ExtractLastLane: {
     // Add on the cost of extracting the element.
     auto *VecTy = toVectorTy(getOperand(0)->getScalarType(), VF);
@@ -2435,7 +2442,7 @@ void VPWidenIntrinsicRecipe::printRecipe(raw_ostream &O, const Twine &Indent,
 
 void VPWidenMemIntrinsicRecipe::execute(VPTransformState &State) {
   CallInst *MemI = createVectorCall(State);
-  auto PtrPos = VPIntrinsic::getMemoryPointerParamPos(getVectorIntrinsicID());
+  auto PtrPos = getVectorMemoryIntrinsicPointerArgIdx(getVectorIntrinsicID());
   assert(PtrPos && "Expected a memory intrinsic with a valid pointer position");
   MemI->addParamAttr(
       *PtrPos, Attribute::getWithAlignment(MemI->getContext(), Alignment));
@@ -2443,6 +2450,12 @@ void VPWidenMemIntrinsicRecipe::execute(VPTransformState &State) {
     State.set(this, MemI);
 }
 
+VPValue *VPWidenMemIntrinsicRecipe::getMask() const {
+  auto MaskPos = getVectorIntrinsicMaskArgIdx(getVectorIntrinsicID());
+  assert(MaskPos && "Expected a memory intrinsic with a valid mask position");
+  return getOperand(*MaskPos);
+}
+
 InstructionCost VPWidenMemIntrinsicRecipe::computeMemIntrinsicCost(
     Intrinsic::ID IID, Type *Ty, bool IsMasked, Align Alignment,
     VPCostContext &Ctx) {
@@ -2455,17 +2468,14 @@ InstructionCost
 VPWidenMemIntrinsicRecipe::computeCost(ElementCount VF,
                                        VPCostContext &Ctx) const {
   Type *DataTy;
-  if (auto DataPos = VPIntrinsic::getMemoryDataParamPos(getVectorIntrinsicID()))
+  if (auto DataPos = getVectorStoreIntrinsicDataArgIdx(getVectorIntrinsicID()))
     DataTy = getOperand(*DataPos)->getScalarType();
   else
     DataTy = getScalarType();
   assert(!DataTy->isVoidTy() && "Expected a non-void data type");
   Type *Ty = toVectorTy(DataTy, VF);
-  auto MaskPos = VPIntrinsic::getMaskParamPos(getVectorIntrinsicID());
-  assert(MaskPos && "Expected a memory intrinsic with a valid mask position");
   return computeMemIntrinsicCost(getVectorIntrinsicID(), Ty,
-                                 !match(getOperand(*MaskPos), m_True()),
-                                 Alignment, Ctx);
+                                 !match(getMask(), m_True()), Alignment, Ctx);
 }
 
 void VPHistogramRecipe::execute(VPTransformState &State) {
@@ -5134,6 +5144,21 @@ bool VPBlendRecipe::usesFirstLaneOnly(const VPValue *Op) const {
   return vputils::onlyFirstLaneUsed(this);
 }
 
+void VPConditionalInductionPHIRecipe::execute(VPTransformState &State) {
+  executePhiRecipe(this, *this, State, /*IsScalar=*/true, "conditional.iv");
+}
+
+#if !defined(NDEBUG) || defined(LLVM_ENABLE_DUMP)
+void VPConditionalInductionPHIRecipe::printRecipe(
+    raw_ostream &O, const Twine &Indent, VPSlotTracker &SlotTracker) const {
+  O << Indent << "CONDITIONAL-INDUCTION-PHI ";
+
+  printAsOperand(O, SlotTracker);
+  O << " = phi ";
+  printOperands(O, SlotTracker);
+}
+#endif
+
 void VPWidenPHIRecipe::execute(VPTransformState &State) {
   executePhiRecipe(this, *this, State, /*IsScalar=*/false, Name);
 }
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
index 09fdaa63bab99..0cf3d5973dba2 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.cpp
@@ -6121,3 +6121,94 @@ void VPlanTransforms::convertToStridedAccesses(VPlan &Plan,
     }
   }
 }
+
+bool VPlanTransforms::handleCompressingPatterns(
+    VPlan &Plan, VPBasicBlock *HeaderVPBB, VPRecipeBuilder &RecipeBuilder) {
+  SmallVector<VPInstruction *> MemOps;
+  for (VPBasicBlock *VPBB :
+       VPBlockUtils::blocksOnly<VPBasicBlock>(vp_depth_first_shallow(
+           Plan.getVectorLoopRegion()->getEntryBasicBlock()))) {
+    for (VPRecipeBase &R : *VPBB) {
+      auto *VPI = dyn_cast<VPInstruction>(&R);
+      if (VPI && VPI->getUnderlyingValue() &&
+          is_contained({Instruction::Load, Instruction::Store},
+                       VPI->getOpcode()))
+        MemOps.push_back(VPI);
+    }
+  }
+
+  VPBuilder Builder;
+  for (VPRecipeBase &R : HeaderVPBB->phis()) {
+    auto *ConditionalInductionPhi =
+        dyn_cast<VPConditionalInductionPHIRecipe>(&R);
+    if (!ConditionalInductionPhi)
+      continue;
+
+    // Obtain the mask for the conditional induction update from the
+    // VPBlendRecipe.
+    auto *BlendR =
+        cast<VPBlendRecipe>(ConditionalInductionPhi->getBackedgeValue());
+    VPValue *Mask = nullptr;
+    for (unsigned I = 0, E = BlendR->getNumIncomingValues(); I != E; ++I)
+      if (auto *IncomingVal = BlendR->getIncomingValue(I);
+          IncomingVal != ConditionalInductionPhi) {
+        Mask = BlendR->getMask(I);
+        break;
+      }
+    assert(Mask);
+
+    // Replace all "compressed" loads and stores with expandload and
+    // compressstore respectively.
+    for (VPInstruction *&VPI : MemOps) {
+      auto *CompressedMemOp = RecipeBuilder.widenIfCompressedLoadOrStore(
+          VPI, ConditionalInductionPhi);
+      if (!CompressedMemOp)
+        continue;
+
+      Builder.setInsertPoint(VPI);
+      Builder.insert(CompressedMemOp);
+
+      // Bail out if the mask for the memory op does not match the condition
+      // used to update the conditional induction.
+      VPValue *MemOpMask = CompressedMemOp->getMask();
+      if (MemOpMask != Mask)
+        return false;
+
+      if (VPI->getOpcode() == Instruction::Load)
+        VPI->replaceAllUsesWith(CompressedMemOp->getVPSingleValue());
+      VPI->eraseFromParent();
+      VPI = nullptr; // Mark handled instructions with a nullptr.
+    }
+
+    // Remove all memory operations we've handled.
+    MemOps.erase(
+        remove_if(MemOps, [](VPInstruction *VPI) { return VPI == nullptr; }),
+        MemOps.end());
+
+    // Update the conditional induction to increment by the number of active
+    // lanes in the mask.
+    auto *BackedgeVal = ConditionalInductionPhi->getBackedgeValue();
+    auto *InsertBlock = BackedgeVal->getDefiningRecipe()->getParent();
+    Builder.setInsertPoint(InsertBlock, InsertBlock->getFirstNonPhi());
+
+    Type *UpdateType = ConditionalInductionPhi->getScalarType();
+    if (UpdateType->isPointerTy())
+      UpdateType = Plan.getDataLayout().getIndexType(UpdateType);
+
+    auto *HandledLanes = Builder.createNaryOp(
+        VPInstruction::NumActiveLanes, {Mask}, nullptr, {}, {},
+        DebugLoc::getUnknown(), "handled.lanes", UpdateType);
+    VPValue *Offset = Builder.createOverflowingOp(
+        Instruction::Mul, {ConditionalInductionPhi->getStep(), HandledLanes});
+    VPValue *Update;
+    if (ConditionalInductionPhi->getScalarType()->isPointerTy())
+      Update = Builder.createPtrAdd(ConditionalInductionPhi, Offset);
+    else
+      Update = Builder.createAdd(ConditionalInductionPhi, Offset, {},
+                                 "conditional.step");
+
+    BackedgeVal->replaceAllUsesWith(Update);
+  }
+
+  return true;
+}
diff --git a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
index fd62693d6068b..c8b6a7c8e42fc 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanTransforms.h
@@ -175,6 +175,8 @@ struct VPlanTransforms {
       const VPDominatorTree &VPDT,
       const MapVector<PHINode *, InductionDescriptor> &Inductions,
       const MapVector<PHINode *, RecurrenceDescriptor> &Reductions,
+      const MapVector<PHINode *, ConditionalInductionDescriptor>
+          &ConditionalInductions,
       const SmallPtrSetImpl<const PHINode *> &FixedOrderRecurrences,
       const SmallPtrSetImpl<PHINode *> &InLoopReductions, bool AllowReordering);
 
@@ -272,6 +274,16 @@ struct VPlanTransforms {
   /// was unsuccessful.
   static bool handleFindLastReductions(VPlan &Plan);
 
+  /// Handles compressing memory loads/stores. Loads/stores where the pointer
+  /// is derived from a conditional induction PHI are replaced with expandloads
+  /// or compressstores respectively. The backedge value of the conditional
+  /// induction PHI is updated to increment by the number of active lanes of
+  /// the block mask.
+  /// Returns false if any memory operation could not be updated (e.g., due to
+  /// having a mask that does not match the PHI).
+  static bool handleCompressingPatterns(VPlan &Plan, VPBasicBlock *HeaderVPBB,
+                                        VPRecipeBuilder &RecipeBuilder);
+
   /// Clear NSW/NUW flags from reduction instructions if necessary.
   static void clearReductionWrapFlags(VPlan &Plan);
 
diff --git a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
index 9a07697f66762..22ec3c69f77ae 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanUtils.cpp
@@ -475,8 +475,8 @@ bool vputils::isSingleScalar(const VPValue *VPV) {
             all_of(VPI->operands(), isSingleScalar));
   if (auto *RR = dyn_cast<VPReductionRecipe>(VPV))
     return !RR->isPartialReduction();
-  if (isa<VPVectorPointerRecipe, VPVectorEndPointerRecipe, VPDerivedIVRecipe>(
-          VPV))
+  if (isa<VPVectorPointerRecipe, VPVectorEndPointerRecipe, VPDerivedIVRecipe,
+          VPConditionalInductionPHIRecipe>(VPV))
     return true;
   if (auto *Expr = dyn_cast<VPExpressionRecipe>(VPV))
     return Expr->isVectorToScalar();
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
new file mode 100644
index 0000000000000..52ecb9d954de0
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/compress-idioms.ll
@@ -0,0 +1,153 @@
+; NOTE: Assertions have been autogenerated by utils/update_analyze_test_checks.py UTC_ARGS: --filter-out-after "^scalar.ph:" --version 6
+; RUN: opt -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -vplan-print-after=printOptimizedVPlan -disable-output %s -S 2>&1 | FileCheck %s
+
+define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: VPlan for loop in 'compress_store'
+; CHECK:  VPlan 'Initial VPlan for VF={4},UF>=1' {
+; CHECK-NEXT:  Live-in vp<[[VP0:%[0-9]+]]> = VF
+; CHECK-NEXT:  Live-in vp<[[VP1:%[0-9]+]]> = VF * UF
+; CHECK-NEXT:  Live-in vp<[[VP2:%[0-9]+]]> = vector-trip-count
+; CHECK-NEXT:  Live-in ir<%n> = original trip-count
+; CHECK-EMPTY:
+; CHECK-NEXT:  ir-bb<entry>:
+; CHECK-NEXT:  Successor(s): scalar.ph, vector.ph
+; CHECK-EMPTY:
+; CHECK-NEXT:  vector.ph:
+; CHECK-NEXT:  Successor(s): vector loop
+; CHECK-EMPTY:
+; CHECK-NEXT:  <x1> vector loop: {
+; CHECK-NEXT:  vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
+; CHECK-EMPTY:
+; CHECK-NEXT:    vector.body:
+; CHECK-NEXT:      CONDITIONAL-INDUCTION-PHI ir<%idx> = phi ir<0>, vp<%conditional.step>, ir<1>
+; CHECK-NEXT:      vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
+; CHECK-NEXT:      CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, vp<[[VP4]]>
+; CHECK-NEXT:      vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
+; CHECK-NEXT:      WIDEN ir<%load.src> = load vp<[[VP5]]>
+; CHECK-NEXT:      WIDEN ir<%cmp> = icmp slt ir<%load.src>, ir<%c>
+; CHECK-NEXT:      CLONE ir<%dst.ptr> = getelementptr inbounds ir<%dst>, ir<%idx>
+; CHECK-NEXT:      vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT:      WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.compressstore(ir<%load.src>, vp<[[VP6]]>, ir<%cmp>) (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
+; CHECK-NEXT:      EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
+; CHECK-NEXT:      EMIT vp<%conditional.step> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT:      EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
+; CHECK-NEXT:      EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
+; CHECK-NEXT:    No successors
+; CHECK-NEXT:  }
+; CHECK-NEXT:  Successor(s): middle.block
+; CHECK-EMPTY:
+; CHECK-NEXT:  middle.block:
+; CHECK-NEXT:    EMIT vp<[[VP9:%[0-9]+]]> = extract-last-part vp<%conditional.step>
+; CHECK-NEXT:    EMIT vp<[[VP10:%[0-9]+]]> = extract-last-lane vp<[[VP9]]>
+; CHECK-NEXT:    EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
+; CHECK-NEXT:    EMIT branch-on-cond vp<%cmp.n>
+; CHECK-NEXT:  Successor(s): ir-bb<exit>, scalar.ph
+; CHECK-EMPTY:
+; CHECK-NEXT:  ir-bb<exit>:
+; CHECK-NEXT:  No successors
+; CHECK-EMPTY:
+; CHECK-NEXT:  scalar.ph:
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+  %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+  %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+  %load.src = load i32, ptr %src.ptr, align 4
+  %cmp = icmp slt i32 %load.src, %c
+  br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+  %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
+  store i32 %load.src, ptr %dst.ptr, align 4
+  %idx.next = add nsw i64 %idx, 1
+  br label %for.inc
+
+for.inc:
+  %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+  %iv.next = add nuw nsw i64 %iv, 1
+  %exitcond.not = icmp eq i64 %iv.next, %n
+  br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+  ret void
+}
+
+define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-LABEL: VPlan for loop in 'expand_load'
+; CHECK:  VPlan 'Initial VPlan for VF={4},UF>=1' {
+; CHECK-NEXT:  Live-in vp<[[VP0:%[0-9]+]]> = VF
+; CHECK-NEXT:  Live-in vp<[[VP1:%[0-9]+]]> = VF * UF
+; CHECK-NEXT:  Live-in vp<[[VP2:%[0-9]+]]> = vector-trip-count
+; CHECK-NEXT:  Live-in ir<%n> = original trip-count
+; CHECK-EMPTY:
+; CHECK-NEXT:  ir-bb<entry>:
+; CHECK-NEXT:  Successor(s): scalar.ph, vector.ph
+; CHECK-EMPTY:
+; CHECK-NEXT:  vector.ph:
+; CHECK-NEXT:  Successor(s): vector loop
+; CHECK-EMPTY:
+; CHECK-NEXT:  <x1> vector loop: {
+; CHECK-NEXT:  vp<[[VP3:%[0-9]+]]> = CANONICAL-IV
+; CHECK-EMPTY:
+; CHECK-NEXT:    vector.body:
+; CHECK-NEXT:      CONDITIONAL-INDUCTION-PHI ir<%idx> = phi ir<0>, vp<%conditional.step>, ir<1>
+; CHECK-NEXT:      vp<[[VP4:%[0-9]+]]> = SCALAR-STEPS vp<[[VP3]]>, ir<1>, vp<[[VP0]]>
+; CHECK-NEXT:      CLONE ir<%dst.ptr> = getelementptr ir<%dst>, vp<[[VP4]]>
+; CHECK-NEXT:      vp<[[VP5:%[0-9]+]]> = vector-pointer inbounds i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT:      WIDEN ir<%load.dst> = load vp<[[VP5]]>
+; CHECK-NEXT:      WIDEN ir<%cmp> = icmp slt ir<%load.dst>, ir<%c>
+; CHECK-NEXT:      CLONE ir<%src.ptr> = getelementptr inbounds ir<%src>, ir<%idx>
+; CHECK-NEXT:      vp<[[VP6:%[0-9]+]]> = vector-pointer inbounds i32, ir<%src.ptr>, ir<1>
+; CHECK-NEXT:      WIDEN-INTRINSIC vp<[[VP7:%[0-9]+]]> = call llvm.masked.expandload(vp<[[VP6]]>, ir<%cmp>, ir<poison>) (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
+; CHECK-NEXT:      vp<[[VP8:%[0-9]+]]> = vector-pointer i32, ir<%dst.ptr>, ir<1>
+; CHECK-NEXT:      WIDEN store vp<[[VP8]]>, vp<[[VP7]]>, ir<%cmp> (!vplan.execution.frequency 4611686018427387904 (50%, estimated))
+; CHECK-NEXT:      EMIT vp<%handled.lanes> = num-active-lanes ir<%cmp>
+; CHECK-NEXT:      EMIT vp<%conditional.step> = add ir<%idx>, vp<%handled.lanes>
+; CHECK-NEXT:      EMIT vp<%index.next> = add nuw vp<[[VP3]]>, vp<[[VP1]]>
+; CHECK-NEXT:      EMIT branch-on-count vp<%index.next>, vp<[[VP2]]>
+; CHECK-NEXT:    No successors
+; CHECK-NEXT:  }
+; CHECK-NEXT:  Successor(s): middle.block
+; CHECK-EMPTY:
+; CHECK-NEXT:  middle.block:
+; CHECK-NEXT:    EMIT vp<[[VP10:%[0-9]+]]> = extract-last-part vp<%conditional.step>
+; CHECK-NEXT:    EMIT vp<[[VP11:%[0-9]+]]> = extract-last-lane vp<[[VP10]]>
+; CHECK-NEXT:    EMIT vp<%cmp.n> = icmp eq ir<%n>, vp<[[VP2]]>
+; CHECK-NEXT:    EMIT branch-on-cond vp<%cmp.n>
+; CHECK-NEXT:  Successor(s): ir-bb<exit>, scalar.ph
+; CHECK-EMPTY:
+; CHECK-NEXT:  ir-bb<exit>:
+; CHECK-NEXT:  No successors
+; CHECK-EMPTY:
+; CHECK-NEXT:  scalar.ph:
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+  %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+  %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+  %load.dst = load i32, ptr %dst.ptr, align 4
+  %cmp = icmp slt i32 %load.dst, %c
+  br i1 %cmp, label %if.then, label %for.inc
+
+if.then:
+  %src.ptr = getelementptr inbounds i32, ptr %src, i64 %idx
+  %load.src = load i32, ptr %src.ptr, align 4
+  store i32 %load.src, ptr %dst.ptr, align 4
+  %idx.next = add nsw i64 %idx, 1
+  br label %for.inc
+
+for.inc:
+  %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
+  %iv.next = add nuw nsw i64 %iv, 1
+  %exitcond.not = icmp eq i64 %iv.next, %n
+  br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+  ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
index 5fd844e186e44..3868bf73e407f 100644
--- a/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
+++ b/llvm/test/Transforms/LoopVectorize/VPlan/vplan-print-before-after-all.ll
@@ -22,6 +22,7 @@
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::handleCountableEarlyExits
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::createLoopRegions
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::introduceMasksAndLinearize
+; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::handleCompressingPatterns
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::createInLoopReductionRecipes
 ; CHECK-BEFORE: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] VPlanTransforms::makeMemOpWideningDecisions
 ; CHECK: VPlan for loop in 'foo' [[BEFORE_OR_AFTER]] lowerMemoryIdioms
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
index da3d76cdadb1a..58ac666ae9157 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms-negative-tests.ll
@@ -94,7 +94,7 @@ exit:
   ret void
 }
 
-; CHECK: loop not vectorized
+; CHECK: the cost-model indicates that vectorization is not beneficial
 
 ; Negative test: In this case the %idx is incremented when %cond.val != 0,
 ; but the store occurs when %cond.val > 100. The store mask does not match the
@@ -136,7 +136,7 @@ exit:
   ret void
 }
 
-; CHECK: loop not vectorized
+; CHECK: the cost-model indicates that vectorization is not beneficial
 
 ; Negative test: Simple early exit loop with a compressstore. This fails in VPlan handling for early exits.
 define i32 @compress_store_with_early_exit(ptr dereferenceable(1024) %dst, ptr noalias dereferenceable(1024) %src, ptr noalias dereferenceable(1024) %cond, ptr noalias dereferenceable(1024) %exit_cond) {
diff --git a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
index 6f6200e8e2074..592fc55169397 100644
--- a/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
+++ b/llvm/unittests/Transforms/Vectorize/VPlanTestBase.h
@@ -101,6 +101,7 @@ class VPlanTestIRBase : public testing::Test {
       VPlanTransforms::createHeaderPhiRecipes(
           *Plan, PSE, *L, VPDT, Inductions,
           MapVector<PHINode *, RecurrenceDescriptor>(),
+          MapVector<PHINode *, ConditionalInductionDescriptor>(),
           SmallPtrSet<const PHINode *, 1>(), SmallPtrSet<PHINode *, 1>(),
           /*AllowReordering=*/false);
     }

>From a79ad539bf659e1a8b9f6bcc732d5090ebad0397 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Tue, 22 Sep 2026 10:12:27 +0000
Subject: [PATCH 2/6] Update checks

---
 .../LoopVectorize/AArch64/compress-idioms.ll  | 102 ++-
 .../LoopVectorize/X86/compress-idioms.ll      | 185 +++--
 .../LoopVectorize/compress-idioms.ll          | 747 ++++++++----------
 .../compress-store-vec-epilogue.ll            |  58 +-
 4 files changed, 576 insertions(+), 516 deletions(-)

diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
index ad8878ca635eb..f65e14dfd79cd 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/compress-idioms.ll
@@ -6,27 +6,35 @@
 define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-LABEL: define void @compress_store(
 ; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
-; CHECK-NEXT:  [[ENTRY:.*]]:
-; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
-; CHECK:       [[FOR_BODY]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
-; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
-; CHECK-NEXT:    [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
-; CHECK-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
-; CHECK-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
-; CHECK:       [[IF_THEN]]:
-; CHECK-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-NEXT:    br label %[[FOR_INC]]
-; CHECK:       [[FOR_INC]]:
-; CHECK-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
-; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
-; CHECK:       [[EXIT]]:
-; CHECK-NEXT:    ret void
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT:    [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
+; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP2]], align 4
+; CHECK-NEXT:    [[TMP3:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.nxv4i32.p0(<vscale x 4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP4]], <vscale x 4 x i1> [[TMP3]])
+; CHECK-NEXT:    [[TMP5:%.*]] = zext <vscale x 4 x i1> [[TMP3]] to <vscale x 4 x i64>
+; CHECK-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP5]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT:    [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK:       [[SCALAR_PH]]:
 ;
 entry:
   br label %for.body
@@ -58,28 +66,36 @@ exit:
 define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-LABEL: define void @expand_load(
 ; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-NEXT:  [[ENTRY:.*]]:
-; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
-; CHECK:       [[FOR_BODY]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
-; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
-; CHECK-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
-; CHECK:       [[IF_THEN]]:
-; CHECK-NEXT:    [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IDX]]
-; CHECK-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
-; CHECK-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-NEXT:    br label %[[FOR_INC]]
-; CHECK:       [[FOR_INC]]:
-; CHECK-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
-; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
-; CHECK:       [[EXIT]]:
-; CHECK-NEXT:    ret void
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    [[TMP0:%.*]] = call i64 @llvm.vscale.i64()
+; CHECK-NEXT:    [[TMP1:%.*]] = shl nuw i64 [[TMP0]], 2
+; CHECK-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], [[TMP1]]
+; CHECK-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[N_MOD_VF:%.*]] = urem i64 [[N]], [[TMP1]]
+; CHECK-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[N_MOD_VF]]
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <vscale x 4 x i32>, ptr [[TMP2]], align 4
+; CHECK-NEXT:    [[TMP3:%.*]] = icmp slt <vscale x 4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP4:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    [[TMP5:%.*]] = call <vscale x 4 x i32> @llvm.masked.expandload.nxv4i32.p0(ptr align 4 [[TMP4]], <vscale x 4 x i1> [[TMP3]], <vscale x 4 x i32> poison)
+; CHECK-NEXT:    call void @llvm.masked.store.nxv4i32.p0(<vscale x 4 x i32> [[TMP5]], ptr align 4 [[TMP2]], <vscale x 4 x i1> [[TMP3]])
+; CHECK-NEXT:    [[TMP6:%.*]] = zext <vscale x 4 x i1> [[TMP3]] to <vscale x 4 x i64>
+; CHECK-NEXT:    [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP6]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], [[TMP1]]
+; CHECK-NEXT:    [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-NEXT:    br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
+; CHECK:       [[SCALAR_PH]]:
 ;
 entry:
   br label %for.body
diff --git a/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
index 99b5c4736dad6..cb905357816f9 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
@@ -54,27 +54,33 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
 ;
 ; CHECK-ZNVER4-LABEL: define void @compress_store(
 ; CHECK-ZNVER4-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
-; CHECK-ZNVER4-NEXT:  [[VECTOR_PH:.*]]:
+; CHECK-ZNVER4-NEXT:  [[ENTRY:.*:]]
+; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 48
+; CHECK-ZNVER4-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ZNVER4:       [[VECTOR_PH]]:
+; CHECK-ZNVER4-NEXT:    [[TMP0:%.*]] = and i64 [[N]], 15
+; CHECK-ZNVER4-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
 ; CHECK-ZNVER4-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK-ZNVER4:       [[VECTOR_BODY]]:
-; CHECK-ZNVER4-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[MIDDLE_BLOCK:.*]] ]
-; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[MIDDLE_BLOCK]] ]
+; CHECK-ZNVER4-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
 ; CHECK-ZNVER4-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-ZNVER4-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP1]], align 4
-; CHECK-ZNVER4-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-ZNVER4-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[MIDDLE_BLOCK]]
-; CHECK-ZNVER4:       [[IF_THEN]]:
+; CHECK-ZNVER4-NEXT:    [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-ZNVER4-NEXT:    [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
 ; CHECK-ZNVER4-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
-; CHECK-ZNVER4-NEXT:    store i32 [[LOAD_SRC]], ptr [[TMP3]], align 4
-; CHECK-ZNVER4-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-ZNVER4-NEXT:    br label %[[MIDDLE_BLOCK]]
+; CHECK-ZNVER4-NEXT:    call void @llvm.masked.compressstore.v16i32.p0(<16 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <16 x i1> [[TMP2]])
+; CHECK-ZNVER4-NEXT:    [[TMP4:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
+; CHECK-ZNVER4-NEXT:    [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP4]])
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
+; CHECK-ZNVER4-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-ZNVER4-NEXT:    [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT:    br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
 ; CHECK-ZNVER4:       [[MIDDLE_BLOCK]]:
-; CHECK-ZNVER4-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-ZNVER4-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-ZNVER4-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-ZNVER4-NEXT:    br i1 [[EXITCOND_NOT]], label %[[SCALAR_PH:.*]], label %[[VECTOR_BODY]]
+; CHECK-ZNVER4-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT:    br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
 ; CHECK-ZNVER4:       [[SCALAR_PH]]:
-; CHECK-ZNVER4-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -106,61 +112,134 @@ exit:
 define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-V4-LABEL: define void @expand_load(
 ; CHECK-V4-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-V4-NEXT:  [[VECTOR_PH:.*]]:
+; CHECK-V4-NEXT:  [[ENTRY:.*:]]
+; CHECK-V4-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-V4-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-V4:       [[VECTOR_PH]]:
+; CHECK-V4-NEXT:    [[TMP0:%.*]] = and i64 [[N]], 7
+; CHECK-V4-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-V4-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-V4-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT]], <8 x i32> poison, <8 x i32> zeroinitializer
 ; CHECK-V4-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK-V4:       [[VECTOR_BODY]]:
-; CHECK-V4-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[MIDDLE_BLOCK:.*]] ]
-; CHECK-V4-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[MIDDLE_BLOCK]] ]
-; CHECK-V4-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-V4-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-V4-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-V4-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[MIDDLE_BLOCK]]
-; CHECK-V4:       [[IF_THEN]]:
+; CHECK-V4-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-V4-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-V4-NEXT:    [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-V4-NEXT:    [[WIDE_LOAD:%.*]] = load <8 x i32>, ptr [[TMP1]], align 4
+; CHECK-V4-NEXT:    [[TMP2:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
 ; CHECK-V4-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
-; CHECK-V4-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP3]], align 4
-; CHECK-V4-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-V4-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-V4-NEXT:    br label %[[MIDDLE_BLOCK]]
+; CHECK-V4-NEXT:    [[TMP4:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP3]], <8 x i1> [[TMP2]], <8 x i32> poison)
+; CHECK-V4-NEXT:    call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP4]], ptr align 4 [[TMP1]], <8 x i1> [[TMP2]])
+; CHECK-V4-NEXT:    [[TMP5:%.*]] = zext <8 x i1> [[TMP2]] to <8 x i64>
+; CHECK-V4-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP5]])
+; CHECK-V4-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-V4-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; CHECK-V4-NEXT:    [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-V4-NEXT:    br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
 ; CHECK-V4:       [[MIDDLE_BLOCK]]:
-; CHECK-V4-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-V4-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-V4-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-V4-NEXT:    br i1 [[EXITCOND_NOT]], label %[[SCALAR_PH:.*]], label %[[VECTOR_BODY]]
+; CHECK-V4-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-V4-NEXT:    br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
 ; CHECK-V4:       [[SCALAR_PH]]:
-; CHECK-V4-NEXT:    ret void
 ;
 ; CHECK-ICELAKE-LABEL: define void @expand_load(
 ; CHECK-ICELAKE-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-ICELAKE-NEXT:  [[VECTOR_PH:.*]]:
+; CHECK-ICELAKE-NEXT:  [[ENTRY:.*:]]
+; CHECK-ICELAKE-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-ICELAKE-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ICELAKE:       [[VECTOR_PH]]:
+; CHECK-ICELAKE-NEXT:    [[TMP0:%.*]] = and i64 [[N]], 7
+; CHECK-ICELAKE-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ICELAKE-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-ICELAKE-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT]], <8 x i32> poison, <8 x i32> zeroinitializer
 ; CHECK-ICELAKE-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK-ICELAKE:       [[VECTOR_BODY]]:
-; CHECK-ICELAKE-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[MIDDLE_BLOCK:.*]] ]
-; CHECK-ICELAKE-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[MIDDLE_BLOCK]] ]
-; CHECK-ICELAKE-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-ICELAKE-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-ICELAKE-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-ICELAKE-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[MIDDLE_BLOCK]]
-; CHECK-ICELAKE:       [[IF_THEN]]:
+; CHECK-ICELAKE-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ICELAKE-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ICELAKE-NEXT:    [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-ICELAKE-NEXT:    [[WIDE_LOAD:%.*]] = load <8 x i32>, ptr [[TMP1]], align 4
+; CHECK-ICELAKE-NEXT:    [[TMP2:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
 ; CHECK-ICELAKE-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
-; CHECK-ICELAKE-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP3]], align 4
-; CHECK-ICELAKE-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-ICELAKE-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-ICELAKE-NEXT:    br label %[[MIDDLE_BLOCK]]
+; CHECK-ICELAKE-NEXT:    [[TMP4:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP3]], <8 x i1> [[TMP2]], <8 x i32> poison)
+; CHECK-ICELAKE-NEXT:    call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP4]], ptr align 4 [[TMP1]], <8 x i1> [[TMP2]])
+; CHECK-ICELAKE-NEXT:    [[TMP5:%.*]] = zext <8 x i1> [[TMP2]] to <8 x i64>
+; CHECK-ICELAKE-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP5]])
+; CHECK-ICELAKE-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-ICELAKE-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 8
+; CHECK-ICELAKE-NEXT:    [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ICELAKE-NEXT:    br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
 ; CHECK-ICELAKE:       [[MIDDLE_BLOCK]]:
-; CHECK-ICELAKE-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-ICELAKE-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-ICELAKE-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-ICELAKE-NEXT:    br i1 [[EXITCOND_NOT]], label %[[SCALAR_PH:.*]], label %[[VECTOR_BODY]]
+; CHECK-ICELAKE-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ICELAKE-NEXT:    br i1 [[CMP_N]], [[EXIT:label %.*]], label %[[SCALAR_PH]]
 ; CHECK-ICELAKE:       [[SCALAR_PH]]:
-; CHECK-ICELAKE-NEXT:    ret void
 ;
 ; CHECK-ZNVER4-LABEL: define void @expand_load(
 ; CHECK-ZNVER4-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-ZNVER4-NEXT:  [[VEC_EPILOG_SCALAR_PH:.*]]:
+; CHECK-ZNVER4-NEXT:  [[ITER_CHECK:.*]]:
+; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-ZNVER4-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK-ZNVER4:       [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[N]], 16
+; CHECK-ZNVER4-NEXT:    br i1 [[MIN_ITERS_CHECK1]], label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK-ZNVER4:       [[VECTOR_PH]]:
+; CHECK-ZNVER4-NEXT:    [[TMP0:%.*]] = and i64 [[N]], 15
+; CHECK-ZNVER4-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK-ZNVER4:       [[VECTOR_BODY]]:
+; CHECK-ZNVER4-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT:    [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-ZNVER4-NEXT:    [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP1]], align 4
+; CHECK-ZNVER4-NEXT:    [[TMP2:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-ZNVER4-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-ZNVER4-NEXT:    [[TMP4:%.*]] = call <16 x i32> @llvm.masked.expandload.v16i32.p0(ptr align 4 [[TMP3]], <16 x i1> [[TMP2]], <16 x i32> poison)
+; CHECK-ZNVER4-NEXT:    call void @llvm.masked.store.v16i32.p0(<16 x i32> [[TMP4]], ptr align 4 [[TMP1]], <16 x i1> [[TMP2]])
+; CHECK-ZNVER4-NEXT:    [[TMP5:%.*]] = zext <16 x i1> [[TMP2]] to <16 x i64>
+; CHECK-ZNVER4-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP5]])
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-ZNVER4-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-ZNVER4-NEXT:    [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT:    br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK-ZNVER4:       [[MIDDLE_BLOCK]]:
+; CHECK-ZNVER4-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-ZNVER4-NEXT:    br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK-ZNVER4:       [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-ZNVER4-NEXT:    [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP0]], 8
+; CHECK-ZNVER4-NEXT:    br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF5:![0-9]+]]
+; CHECK-ZNVER4:       [[VEC_EPILOG_PH]]:
+; CHECK-ZNVER4-NEXT:    [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT:    [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT:    [[TMP8:%.*]] = and i64 [[N]], 7
+; CHECK-ZNVER4-NEXT:    [[N_VEC2:%.*]] = sub i64 [[N]], [[TMP8]]
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLAT4:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT3]], <8 x i32> poison, <8 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT:    br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
+; CHECK-ZNVER4:       [[VEC_EPILOG_VECTOR_BODY]]:
+; CHECK-ZNVER4-NEXT:    [[INDEX5:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT9:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[CONDITIONAL_STEP8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-ZNVER4-NEXT:    [[TMP9:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX5]]
+; CHECK-ZNVER4-NEXT:    [[WIDE_LOAD7:%.*]] = load <8 x i32>, ptr [[TMP9]], align 4
+; CHECK-ZNVER4-NEXT:    [[TMP10:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
+; CHECK-ZNVER4-NEXT:    [[TMP11:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV6]]
+; CHECK-ZNVER4-NEXT:    [[TMP12:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP11]], <8 x i1> [[TMP10]], <8 x i32> poison)
+; CHECK-ZNVER4-NEXT:    call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP12]], ptr align 4 [[TMP9]], <8 x i1> [[TMP10]])
+; CHECK-ZNVER4-NEXT:    [[TMP13:%.*]] = zext <8 x i1> [[TMP10]] to <8 x i64>
+; CHECK-ZNVER4-NEXT:    [[TMP14:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP13]])
+; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_STEP8]] = add i64 [[CONDITIONAL_IV6]], [[TMP14]]
+; CHECK-ZNVER4-NEXT:    [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 8
+; CHECK-ZNVER4-NEXT:    [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT9]], [[N_VEC2]]
+; CHECK-ZNVER4-NEXT:    br i1 [[TMP15]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK-ZNVER4:       [[VEC_EPILOG_MIDDLE_BLOCK]]:
+; CHECK-ZNVER4-NEXT:    [[CMP_N10:%.*]] = icmp eq i64 [[N]], [[N_VEC2]]
+; CHECK-ZNVER4-NEXT:    br i1 [[CMP_N10]], label %[[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
+; CHECK-ZNVER4:       [[VEC_EPILOG_SCALAR_PH]]:
+; CHECK-ZNVER4-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC2]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-ZNVER4-NEXT:    [[BC_MERGE_RDX11:%.*]] = phi i64 [ [[CONDITIONAL_STEP8]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
 ; CHECK-ZNVER4-NEXT:    br label %[[FOR_BODY:.*]]
 ; CHECK-ZNVER4:       [[FOR_BODY]]:
-; CHECK-ZNVER4-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
-; CHECK-ZNVER4-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-ZNVER4-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-ZNVER4-NEXT:    [[IDX:%.*]] = phi i64 [ [[BC_MERGE_RDX11]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
 ; CHECK-ZNVER4-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
 ; CHECK-ZNVER4-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
 ; CHECK-ZNVER4-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
@@ -175,7 +254,7 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-ZNVER4-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
 ; CHECK-ZNVER4-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
 ; CHECK-ZNVER4-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-ZNVER4-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
+; CHECK-ZNVER4-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP7:![0-9]+]]
 ; CHECK-ZNVER4:       [[EXIT]]:
 ; CHECK-ZNVER4-NEXT:    ret void
 ;
diff --git a/llvm/test/Transforms/LoopVectorize/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
index c0dbd746fc997..c634cc83429e9 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-idioms.ll
@@ -1,55 +1,34 @@
 ; NOTE: Assertions have been autogenerated by utils/update_test_checks.py UTC_ARGS: --check-globals none --version 6
-; RUN: opt < %s -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -S | FileCheck %s -check-prefix=CHECK-IC1
-; RUN: opt < %s -force-target-supports-masked-memory-ops -force-vector-width=4 -tail-folding-policy=must-fold-tail -passes=loop-vectorize -S | FileCheck %s -check-prefix=CHECK-TF
+; RUN: opt < %s -force-target-supports-masked-memory-ops -force-vector-width=4 -passes=loop-vectorize -S | FileCheck %s -check-prefixes=CHECK,CHECK-IC1
+; RUN: opt < %s -force-target-supports-masked-memory-ops -force-vector-width=4 -tail-folding-policy=must-fold-tail -passes=loop-vectorize -S | FileCheck %s -check-prefixes=CHECK,CHECK-TF
 
 define void @test_compress_store_with_index(ptr writeonly noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_compress_store_with_index(
-; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[MIDDLE_BLOCK]]:
-; CHECK-IC1-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[IDX]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_compress_store_with_index(
-; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-TF-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-TF-NEXT:    br i1 [[CMP]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-TF:       [[MIDDLE_BLOCK]]:
-; CHECK-TF-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[IDX]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_compress_store_with_index(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP2]], <4 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP3:%.*]] = zext <4 x i1> [[TMP1]] to <4 x i64>
+; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP3]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -79,55 +58,33 @@ exit:
 }
 
 define void @test_expand_load_with_index(ptr noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_expand_load_with_index(
-; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[ENTRY:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_PH:.*]]
-; CHECK-IC1:       [[VECTOR_PH]]:
-; CHECK-IC1-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-IC1-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[VECTOR_BODY:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP2]], align 4
-; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_PH]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_PH]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_expand_load_with_index(
-; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[ENTRY:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_PH:.*]]
-; CHECK-TF:       [[VECTOR_PH]]:
-; CHECK-TF-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-TF-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-TF-NEXT:    br i1 [[CMP]], label %[[VECTOR_BODY:.*]], label %[[EXIT]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP2]], align 4
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_PH]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_PH]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_expand_load_with_index(
+; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    [[TMP3:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-NEXT:    call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP3]], ptr align 4 [[TMP0]], <4 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP4:%.*]] = zext <4 x i1> [[TMP1]] to <4 x i64>
+; CHECK-NEXT:    [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP3:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -158,55 +115,32 @@ exit:
 }
 
 define i64 @test_conditionally_incremented_phi_liveout(ptr writeonly noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
-; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[MIDDLE_BLOCK]]:
-; CHECK-IC1-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[IDX]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    [[CONDITIONAL_STEP:%.*]] = phi i64 [ [[IDX_1]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    ret i64 [[CONDITIONAL_STEP]]
-;
-; CHECK-TF-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
-; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-TF-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-TF-NEXT:    br i1 [[CMP]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-TF:       [[MIDDLE_BLOCK]]:
-; CHECK-TF-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[IDX]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    [[CONDITIONAL_STEP:%.*]] = phi i64 [ [[IDX_1]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    ret i64 [[CONDITIONAL_STEP]]
+; CHECK-LABEL: define i64 @test_conditionally_incremented_phi_liveout(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP2]], <4 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP3:%.*]] = zext <4 x i1> [[TMP1]] to <4 x i64>
+; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP3]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret i64 [[CONDITIONAL_STEP]]
 ;
 entry:
   br label %for.body
@@ -236,55 +170,33 @@ exit:
 }
 
 define void @test_compress_store_with_scaled_pointer(ptr writeonly noalias %dst.bytes, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_compress_store_with_scaled_pointer(
-; CHECK-IC1-SAME: ptr noalias writeonly [[DST_BYTES:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[IF_THEN]]:
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = shl nsw i64 [[CONDITIONAL_IV]], 2
-; CHECK-IC1-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP2]]
-; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[TMP3]], align 4
-; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_compress_store_with_scaled_pointer(
-; CHECK-TF-SAME: ptr noalias writeonly [[DST_BYTES:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-TF-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-TF-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[EXIT]]
-; CHECK-TF:       [[IF_THEN]]:
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = shl nsw i64 [[CONDITIONAL_IV]], 2
-; CHECK-TF-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP2]]
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[TMP3]], align 4
-; CHECK-TF-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_compress_store_with_scaled_pointer(
+; CHECK-SAME: ptr noalias writeonly [[DST_BYTES:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = shl nsw i64 [[CONDITIONAL_IV]], 2
+; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST_BYTES]], i64 [[TMP2]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP4:%.*]] = zext <4 x i1> [[TMP1]] to <4 x i64>
+; CHECK-NEXT:    [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -316,59 +228,34 @@ exit:
 
 ; Test a nested conditional compress store, where the phi is only updated on iterations where the store takes place.
 define void @test_nested_conditional_compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_nested_conditional_compress_store(
-; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[PHI:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[UPDATE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[UPDATE_BLOCK]]:
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
-; CHECK-IC1-NEXT:    [[UPDATE:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-IC1-NEXT:    [[CMP2:%.*]] = icmp sgt i32 [[LOAD_SRC]], 0
-; CHECK-IC1-NEXT:    br i1 [[CMP2]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-IC1:       [[MIDDLE_BLOCK]]:
-; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[TMP2]], align 4
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[PHI]] = phi i64 [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[UPDATE_BLOCK]] ], [ [[UPDATE]], %[[MIDDLE_BLOCK]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_nested_conditional_compress_store(
-; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[PHI:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-TF-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-TF-NEXT:    br i1 [[CMP]], label %[[UPDATE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-TF:       [[UPDATE_BLOCK]]:
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
-; CHECK-TF-NEXT:    [[UPDATE:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-TF-NEXT:    [[CMP2:%.*]] = icmp sgt i32 [[LOAD_SRC]], 0
-; CHECK-TF-NEXT:    br i1 [[CMP2]], label %[[MIDDLE_BLOCK:.*]], label %[[EXIT]]
-; CHECK-TF:       [[MIDDLE_BLOCK]]:
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[TMP2]], align 4
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[PHI]] = phi i64 [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[UPDATE_BLOCK]] ], [ [[UPDATE]], %[[MIDDLE_BLOCK]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_nested_conditional_compress_store(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    [[TMP3:%.*]] = icmp sgt <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT:    [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP2]], <4 x i1> [[TMP4]])
+; CHECK-NEXT:    [[TMP5:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP7:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP7]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -403,47 +290,26 @@ exit:
 
 ; An unconditional increment should lower as a simple induction (not a conditional induction).
 define void @test_unconditional_increment(ptr writeonly noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_unconditional_increment(
-; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-IC1-NEXT:  [[ENTRY:.*:]]
-; CHECK-IC1-NEXT:    br label %[[VECTOR_PH:.*]]
-; CHECK-IC1:       [[VECTOR_PH]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = add i64 15, [[INDEX]]
-; CHECK-IC1-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP0]]
-; CHECK-IC1-NEXT:    store <4 x i32> [[WIDE_LOAD]], ptr [[TMP2]], align 4
-; CHECK-IC1-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-IC1-NEXT:    [[TMP3:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[TMP3]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
-; CHECK-IC1:       [[MIDDLE_BLOCK]]:
-; CHECK-IC1-NEXT:    br label %[[EXIT:.*]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_unconditional_increment(
-; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[ENTRY:.*:]]
-; CHECK-TF-NEXT:    br label %[[VECTOR_PH:.*]]
-; CHECK-TF:       [[VECTOR_PH]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = add i64 15, [[INDEX]]
-; CHECK-TF-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP0]]
-; CHECK-TF-NEXT:    store <4 x i32> [[WIDE_LOAD]], ptr [[TMP2]], align 4
-; CHECK-TF-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
-; CHECK-TF-NEXT:    [[TMP3:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[TMP3]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
-; CHECK-TF:       [[MIDDLE_BLOCK]]:
-; CHECK-TF-NEXT:    br label %[[EXIT:.*]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_unconditional_increment(
+; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = add i64 15, [[INDEX]]
+; CHECK-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[TMP0]]
+; CHECK-NEXT:    store <4 x i32> [[WIDE_LOAD]], ptr [[TMP2]], align 4
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP3:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP3]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP7:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -473,85 +339,43 @@ exit:
 }
 
 define void @test_multiple_conditional_inductions(ptr %dst, ptr noalias %dst2, ptr noalias %src, ptr noalias %cond, ptr noalias %cond2) {
-; CHECK-IC1-LABEL: define void @test_multiple_conditional_inductions(
-; CHECK-IC1-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]]) {
-; CHECK-IC1-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[DST_INC:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[DST2_IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[DST2_INC:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[COND_VAL:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-IC1-NEXT:    [[COND_IS_ZERO:%.*]] = icmp eq i32 [[COND_VAL]], 0
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[SRC_VAL:%.*]] = load i32, ptr [[TMP2]], align 4
-; CHECK-IC1-NEXT:    br i1 [[COND_IS_ZERO]], label %[[IF_END:.*]], label %[[IF_THEN0:.*]]
-; CHECK-IC1:       [[IF_THEN0]]:
-; CHECK-IC1-NEXT:    [[DST_IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-IC1-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[CONDITIONAL_IV]]
-; CHECK-IC1-NEXT:    [[DST_VAL_TRUNC:%.*]] = trunc i32 [[SRC_VAL]] to i8
-; CHECK-IC1-NEXT:    store i8 [[DST_VAL_TRUNC]], ptr [[TMP3]], align 1
-; CHECK-IC1-NEXT:    br label %[[IF_END]]
-; CHECK-IC1:       [[IF_END]]:
-; CHECK-IC1-NEXT:    [[DST_INC]] = phi i64 [ [[DST_IDX_NEXT]], %[[IF_THEN0]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-IC1-NEXT:    [[TMP7:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
-; CHECK-IC1-NEXT:    [[COND2_VAL:%.*]] = load i32, ptr [[TMP7]], align 4
-; CHECK-IC1-NEXT:    [[COND2_IS_ZERO:%.*]] = icmp eq i32 [[COND2_VAL]], 0
-; CHECK-IC1-NEXT:    br i1 [[COND2_IS_ZERO]], label %[[EXIT]], label %[[MIDDLE_BLOCK:.*]]
-; CHECK-IC1:       [[MIDDLE_BLOCK]]:
-; CHECK-IC1-NEXT:    [[DST2_VAL_TRUNC:%.*]] = trunc i32 [[SRC_VAL]] to i16
-; CHECK-IC1-NEXT:    [[DST2_IDX_NEXT:%.*]] = add nsw i64 [[DST2_IDX]], 1
-; CHECK-IC1-NEXT:    [[DST2_GEP:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[DST2_IDX]]
-; CHECK-IC1-NEXT:    store i16 [[DST2_VAL_TRUNC]], ptr [[DST2_GEP]], align 2
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[DST2_INC]] = phi i64 [ [[DST2_IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[DST2_IDX]], %[[IF_END]] ]
-; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-IC1:       [[EXIT1]]:
-; CHECK-IC1-NEXT:    ret void
-;
-; CHECK-TF-LABEL: define void @test_multiple_conditional_inductions(
-; CHECK-TF-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]]) {
-; CHECK-TF-NEXT:  [[VECTOR_PH:.*]]:
-; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
-; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[DST_INC:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[DST2_IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[DST2_INC:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[COND_VAL:%.*]] = load i32, ptr [[TMP0]], align 4
-; CHECK-TF-NEXT:    [[COND_IS_ZERO:%.*]] = icmp eq i32 [[COND_VAL]], 0
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[SRC_VAL:%.*]] = load i32, ptr [[TMP2]], align 4
-; CHECK-TF-NEXT:    br i1 [[COND_IS_ZERO]], label %[[IF_END:.*]], label %[[IF_THEN0:.*]]
-; CHECK-TF:       [[IF_THEN0]]:
-; CHECK-TF-NEXT:    [[DST_IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-TF-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[CONDITIONAL_IV]]
-; CHECK-TF-NEXT:    [[DST_VAL_TRUNC:%.*]] = trunc i32 [[SRC_VAL]] to i8
-; CHECK-TF-NEXT:    store i8 [[DST_VAL_TRUNC]], ptr [[TMP3]], align 1
-; CHECK-TF-NEXT:    br label %[[IF_END]]
-; CHECK-TF:       [[IF_END]]:
-; CHECK-TF-NEXT:    [[DST_INC]] = phi i64 [ [[DST_IDX_NEXT]], %[[IF_THEN0]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_BODY]] ]
-; CHECK-TF-NEXT:    [[TMP7:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
-; CHECK-TF-NEXT:    [[COND2_VAL:%.*]] = load i32, ptr [[TMP7]], align 4
-; CHECK-TF-NEXT:    [[COND2_IS_ZERO:%.*]] = icmp eq i32 [[COND2_VAL]], 0
-; CHECK-TF-NEXT:    br i1 [[COND2_IS_ZERO]], label %[[EXIT]], label %[[MIDDLE_BLOCK:.*]]
-; CHECK-TF:       [[MIDDLE_BLOCK]]:
-; CHECK-TF-NEXT:    [[DST2_VAL_TRUNC:%.*]] = trunc i32 [[SRC_VAL]] to i16
-; CHECK-TF-NEXT:    [[DST2_IDX_NEXT:%.*]] = add nsw i64 [[DST2_IDX]], 1
-; CHECK-TF-NEXT:    [[DST2_GEP:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[DST2_IDX]]
-; CHECK-TF-NEXT:    store i16 [[DST2_VAL_TRUNC]], ptr [[DST2_GEP]], align 2
-; CHECK-TF-NEXT:    br label %[[EXIT]]
-; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[DST2_INC]] = phi i64 [ [[DST2_IDX_NEXT]], %[[MIDDLE_BLOCK]] ], [ [[DST2_IDX]], %[[IF_END]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[INDEX]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_BODY]]
-; CHECK-TF:       [[EXIT1]]:
-; CHECK-TF-NEXT:    ret void
+; CHECK-LABEL: define void @test_multiple_conditional_inductions(
+; CHECK-SAME: ptr [[DST:%.*]], ptr noalias [[DST2:%.*]], ptr noalias [[SRC:%.*]], ptr noalias [[COND:%.*]], ptr noalias [[COND2:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV1:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP4:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[COND]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD]], zeroinitializer
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD2:%.*]] = load <4 x i32>, ptr [[TMP2]], align 4
+; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i8, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    [[TMP4:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i8>
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i8.p0(<4 x i8> [[TMP4]], ptr align 1 [[TMP3]], <4 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP5:%.*]] = zext <4 x i1> [[TMP1]] to <4 x i64>
+; CHECK-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP5]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP6]]
+; CHECK-NEXT:    [[TMP7:%.*]] = getelementptr inbounds i32, ptr [[COND2]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD3:%.*]] = load <4 x i32>, ptr [[TMP7]], align 4
+; CHECK-NEXT:    [[TMP8:%.*]] = icmp ne <4 x i32> [[WIDE_LOAD3]], zeroinitializer
+; CHECK-NEXT:    [[TMP9:%.*]] = trunc <4 x i32> [[WIDE_LOAD2]] to <4 x i16>
+; CHECK-NEXT:    [[TMP10:%.*]] = getelementptr inbounds i16, ptr [[DST2]], i64 [[CONDITIONAL_IV1]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i16.p0(<4 x i16> [[TMP9]], ptr align 2 [[TMP10]], <4 x i1> [[TMP8]])
+; CHECK-NEXT:    [[TMP11:%.*]] = zext <4 x i1> [[TMP8]] to <4 x i64>
+; CHECK-NEXT:    [[TMP12:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP11]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP4]] = add i64 [[CONDITIONAL_IV1]], [[TMP12]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP13:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP13]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP8:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
 ;
 entry:
   br label %for.body
@@ -600,54 +424,146 @@ exit:
 
 ; Test expand load with always false update (this probably can be simplified).
 define void @test_expand_load_always_false_cond(ptr noalias %dst, ptr readonly %src, i32 %c) {
-; CHECK-IC1-LABEL: define void @test_expand_load_always_false_cond(
-; CHECK-IC1-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-LABEL: define void @test_expand_load_always_false_cond(
+; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
+; CHECK-NEXT:  [[ENTRY:.*:]]
+; CHECK-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i32 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[CONDITIONAL_IV]] to i64
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[TMP1]]
+; CHECK-NEXT:    [[TMP3:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> zeroinitializer, <4 x i32> poison)
+; CHECK-NEXT:    call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP3]], ptr align 4 [[TMP0]], <4 x i1> zeroinitializer)
+; CHECK-NEXT:    [[TMP4:%.*]] = call i32 @llvm.vector.reduce.add.v4i32(<4 x i32> zeroinitializer)
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i32 [[CONDITIONAL_IV]], [[TMP4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-NEXT:    [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP9:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br label %[[EXIT:.*]]
+; CHECK:       [[EXIT]]:
+; CHECK-NEXT:    ret void
+;
+entry:
+  br label %for.body
+
+for.body:
+  %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
+  %idx = phi i32 [ 0, %entry ], [ %idx.1, %for.inc ]
+  %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
+  %load.dst = load i32, ptr %dst.ptr, align 4
+  br i1 0, label %if.then, label %for.inc
+
+if.then:
+  %src.idx = sext i32 %idx to i64
+  %src.ptr = getelementptr inbounds i32, ptr %src, i64 %src.idx
+  %load.src = load i32, ptr %src.ptr, align 4
+  store i32 %load.src, ptr %dst.ptr, align 4
+  %idx.next = add nsw i32 %idx, 1
+  br label %for.inc
+
+for.inc:
+  %idx.1 = phi i32 [ %idx.next, %if.then ], [ %idx, %for.body ]
+  %iv.next = add nuw nsw i64 %iv, 1
+  %exitcond.not = icmp eq i64 %iv.next, 4096
+  br i1 %exitcond.not, label %exit, label %for.body
+
+exit:
+  ret void
+}
+
+define void @test_compress_with_unknown_tc(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
+; CHECK-IC1-LABEL: define void @test_compress_with_unknown_tc(
+; CHECK-IC1-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
 ; CHECK-IC1-NEXT:  [[ENTRY:.*]]:
-; CHECK-IC1-NEXT:    br label %[[VECTOR_PH:.*]]
+; CHECK-IC1-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
+; CHECK-IC1-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[SCALAR_PH:.*]], label %[[VECTOR_PH:.*]]
 ; CHECK-IC1:       [[VECTOR_PH]]:
-; CHECK-IC1-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i32 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-IC1-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-IC1-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    br i1 false, label %[[VECTOR_BODY:.*]], label %[[EXIT]]
+; CHECK-IC1-NEXT:    [[TMP0:%.*]] = and i64 [[N]], 3
+; CHECK-IC1-NEXT:    [[N_VEC:%.*]] = sub i64 [[N]], [[TMP0]]
+; CHECK-IC1-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-IC1-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-IC1-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK-IC1:       [[VECTOR_BODY]]:
-; CHECK-IC1-NEXT:    [[TMP1:%.*]] = sext i32 [[CONDITIONAL_IV]] to i64
-; CHECK-IC1-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[TMP1]]
-; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP2]], align 4
+; CHECK-IC1-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-IC1-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-IC1-NEXT:    [[WIDE_LOAD:%.*]] = load <4 x i32>, ptr [[TMP1]], align 4
+; CHECK-IC1-NEXT:    [[TMP2:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-IC1-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-IC1-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP3]], <4 x i1> [[TMP2]])
+; CHECK-IC1-NEXT:    [[TMP4:%.*]] = zext <4 x i1> [[TMP2]] to <4 x i64>
+; CHECK-IC1-NEXT:    [[TMP5:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP4]])
+; CHECK-IC1-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP5]]
+; CHECK-IC1-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 4
+; CHECK-IC1-NEXT:    [[TMP6:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-IC1-NEXT:    br i1 [[TMP6]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
+; CHECK-IC1:       [[MIDDLE_BLOCK]]:
+; CHECK-IC1-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
+; CHECK-IC1-NEXT:    br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[SCALAR_PH]]
+; CHECK-IC1:       [[SCALAR_PH]]:
+; CHECK-IC1-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT:    [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[MIDDLE_BLOCK]] ], [ 0, %[[ENTRY]] ]
+; CHECK-IC1-NEXT:    br label %[[FOR_BODY:.*]]
+; CHECK-IC1:       [[FOR_BODY]]:
+; CHECK-IC1-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-IC1-NEXT:    [[IDX:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-IC1-NEXT:    [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
+; CHECK-IC1-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
+; CHECK-IC1-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
+; CHECK-IC1-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
+; CHECK-IC1:       [[IF_THEN]]:
+; CHECK-IC1-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
 ; CHECK-IC1-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i32 [[CONDITIONAL_IV]], 1
-; CHECK-IC1-NEXT:    br label %[[EXIT]]
-; CHECK-IC1:       [[EXIT]]:
-; CHECK-IC1-NEXT:    [[IDX_1]] = phi i32 [ [[IDX_NEXT]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_PH]] ]
+; CHECK-IC1-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
+; CHECK-IC1-NEXT:    br label %[[FOR_INC]]
+; CHECK-IC1:       [[FOR_INC]]:
+; CHECK-IC1-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
 ; CHECK-IC1-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_PH]]
-; CHECK-IC1:       [[EXIT1]]:
+; CHECK-IC1-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
+; CHECK-IC1-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP11:![0-9]+]]
+; CHECK-IC1:       [[EXIT]]:
 ; CHECK-IC1-NEXT:    ret void
 ;
-; CHECK-TF-LABEL: define void @test_expand_load_always_false_cond(
-; CHECK-TF-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-TF-NEXT:  [[ENTRY:.*]]:
+; CHECK-TF-LABEL: define void @test_compress_with_unknown_tc(
+; CHECK-TF-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) {
+; CHECK-TF-NEXT:  [[ENTRY:.*:]]
 ; CHECK-TF-NEXT:    br label %[[VECTOR_PH:.*]]
 ; CHECK-TF:       [[VECTOR_PH]]:
-; CHECK-TF-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[ENTRY]] ], [ [[IV_NEXT:%.*]], %[[EXIT:.*]] ]
-; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i32 [ 0, %[[ENTRY]] ], [ [[IDX_1:%.*]], %[[EXIT]] ]
-; CHECK-TF-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-TF-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    br i1 false, label %[[VECTOR_BODY:.*]], label %[[EXIT]]
+; CHECK-TF-NEXT:    [[N_RND_UP:%.*]] = add i64 [[N]], 3
+; CHECK-TF-NEXT:    [[TMP0:%.*]] = and i64 [[N_RND_UP]], 3
+; CHECK-TF-NEXT:    [[N_VEC:%.*]] = sub i64 [[N_RND_UP]], [[TMP0]]
+; CHECK-TF-NEXT:    [[TRIP_COUNT_MINUS_1:%.*]] = sub i64 [[N]], 1
+; CHECK-TF-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <4 x i64> poison, i64 [[TRIP_COUNT_MINUS_1]], i64 0
+; CHECK-TF-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <4 x i64> [[BROADCAST_SPLATINSERT]], <4 x i64> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT:    [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-TF-NEXT:    [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-TF-NEXT:    br label %[[VECTOR_BODY:.*]]
 ; CHECK-TF:       [[VECTOR_BODY]]:
-; CHECK-TF-NEXT:    [[TMP1:%.*]] = sext i32 [[CONDITIONAL_IV]] to i64
-; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[TMP1]]
-; CHECK-TF-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP2]], align 4
-; CHECK-TF-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-TF-NEXT:    [[IDX_NEXT:%.*]] = add nsw i32 [[CONDITIONAL_IV]], 1
-; CHECK-TF-NEXT:    br label %[[EXIT]]
+; CHECK-TF-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT:    [[VEC_IND:%.*]] = phi <4 x i64> [ <i64 0, i64 1, i64 2, i64 3>, %[[VECTOR_PH]] ], [ [[VEC_IND_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-TF-NEXT:    [[TMP1:%.*]] = icmp ule <4 x i64> [[VEC_IND]], [[BROADCAST_SPLAT]]
+; CHECK-TF-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-TF-NEXT:    [[WIDE_MASKED_LOAD:%.*]] = call <4 x i32> @llvm.masked.load.v4i32.p0(ptr align 4 [[TMP2]], <4 x i1> [[TMP1]], <4 x i32> poison)
+; CHECK-TF-NEXT:    [[TMP3:%.*]] = icmp slt <4 x i32> [[WIDE_MASKED_LOAD]], [[BROADCAST_SPLAT2]]
+; CHECK-TF-NEXT:    [[TMP4:%.*]] = select <4 x i1> [[TMP1]], <4 x i1> [[TMP3]], <4 x i1> zeroinitializer
+; CHECK-TF-NEXT:    [[TMP5:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-TF-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_MASKED_LOAD]], ptr align 4 [[TMP5]], <4 x i1> [[TMP4]])
+; CHECK-TF-NEXT:    [[TMP6:%.*]] = zext <4 x i1> [[TMP4]] to <4 x i64>
+; CHECK-TF-NEXT:    [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP6]])
+; CHECK-TF-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
+; CHECK-TF-NEXT:    [[INDEX_NEXT]] = add i64 [[INDEX]], 4
+; CHECK-TF-NEXT:    [[VEC_IND_NEXT]] = add nuw <4 x i64> [[VEC_IND]], splat (i64 4)
+; CHECK-TF-NEXT:    [[TMP8:%.*]] = icmp eq i64 [[INDEX_NEXT]], [[N_VEC]]
+; CHECK-TF-NEXT:    br i1 [[TMP8]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP10:![0-9]+]]
+; CHECK-TF:       [[MIDDLE_BLOCK]]:
+; CHECK-TF-NEXT:    br label %[[EXIT:.*]]
 ; CHECK-TF:       [[EXIT]]:
-; CHECK-TF-NEXT:    [[IDX_1]] = phi i32 [ [[IDX_NEXT]], %[[VECTOR_BODY]] ], [ [[CONDITIONAL_IV]], %[[VECTOR_PH]] ]
-; CHECK-TF-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-TF-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4096
-; CHECK-TF-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT1:.*]], label %[[VECTOR_PH]]
-; CHECK-TF:       [[EXIT1]]:
 ; CHECK-TF-NEXT:    ret void
 ;
 entry:
@@ -655,23 +571,22 @@ entry:
 
 for.body:
   %iv = phi i64 [ 0, %entry ], [ %iv.next, %for.inc ]
-  %idx = phi i32 [ 0, %entry ], [ %idx.1, %for.inc ]
-  %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %iv
-  %load.dst = load i32, ptr %dst.ptr, align 4
-  br i1 0, label %if.then, label %for.inc
+  %idx = phi i64 [ 0, %entry ], [ %idx.1, %for.inc ]
+  %src.ptr = getelementptr inbounds i32, ptr %src, i64 %iv
+  %load.src = load i32, ptr %src.ptr, align 4
+  %cmp = icmp slt i32 %load.src, %c
+  br i1 %cmp, label %if.then, label %for.inc
 
 if.then:
-  %src.idx = sext i32 %idx to i64
-  %src.ptr = getelementptr inbounds i32, ptr %src, i64 %src.idx
-  %load.src = load i32, ptr %src.ptr, align 4
+  %dst.ptr = getelementptr inbounds i32, ptr %dst, i64 %idx
   store i32 %load.src, ptr %dst.ptr, align 4
-  %idx.next = add nsw i32 %idx, 1
+  %idx.next = add nsw i64 %idx, 1
   br label %for.inc
 
 for.inc:
-  %idx.1 = phi i32 [ %idx.next, %if.then ], [ %idx, %for.body ]
+  %idx.1 = phi i64 [ %idx.next, %if.then ], [ %idx, %for.body ]
   %iv.next = add nuw nsw i64 %iv, 1
-  %exitcond.not = icmp eq i64 %iv.next, 4096
+  %exitcond.not = icmp eq i64 %iv.next, %n
   br i1 %exitcond.not, label %exit, label %for.body
 
 exit:
diff --git a/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
index 3932acd2f2a3f..0129d5a8b673e 100644
--- a/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
+++ b/llvm/test/Transforms/LoopVectorize/compress-store-vec-epilogue.ll
@@ -4,11 +4,61 @@
 define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c) {
 ; CHECK-LABEL: define void @compress_store(
 ; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]]) {
-; CHECK-NEXT:  [[VEC_EPILOG_SCALAR_PH:.*]]:
+; CHECK-NEXT:  [[ITER_CHECK:.*]]:
+; CHECK-NEXT:    br i1 false, label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
+; CHECK:       [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
+; CHECK-NEXT:    br i1 false, label %[[VEC_EPILOG_PH:.*]], label %[[VECTOR_PH:.*]]
+; CHECK:       [[VECTOR_PH]]:
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <16 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <16 x i32> [[BROADCAST_SPLATINSERT]], <16 x i32> poison, <16 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VECTOR_BODY:.*]]
+; CHECK:       [[VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[INDEX_NEXT:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 42, %[[VECTOR_PH]] ], [ [[CONDITIONAL_STEP:%.*]], %[[VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[WIDE_LOAD:%.*]] = load <16 x i32>, ptr [[TMP0]], align 4
+; CHECK-NEXT:    [[TMP1:%.*]] = icmp slt <16 x i32> [[WIDE_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP2:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v16i32.p0(<16 x i32> [[WIDE_LOAD]], ptr align 4 [[TMP2]], <16 x i1> [[TMP1]])
+; CHECK-NEXT:    [[TMP3:%.*]] = zext <16 x i1> [[TMP1]] to <16 x i64>
+; CHECK-NEXT:    [[TMP4:%.*]] = call i64 @llvm.vector.reduce.add.v16i64(<16 x i64> [[TMP3]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP4]]
+; CHECK-NEXT:    [[INDEX_NEXT]] = add nuw i64 [[INDEX]], 16
+; CHECK-NEXT:    [[TMP5:%.*]] = icmp eq i64 [[INDEX_NEXT]], 4096
+; CHECK-NEXT:    br i1 [[TMP5]], label %[[MIDDLE_BLOCK:.*]], label %[[VECTOR_BODY]], !llvm.loop [[LOOP0:![0-9]+]]
+; CHECK:       [[MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br i1 false, label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
+; CHECK:       [[VEC_EPILOG_ITER_CHECK]]:
+; CHECK-NEXT:    br i1 false, label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF3:![0-9]+]]
+; CHECK:       [[VEC_EPILOG_PH]]:
+; CHECK-NEXT:    [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ 4096, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT:    [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT1:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT2:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT1]], <4 x i32> poison, <4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
+; CHECK:       [[VEC_EPILOG_VECTOR_BODY]]:
+; CHECK-NEXT:    [[INDEX3:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT7:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV4:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[CONDITIONAL_STEP6:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
+; CHECK-NEXT:    [[TMP6:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX3]]
+; CHECK-NEXT:    [[WIDE_LOAD5:%.*]] = load <4 x i32>, ptr [[TMP6]], align 4
+; CHECK-NEXT:    [[TMP7:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD5]], [[BROADCAST_SPLAT2]]
+; CHECK-NEXT:    [[TMP8:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[CONDITIONAL_IV4]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.v4i32.p0(<4 x i32> [[WIDE_LOAD5]], ptr align 4 [[TMP8]], <4 x i1> [[TMP7]])
+; CHECK-NEXT:    [[TMP9:%.*]] = zext <4 x i1> [[TMP7]] to <4 x i64>
+; CHECK-NEXT:    [[TMP10:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP9]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP6]] = add i64 [[CONDITIONAL_IV4]], [[TMP10]]
+; CHECK-NEXT:    [[INDEX_NEXT7]] = add nuw i64 [[INDEX3]], 4
+; CHECK-NEXT:    [[TMP11:%.*]] = icmp eq i64 [[INDEX_NEXT7]], 4108
+; CHECK-NEXT:    br i1 [[TMP11]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP4:![0-9]+]]
+; CHECK:       [[VEC_EPILOG_MIDDLE_BLOCK]]:
+; CHECK-NEXT:    br i1 false, label %[[EXIT]], label %[[VEC_EPILOG_SCALAR_PH]]
+; CHECK:       [[VEC_EPILOG_SCALAR_PH]]:
+; CHECK-NEXT:    [[BC_RESUME_VAL:%.*]] = phi i64 [ 4108, %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ 4096, %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[ITER_CHECK]] ]
+; CHECK-NEXT:    [[BC_MERGE_RDX8:%.*]] = phi i64 [ [[CONDITIONAL_STEP6]], %[[VEC_EPILOG_MIDDLE_BLOCK]] ], [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 42, %[[ITER_CHECK]] ]
 ; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
 ; CHECK:       [[FOR_BODY]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
-; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ 42, %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
+; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ [[BC_RESUME_VAL]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
+; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ [[BC_MERGE_RDX8]], %[[VEC_EPILOG_SCALAR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
 ; CHECK-NEXT:    [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
 ; CHECK-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
 ; CHECK-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
@@ -22,7 +72,7 @@ define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %
 ; CHECK-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
 ; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
 ; CHECK-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], 4110
-; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
+; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT]], label %[[FOR_BODY]], !llvm.loop [[LOOP5:![0-9]+]]
 ; CHECK:       [[EXIT]]:
 ; CHECK-NEXT:    ret void
 ;

>From ed189ba949de84173e1d03969719701adfc229d3 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Tue, 22 Sep 2026 15:15:24 +0000
Subject: [PATCH 3/6] Always pass a vector type to
 TTI.isLegalMaskedCompressStore/ExpandLoad

---
 .../lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp | 7 ++++---
 llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h  | 2 +-
 llvm/lib/Transforms/Vectorize/LoopVectorize.cpp           | 8 ++++----
 3 files changed, 9 insertions(+), 8 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
index 26d92284d31b0..ff34beae3ca33 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.cpp
@@ -160,10 +160,11 @@ bool VFSelectionContext::isLegalGatherOrScatter(bool IsLoad, Type *ScalarTy,
 }
 
 bool VFSelectionContext::isLegalExpandLoadOrCompressStore(
-    bool IsLoad, Type *ScalarTy, Align Alignment) const {
+    bool IsLoad, Type *ScalarTy, Align Alignment, ElementCount VF) const {
+  Type *VectorTy = toVectorTy(ScalarTy, VF);
   return ForceTargetSupportsMaskedMemoryOps ||
-         (IsLoad ? TTI.isLegalMaskedExpandLoad(ScalarTy, Alignment)
-                 : TTI.isLegalMaskedCompressStore(ScalarTy, Alignment));
+         (IsLoad ? TTI.isLegalMaskedExpandLoad(VectorTy, Alignment)
+                 : TTI.isLegalMaskedCompressStore(VectorTy, Alignment));
 }
 
 bool VFSelectionContext::supportsScalableVectors() const {
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
index 4fe32f4467104..e803b79b597bd 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
@@ -843,7 +843,7 @@ class VFSelectionContext {
   /// IsLoad) or masked compress store of scalar type \p ScalarTy with \p
   /// Alignment.
   bool isLegalExpandLoadOrCompressStore(bool IsLoad, Type *ScalarTy,
-                                        Align Alignment) const;
+                                        Align Alignment, ElementCount VF) const;
 
   /// Split reductions into those that happen in the loop, and those that
   /// happen outside. In-loop reductions are collected into InLoopReductions.
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 2078473933465..fd28d6950f62c 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -1075,7 +1075,7 @@ class LoopVectorizationCostModel {
 
   /// Returns true if the target machine supports a masked expand load or masked
   /// compress store for \p I's data type and alignment.
-  bool isLegalExpandLoadOrCompressStore(Instruction *I) const;
+  bool isLegalExpandLoadOrCompressStore(Instruction *I, ElementCount VF) const;
 
   /// Check if \p Instr belongs to any interleaved access group.
   bool isAccessInterleaved(Instruction *Instr) const {
@@ -2391,10 +2391,10 @@ bool LoopVectorizationCostModel::isLegalGatherOrScatter(Instruction *I,
 }
 
 bool LoopVectorizationCostModel::isLegalExpandLoadOrCompressStore(
-    Instruction *I) const {
+    Instruction *I, ElementCount VF) const {
   assert(isa<LoadInst>(I) || isa<StoreInst>(I));
   return Config.isLegalExpandLoadOrCompressStore(
-      isa<LoadInst>(I), getLoadStoreType(I), getLoadStoreAlignment(I));
+      isa<LoadInst>(I), getLoadStoreType(I), getLoadStoreAlignment(I), VF);
 }
 
 bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
@@ -2423,7 +2423,7 @@ bool LoopVectorizationCostModel::isScalarWithPredication(Instruction *I,
     bool IsConsecutive = Legal->isConsecutivePtr(ScalarTy, Ptr);
     return !(IsConsecutive && !IsCompressed &&
              isLegalMaskedLoadOrStore(I, VF)) &&
-           !(IsCompressed && isLegalExpandLoadOrCompressStore(I)) &&
+           !(IsCompressed && isLegalExpandLoadOrCompressStore(I, VF)) &&
            !isLegalGatherOrScatter(I, VF);
   }
   case Instruction::UDiv:

>From f7f6d06ce3a2b2858011525f49226dc25c880d39 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Wed, 23 Sep 2026 11:09:56 +0000
Subject: [PATCH 4/6] Rebase fixups

---
 llvm/lib/Transforms/Utils/LoopUtils.cpp | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/llvm/lib/Transforms/Utils/LoopUtils.cpp b/llvm/lib/Transforms/Utils/LoopUtils.cpp
index a981dcf6e4b40..63935572e6885 100644
--- a/llvm/lib/Transforms/Utils/LoopUtils.cpp
+++ b/llvm/lib/Transforms/Utils/LoopUtils.cpp
@@ -2530,7 +2530,7 @@ bool llvm::collectCompressedPtrs(
   // Over-approximates the conditional induction as a SCEVAddRec assuming the
   // condition is always true.
   const SCEV *ApproximatePhiSCEV = SE.getAddRecExpr(
-      CondID.getStartSCEV(), CondID.getStepSCEV(), &L, SCEV::FlagAnyWrap);
+      CondID.getStartSCEV(), CondID.getStepSCEV(), &L, SCEV::FlagNone);
 
   // TODO: Take into account the non-wrap flags of the MD when rewriting the
   // SCEV expressions for pointers. This should allow folding away zext/sext

>From ce28845d39002e1dde040837cd65900067f7bdde Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Wed, 23 Sep 2026 12:00:49 +0000
Subject: [PATCH 5/6] Add LLVM_ABI marker

---
 llvm/include/llvm/Transforms/Utils/LoopUtils.h | 7 +++----
 1 file changed, 3 insertions(+), 4 deletions(-)

diff --git a/llvm/include/llvm/Transforms/Utils/LoopUtils.h b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
index 1245c122e3906..7b49242a23353 100644
--- a/llvm/include/llvm/Transforms/Utils/LoopUtils.h
+++ b/llvm/include/llvm/Transforms/Utils/LoopUtils.h
@@ -714,10 +714,9 @@ hasPartialIVCondition(const Loop &L, unsigned MSSAThreshold,
 /// approximate SCEV expressions (assuming the monotonic PHI always increments)
 /// for the pointers are placed in \p CompressedPtrs. Returns true if all
 /// in-loop users of the conditional induction are loads/stores.
-bool collectCompressedPtrs(DenseMap<Value *, const SCEV *> &CompressedPtrs,
-                           const Loop &L,
-                           const ConditionalInductionDescriptor &CondID,
-                           ScalarEvolution &SE);
+LLVM_ABI bool collectCompressedPtrs(
+    DenseMap<Value *, const SCEV *> &CompressedPtrs, const Loop &L,
+    const ConditionalInductionDescriptor &CondID, ScalarEvolution &SE);
 
 } // end namespace llvm
 

>From 8440f58fd2b4d862c7b7f4c025a65ef81b0561e4 Mon Sep 17 00:00:00 2001
From: Benjamin Maxwell <benjamin.maxwell at arm.com>
Date: Fri, 25 Sep 2026 15:12:38 +0000
Subject: [PATCH 6/6] Rebase tests

---
 .../LoopVectorize/RISCV/compress-idioms.ll    | 74 +++++++++++--------
 .../LoopVectorize/X86/compress-idioms.ll      | 24 +++---
 2 files changed, 57 insertions(+), 41 deletions(-)

diff --git a/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
index 3bc2fe5804c74..25d0560a6441c 100644
--- a/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/RISCV/compress-idioms.ll
@@ -4,25 +4,33 @@
 define void @compress_store(ptr writeonly noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-LABEL: define void @compress_store(
 ; CHECK-SAME: ptr noalias writeonly [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0:[0-9]+]] {
-; CHECK-NEXT:  [[VECTOR_PH:.*]]:
+; CHECK-NEXT:  [[VECTOR_PH:.*:]]
 ; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
 ; CHECK:       [[FOR_BODY]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IV_NEXT:%.*]], %[[FOR_INC:.*]] ]
-; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[VECTOR_PH]] ], [ [[IDX_1:%.*]], %[[FOR_INC]] ]
-; CHECK-NEXT:    [[SRC_PTR:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[IV]]
-; CHECK-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[SRC_PTR]], align 4
-; CHECK-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_SRC]], [[C]]
-; CHECK-NEXT:    br i1 [[CMP]], label %[[IF_THEN:.*]], label %[[FOR_INC]]
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[IF_THEN:.*]]
 ; CHECK:       [[IF_THEN]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[FOR_BODY]] ], [ [[CURRENT_ITERATION_NEXT:%.*]], %[[IF_THEN]] ]
+; CHECK-NEXT:    [[IDX:%.*]] = phi i64 [ 0, %[[FOR_BODY]] ], [ [[CONDITIONAL_STEP:%.*]], %[[IF_THEN]] ]
+; CHECK-NEXT:    [[AVL:%.*]] = phi i64 [ [[N]], %[[FOR_BODY]] ], [ [[AVL_NEXT:%.*]], %[[IF_THEN]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = call i32 @llvm.experimental.get.vector.length.i64(i64 [[AVL]], i32 4, i1 true)
+; CHECK-NEXT:    [[TMP1:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[INDEX]]
+; CHECK-NEXT:    [[VP_OP_LOAD:%.*]] = call <vscale x 4 x i32> @llvm.vp.load.nxv4i32.p0(ptr align 4 [[TMP1]], <vscale x 4 x i1> splat (i1 true), i32 [[TMP0]])
+; CHECK-NEXT:    [[TMP2:%.*]] = icmp slt <vscale x 4 x i32> [[VP_OP_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP3:%.*]] = call <vscale x 4 x i1> @llvm.vp.merge.nxv4i1(<vscale x 4 x i1> splat (i1 true), <vscale x 4 x i1> [[TMP2]], <vscale x 4 x i1> zeroinitializer, i32 [[TMP0]])
 ; CHECK-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IDX]]
-; CHECK-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[IDX]], 1
-; CHECK-NEXT:    br label %[[FOR_INC]]
+; CHECK-NEXT:    call void @llvm.masked.compressstore.nxv4i32.p0(<vscale x 4 x i32> [[VP_OP_LOAD]], ptr align 4 [[DST_PTR]], <vscale x 4 x i1> [[TMP3]])
+; CHECK-NEXT:    [[TMP5:%.*]] = zext <vscale x 4 x i1> [[TMP3]] to <vscale x 4 x i64>
+; CHECK-NEXT:    [[TMP6:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP5]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[IDX]], [[TMP6]]
+; CHECK-NEXT:    [[TMP7:%.*]] = zext i32 [[TMP0]] to i64
+; CHECK-NEXT:    [[CURRENT_ITERATION_NEXT]] = add i64 [[TMP7]], [[INDEX]]
+; CHECK-NEXT:    [[AVL_NEXT]] = sub nuw i64 [[AVL]], [[TMP7]]
+; CHECK-NEXT:    [[TMP8:%.*]] = icmp eq i64 [[AVL_NEXT]], 0
+; CHECK-NEXT:    br i1 [[TMP8]], label %[[FOR_INC:.*]], label %[[IF_THEN]], !llvm.loop [[LOOP0:![0-9]+]]
 ; CHECK:       [[FOR_INC]]:
-; CHECK-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN]] ], [ [[IDX]], %[[FOR_BODY]] ]
-; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[EXIT:.*]], label %[[FOR_BODY]]
+; CHECK-NEXT:    br label %[[EXIT:.*]]
 ; CHECK:       [[EXIT]]:
 ; CHECK-NEXT:    ret void
 ;
@@ -56,26 +64,34 @@ exit:
 define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-LABEL: define void @expand_load(
 ; CHECK-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
-; CHECK-NEXT:  [[IF_THEN:.*]]:
+; CHECK-NEXT:  [[IF_THEN:.*:]]
 ; CHECK-NEXT:    br label %[[FOR_INC:.*]]
 ; CHECK:       [[FOR_INC]]:
-; CHECK-NEXT:    [[IV:%.*]] = phi i64 [ 0, %[[IF_THEN]] ], [ [[IV_NEXT:%.*]], %[[MIDDLE_BLOCK:.*]] ]
-; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[IF_THEN]] ], [ [[IDX_1:%.*]], %[[MIDDLE_BLOCK]] ]
-; CHECK-NEXT:    [[DST_PTR:%.*]] = getelementptr inbounds i32, ptr [[DST]], i64 [[IV]]
-; CHECK-NEXT:    [[LOAD_DST:%.*]] = load i32, ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[CMP:%.*]] = icmp slt i32 [[LOAD_DST]], [[C]]
-; CHECK-NEXT:    br i1 [[CMP]], label %[[IF_THEN1:.*]], label %[[MIDDLE_BLOCK]]
+; CHECK-NEXT:    [[BROADCAST_SPLATINSERT:%.*]] = insertelement <vscale x 4 x i32> poison, i32 [[C]], i64 0
+; CHECK-NEXT:    [[BROADCAST_SPLAT:%.*]] = shufflevector <vscale x 4 x i32> [[BROADCAST_SPLATINSERT]], <vscale x 4 x i32> poison, <vscale x 4 x i32> zeroinitializer
+; CHECK-NEXT:    br label %[[IF_THEN1:.*]]
 ; CHECK:       [[IF_THEN1]]:
+; CHECK-NEXT:    [[INDEX:%.*]] = phi i64 [ 0, %[[FOR_INC]] ], [ [[CURRENT_ITERATION_NEXT:%.*]], %[[IF_THEN1]] ]
+; CHECK-NEXT:    [[CONDITIONAL_IV:%.*]] = phi i64 [ 0, %[[FOR_INC]] ], [ [[CONDITIONAL_STEP:%.*]], %[[IF_THEN1]] ]
+; CHECK-NEXT:    [[AVL:%.*]] = phi i64 [ [[N]], %[[FOR_INC]] ], [ [[AVL_NEXT:%.*]], %[[IF_THEN1]] ]
+; CHECK-NEXT:    [[TMP0:%.*]] = call i32 @llvm.experimental.get.vector.length.i64(i64 [[AVL]], i32 4, i1 true)
+; CHECK-NEXT:    [[TMP1:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX]]
+; CHECK-NEXT:    [[VP_OP_LOAD:%.*]] = call <vscale x 4 x i32> @llvm.vp.load.nxv4i32.p0(ptr align 4 [[TMP1]], <vscale x 4 x i1> splat (i1 true), i32 [[TMP0]])
+; CHECK-NEXT:    [[TMP2:%.*]] = icmp slt <vscale x 4 x i32> [[VP_OP_LOAD]], [[BROADCAST_SPLAT]]
+; CHECK-NEXT:    [[TMP4:%.*]] = call <vscale x 4 x i1> @llvm.vp.merge.nxv4i1(<vscale x 4 x i1> splat (i1 true), <vscale x 4 x i1> [[TMP2]], <vscale x 4 x i1> zeroinitializer, i32 [[TMP0]])
 ; CHECK-NEXT:    [[TMP3:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV]]
-; CHECK-NEXT:    [[LOAD_SRC:%.*]] = load i32, ptr [[TMP3]], align 4
-; CHECK-NEXT:    store i32 [[LOAD_SRC]], ptr [[DST_PTR]], align 4
-; CHECK-NEXT:    [[IDX_NEXT:%.*]] = add nsw i64 [[CONDITIONAL_IV]], 1
-; CHECK-NEXT:    br label %[[MIDDLE_BLOCK]]
+; CHECK-NEXT:    [[TMP5:%.*]] = call <vscale x 4 x i32> @llvm.masked.expandload.nxv4i32.p0(ptr align 4 [[TMP3]], <vscale x 4 x i1> [[TMP4]], <vscale x 4 x i32> poison)
+; CHECK-NEXT:    call void @llvm.vp.store.nxv4i32.p0(<vscale x 4 x i32> [[TMP5]], ptr align 4 [[TMP1]], <vscale x 4 x i1> [[TMP2]], i32 [[TMP0]])
+; CHECK-NEXT:    [[TMP6:%.*]] = zext <vscale x 4 x i1> [[TMP4]] to <vscale x 4 x i64>
+; CHECK-NEXT:    [[TMP7:%.*]] = call i64 @llvm.vector.reduce.add.nxv4i64(<vscale x 4 x i64> [[TMP6]])
+; CHECK-NEXT:    [[CONDITIONAL_STEP]] = add i64 [[CONDITIONAL_IV]], [[TMP7]]
+; CHECK-NEXT:    [[TMP8:%.*]] = zext i32 [[TMP0]] to i64
+; CHECK-NEXT:    [[CURRENT_ITERATION_NEXT]] = add i64 [[TMP8]], [[INDEX]]
+; CHECK-NEXT:    [[AVL_NEXT]] = sub nuw i64 [[AVL]], [[TMP8]]
+; CHECK-NEXT:    [[TMP9:%.*]] = icmp eq i64 [[AVL_NEXT]], 0
+; CHECK-NEXT:    br i1 [[TMP9]], label %[[MIDDLE_BLOCK:.*]], label %[[IF_THEN1]], !llvm.loop [[LOOP3:![0-9]+]]
 ; CHECK:       [[MIDDLE_BLOCK]]:
-; CHECK-NEXT:    [[IDX_1]] = phi i64 [ [[IDX_NEXT]], %[[IF_THEN1]] ], [ [[CONDITIONAL_IV]], %[[FOR_INC]] ]
-; CHECK-NEXT:    [[IV_NEXT]] = add nuw nsw i64 [[IV]], 1
-; CHECK-NEXT:    [[EXITCOND_NOT:%.*]] = icmp eq i64 [[IV_NEXT]], [[N]]
-; CHECK-NEXT:    br i1 [[EXITCOND_NOT]], label %[[FOR_BODY:.*]], label %[[FOR_INC]]
+; CHECK-NEXT:    br label %[[FOR_BODY:.*]]
 ; CHECK:       [[FOR_BODY]]:
 ; CHECK-NEXT:    ret void
 ;
diff --git a/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
index cb905357816f9..36227ac4ab7e6 100644
--- a/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
+++ b/llvm/test/Transforms/LoopVectorize/X86/compress-idioms.ll
@@ -175,7 +175,7 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-ZNVER4-LABEL: define void @expand_load(
 ; CHECK-ZNVER4-SAME: ptr noalias [[DST:%.*]], ptr readonly [[SRC:%.*]], i32 [[C:%.*]], i64 [[N:%.*]]) #[[ATTR0]] {
 ; CHECK-ZNVER4-NEXT:  [[ITER_CHECK:.*]]:
-; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 8
+; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK:%.*]] = icmp ult i64 [[N]], 4
 ; CHECK-ZNVER4-NEXT:    br i1 [[MIN_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH:.*]], label %[[VECTOR_MAIN_LOOP_ITER_CHECK:.*]]
 ; CHECK-ZNVER4:       [[VECTOR_MAIN_LOOP_ITER_CHECK]]:
 ; CHECK-ZNVER4-NEXT:    [[MIN_ITERS_CHECK1:%.*]] = icmp ult i64 [[N]], 16
@@ -205,29 +205,29 @@ define void @expand_load(ptr noalias %dst, ptr readonly %src, i32 %c, i64 %n) {
 ; CHECK-ZNVER4-NEXT:    [[CMP_N:%.*]] = icmp eq i64 [[N]], [[N_VEC]]
 ; CHECK-ZNVER4-NEXT:    br i1 [[CMP_N]], label %[[EXIT:.*]], label %[[VEC_EPILOG_ITER_CHECK:.*]]
 ; CHECK-ZNVER4:       [[VEC_EPILOG_ITER_CHECK]]:
-; CHECK-ZNVER4-NEXT:    [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP0]], 8
+; CHECK-ZNVER4-NEXT:    [[MIN_EPILOG_ITERS_CHECK:%.*]] = icmp ult i64 [[TMP0]], 4
 ; CHECK-ZNVER4-NEXT:    br i1 [[MIN_EPILOG_ITERS_CHECK]], label %[[VEC_EPILOG_SCALAR_PH]], label %[[VEC_EPILOG_PH]], !prof [[PROF5:![0-9]+]]
 ; CHECK-ZNVER4:       [[VEC_EPILOG_PH]]:
 ; CHECK-ZNVER4-NEXT:    [[VEC_EPILOG_RESUME_VAL:%.*]] = phi i64 [ [[N_VEC]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
 ; CHECK-ZNVER4-NEXT:    [[BC_MERGE_RDX:%.*]] = phi i64 [ [[CONDITIONAL_STEP]], %[[VEC_EPILOG_ITER_CHECK]] ], [ 0, %[[VECTOR_MAIN_LOOP_ITER_CHECK]] ]
-; CHECK-ZNVER4-NEXT:    [[TMP8:%.*]] = and i64 [[N]], 7
+; CHECK-ZNVER4-NEXT:    [[TMP8:%.*]] = and i64 [[N]], 3
 ; CHECK-ZNVER4-NEXT:    [[N_VEC2:%.*]] = sub i64 [[N]], [[TMP8]]
-; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <8 x i32> poison, i32 [[C]], i64 0
-; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLAT4:%.*]] = shufflevector <8 x i32> [[BROADCAST_SPLATINSERT3]], <8 x i32> poison, <8 x i32> zeroinitializer
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLATINSERT3:%.*]] = insertelement <4 x i32> poison, i32 [[C]], i64 0
+; CHECK-ZNVER4-NEXT:    [[BROADCAST_SPLAT4:%.*]] = shufflevector <4 x i32> [[BROADCAST_SPLATINSERT3]], <4 x i32> poison, <4 x i32> zeroinitializer
 ; CHECK-ZNVER4-NEXT:    br label %[[VEC_EPILOG_VECTOR_BODY:.*]]
 ; CHECK-ZNVER4:       [[VEC_EPILOG_VECTOR_BODY]]:
 ; CHECK-ZNVER4-NEXT:    [[INDEX5:%.*]] = phi i64 [ [[VEC_EPILOG_RESUME_VAL]], %[[VEC_EPILOG_PH]] ], [ [[INDEX_NEXT9:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
 ; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_IV6:%.*]] = phi i64 [ [[BC_MERGE_RDX]], %[[VEC_EPILOG_PH]] ], [ [[CONDITIONAL_STEP8:%.*]], %[[VEC_EPILOG_VECTOR_BODY]] ]
 ; CHECK-ZNVER4-NEXT:    [[TMP9:%.*]] = getelementptr i32, ptr [[DST]], i64 [[INDEX5]]
-; CHECK-ZNVER4-NEXT:    [[WIDE_LOAD7:%.*]] = load <8 x i32>, ptr [[TMP9]], align 4
-; CHECK-ZNVER4-NEXT:    [[TMP10:%.*]] = icmp slt <8 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
+; CHECK-ZNVER4-NEXT:    [[WIDE_LOAD7:%.*]] = load <4 x i32>, ptr [[TMP9]], align 4
+; CHECK-ZNVER4-NEXT:    [[TMP10:%.*]] = icmp slt <4 x i32> [[WIDE_LOAD7]], [[BROADCAST_SPLAT4]]
 ; CHECK-ZNVER4-NEXT:    [[TMP11:%.*]] = getelementptr inbounds i32, ptr [[SRC]], i64 [[CONDITIONAL_IV6]]
-; CHECK-ZNVER4-NEXT:    [[TMP12:%.*]] = call <8 x i32> @llvm.masked.expandload.v8i32.p0(ptr align 4 [[TMP11]], <8 x i1> [[TMP10]], <8 x i32> poison)
-; CHECK-ZNVER4-NEXT:    call void @llvm.masked.store.v8i32.p0(<8 x i32> [[TMP12]], ptr align 4 [[TMP9]], <8 x i1> [[TMP10]])
-; CHECK-ZNVER4-NEXT:    [[TMP13:%.*]] = zext <8 x i1> [[TMP10]] to <8 x i64>
-; CHECK-ZNVER4-NEXT:    [[TMP14:%.*]] = call i64 @llvm.vector.reduce.add.v8i64(<8 x i64> [[TMP13]])
+; CHECK-ZNVER4-NEXT:    [[TMP12:%.*]] = call <4 x i32> @llvm.masked.expandload.v4i32.p0(ptr align 4 [[TMP11]], <4 x i1> [[TMP10]], <4 x i32> poison)
+; CHECK-ZNVER4-NEXT:    call void @llvm.masked.store.v4i32.p0(<4 x i32> [[TMP12]], ptr align 4 [[TMP9]], <4 x i1> [[TMP10]])
+; CHECK-ZNVER4-NEXT:    [[TMP13:%.*]] = zext <4 x i1> [[TMP10]] to <4 x i64>
+; CHECK-ZNVER4-NEXT:    [[TMP14:%.*]] = call i64 @llvm.vector.reduce.add.v4i64(<4 x i64> [[TMP13]])
 ; CHECK-ZNVER4-NEXT:    [[CONDITIONAL_STEP8]] = add i64 [[CONDITIONAL_IV6]], [[TMP14]]
-; CHECK-ZNVER4-NEXT:    [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 8
+; CHECK-ZNVER4-NEXT:    [[INDEX_NEXT9]] = add nuw i64 [[INDEX5]], 4
 ; CHECK-ZNVER4-NEXT:    [[TMP15:%.*]] = icmp eq i64 [[INDEX_NEXT9]], [[N_VEC2]]
 ; CHECK-ZNVER4-NEXT:    br i1 [[TMP15]], label %[[VEC_EPILOG_MIDDLE_BLOCK:.*]], label %[[VEC_EPILOG_VECTOR_BODY]], !llvm.loop [[LOOP6:![0-9]+]]
 ; CHECK-ZNVER4:       [[VEC_EPILOG_MIDDLE_BLOCK]]:



More information about the llvm-commits mailing list