[llvm] [LV] Add cross-part load overlap heuristic for interleaving (PR #214500)

Sergey Shcherbinin via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 12 05:56:14 PDT 2026


https://github.com/SergeyShch01 updated https://github.com/llvm/llvm-project/pull/214500

>From 932e6af06f576da563c047b99d2a0f04555ac462 Mon Sep 17 00:00:00 2001
From: Sergey Shcherbinin <sscherbinin at nvidia.com>
Date: Thu, 6 Aug 2026 18:25:49 +0400
Subject: [PATCH 1/3] [LV] Add cross-part load overlap heuristic for
 interleaving

The existing interleave heuristics can miss redundancies exposed only
between logical unrolled parts and may select IC=1 when IC=2 would enable
downstream CSE.

Add a disabled-by-default heuristic for eligible fixed-width,
single-block VPlans. The analysis models two logical parts, matches
simple widened loads using exact part-aware SCEV addresses, and rejects
a pair if any write occurs between its loads. It uses retained
per-recipe costs to raise a heuristic IC=1 to IC=2 when the predicted
saving meets the configured opportunity and percentage thresholds.

This analysis is prediction-only: it neither mutates VPlan nor removes
loads. It implements the minimal first step of the upstream sequence
described in the RFC.

RFC: https://discourse.llvm.org/t/rfc-using-cross-part-cse-to-guide-loop-interleaving/91438
---
 llvm/lib/Transforms/Vectorize/CMakeLists.txt  |   1 +
 .../Vectorize/LoopVectorizationPlanner.h      |  43 ++
 .../Transforms/Vectorize/LoopVectorize.cpp    | 140 ++++-
 .../Vectorize/VPlanCrossPartCSE.cpp           | 288 +++++++++++
 .../Transforms/Vectorize/VPlanCrossPartCSE.h  |  61 +++
 llvm/lib/Transforms/Vectorize/VPlanHelpers.h  |  16 +
 .../lib/Transforms/Vectorize/VPlanRecipes.cpp |   2 +
 .../AArch64/cross-part-load-cse-debug.ll      | 146 ++++++
 .../AArch64/cross-part-load-cse.ll            | 479 ++++++++++++++++++
 .../llvm/lib/Transforms/Vectorize/BUILD.gn    |   1 +
 10 files changed, 1174 insertions(+), 3 deletions(-)
 create mode 100644 llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.cpp
 create mode 100644 llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.h
 create mode 100644 llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
 create mode 100644 llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse.ll

diff --git a/llvm/lib/Transforms/Vectorize/CMakeLists.txt b/llvm/lib/Transforms/Vectorize/CMakeLists.txt
index 9073211280886..f459cbc618557 100644
--- a/llvm/lib/Transforms/Vectorize/CMakeLists.txt
+++ b/llvm/lib/Transforms/Vectorize/CMakeLists.txt
@@ -35,6 +35,7 @@ add_llvm_component_library(LLVMVectorize
   VPlan.cpp
   VPlanAnalysis.cpp
   VPlanConstruction.cpp
+  VPlanCrossPartCSE.cpp
   VPlanDominatorTree.cpp
   VPlanEVLTailFolding.cpp
   VPlanLowering.cpp
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
index 0d43d7a8f0cca..f17de81888495 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorizationPlanner.h
@@ -25,6 +25,8 @@
 #define LLVM_TRANSFORMS_VECTORIZE_LOOPVECTORIZATIONPLANNER_H
 
 #include "VPlan.h"
+#include "VPlanHelpers.h"
+#include "llvm/ADT/DenseMap.h"
 #include "llvm/ADT/SmallSet.h"
 #include "llvm/Analysis/TargetTransformInfo.h"
 #include "llvm/Support/InstructionCost.h"
@@ -894,6 +896,17 @@ class LoopVectorizationPlanner {
   /// A builder used to construct the current plan.
   VPBuilder Builder;
 
+  /// Cost-traversal results retained for eligible fixed-VF Plan/VF candidates
+  /// until IC selection decides whether cross-part CSE can affect interleaving.
+  ///
+  /// Each map contains the recipes visited while costing its exact candidate,
+  /// including zero entries for recipes whose direct cost was skipped. Recipes
+  /// outside that traversal are absent; cross-part analysis queries only its
+  /// supported widened-load recipes.
+  mutable DenseMap<std::pair<const VPlan *, ElementCount>,
+                   std::unique_ptr<VPRecipeCostMap>>
+      CrossPartCSERecipeCosts;
+
   /// Computes the cost of \p Plan for vectorization factor \p VF.
   ///
   /// The current implementation requires access to the
@@ -904,6 +917,36 @@ class LoopVectorizationPlanner {
   /// been retired.
   InstructionCost cost(VPlan &Plan, ElementCount VF, VPRegisterUsage *RU) const;
 
+  /// Return whether cross-part CSE may participate in IC selection for this
+  /// loop.
+  ///
+  /// This is the loop-level policy gate: it checks the feature flag and
+  /// conditions independent of a particular VPlan/VF candidate, including
+  /// whether the loop is innermost, has no explicit interleave count, and does
+  /// not require partial-alias masking. A true result does not mean that a
+  /// supported or profitable cross-part opportunity has been found.
+  bool shouldUseCrossPartCSE() const;
+
+  /// Return whether to retain per-recipe costs for cross-part CSE analysis of
+  /// this exact \p Plan and \p VF candidate.
+  ///
+  /// Unlike shouldUseCrossPartCSE(), this candidate-level gate additionally
+  /// requires a fixed vector VF and a plan capable of the required UF=2. It
+  /// controls transient cost-map population during candidate costing; the
+  /// retained costs are consumed only if this Plan/VF pair reaches IC
+  /// selection.
+  bool shouldCollectCrossPartCSECosts(VPlan &Plan, ElementCount VF) const;
+
+  /// Return whether cross-part CSE opportunities in \p Plan for \p VF justify
+  /// requesting the interleave count required by the analysis.
+  ///
+  /// This supplements the ordinary interleave heuristics with VPlan-specific
+  /// analysis across logical parts. A false result leaves the ordinary
+  /// interleave-count decision unchanged.
+  bool shouldInterleaveForCrossPartCSE(VPlan &Plan, ElementCount VF,
+                                       InstructionCost LoopCost,
+                                       unsigned MaxIC);
+
   /// Precompute costs for certain instructions using the legacy cost model. The
   /// function is used to bring up the VPlan-based cost model to initially avoid
   /// taking different decisions due to inaccuracies in the legacy cost model.
diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 04abdac5ee680..9301e0a882ce6 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -59,6 +59,7 @@
 #include "VPlan.h"
 #include "VPlanAnalysis.h"
 #include "VPlanCFG.h"
+#include "VPlanCrossPartCSE.h"
 #include "VPlanHelpers.h"
 #include "VPlanPatternMatch.h"
 #include "VPlanTransforms.h"
@@ -70,6 +71,7 @@
 #include "llvm/ADT/Hashing.h"
 #include "llvm/ADT/MapVector.h"
 #include "llvm/ADT/STLExtras.h"
+#include "llvm/ADT/ScopeExit.h"
 #include "llvm/ADT/SmallPtrSet.h"
 #include "llvm/ADT/SmallVector.h"
 #include "llvm/ADT/Statistic.h"
@@ -300,6 +302,23 @@ static cl::opt<bool> EnableLoadStoreRuntimeInterleave(
     cl::desc(
         "Enable runtime interleaving until load/store ports are saturated"));
 
+/// Enable cross-part load-overlap analysis during IC selection.
+static cl::opt<bool> EnableInterleaveCSE(
+    "enable-interleave-cse", cl::init(false), cl::Hidden,
+    cl::desc("Raise heuristic IC=1 to IC=2 when exact cross-part load overlap "
+             "predicts a downstream saving"));
+
+/// Minimum percentage of the modeled UF=2 body predicted to be saved.
+static cl::opt<unsigned> InterleaveCSEMinSavingPct(
+    "interleave-cse-min-pct", cl::init(5), cl::Hidden,
+    cl::desc("Minimum predicted downstream load saving as a percentage of the "
+             "modeled UF=2 vector loop body"));
+
+/// Minimum number of exact load overlaps required to request IC=2.
+static cl::opt<unsigned> InterleaveCSEMinOpportunities(
+    "interleave-cse-min-ops", cl::init(2), cl::Hidden,
+    cl::desc("Minimum number of exact cross-part load-overlap opportunities"));
+
 // TODO: Move size-based thresholds out of legality checking, make cost based
 // decisions instead of hard thresholds.
 static cl::opt<unsigned> VectorizeSCEVCheckThreshold(
@@ -3630,9 +3649,53 @@ std::unique_ptr<VPlan> LoopVectorizationPlanner::selectBestEpiloguePlan(
   return Clone;
 }
 
+bool LoopVectorizationPlanner::shouldUseCrossPartCSE() const {
+  // Keep the disabled path free of policy queries and cost-map allocation.
+  if (!EnableInterleaveCSE)
+    return false;
+  return OrigLoop->isInnermost() && Hints.getInterleave() == 0 &&
+         !CM.maskPartialAliasing();
+}
+
+bool LoopVectorizationPlanner::shouldCollectCrossPartCSECosts(
+    VPlan &Plan, ElementCount VF) const {
+  return shouldUseCrossPartCSE() && VF.isVector() && !VF.isScalable() &&
+         Plan.hasUF(CrossPartCSERequiredInterleaveCount);
+}
+
+bool LoopVectorizationPlanner::shouldInterleaveForCrossPartCSE(
+    VPlan &Plan, ElementCount VF, InstructionCost LoopCost, unsigned MaxIC) {
+  // Cross-part CSE only augments ordinary heuristic selection. The policy gate
+  // excludes explicit user counts and loops requiring partial-alias masking,
+  // while MaxIC preserves target, trip-count, and register-pressure limits.
+  if (!shouldUseCrossPartCSE() || MaxIC < CrossPartCSERequiredInterleaveCount ||
+      !VF.isVector() || VF.isScalable())
+    return false;
+
+  auto CostIt = CrossPartCSERecipeCosts.find({&Plan, VF});
+  if (CostIt == CrossPartCSERecipeCosts.end())
+    return false;
+  if (!CostIt->second)
+    return false;
+
+  CrossPartCSEOptions Options;
+  Options.MinSavingPct = InterleaveCSEMinSavingPct;
+  Options.MinOpportunities = InterleaveCSEMinOpportunities;
+  if (!isCrossPartCSEProfitable(Plan, VF, LoopCost, OrigLoop, PSE,
+                                *CostIt->second, Options))
+    return false;
+
+  LLVM_DEBUG(dbgs() << "LV: Exact cross-part load overlap predicts a "
+                       "downstream saving; raising IC to 2.\n");
+  return true;
+}
+
 unsigned
 LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
                                                 InstructionCost LoopCost) {
+  // Recipe keys borrow VPlan storage and are useful only during this selection.
+  scope_exit ClearCrossPartCSECosts([&] { CrossPartCSERecipeCosts.clear(); });
+
   // -- The interleave heuristics --
   // We interleave the loop in order to expose ILP and reduce the loop overhead.
   // There are many micro-architectural considerations that we can't predict
@@ -3967,6 +4030,9 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
       return std::max(IC / 2, SmallIC);
     }
 
+    if (SmallIC == 1 && shouldInterleaveForCrossPartCSE(Plan, VF, LoopCost, IC))
+      return CrossPartCSERequiredInterleaveCount;
+
     LLVM_DEBUG(dbgs() << "LV: Interleaving to reduce branch cost.\n");
     return SmallIC;
   }
@@ -3978,6 +4044,9 @@ LoopVectorizationPlanner::selectInterleaveCount(VPlan &Plan, ElementCount VF,
     return IC;
   }
 
+  if (shouldInterleaveForCrossPartCSE(Plan, VF, LoopCost, IC))
+    return CrossPartCSERequiredInterleaveCount;
+
   LLVM_DEBUG(dbgs() << "LV: Not Interleaving.\n");
   return 1;
 }
@@ -5357,8 +5426,15 @@ void LoopVectorizationPlanner::plan(ElementCount UserVF, unsigned UserIC) {
       if (!VPlans.empty() && VPlans.front()->getSingleVF() == UserVF) {
         // For scalar VF, skip VPlan cost check as VPlan cost is designed for
         // vector VFs only.
-        if (UserVF.isScalar() ||
-            cost(*VPlans.front(), UserVF, /*RU=*/nullptr).isValid()) {
+        InstructionCost UserVFCost = 0;
+        if (!UserVF.isScalar()) {
+          UserVFCost = cost(*VPlans.front(), UserVF, /*RU=*/nullptr);
+          // Validation must not retain borrowed VPlan recipe keys. A valid
+          // forced plan is costed lazily during IC selection, while an invalid
+          // plan may be destroyed immediately below.
+          CrossPartCSERecipeCosts.clear();
+        }
+        if (UserVF.isScalar() || UserVFCost.isValid()) {
           LLVM_DEBUG(dbgs() << "LV: Using user VF " << UserVF << ".\n");
           LLVM_DEBUG(printPlans(dbgs()));
           return;
@@ -5417,6 +5493,12 @@ bool VPCostContext::skipCostComputation(Instruction *UI, bool IsVector) const {
          SkipCostComputation.contains(UI);
 }
 
+void VPCostContext::recordRecipeCost(const VPRecipeBase *R,
+                                     InstructionCost Cost) {
+  assert(RecipeCosts && "recipe costs must be enabled by the caller");
+  (*RecipeCosts)[R] = Cost;
+}
+
 void VPCostContext::invalidateWideningDecision(Instruction *I,
                                                ElementCount VF) {
   CM.setWideningDecision(I, VF,
@@ -5580,6 +5662,15 @@ InstructionCost LoopVectorizationPlanner::cost(VPlan &Plan, ElementCount VF,
                                                VPRegisterUsage *RU) const {
   VPCostContext CostCtx(*TLI, Plan, *CM, Config,
                         /*ReusePrintingSlotTracker=*/true);
+  if (shouldCollectCrossPartCSECosts(Plan, VF)) {
+    std::unique_ptr<VPRecipeCostMap> &Costs =
+        CrossPartCSERecipeCosts[{&Plan, VF}];
+    if (!Costs)
+      Costs = std::make_unique<VPRecipeCostMap>();
+    Costs->clear();
+    // The pointee remains stable if the owning DenseMap rehashes.
+    CostCtx.RecipeCosts = Costs.get();
+  }
   InstructionCost Cost = precomputeCosts(Plan, VF, CostCtx);
 
   // Now compute and add the VPlan-based cost.
@@ -5615,6 +5706,11 @@ InstructionCost LoopVectorizationPlanner::cost(VPlan &Plan, ElementCount VF,
 
 std::pair<VectorizationFactor, VPlan *>
 LoopVectorizationPlanner::computeBestVF() {
+  // Discard maps produced while validating a user VF; the candidate loop below
+  // retains costs only for the actual winning Plan/VF pair.
+  CrossPartCSERecipeCosts.clear();
+  bool TrackRecipeCosts = shouldUseCrossPartCSE();
+
   if (VPlans.empty())
     return {VectorizationFactor::Disabled(), nullptr};
   // If there is a single VPlan with a single VF, return it directly.
@@ -5668,6 +5764,10 @@ LoopVectorizationPlanner::computeBestVF() {
   }
 
   VPlan *PlanForBestVF = &FirstPlan;
+  // Identify the retained recipe-cost map for the current winning Plan/VF
+  // pair. The key is absent while the winner is scalar or otherwise ineligible
+  // for cross-part analysis.
+  std::optional<std::pair<const VPlan *, ElementCount>> BestRecipeCostKey;
 
   for (auto &P : VPlans) {
     ArrayRef<ElementCount> VFs(P->vectorFactors().begin(),
@@ -5704,11 +5804,36 @@ LoopVectorizationPlanner::computeBestVF() {
           cost(*P, VF, ConsiderRegPressure ? &RUs[I] : nullptr);
       VectorizationFactor CurrentFactor(VF, Cost, ScalarCost);
 
-      if (isMoreProfitable(CurrentFactor, BestFactor, P->hasScalarTail())) {
+      bool IsMoreProfitable =
+          isMoreProfitable(CurrentFactor, BestFactor, P->hasScalarTail());
+      if (IsMoreProfitable) {
         BestFactor = CurrentFactor;
         PlanForBestVF = P.get();
       }
 
+      // Costing may retain a recipe-cost map for each eligible candidate. Keep
+      // only the map associated with the best candidate seen so far, because
+      // IC selection consumes costs exclusively for the selected Plan/VF pair.
+      if (TrackRecipeCosts) {
+        std::pair<const VPlan *, ElementCount> CurrentKey = {P.get(), VF};
+        bool HasCurrentRecipeCosts =
+            CrossPartCSERecipeCosts.contains(CurrentKey);
+        if (IsMoreProfitable) {
+          // The current candidate replaces the previous winner, so its map
+          // also replaces any recipe costs retained for that winner.
+          if (BestRecipeCostKey)
+            CrossPartCSERecipeCosts.erase(*BestRecipeCostKey);
+          if (HasCurrentRecipeCosts)
+            BestRecipeCostKey = CurrentKey;
+          else
+            BestRecipeCostKey.reset();
+        } else if (HasCurrentRecipeCosts) {
+          // Discard costs for a losing candidate immediately to prevent recipe
+          // pointers from outliving a plan that is not selected.
+          CrossPartCSERecipeCosts.erase(CurrentKey);
+        }
+      }
+
       // If profitable add it to ProfitableVF list.
       if (isMoreProfitable(CurrentFactor, ScalarFactor, P->hasScalarTail()))
         ProfitableVFs.push_back(CurrentFactor);
@@ -5716,6 +5841,15 @@ LoopVectorizationPlanner::computeBestVF() {
   }
 
   VPlan &BestPlan = *PlanForBestVF;
+  assert((!TrackRecipeCosts || CrossPartCSERecipeCosts.size() <= 1) &&
+         "only the selected VF recipe costs may remain");
+  assert((!shouldCollectCrossPartCSECosts(BestPlan, BestFactor.Width) ||
+          (BestRecipeCostKey &&
+           *BestRecipeCostKey ==
+               std::make_pair(static_cast<const VPlan *>(PlanForBestVF),
+                              BestFactor.Width) &&
+           CrossPartCSERecipeCosts.contains(*BestRecipeCostKey))) &&
+         "selected fixed VF must retain its recipe costs");
 
   assert((BestFactor.Width.isScalar() || BestFactor.ScalarCost > 0) &&
          "when vectorizing, the scalar cost must be computed.");
diff --git a/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.cpp b/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.cpp
new file mode 100644
index 0000000000000..75d826d1d3144
--- /dev/null
+++ b/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.cpp
@@ -0,0 +1,288 @@
+//===- VPlanCrossPartCSE.cpp - Cross-part CSE for VPlan -------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// This file implements exact load-overlap profitability analysis across two
+// logical VPlan parts.
+//
+//===----------------------------------------------------------------------===//
+
+#include "VPlanCrossPartCSE.h"
+#include "VPlan.h"
+#include "VPlanPatternMatch.h"
+#include "VPlanUtils.h"
+#include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/Hashing.h"
+#include "llvm/Analysis/LoopInfo.h"
+#include "llvm/Analysis/ScalarEvolution.h"
+#include "llvm/IR/Instructions.h"
+#include "llvm/Support/Debug.h"
+#include "llvm/Support/raw_ostream.h"
+#include <cstdint>
+#include <optional>
+
+using namespace llvm;
+
+#define DEBUG_TYPE "loop-vectorize"
+
+namespace {
+
+/// A simple widened load supported by the prediction-only analysis.
+struct CrossPartSupportedLoad {
+  /// VPlan recipe whose logical UF=2 instances are modeled.
+  VPWidenLoadRecipe *Recipe;
+  /// Underlying scalar load used for type, SCEV, and poison identity.
+  LoadInst *Load;
+};
+
+/// Return an unmasked, non-EVL, consecutive simple widened load.
+static std::optional<CrossPartSupportedLoad>
+getCrossPartSupportedLoad(VPRecipeBase &R) {
+  auto *Widen = dyn_cast<VPWidenLoadRecipe>(&R);
+  if (!Widen || Widen->isMasked() || !Widen->isConsecutive())
+    return std::nullopt;
+
+  auto *Load = dyn_cast<LoadInst>(&Widen->getIngredient());
+  if (!Load || !Load->isSimple())
+    return std::nullopt;
+  return CrossPartSupportedLoad{Widen, Load};
+}
+
+/// Build addresses only for provenance whose physical UF mapping is explicit.
+class CrossPartAddressBuilder {
+  /// Return the fixed width after validating the address model's precondition.
+  static unsigned getFixedVF(ElementCount VF) {
+    assert(!VF.isScalable() && "cross-part analysis requires a fixed VF");
+    return VF.getFixedValue();
+  }
+
+  /// Predicated SCEV state carrying vectorization assumptions.
+  PredicatedScalarEvolution &PSE;
+  /// ScalarEvolution used for canonical exact identities.
+  ScalarEvolution &SE;
+  /// Original loop used to interpret loop-varying VPlan values.
+  const Loop *OrigLoop;
+  /// Exact fixed vector width used for Part * VF.
+  const unsigned FixedVF;
+  /// Base SCEVs cached by VPlan value for reuse across loads and parts.
+  DenseMap<const VPValue *, const SCEV *> BaseSCEVs;
+
+  /// Return the SCEV represented by \p V, caching it after first construction.
+  const SCEV *getBaseSCEV(const VPValue *V) {
+    auto It = BaseSCEVs.find(V);
+    if (It != BaseSCEVs.end())
+      return It->second;
+
+    const SCEV *S = vputils::getSCEVExprForVPValue(V, PSE, OrigLoop);
+    BaseSCEVs.try_emplace(V, S);
+    return S;
+  }
+
+public:
+  /// Bind the fixed VF, original loop, and predicated SCEV state.
+  CrossPartAddressBuilder(ElementCount VF, PredicatedScalarEvolution &PSE,
+                          const Loop *OrigLoop)
+      : PSE(PSE), SE(*PSE.getSE()), OrigLoop(OrigLoop),
+        FixedVF(getFixedVF(VF)) {}
+
+  /// Return the exact address used by \p Load in logical part \p Part.
+  const SCEV *getAddress(const CrossPartSupportedLoad &Load, unsigned Part) {
+    assert(Part < CrossPartCSERequiredInterleaveCount &&
+           "logical part must be zero or one");
+    VPValue *Addr = Load.Recipe->getAddr();
+
+    // VPVectorPointerRecipe is explicitly cloned with a Part * VF offset by
+    // VPlanUnroll. Reproduce only that recipe-specific physical rewrite.
+    auto *VectorPtr = dyn_cast<VPVectorPointerRecipe>(Addr);
+    if (!VectorPtr)
+      return SE.getCouldNotCompute();
+
+    // TODO: Generalize the part offset to VFxPart * Stride when adding complete
+    // supported load shapes. Until then, accept only a literal unit stride so
+    // the simplified Part * VF offset remains exact.
+    using namespace VPlanPatternMatch;
+    if (!match(VectorPtr->getStride(), m_One()))
+      return SE.getCouldNotCompute();
+
+    const SCEV *Base = getBaseSCEV(VectorPtr->getOperand(0));
+    if (isa<SCEVCouldNotCompute>(Base))
+      return SE.getCouldNotCompute();
+    if (Part == 0)
+      return Base;
+
+    Type *IndexTy = SE.getDataLayout().getIndexType(VectorPtr->getScalarType());
+    const SCEV *Offset = SE.getConstant(IndexTy, uint64_t(Part) * FixedVF);
+    // Keep the synthetic address conservative: ScalarEvolution imports GEP
+    // nowrap facts only after accounting for their poison semantics. AddExpr
+    // uniquing still recognizes equal operands without those facts.
+    return SE.getGEPExpr(Base, {Offset}, VectorPtr->getSourceElementType());
+  }
+};
+
+/// Key for exact load equality after annotated loads have been excluded.
+struct CrossPartLoadKey {
+  /// Canonical SCEV address for this logical load instance.
+  const SCEV *Address;
+  /// Loaded scalar type required for value compatibility.
+  Type *ValueType;
+};
+
+/// DenseMap policy for exact canonical load keys.
+struct CrossPartLoadKeyInfo {
+  /// Hash every property required by exact load equality.
+  static unsigned getHashValue(const CrossPartLoadKey &Key) {
+    return hash_combine(Key.Address, Key.ValueType);
+  }
+
+  /// Compare every property required by exact load equality.
+  static bool isEqual(const CrossPartLoadKey &A, const CrossPartLoadKey &B) {
+    return A.Address == B.Address && A.ValueType == B.ValueType;
+  }
+};
+
+/// Return whether \p R may write memory during VPlan execution.
+static bool isCrossPartWrite(const VPRecipeBase &R) {
+  // VPVectorEndPointerRecipe is pure but inherits the conservative memory
+  // default. This local exception prevents its address computation from being
+  // mistaken for a write without changing global recipe memory behavior.
+  switch (R.getVPRecipeID()) {
+  case VPRecipeBase::VPVectorEndPointerSC:
+    return false;
+  default:
+    return R.mayWriteToMemory();
+  }
+}
+
+/// Return whether \p Plan keeps the canonical IV increment in the symbolic
+/// VF * UF form required to model consecutive logical parts.
+static bool hasCanonicalIVIncrementForCrossPartCSE(VPlan &Plan) {
+  return !Plan.getVFxUF().isMaterialized() &&
+         vputils::findCanonicalIVIncrement(Plan);
+}
+
+} // namespace
+
+bool llvm::isCrossPartCSEProfitable(VPlan &Plan, ElementCount VF,
+                                    InstructionCost LoopCost,
+                                    const Loop *OrigLoop,
+                                    PredicatedScalarEvolution &PSE,
+                                    const VPRecipeCostMap &RecipeCosts,
+                                    const CrossPartCSEOptions &Options) {
+  // Reject a partially initialized policy before its sentinel values can
+  // participate in saturating cost arithmetic.
+  if (Options.MinSavingPct == CrossPartCSEOptions::Unspecified ||
+      Options.MinOpportunities == CrossPartCSEOptions::Unspecified)
+    return false;
+
+  VPRegionBlock *LoopRegion = Plan.getVectorLoopRegion();
+  if (!LoopRegion)
+    return false;
+
+  // Fail closed for every shape outside the exact fixed-width, single-block
+  // UF=2 model.
+  // TODO: Expand coverage by accepting additional plan shapes once their
+  // cross-part semantics can be modeled exactly.
+  if (VF.isScalable() || !VF.isVector() || !LoopCost.isValid() ||
+      LoopCost <= 0 || !OrigLoop->isInnermost() ||
+      OrigLoop->getNumBlocks() != 1 ||
+      LoopRegion->getEntryBasicBlock() != LoopRegion->getExitingBasicBlock() ||
+      !Plan.hasUF(CrossPartCSERequiredInterleaveCount) || Plan.isUnrolled() ||
+      !hasCanonicalIVIncrementForCrossPartCSE(Plan))
+    return false;
+
+  using AvailableLoadMap =
+      DenseMap<CrossPartLoadKey, unsigned, CrossPartLoadKeyInfo>;
+  AvailableLoadMap AvailableLoadParts;
+  CrossPartAddressBuilder Addresses(VF, PSE, OrigLoop);
+  unsigned NumOpportunities = 0;
+  InstructionCost SavedCost = 0;
+
+  // Match VPlanUnroll's recipe-major UF=2 order. Clearing on every write
+  // enforces a strict no-write interval without alias disambiguation.
+  for (VPRecipeBase &R : *LoopRegion->getEntryBasicBlock()) {
+    if (isCrossPartWrite(R)) {
+      AvailableLoadParts.clear();
+      continue;
+    }
+
+    std::optional<CrossPartSupportedLoad> Load = getCrossPartSupportedLoad(R);
+    if (!Load)
+      continue;
+
+    // Conservatively reject annotations whose poison behavior requires the
+    // downstream realization to combine metadata across the load pair.
+    // TODO: Accept compatible annotated pairs once realization supports the
+    // required CSE metadata merge.
+    if (Load->Load->hasPoisonGeneratingAnnotations())
+      continue;
+
+    for (unsigned Part = 0; Part != CrossPartCSERequiredInterleaveCount;
+         ++Part) {
+      const SCEV *Address = Addresses.getAddress(*Load, Part);
+      if (isa<SCEVCouldNotCompute>(Address))
+        continue;
+
+      CrossPartLoadKey Key = {Address, Load->Load->getType()};
+      // Only reuse between different logical parts can justify raising IC from
+      // 1 to 2. A duplicate already seen in the same part also exists at IC=1
+      // and therefore provides no interleaving-specific saving.
+      unsigned PartBit = 1U << Part;
+      unsigned &AvailableParts = AvailableLoadParts[Key];
+      if (AvailableParts & PartBit)
+        continue;
+
+      bool HasOppositePart = (AvailableParts & ~PartBit) != 0;
+      AvailableParts |= PartBit;
+      if (!HasOppositePart) {
+        // Record the first occurrence in this part without assigning
+        // cross-part credit.
+        continue;
+      }
+
+      auto CostIt = RecipeCosts.find(Load->Recipe);
+      // RecipeCosts records direct VPlan-owned costs only. A zero entry may
+      // mean that the underlying instruction was charged or ignored elsewhere,
+      // so it cannot establish independently removable work.
+      // TODO: Before crediting non-load congruence, model legacy-owned costs
+      // with explicit per-occurrence ownership and deduplication.
+      if (CostIt == RecipeCosts.end() || !CostIt->second.isValid() ||
+          CostIt->second <= 0)
+        continue;
+      SavedCost += CostIt->second;
+      ++NumOpportunities;
+    }
+  }
+
+  bool MeetsOpportunityThreshold =
+      NumOpportunities >= Options.MinOpportunities && SavedCost > 0;
+  using CostType = InstructionCost::CostType;
+  bool Select = false;
+  if (MeetsOpportunityThreshold) {
+    // Use InstructionCost arithmetic to preserve fractional cost units.
+    InstructionCost ScaledSavedCost = SavedCost * CostType(100);
+    InstructionCost RequiredCost =
+        LoopCost * CostType(CrossPartCSERequiredInterleaveCount);
+    RequiredCost *= CostType(Options.MinSavingPct);
+    Select = ScaledSavedCost >= RequiredCost;
+  }
+
+  LLVM_DEBUG({
+    CostType SavingPct = 0;
+    if (SavedCost.isValid() && SavedCost > 0 && LoopCost.isValid() &&
+        LoopCost > 0)
+      SavingPct = ((SavedCost * CostType(100)) /
+                   (LoopCost * CostType(CrossPartCSERequiredInterleaveCount)))
+                      .getValue();
+    dbgs() << "LV: Cross-part load overlap estimate: ops=" << NumOpportunities
+           << ", required-ops=" << Options.MinOpportunities
+           << ", predicted-saved-cost=" << SavedCost
+           << ", loop-cost=" << LoopCost << ", saving=" << SavingPct
+           << "%, required=" << Options.MinSavingPct << "%; "
+           << (Select ? "selecting IC=2" : "skipping") << ".\n";
+  });
+  return Select;
+}
diff --git a/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.h b/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.h
new file mode 100644
index 0000000000000..d3a7dcae27e94
--- /dev/null
+++ b/llvm/lib/Transforms/Vectorize/VPlanCrossPartCSE.h
@@ -0,0 +1,61 @@
+//===- VPlanCrossPartCSE.h - Cross-part CSE for VPlan -----------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// This file declares prediction-only profitability analysis for exact load
+// overlap across two modeled logical VPlan parts. It does not transform VPlan.
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_TRANSFORMS_VECTORIZE_VPLANCROSSPARTCSE_H
+#define LLVM_TRANSFORMS_VECTORIZE_VPLANCROSSPARTCSE_H
+
+#include "VPlanHelpers.h"
+#include "llvm/Support/InstructionCost.h"
+#include "llvm/Support/TypeSize.h"
+#include <limits>
+
+namespace llvm {
+
+class Loop;
+class PredicatedScalarEvolution;
+class VPlan;
+
+/// The interleave count and logical unroll factor modeled by the analysis.
+constexpr unsigned CrossPartCSERequiredInterleaveCount = 2;
+
+/// Profitability criteria supplied by the caller.
+///
+/// Fail-closed defaults require callers to provide both criteria explicitly.
+struct CrossPartCSEOptions {
+  /// Sentinel used until the caller supplies an explicit policy value.
+  static constexpr unsigned Unspecified = std::numeric_limits<unsigned>::max();
+
+  /// Minimum saving; the default rejects analysis until policy supplies it.
+  unsigned MinSavingPct = Unspecified;
+  /// Minimum opportunities; the default likewise keeps the API fail-closed.
+  unsigned MinOpportunities = Unspecified;
+};
+
+/// Return whether predicted exact load overlap is profitable for \p Plan and
+/// \p VF under \p Options.
+///
+/// The caller must establish that interleaving \p Plan is legal before using
+/// this profitability result to raise its interleave count. The analysis reads
+/// \p Plan but takes a non-const reference because the VPlan query APIs it uses
+/// are not const-qualified.
+///
+/// \p RecipeCosts is borrowed for this call and is neither copied nor retained.
+bool isCrossPartCSEProfitable(VPlan &Plan, ElementCount VF,
+                              InstructionCost LoopCost, const Loop *OrigLoop,
+                              PredicatedScalarEvolution &PSE,
+                              const VPRecipeCostMap &RecipeCosts,
+                              const CrossPartCSEOptions &Options);
+
+} // namespace llvm
+
+#endif // LLVM_TRANSFORMS_VECTORIZE_VPLANCROSSPARTCSE_H
diff --git a/llvm/lib/Transforms/Vectorize/VPlanHelpers.h b/llvm/lib/Transforms/Vectorize/VPlanHelpers.h
index d31c5acdf69a9..51eec5beffb56 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanHelpers.h
+++ b/llvm/lib/Transforms/Vectorize/VPlanHelpers.h
@@ -325,6 +325,13 @@ struct VPTransformState {
   VPDominatorTree VPDT;
 };
 
+/// Effective costs recorded for recipes visited by one VPlan cost traversal.
+///
+/// A visited recipe whose direct computation is skipped, for example because
+/// the legacy model precomputed its cost, has a zero entry. Recipes not visited
+/// by that traversal have no entry.
+using VPRecipeCostMap = DenseMap<const VPRecipeBase *, InstructionCost>;
+
 /// Struct to hold various analysis needed for cost computations.
 struct VPCostContext {
   const TargetTransformInfo &TTI;
@@ -337,6 +344,12 @@ struct VPCostContext {
   PredicatedScalarEvolution &PSE;
   const Loop *L;
 
+  /// Optional destination for the recipe costs produced by this traversal.
+  ///
+  /// It is non-null only for an eligible fixed-VF cost run while cross-part CSE
+  /// for interleaving is enabled.
+  VPRecipeCostMap *RecipeCosts = nullptr;
+
   /// Number of predicated stores in the VPlan, computed on demand.
   std::optional<unsigned> NumPredStores;
 
@@ -352,6 +365,9 @@ struct VPCostContext {
   /// has already been pre-computed.
   bool skipCostComputation(Instruction *UI, bool IsVector) const;
 
+  /// Record one recipe's directly computed effective cost.
+  void recordRecipeCost(const VPRecipeBase *R, InstructionCost Cost);
+
   /// Mark the widening decision for \p I at \p VF as invalidated since a VPlan
   /// transform replaced the original recipe.
   void invalidateWideningDecision(Instruction *I, ElementCount VF);
diff --git a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
index 02c570d0437bb..1f0ed24760adc 100644
--- a/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
+++ b/llvm/lib/Transforms/Vectorize/VPlanRecipes.cpp
@@ -347,6 +347,8 @@ InstructionCost VPRecipeBase::cost(ElementCount VF, VPCostContext &Ctx) {
       dump();
     }
   });
+  if (Ctx.RecipeCosts)
+    Ctx.recordRecipeCost(this, RecipeCost);
   return RecipeCost;
 }
 
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
new file mode 100644
index 0000000000000..1c239f95ba2d2
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
@@ -0,0 +1,146 @@
+; REQUIRES: asserts
+; RUN: split-file %s %t
+;
+; When cross-part analysis selects IC=2 after the ordinary heuristics decline
+; interleaving, it must not emit a contradictory non-interleaving diagnostic.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 \
+; RUN:     -debug-only=loop-vectorize -disable-output %t/success.ll 2>&1 \
+; RUN:     | FileCheck %t/success.ll --check-prefix=SUCCESS
+;
+; A fixed-VF, wide-lane-mask tail-folded plan reaches IC selection, but its
+; masked widened loads remain outside the exact unmasked-load model.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -force-target-supports-masked-memory-ops \
+; RUN:     -force-tail-folding-style=data-and-control \
+; RUN:     -tail-folding-policy=must-fold-tail -enable-wide-lane-mask \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -debug-only=loop-vectorize \
+; RUN:     -disable-output %t/success.ll 2>&1 \
+; RUN:     | FileCheck %t/success.ll --check-prefix=MASKED
+;
+; When the ordinary branch-cost heuristic recommends IC=1, a successful
+; cross-part selection must return before emitting its baseline diagnostic.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 \
+; RUN:     -force-target-instruction-cost=1 -small-loop-cost=12 \
+; RUN:     -enable-loadstore-runtime-interleave=false \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 \
+; RUN:     -debug-only=loop-vectorize -disable-output %t/success.ll 2>&1 \
+; RUN:     | FileCheck %t/success.ll --check-prefix=SUCCESS-SMALL
+;
+; The same-part duplicate after a genuine cross-part match must report exactly
+; one opportunity and fail the requested minimum of two.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=2 \
+; RUN:     -debug-only=loop-vectorize \
+; RUN:     -disable-output %t/duplicate.ll 2>&1 \
+; RUN:     | FileCheck %t/duplicate.ll
+;
+; With cross-part analysis disabled, the ordinary branch-cost diagnostic
+; remains unchanged.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 \
+; RUN:     -force-target-instruction-cost=1 -small-loop-cost=12 \
+; RUN:     -enable-loadstore-runtime-interleave=false \
+; RUN:     -debug-only=loop-vectorize -disable-output %t/success.ll 2>&1 \
+; RUN:     | FileCheck %t/success.ll --check-prefix=DISABLED-SMALL
+;
+; A forced scalable VF reaches vectorization but remains outside the fixed-VF
+; heuristic and executes one scalable part.
+; RUN: opt -passes=loop-vectorize -mtriple=aarch64-none-linux-gnu -mattr=+sve \
+; RUN:     -force-vector-width="vscale x 2" \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -debug-only=loop-vectorize \
+; RUN:     -disable-output %t/success.ll 2>&1 \
+; RUN:     | FileCheck %t/success.ll --check-prefix=SCALABLE
+
+;--- success.ll
+; SUCCESS-LABEL: LV: Checking a loop in 'positive'
+; SUCCESS: LV: Cross-part load overlap estimate: ops=1, required-ops=1, predicted-saved-cost={{[^,]+}}, loop-cost={{[^,]+}}, saving={{[0-9]+}}%, required=1%; selecting IC=2.
+; SUCCESS-NEXT: LV: Exact cross-part load overlap predicts a downstream saving; raising IC to 2.
+; SUCCESS-NOT: LV: Not Interleaving.
+; SUCCESS: LV: Found a vectorizable loop
+; MASKED-LABEL: LV: Checking a loop in 'positive'
+; MASKED: LV: Cross-part load overlap estimate: ops=0,
+; MASKED-NOT: Exact cross-part load overlap predicts a downstream saving
+; MASKED: Executing best plan with VF=4, UF=1
+; SUCCESS-SMALL-LABEL: LV: Checking a loop in 'positive'
+; SUCCESS-SMALL: LV: Cross-part load overlap estimate: ops=1, required-ops=1, predicted-saved-cost={{[^,]+}}, loop-cost={{[^,]+}}, saving={{[0-9]+}}%, required=1%; selecting IC=2.
+; SUCCESS-SMALL-NEXT: LV: Exact cross-part load overlap predicts a downstream saving; raising IC to 2.
+; SUCCESS-SMALL-NOT: LV: Interleaving to reduce branch cost.
+; SUCCESS-SMALL: LV: Found a vectorizable loop
+; DISABLED-SMALL-LABEL: LV: Checking a loop in 'positive'
+; DISABLED-SMALL-NOT: Cross-part load overlap
+; DISABLED-SMALL: LV: Interleaving to reduce branch cost.
+; DISABLED-SMALL-NOT: Cross-part load overlap
+; DISABLED-SMALL: LV: Found a vectorizable loop
+; SCALABLE-LABEL: LV: Checking a loop in 'positive'
+; SCALABLE-NOT: LV: Cross-part load overlap estimate:
+; SCALABLE: LV: VF is vscale x 2
+; SCALABLE-NEXT: LV: Not Interleaving.
+; SCALABLE-NOT: LV: Cross-part load overlap estimate:
+; SCALABLE: LV: Found a vectorizable loop (vscale x 2)
+; SCALABLE: Executing best plan with VF=vscale x 2, UF=1
+
+target triple = "aarch64-unknown-linux-gnu"
+
+define void @positive(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+
+exit:
+  ret void
+}
+
+;--- duplicate.ll
+; CHECK-LABEL: LV: Checking a loop in 'duplicate_after_cross_part'
+; CHECK: LV: Cross-part load overlap estimate: ops=1, required-ops=2, predicted-saved-cost={{[^,]+}}, loop-cost={{[^,]+}}, saving={{[0-9]+}}%, required=5%; skipping.
+; CHECK-NOT: Exact cross-part load overlap predicts a downstream saving
+; CHECK: LV: Found a vectorizable loop
+
+target triple = "aarch64-unknown-linux-gnu"
+
+define void @duplicate_after_cross_part(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %l3 = load i32, ptr %a.i4, align 4
+  %sum.1 = add i32 %l1, %l2
+  %sum.2 = add i32 %sum.1, %l3
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum.2, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+
+exit:
+  ret void
+}
diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse.ll b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse.ll
new file mode 100644
index 0000000000000..c0ad06a15b809
--- /dev/null
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse.ll
@@ -0,0 +1,479 @@
+; RUN: split-file %s %t
+;
+; The analysis only predicts downstream savings to guide IC selection. It does
+; not modify VPlan or eliminate any widened loads.
+;
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 \
+; RUN:     -S %t/positive.ll | FileCheck %t/positive.ll --check-prefix=FORCED
+; RUN: opt -passes=loop-vectorize -force-target-max-vector-interleave=2 \
+; RUN:     -small-loop-cost=0 -enable-interleave-cse \
+; RUN:     -interleave-cse-min-ops=1 -interleave-cse-min-pct=1 \
+; RUN:     -S %t/positive.ll \
+; RUN:     | FileCheck %t/positive.ll --check-prefix=PRODUCTION
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -S %t/positive.ll \
+; RUN:     | FileCheck %t/positive.ll --check-prefix=THRESHOLD
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=100 -S %t/positive.ll \
+; RUN:     | FileCheck %t/positive.ll --check-prefix=PERCENT
+; Force deterministic costs so the ordinary branch-cost heuristic recommends
+; IC=1, allowing this run to verify that cross-part analysis can raise it to 2.
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 \
+; RUN:     -force-target-instruction-cost=1 -small-loop-cost=12 \
+; RUN:     -enable-loadstore-runtime-interleave=false \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 \
+; RUN:     -S %t/positive.ll | FileCheck %t/positive.ll --check-prefix=SMALL-IC
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -S %t/positive.ll | FileCheck %t/positive.ll --check-prefix=DISABLED
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -force-vector-interleave=1 \
+; RUN:     -S %t/positive.ll | FileCheck %t/positive.ll --check-prefix=USERIC
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=1 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 \
+; RUN:     -S %t/positive.ll | FileCheck %t/positive.ll --check-prefix=MAXIC
+; RUN: opt -passes=loop-vectorize -mtriple=aarch64-none-linux-gnu -mattr=+sve \
+; RUN:     -force-vector-width="vscale x 2" \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/positive.ll \
+; RUN:     | FileCheck %t/positive.ll --check-prefix=SCALABLE
+;
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=2 \
+; RUN:     -S %t/duplicate.ll \
+; RUN:     | FileCheck %t/duplicate.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=2 \
+; RUN:     -interleave-cse-min-pct=6 -S %t/two-opportunities.ll \
+; RUN:     | FileCheck %t/two-opportunities.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/type-mismatch.ll \
+; RUN:     | FileCheck %t/type-mismatch.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/simple-negatives.ll \
+; RUN:     | FileCheck %t/simple-negatives.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/provenance.ll \
+; RUN:     | FileCheck %t/provenance.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/poison.ll \
+; RUN:     | FileCheck %t/poison.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/multi-block.ll \
+; RUN:     | FileCheck %t/multi-block.ll
+; RUN: opt -passes=loop-vectorize -force-vector-width=4 \
+; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
+; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
+; RUN:     -interleave-cse-min-pct=1 -S %t/reverse.ll \
+; RUN:     | FileCheck %t/reverse.ll
+
+;--- positive.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; With VF=4, a[i+4] in part 0 has the same modeled vector address as a[i] in
+; part 1. The predicted downstream saving raises IC to 2, while all four
+; widened loads remain because the analysis does not realize the overlap.
+;
+; FORCED-LABEL: @positive(
+; FORCED:       vector.body:
+; FORCED-COUNT-4: load <4 x i32>
+; FORCED-NOT:   load <4 x i32>
+; FORCED:       add nuw i64 %index, 8
+;
+; PRODUCTION-LABEL: @positive(
+; PRODUCTION:       vector.body:
+; PRODUCTION-COUNT-4: load <4 x i32>
+; PRODUCTION-NOT:   load <4 x i32>
+; PRODUCTION:       add nuw i64 %index, 8
+;
+; THRESHOLD-LABEL: @positive(
+; THRESHOLD:       add nuw i64 %index, 4
+;
+; PERCENT-LABEL: @positive(
+; PERCENT:       add nuw i64 %index, 4
+;
+; SMALL-IC-LABEL: @positive(
+; SMALL-IC:       add nuw i64 %index, 8
+;
+; DISABLED-LABEL: @positive(
+; DISABLED:       add nuw i64 %index, 4
+;
+; USERIC-LABEL: @positive(
+; USERIC:       add nuw i64 %index, 4
+;
+; MAXIC-LABEL: @positive(
+; MAXIC:       add nuw i64 %index, 4
+;
+; SCALABLE-LABEL: @positive(
+; SCALABLE:       call i64 @llvm.vscale.i64()
+; SCALABLE:       vector.body:
+; SCALABLE-COUNT-2: load <vscale x 2 x i32>
+; SCALABLE-NOT:   load <vscale x 2 x i32>
+; SCALABLE:       %index.next = add nuw i64 %index, %
+define void @positive(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- duplicate.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; The first a[i+4] forms a predicted cross-part overlap with a[i] in part 1.
+; The same-part duplicate must not count as a second opportunity.
+;
+; CHECK-LABEL: @duplicate_after_cross_part(
+; CHECK:       add nuw i64 %index, 4
+define void @duplicate_after_cross_part(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %l3 = load i32, ptr %a.i4, align 4
+  %sum.1 = add i32 %l1, %l2
+  %sum.2 = add i32 %sum.1, %l3
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum.2, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- two-opportunities.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; At VF=4, a[i+4] in part 0 overlaps a[i] in part 1, and a[i+8] in
+; part 0 independently overlaps a[i+4] in part 1. Both opportunities are
+; credited, satisfying the requested minimum of two and raising IC to 2.
+;
+; CHECK-LABEL: @two_opportunities(
+; CHECK:       vector.body:
+; CHECK-COUNT-6: load <4 x i32>
+; CHECK-NOT:   load <4 x i32>
+; CHECK:       add nuw i64 %index, 8
+define void @two_opportunities(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %i8 = add nuw nsw i64 %i, 8
+  %a.i8 = getelementptr inbounds i32, ptr %a, i64 %i8
+  %l3 = load i32, ptr %a.i8, align 4
+  %sum.1 = add i32 %l1, %l2
+  %sum.2 = add i32 %sum.1, %l3
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum.2, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- type-mismatch.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; The part-shifted addresses are equal, but loads of different value types
+; cannot share a result. The type component of the key keeps IC at 1.
+;
+; CHECK-LABEL: @different_types(
+; CHECK:       vector.body:
+; CHECK-COUNT-1: load <4 x i32>
+; CHECK-COUNT-1: load <4 x float>
+; CHECK-NOT:   load <
+; CHECK:       add nuw i64 %index, 4
+define void @different_types(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds float, ptr %a, i64 %i4
+  %l2 = load float, ptr %a.i4, align 4
+  %l2.bits = bitcast float %l2 to i32
+  %sum = add i32 %l1, %l2.bits
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- simple-negatives.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; a[i] + a[i+3]: the offset 3 is not a multiple of VF=4, so no part-shifted
+; address ever matches exactly. Interleave count stays 1.
+;
+; CHECK-LABEL: @inequality(
+; CHECK:       vector.body:
+; CHECK-COUNT-2: load <4 x i32>
+; CHECK-NOT:   load <4 x i32>
+; CHECK:       add nuw i64 %index, 4
+define void @inequality(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i3 = add nuw nsw i64 %i, 3
+  %a.i3 = getelementptr inbounds i32, ptr %a, i64 %i3
+  %l2 = load i32, ptr %a.i3, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+; A store occurs between the two matching logical-part loads, so the analysis
+; must not credit their overlap and IC remains 1.
+;
+; CHECK-LABEL: @write_between(
+; CHECK:       vector.body:
+; CHECK-COUNT-2: load <4 x i32>
+; CHECK-NOT:   load <4 x i32>
+; CHECK:       add nuw i64 %index, 4
+define void @write_between(ptr noalias %a, ptr noalias %b, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %b.i = getelementptr inbounds i32, ptr %b, i64 %i
+  store i32 %l1, ptr %b.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+; Stride-2 accesses are represented by interleave recipes rather than supported
+; simple consecutive widened loads, so the analysis fails closed and leaves IC
+; at 1.
+;
+; CHECK-LABEL: @non_unit_stride(
+; CHECK:       vector.body:
+; CHECK-COUNT-2: load <8 x i32>
+; CHECK-NOT:   load <8 x i32>
+; CHECK:       add nuw i64 %index, 4
+define void @non_unit_stride(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %twice.i = shl nuw nsw i64 %i, 1
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %twice.i
+  %l1 = load i32, ptr %a.i, align 4
+  %twice.i4 = add nuw nsw i64 %twice.i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %twice.i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- provenance.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; VPlan folds each identical-arm select to its underlying GEP. The constant
+; condition also lets ScalarEvolution canonicalize the original select to that
+; same GEP. Deriving the address from the folded VPValue therefore exposes the
+; exact cross-part overlap and raises IC to 2.
+;
+; CHECK-LABEL: @folded_provenance(
+; CHECK:       vector.body:
+; CHECK:       add nuw i64 %index, 8
+define void @folded_provenance(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %p1 = select i1 true, ptr %a.i, ptr %a.i
+  %l1 = load i32, ptr %p1, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %p2 = select i1 true, ptr %a.i4, ptr %a.i4
+  %l2 = load i32, ptr %p2, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+;--- poison.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; The addresses overlap, but poison-generating annotations require metadata
+; intersection when the overlap is realized. The analysis rejects these loads
+; rather than predicting reuse it cannot preserve directly.
+;
+; CHECK-LABEL: @poison_annotations(
+; CHECK:       vector.body:
+; CHECK:       add nuw i64 %index, 4
+define void @poison_annotations(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4, !range !0
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4, !range !1
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
+
+!0 = !{i32 0, i32 100}
+!1 = !{i32 0, i32 101}
+
+;--- multi-block.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; The matching forward loads are present, but the separate latch places this
+; loop outside the analysis's single-block scope.
+;
+; CHECK-LABEL: @multi_block(
+; CHECK:       vector.body:
+; CHECK:       add nuw i64 %index, 4
+define void @multi_block(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  br label %loop.header
+loop.header:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop.latch ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum = add i32 %l1, %l2
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  br label %loop.latch
+loop.latch:
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop.header, label %exit
+exit:
+  ret void
+}
+
+;--- reverse.ll
+target triple = "aarch64-unknown-linux-gnu"
+
+; The reverse access is physically between the two forward loads. Its load is
+; not credited, while its pure end-pointer address recipe must not be treated as
+; a write barrier. The one forward equality therefore raises IC to 2.
+;
+; CHECK-LABEL: @reverse_between(
+; CHECK:       vector.body:
+; CHECK-COUNT-6: load <4 x i32>
+; CHECK-NOT:   load <4 x i32>
+; CHECK:       add nuw i64 %index, 8
+define void @reverse_between(ptr noalias %a, ptr noalias %c, i64 %n) {
+entry:
+  %last = add i64 %n, -1
+  br label %loop
+loop:
+  %i = phi i64 [ 0, %entry ], [ %i.next, %loop ]
+  %a.i = getelementptr inbounds i32, ptr %a, i64 %i
+  %l1 = load i32, ptr %a.i, align 4
+  %reverse.i = sub i64 %last, %i
+  %a.reverse = getelementptr inbounds i32, ptr %a, i64 %reverse.i
+  %reverse = load i32, ptr %a.reverse, align 4
+  %i4 = add nuw nsw i64 %i, 4
+  %a.i4 = getelementptr inbounds i32, ptr %a, i64 %i4
+  %l2 = load i32, ptr %a.i4, align 4
+  %sum.forward = add i32 %l1, %l2
+  %sum = add i32 %sum.forward, %reverse
+  %c.i = getelementptr inbounds i32, ptr %c, i64 %i
+  store i32 %sum, ptr %c.i, align 4
+  %i.next = add nuw nsw i64 %i, 1
+  %cmp = icmp slt i64 %i.next, %n
+  br i1 %cmp, label %loop, label %exit
+exit:
+  ret void
+}
diff --git a/llvm/utils/gn/secondary/llvm/lib/Transforms/Vectorize/BUILD.gn b/llvm/utils/gn/secondary/llvm/lib/Transforms/Vectorize/BUILD.gn
index 0ffc24ec855f7..6ceb8ce67b207 100644
--- a/llvm/utils/gn/secondary/llvm/lib/Transforms/Vectorize/BUILD.gn
+++ b/llvm/utils/gn/secondary/llvm/lib/Transforms/Vectorize/BUILD.gn
@@ -42,6 +42,7 @@ static_library("Vectorize") {
     "VPlan.cpp",
     "VPlanAnalysis.cpp",
     "VPlanConstruction.cpp",
+    "VPlanCrossPartCSE.cpp",
     "VPlanDominatorTree.cpp",
     "VPlanEVLTailFolding.cpp",
     "VPlanLowering.cpp",

>From 073a0d5523ad018e019064b10bae487a9702ff4b Mon Sep 17 00:00:00 2001
From: Sergey Shcherbinin <sscherbinin at nvidia.com>
Date: Sat, 12 Sep 2026 14:55:30 +0400
Subject: [PATCH 2/3] [LV] Adapt cross-part CSE to planner API changes

---
 llvm/lib/Transforms/Vectorize/LoopVectorize.cpp | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
index 9301e0a882ce6..9fc2b101eeed0 100644
--- a/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
+++ b/llvm/lib/Transforms/Vectorize/LoopVectorize.cpp
@@ -3653,8 +3653,8 @@ bool LoopVectorizationPlanner::shouldUseCrossPartCSE() const {
   // Keep the disabled path free of policy queries and cost-map allocation.
   if (!EnableInterleaveCSE)
     return false;
-  return OrigLoop->isInnermost() && Hints.getInterleave() == 0 &&
-         !CM.maskPartialAliasing();
+  return OrigLoop->isInnermost() && Config.getHints().getInterleave() == 0 &&
+         !CM->maskPartialAliasing();
 }
 
 bool LoopVectorizationPlanner::shouldCollectCrossPartCSECosts(

>From 356a054766ba4bdf500bf540ab954ba1dff1d0fa Mon Sep 17 00:00:00 2001
From: Sergey Shcherbinin <sscherbinin at nvidia.com>
Date: Sat, 12 Sep 2026 16:55:23 +0400
Subject: [PATCH 3/3] [LV] Remove obsolete wide lane mask test option

---
 .../LoopVectorize/AArch64/cross-part-load-cse-debug.ll          | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
index 1c239f95ba2d2..4be01f517abc9 100644
--- a/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
+++ b/llvm/test/Transforms/LoopVectorize/AArch64/cross-part-load-cse-debug.ll
@@ -16,7 +16,7 @@
 ; RUN:     -force-target-max-vector-interleave=2 -small-loop-cost=0 \
 ; RUN:     -force-target-supports-masked-memory-ops \
 ; RUN:     -force-tail-folding-style=data-and-control \
-; RUN:     -tail-folding-policy=must-fold-tail -enable-wide-lane-mask \
+; RUN:     -tail-folding-policy=must-fold-tail \
 ; RUN:     -enable-interleave-cse -interleave-cse-min-ops=1 \
 ; RUN:     -interleave-cse-min-pct=1 -debug-only=loop-vectorize \
 ; RUN:     -disable-output %t/success.ll 2>&1 \



More information about the llvm-commits mailing list