[llvm] [SLP] Add store-to-load forwarding conflict cost for widened store chains (PR #199606)

Milin Bhade via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 28 23:25:36 PDT 2026


================
@@ -18049,6 +18186,72 @@ BoUpSLP::getEntryCost(const TreeEntry *E, ArrayRef<Value *> VectorizedVals,
               BaseSI->getPointerAddressSpace(), CostKind, OpInfo);
         }
       }
+      // Widening this store chain can break store-to-load forwarding for a
+      // nearby loop-carried load. Rather than reject the tree outright, add
+      // the target's modeled STLF penalty so a chain that is still profitable
+      // after paying it can vectorize. The penalty is a throughput/latency
+      // hazard, so only account for it under those cost kinds. Collect every
+      // conflicting load (not just the first) so a load already paid for by
+      // an earlier committed tree does not, on its own, cause a second
+      // charge here: only a load this store is the *first* committed
+      // conflict for should add the penalty and be staged for promotion.
+      //
+      // Covers Vectorize (contiguous window) and, when the intra-vector lane
+      // stride resolves to a compile-time constant, StridedVectorize (real
+      // window `(VF-1)*|byteStride| + elementSize`, not `VF*elementSize`).
+      // ExpandVectorize (narrow/sparse masked writes: no single contiguous
+      // window) is intentionally excluded; a StridedVectorize entry whose
+      // stride is only known at runtime (a general SCEV/Value, not a
+      // constant) is also excluded, since no single fixed byte window can be
+      // derived for it at compile time.
+      bool IsStoreStateSupported = E->State == TreeEntry::Vectorize;
+      unsigned StoreSTLFVF =
+          E->Scalars.size() * std::max(1u, E->getInterleaveFactor());
+      std::optional<uint64_t> StoreSizeOverride;
+      if (E->State == TreeEntry::StridedVectorize) {
+        const StridedPtrInfo &SPtrInfo = TreeEntryToStridedPtrInfoMap.at(E);
+        std::optional<int64_t> StrideUnits;
+        if (auto *CI = dyn_cast_or_null<ConstantInt>(SPtrInfo.StrideVal))
+          StrideUnits = CI->getSExtValue();
+        else if (auto *SC = dyn_cast_or_null<SCEVConstant>(SPtrInfo.StrideSCEV))
+          StrideUnits = SC->getAPInt().getSExtValue();
+        TypeSize StoreScalarSize =
+            DL->getTypeStoreSize(BaseSI->getValueOperand()->getType());
+        if (StrideUnits && SPtrInfo.Ty && !StoreScalarSize.isScalable() &&
+            StoreScalarSize.getFixedValue() != 0) {
+          uint64_t ScalarBytes = StoreScalarSize.getFixedValue();
+          uint64_t ByteStride =
+              static_cast<uint64_t>(std::abs(*StrideUnits)) *
+              DL->getTypeAllocSize(BaseSI->getValueOperand()->getType())
+                  .getFixedValue();
+          unsigned StridedVF = SPtrInfo.Ty->getNumElements();
+          if (StridedVF > 0) {
+            StoreSizeOverride = (StridedVF - 1) * ByteStride + ScalarBytes;
+            StoreSTLFVF = StridedVF;
+            IsStoreStateSupported = true;
+          }
+        }
+      }
+      if (EnableSLPStoreLoadForwardCheck && IsStoreStateSupported &&
+          (CostKind == TTI::TCK_RecipThroughput ||
+           CostKind == TTI::TCK_Latency)) {
+        SmallVector<LoadInst *> ConflictingLoads;
+        if (findStoreLoadForwardingConflict(
+                BaseSI, StoreSTLFVF,
----------------
mbhade-amd wrote:

Addressed for contiguous stores by using VL0. I retained BaseSI only for reverse StridedVectorize, because emission rebinds the store pointer to Scalars[ReorderIndices.front()] in that case. This keeps costing aligned with the emitted store.

https://github.com/llvm/llvm-project/pull/199606


More information about the llvm-commits mailing list