[llvm] [SLP] Add store-to-load forwarding conflict cost for widened store chains (PR #199606)
Milin Bhade via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 28 23:25:36 PDT 2026
================
@@ -18049,6 +18186,72 @@ BoUpSLP::getEntryCost(const TreeEntry *E, ArrayRef<Value *> VectorizedVals,
BaseSI->getPointerAddressSpace(), CostKind, OpInfo);
}
}
+ // Widening this store chain can break store-to-load forwarding for a
+ // nearby loop-carried load. Rather than reject the tree outright, add
+ // the target's modeled STLF penalty so a chain that is still profitable
+ // after paying it can vectorize. The penalty is a throughput/latency
+ // hazard, so only account for it under those cost kinds. Collect every
+ // conflicting load (not just the first) so a load already paid for by
+ // an earlier committed tree does not, on its own, cause a second
+ // charge here: only a load this store is the *first* committed
+ // conflict for should add the penalty and be staged for promotion.
+ //
+ // Covers Vectorize (contiguous window) and, when the intra-vector lane
+ // stride resolves to a compile-time constant, StridedVectorize (real
+ // window `(VF-1)*|byteStride| + elementSize`, not `VF*elementSize`).
+ // ExpandVectorize (narrow/sparse masked writes: no single contiguous
+ // window) is intentionally excluded; a StridedVectorize entry whose
+ // stride is only known at runtime (a general SCEV/Value, not a
+ // constant) is also excluded, since no single fixed byte window can be
+ // derived for it at compile time.
+ bool IsStoreStateSupported = E->State == TreeEntry::Vectorize;
+ unsigned StoreSTLFVF =
+ E->Scalars.size() * std::max(1u, E->getInterleaveFactor());
+ std::optional<uint64_t> StoreSizeOverride;
+ if (E->State == TreeEntry::StridedVectorize) {
+ const StridedPtrInfo &SPtrInfo = TreeEntryToStridedPtrInfoMap.at(E);
+ std::optional<int64_t> StrideUnits;
+ if (auto *CI = dyn_cast_or_null<ConstantInt>(SPtrInfo.StrideVal))
+ StrideUnits = CI->getSExtValue();
+ else if (auto *SC = dyn_cast_or_null<SCEVConstant>(SPtrInfo.StrideSCEV))
+ StrideUnits = SC->getAPInt().getSExtValue();
+ TypeSize StoreScalarSize =
+ DL->getTypeStoreSize(BaseSI->getValueOperand()->getType());
+ if (StrideUnits && SPtrInfo.Ty && !StoreScalarSize.isScalable() &&
+ StoreScalarSize.getFixedValue() != 0) {
+ uint64_t ScalarBytes = StoreScalarSize.getFixedValue();
+ uint64_t ByteStride =
+ static_cast<uint64_t>(std::abs(*StrideUnits)) *
+ DL->getTypeAllocSize(BaseSI->getValueOperand()->getType())
+ .getFixedValue();
+ unsigned StridedVF = SPtrInfo.Ty->getNumElements();
+ if (StridedVF > 0) {
+ StoreSizeOverride = (StridedVF - 1) * ByteStride + ScalarBytes;
+ StoreSTLFVF = StridedVF;
+ IsStoreStateSupported = true;
+ }
+ }
+ }
+ if (EnableSLPStoreLoadForwardCheck && IsStoreStateSupported &&
+ (CostKind == TTI::TCK_RecipThroughput ||
+ CostKind == TTI::TCK_Latency)) {
+ SmallVector<LoadInst *> ConflictingLoads;
+ if (findStoreLoadForwardingConflict(
+ BaseSI, StoreSTLFVF,
----------------
mbhade-amd wrote:
Addressed for contiguous stores by using VL0. I retained BaseSI only for reverse StridedVectorize, because emission rebinds the store pointer to Scalars[ReorderIndices.front()] in that case. This keeps costing aligned with the emitted store.
https://github.com/llvm/llvm-project/pull/199606
More information about the llvm-commits
mailing list