[llvm] [SLP] Bail out on store-to-load forwarding hazards (PR #199606)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Jul 20 21:54:29 PDT 2026
================
@@ -27384,6 +27418,147 @@ bool SLPVectorizerPass::runImpl(Function &F, ScalarEvolution *SE_,
return Changed;
}
+bool BoUpSLP::findStoreLoadForwardingConflict(StoreInst *BaseStore,
+ unsigned VF) {
+ if (!BaseStore || !LAIs)
+ return false;
+
+ StoreInst *FirstStore = BaseStore;
+
+ // Cache lookup: avoid re-walking the LAA dependence list when the same
+ // chain is retried at multiple vector factors.
+ auto Key = std::make_pair(FirstStore, VF);
+ auto CacheIt = StlfConflictCache.find(Key);
+ if (CacheIt != StlfConflictCache.end())
+ return CacheIt->second;
+
+ auto CacheAndReturn = [&](bool Result) -> bool {
+ StlfConflictCache[Key] = Result;
+ return Result;
+ };
+
+ Loop *L = LI->getLoopFor(FirstStore->getParent());
+ if (!L)
+ return CacheAndReturn(false);
+
+ Type *ValueTy = FirstStore->getValueOperand()->getType();
+ TypeSize StoreSize = DL->getTypeStoreSize(ValueTy);
+ if (StoreSize.isScalable())
+ return CacheAndReturn(false);
+ uint64_t ElementSize = StoreSize.getFixedValue();
+ if (ElementSize == 0)
+ return CacheAndReturn(false);
+ uint64_t VectorStoreBytes = uint64_t(VF) * ElementSize;
+ LLVM_DEBUG(dbgs() << "SLP: STLF check: VF=" << VF
+ << " ElementSize=" << ElementSize
+ << " VectorStoreBytes=" << VectorStoreBytes << "\n");
+
+ // Cheap early-gate: if the loop has no loads from the same underlying
+ // object as the store chain, STLF conflicts are impossible. This avoids
+ // paying for LAA on loops that obviously cannot conflict.
+ Value *StoreBase = getUnderlyingObject(FirstStore->getPointerOperand());
+ bool HasSameBaseLoad = false;
+ for (BasicBlock *BB : L->blocks()) {
+ for (Instruction &I : *BB) {
+ if (auto *LdI = dyn_cast<LoadInst>(&I)) {
+ if (LdI->isSimple() &&
+ getUnderlyingObject(LdI->getPointerOperand()) == StoreBase) {
+ HasSameBaseLoad = true;
+ break;
+ }
+ }
+ }
+ if (HasSameBaseLoad)
+ break;
+ }
+ if (!HasSameBaseLoad)
+ return CacheAndReturn(false);
+
+ const LoopAccessInfo &LAI = LAIs->getInfo(*L);
----------------
mbhade-amd wrote:
Measured **CTMark** (Release, no-asserts) with retired-instruction counts (`perf`, pinned to one core):
- **Base (PR parent `0ef35be`) vs PR tip:** +0.061% instructions (+424M / 696G), −0.26% wall time (within noise), 0% code-size change.
- **Isolating the check on the same binary** (heuristic on vs `-slp-store-load-forward-check=false`): **+0.055%** — i.e. essentially all of the PR's compile-time cost is the STLF check itself, and it's negligible. The non-zero delta also confirms the path is actually exercised on CTMark.
So the added `LoopAccessInfo` query isn't doing measurable unnecessary work: the early bail-out (no same-object load in the loop → skip LAA) and the per-`(chain-base, VF)` cache keep it off the hot path. Happy to share the per-benchmark table or run the full test-suite if useful.
https://github.com/llvm/llvm-project/pull/199606
More information about the llvm-commits
mailing list