[llvm] [SLP] Bail out on store-to-load forwarding hazards (PR #199606)

via llvm-commits llvm-commits at lists.llvm.org
Mon Jul 20 21:54:29 PDT 2026


================
@@ -27384,6 +27418,147 @@ bool SLPVectorizerPass::runImpl(Function &F, ScalarEvolution *SE_,
   return Changed;
 }
 
+bool BoUpSLP::findStoreLoadForwardingConflict(StoreInst *BaseStore,
+                                              unsigned VF) {
+  if (!BaseStore || !LAIs)
+    return false;
+
+  StoreInst *FirstStore = BaseStore;
+
+  // Cache lookup: avoid re-walking the LAA dependence list when the same
+  // chain is retried at multiple vector factors.
+  auto Key = std::make_pair(FirstStore, VF);
+  auto CacheIt = StlfConflictCache.find(Key);
+  if (CacheIt != StlfConflictCache.end())
+    return CacheIt->second;
+
+  auto CacheAndReturn = [&](bool Result) -> bool {
+    StlfConflictCache[Key] = Result;
+    return Result;
+  };
+
+  Loop *L = LI->getLoopFor(FirstStore->getParent());
+  if (!L)
+    return CacheAndReturn(false);
+
+  Type *ValueTy = FirstStore->getValueOperand()->getType();
+  TypeSize StoreSize = DL->getTypeStoreSize(ValueTy);
+  if (StoreSize.isScalable())
+    return CacheAndReturn(false);
+  uint64_t ElementSize = StoreSize.getFixedValue();
+  if (ElementSize == 0)
+    return CacheAndReturn(false);
+  uint64_t VectorStoreBytes = uint64_t(VF) * ElementSize;
+  LLVM_DEBUG(dbgs() << "SLP: STLF check: VF=" << VF
+                    << " ElementSize=" << ElementSize
+                    << " VectorStoreBytes=" << VectorStoreBytes << "\n");
+
+  // Cheap early-gate: if the loop has no loads from the same underlying
+  // object as the store chain, STLF conflicts are impossible. This avoids
+  // paying for LAA on loops that obviously cannot conflict.
+  Value *StoreBase = getUnderlyingObject(FirstStore->getPointerOperand());
+  bool HasSameBaseLoad = false;
+  for (BasicBlock *BB : L->blocks()) {
+    for (Instruction &I : *BB) {
+      if (auto *LdI = dyn_cast<LoadInst>(&I)) {
+        if (LdI->isSimple() &&
+            getUnderlyingObject(LdI->getPointerOperand()) == StoreBase) {
+          HasSameBaseLoad = true;
+          break;
+        }
+      }
+    }
+    if (HasSameBaseLoad)
+      break;
+  }
+  if (!HasSameBaseLoad)
+    return CacheAndReturn(false);
+
+  const LoopAccessInfo &LAI = LAIs->getInfo(*L);
----------------
mbhade-amd wrote:

Measured **CTMark** (Release, no-asserts) with retired-instruction counts (`perf`, pinned to one core):

- **Base (PR parent `0ef35be`) vs PR tip:** +0.061% instructions (+424M / 696G), −0.26% wall time (within noise), 0% code-size change.
- **Isolating the check on the same binary** (heuristic on vs `-slp-store-load-forward-check=false`): **+0.055%** — i.e. essentially all of the PR's compile-time cost is the STLF check itself, and it's negligible. The non-zero delta also confirms the path is actually exercised on CTMark.

So the added `LoopAccessInfo` query isn't doing measurable unnecessary work: the early bail-out (no same-object load in the loop → skip LAA) and the per-`(chain-base, VF)` cache keep it off the hot path. Happy to share the per-benchmark table or run the full test-suite if useful.

https://github.com/llvm/llvm-project/pull/199606


More information about the llvm-commits mailing list