[llvm] [LAA] Add stencil group merging to reduce runtime pointer checks (PR #187252)
Florian Hahn via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 7 01:11:06 PDT 2026
================
@@ -747,6 +781,643 @@ void RuntimePointerChecking::groupChecks(
}
}
+/// Result of decomposing a SCEV expression into stencil offset form:
+/// Offset = Constant + sum(Coefficients[stride] * stride)
+/// where each stride is a loop-invariant SCEV expression.
+struct StencilDecomposition {
+ int64_t Constant = 0;
+ /// Map from loop-invariant stride SCEV to its integer coefficient.
+ SmallMapVector<const SCEV *, int64_t, 4> Coefficients;
+};
+
+/// Recursion cap for addScaledStencilTerm. Depth counts how deep a term
+/// sits inside the offset expression. For example, the offset
+/// 8 + (64 * (s1 + s2 + (4 * s3)))
+/// is visited like this:
+/// depth 0: the whole add
+/// depth 1: its operands 8 and (64 * (s1 + s2 + (4 * s3)))
+/// depth 2: (s1 + s2 + (4 * s3)), the operand of the multiply
+/// depth 3: s1, s2 and (4 * s3), the operands of that add
+/// At depth 3 addScaledStencilTerm stops going deeper. s1 and s2 are plain
+/// strides anyway. (4 * s3) is not split into 4 times s3: it becomes one
+/// stride key as it is, with coefficient 64. The result is Constant = 8
+/// and coefficients {s1: 64, s2: 64, (4 * s3): 64}.
+/// Three levels cover the stencil offsets we care about: a top-level add,
+/// a constant times a sum inside it, and the strides in that sum. A deeper
+/// term is kept whole as one stride key. The merge does not care what is
+/// inside a key. It only needs a loop-invariant value with a
+/// positive-stride predicate, and a whole term has both. The only cost is
+/// precision, when another member uses a part of that term, here s3 alone,
+/// as a key of its own. isNeverAbove sees two unrelated keys, so a member
+/// that is in fact always lower or higher may stay a candidate.
+constexpr unsigned MaxStencilDecomposeDepth = 3;
+
+/// Add one term of a stencil offset to \p D. \p Mult is the factor in
+/// front of the term; the top-level call passes 1.
+/// Example: the offset 8 + (-64 * (s1 + s2)) + (-32 * s1), Mult = 1. It is
+/// an add, so each operand is visited in turn with the same Mult = 1:
+/// 8 a constant: D.Constant += 1 * 8
+/// (-64 * (s1 + s2)) a constant times X: visit X = (s1 + s2) with
+/// Mult = 1 * -64. X is an add, so each operand is
+/// visited with Mult = -64:
+/// s1 a stride: D.Coefficients[s1] += -64
+/// s2 a stride: D.Coefficients[s2] += -64
+/// (-32 * s1) a constant times X: visit X = s1 with Mult = -32:
+/// s1 a stride: D.Coefficients[s1] += -32
+/// Result: Constant = 8, Coefficients {s1: -96, s2: -64}. The -64 and the
+/// -32 for s1 come from two different terms and add up in the map.
+/// So, by the kind of term:
+/// constant K D.Constant += Mult * K
+/// (K * X) visit X with Mult * K
+/// (a + b + ...) visit a, b, ... each with this same Mult
+/// anything else a stride key: D.Coefficients[Term] += Mult
+/// The two recursive cases only fire while Depth is below
+/// MaxStencilDecomposeDepth. At the cap, (K * X) and (a + b + ...) are
+/// stride keys like anything else; that is not a bailout.
+/// Returns false when a constant does not fit in int64_t or an update
+/// overflows. The caller then drops the whole decomposition.
+static bool addScaledStencilTerm(const SCEV *Term, int64_t Mult, unsigned Depth,
+ StencilDecomposition &D) {
+ const SCEVConstant *C;
+ // A constant folds into the running constant at any depth.
+ if (match(Term, m_SCEVConstant(C))) {
+ std::optional<int64_t> V = C->getAPInt().trySExtValue();
+ int64_t Scaled;
+ return V && !MulOverflow(Mult, *V, Scaled) &&
+ !AddOverflow(D.Constant, Scaled, D.Constant);
+ }
+
+ if (Depth < MaxStencilDecomposeDepth) {
+ const SCEV *Inner;
+ if (match(Term, m_scev_Mul(m_SCEVConstant(C), m_SCEV(Inner)))) {
+ std::optional<int64_t> V = C->getAPInt().trySExtValue();
+ int64_t NewMult;
+ return V && !MulOverflow(Mult, *V, NewMult) &&
+ addScaledStencilTerm(Inner, NewMult, Depth + 1, D);
+ }
+ if (auto *Add = dyn_cast<SCEVAddExpr>(Term))
+ return all_of(Add->operands(), [&](const SCEV *Op) {
+ return addScaledStencilTerm(Op, Mult, Depth + 1, D);
+ });
+ }
+
+ // Anything else is one stride key.
+ int64_t &Coeff = D.Coefficients[Term];
+ return !AddOverflow(Coeff, Mult, Coeff);
+}
+
+/// Try to decompose \p Expr into a stencil offset function of loop-invariant
+/// strides: C + a1*s1 + a2*s2 + ...
+/// \p Expr is the difference of two access "Start" SCEVs (Start_member -
+/// Start_base). A "Start" is the low bound of a memory access range as computed
+/// by getStartAndEndForAccess: the address of the first byte the access can
+/// touch. The result describes where one member's range sits relative to the
+/// base member's range.
+/// Constant factors are distributed over sums. SCEV can keep a factored form:
+/// -64*s1 + -64*s2 is stored as (-64 * (s1 + s2)). Distributing the -64 gives
+/// the coefficients {s1: -64, s2: -64}, so every member of a group is keyed
+/// on the same base strides.
+/// Relies on SCEV's canonical form: AddExpr operands are flattened (N-ary),
+/// MulExpr has the constant operand first when present.
+/// Returns std::nullopt if a constant does not fit in int64_t or a multiplier
+/// or coefficient update overflows (we commit to the signed interpretation;
+/// values that need more than 64 significant bits are out of scope).
+static std::optional<StencilDecomposition>
+decomposeStencilOffset(const SCEV *Expr, ScalarEvolution &SE, const Loop &L) {
+ // A "Start" is always loop-invariant (getStartAndEndForAccess asserts it), so
+ // the difference Expr passed in by the caller is loop-invariant too, and so
+ // is every term addScaledStencilTerm visits.
+ assert(SE.isLoopInvariant(Expr, &L) && "expected a loop-invariant offset");
+
+ StencilDecomposition D;
+ if (!addScaledStencilTerm(Expr, /*Mult=*/1, /*Depth=*/0, D))
+ return std::nullopt;
+ return D;
+}
+
+/// Return true if offset A is never higher than offset B.
+/// A and B are these sums:
+/// A = A.Constant + CoefA_1 * stride_1 + CoefA_2 * stride_2 + ...
----------------
fhahn wrote:
It is not entirely clear to me how we guard against any term of Coef * stride wrapping? It would be great if the comment could clarify that
https://github.com/llvm/llvm-project/pull/187252
More information about the llvm-commits
mailing list