[flang-commits] [flang] [flang][docs] add design doc for inline runtime checks (PR #228502)
Tom Eccles via flang-commits
flang-commits at lists.llvm.org
Mon Oct 5 03:43:52 PDT 2026
================
@@ -0,0 +1,635 @@
+<!--===- docs/InlineRuntimeChecks.md
+
+ Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+ See https://llvm.org/LICENSE.txt for license information.
+ SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+
+-->
+
+# Inline runtime checks in Flang with `fir.assert`
+
+Status: draft design, for discussion.
+Context: [llvm/llvm-project#223287](https://github.com/llvm/llvm-project/pull/223287).
+Source references are to `llvm-project` at `fe1de64cb3c2` (2026-10-02). Line
+numbers drift, so each reference also names the function or pattern.
+
+Abbreviations used in the tables:
+
+| Short | Path |
+|---|---|
+| `SHI` | `flang/lib/Optimizer/HLFIR/Transforms/SimplifyHLFIRIntrinsics.cpp` |
+| `OB` | `flang/lib/Optimizer/HLFIR/Transforms/OptimizedBufferization.cpp` |
+| `IHA` | `flang/lib/Optimizer/HLFIR/Transforms/InlineHLFIRAssign.cpp` |
+| `SAA` | `flang/lib/Optimizer/HLFIR/Transforms/SeparateAllocatableAssign.cpp` |
+| `C2F` | `flang/lib/Optimizer/HLFIR/Transforms/ConvertToFIR.cpp` |
+| `SI` | `flang/lib/Optimizer/Transforms/SimplifyIntrinsics.cpp` (FIR level) |
+| `IC` | `flang/lib/Optimizer/Builder/IntrinsicCall.cpp` |
+| `X2H` | `flang/lib/Lower/ConvertExprToHLFIR.cpp` |
+| `RT/` | `flang-rt/lib/runtime/` |
+
+## 1. Problem
+
+Flang relies on the Fortran runtime for almost all dynamic checks: shape
+conformance, allocation status, argument consistency of intrinsics. Every
+time an HLFIR or FIR pass replaces a runtime call with inline code, the check
+silently disappears. Three consequences:
+
+1. **Behavior depends on the optimization level.** The inlining passes mostly
+ run at `-O1` and above, so the same non-conforming program stops with a
+ clear message at `-O0` and does something else at `-O2`.
+2. **The inline code is often memory-unsafe on non-conforming input, not just
+ wrong.** Loop bounds are taken from one operand and used to index another
+ (`fir::factory::deduceOptimalExtents`). A size mismatch becomes an
+ out-of-bounds read or write instead of an error.
+3. **Some checks the standard requires are missing on every path.** For
+ example, the inline ALLOCATE/DEALLOCATE of intrinsic-type allocatables
+ does not check the allocation status (F2018 9.7.1.3, 9.7.3.2, 9.7.4).
+
+The few inline checks that exist today (MOD/MODULO `P==0`, NEAREST `S==0`,
+STORAGE_SIZE, the realloc-with-scalar-RHS case) use `fir.if` plus
+`fir.call @_FortranAReportFatalUserError`. That pattern:
+
+- has no data dependence on the operation it protects, so passes may move the
+ protected operation above the check (MLIR models neither "may not return"
+ nor non-local control flow: `mlir/docs/Rationale/SideEffectsAndSpeculation.md`);
+- introduces a call with unknown memory effects, which blocks alias-based
+ optimizations;
+- does not tell LLVM the call never returns (the FIR declaration has no
+ `noreturn`), so LLVM cannot use the checked condition afterwards.
+
+## 2. Goals and non-goals
+
+Goals:
+
+- Keep every user-facing check the runtime performs when the call is inlined,
+ and report it with the runtime's message format:
+ `fatal Fortran runtime error(<file>:<line>): <message>`.
+- Keep the checks cheap and transparent to optimization. They must not block
+ store-to-load forwarding, CSE of loads, LICM of loads, or alias analysis.
+- Let checks be hoisted, merged, folded, and dropped under an option.
+- Provide one mechanism for future opt-in checks (`-fcheck=bounds`,
+ `pointer`, `do`, `bits`).
+
+Non-goals:
+
+- Detecting undefined POINTER association status (only disassociated, that
+ is null, is detectable).
+- Checking explicit-shape dummy size against the actual argument in general
+ (the callee only receives an address).
+- Changing the runtime's own checks, apart from the bugs listed in the
+ appendix.
+
+## 3. The `fir.assert` operation
+
+### 3.1 Semantics and syntax
+
+```
+%g:N = fir.assert %ok, <kind>, "<message>" [values(%v0, ... : T0, ...)] [guard(%x0, ... : U0, ...)]
+```
+
+- `%ok` (`i1`) is the condition that must hold.
+- `<kind>` is a check category (section 4). Options act on categories.
+- `<message>` is the error text, a `printf` format.
+- `values(%v0, ... : T0, ...)` are optional integer or `index` SSA values
+ (`T0`, ... are their types) printed into the message, one per `%jd`
+ directive. They are only read on the failure path. For example, the two
+ extents in `"DOT_PRODUCT: SIZE(VECTOR_A) is %jd but SIZE(VECTOR_B) is %jd"`.
+- `guard(%x0, ... : U0, ...)` are the optional guarded values, of any type.
+ The op has one result per guarded value, with the same type (`%g#i` has
+ type `Ui`).
+- If `%ok` is true, each result equals the corresponding guarded value and
+ nothing else happens. Otherwise the program terminates through the Fortran
+ runtime with the formatted message.
+- Operations that must not run before the check consume the guarded results.
+ The ordering is then enforced by SSA dominance, for every pass, without any
+ pass having to know about `fir.assert`. Asserts without a guard only order
+ themselves with respect to I/O and calls (section 3.2).
+
+What to guard: the value from which the dangerous access is derived. For
+inlined loops this is the **extent used as the loop bound** (or the
+`fir.shape` operands): every access in the loop depends on the induction
+variable, which depends on the guarded extent. Guarding extents rather than
+array entities avoids threading `!hlfir.expr` values through the assert.
+
+### 3.2 Traits
+
+- `MemoryEffects<[MemWrite<FortranRuntimeResource>]>`, where
+ `FortranRuntimeResource` is a non-addressable resource (like the existing
+ `fir::DebuggingResource`).
+ - The op is never trivially dead, so DCE keeps asserts without uses.
+ - It stays ordered with calls and I/O, whose unknown effects cover every
+ resource.
+ - It does not touch user memory. FIR alias analysis skips non-addressable
+ resources (`flang/lib/Optimizer/Analysis/AliasAnalysis.cpp`, `getModRef`),
+ and MLIR CSE ignores writes to resources disjoint from a load's resource.
+ So asserts do not block load CSE, store-to-load forwarding, or FIR LICM of
+ loads. This also means `SimplifyHLFIRIntrinsics` can emit asserts in its
+ first run (`allowNewSideEffects=false`, `SHI:3321-3333`): the concern
+ there is new effects that stop CSE from merging loads, which this op does
+ not have.
+- Not speculatable.
+- **Not transparent to other operations' folds and patterns.** The assert's
+ own canonicalization removes it when `%ok` is a constant true (section 3.4).
+ But no other rewrite may look through an assert to its guarded operand,
+ because that drops the dependence. For example, in
+ `%a = fir.box_addr %b; %a_ok = fir.assert ... guard(%a); %c = fir.embox %a_ok`,
+ a fold that rewrites an `fir.embox` of a `fir.box_addr` into the original
+ box must not replace `%c` with `%b`. In practice:
+ - the op does not implement `ViewLikeOpInterface`;
+ - it is not added to the helpers that skip `fir.declare` or `fir.convert`
+ to find a defining op;
+ - it is not given a folder that returns a guarded operand when `%ok` is
+ unknown.
+
+ Alias analysis may look through the assert, since it only answers
+ aliasing queries and does not rewrite uses.
+
+### 3.3 TableGen sketch
+
+```
+def fir_AssertOp : fir_Op<"assert", [AttrSizedOperandSegments,
+ MemoryEffects<[MemWrite<FortranRuntimeResource>]>]> {
+ let arguments = (ins I1:$condition,
+ fir_RuntimeCheckKindAttr:$kind,
+ StrAttr:$message,
+ Variadic<AnySignlessIntegerOrIndex>:$values,
+ Variadic<AnyType>:$guarded);
+ let results = (outs Variadic<AnyType>:$results); // types == guarded types
+ let hasVerifier = 1;
+ let hasCanonicalizer = 1;
+}
+```
+
+### 3.4 Canonicalization and redundancy elimination
+
+- **Constant true:** erase the assert and replace its results with the guarded
+ operands. This replaces the hand-written `fir::getIntIfConstant` early-outs
+ of today's checks (`IC:6929-6931`). It must be a canonicalization pattern
+ rather than only a `fold`: for an assert with no guard, an empty fold result
+ means an in-place update, not an erasure.
+- **Constant false:** keep it, with no diagnostic. A false condition is common
+ in code that is unreachable but hard to prove unreachable at compile time
+ (for example, after inlining or specialization on constant arguments), so a
+ warning would be noisy.
+
+ A later pass could erase the code strictly dominated by a constant-false
+ assert. This is not proposed for the first version, and it must not run in
+ any pipeline that may drop asserts (for example, device code with checks
+ disabled). There, the assert would disappear together with the pruned
+ code, and a program that should have stopped would silently produce
+ wrong results.
+- **Dominating assert with the same condition:** remove the later one, mapping
+ its results to the earlier one's results. This is valid regardless of what
+ lies between them and regardless of the message, because the earlier check
+ fails first. Only remove it when the earlier assert guards everything the
+ later one guards; otherwise the consumers would lose their dependence.
+ Upstream CSE cannot do this: it skips ops with write effects
+ (`mlir/lib/Transforms/Utils/CSE.cpp`, `simplifyOperation`). So this needs a
+ small dominance-based pass (section 6).
+
+### 3.5 Lowering
+
+Lowering is done in `FIRToLLVMLowering`, where the IR is already a control-flow
+graph:
+
+```
+// fir.assert %ok, conformance, "..." values(%a, %b : index, index) guard(%n : index)
+llvm.cond_br %ok weights([2000, 1]), ^cont, ^fail
+^fail:
+ llvm.call @_FortranAReportFatalUserErrorValues(%msg, %file, %line, %a64, %b64, ...)
+ llvm.unreachable
+^cont: // uses of the result are replaced with %n
+```
+
+- **Runtime entry.** `_FortranAReportFatalUserError(msg, source, line)`
+ (`flang/include/flang/Runtime/stop.h`) is `[[noreturn]]` and device-callable
+ (`RT_API_ATTRS`). It takes no values. Proposed addition:
+ `ReportFatalUserErrorValues(msg, source, line, int64 v0..v3)` with a fixed
+ number of `int64_t` arguments. That keeps the ABI simple on devices, and
+ messages use `%jd` only. The message is passed to `printf`, so a literal
+ `%` in a message must be escaped.
+- **Declaration.** The runtime function declaration must carry `noreturn`,
+ so LLVM can use the checked condition on the continuation path.
+- **Source location.** Take it from the assert's own location. Resolve
+ `FusedLoc` and `CallSiteLoc` to the innermost `FileLineColLoc`: today
+ `fir::factory::locationToFilename` returns null for them, so inlined checks
+ lose file and line.
+- **Device code.** String globals must be created in the nearest symbol table
+ (a `gpu.module` for device code), and the lowering must also exist in any
+ device code generation pipeline. If a device runtime lacks the entry point,
+ fall back to `printf` plus `llvm.trap`. The function is not in the OpenMP
+ offload API group today, unlike `Abort`.
+
+Lowering before `CFGConversion` (to `fir.if` plus call) is also possible, but
+`fir.if` regions cannot end with `fir.unreachable`, and asserts would stop
+being visible to the passes after that point.
+
+### 3.6 Alternatives considered
+
+**`fir.if` plus a runtime call (today's pattern).** Not retained, for the
+reasons in section 1: there is no data dependence on the protected operation,
+the call's unknown memory effects block alias-based optimizations, and LLVM
+is not told that the call never returns.
+
+**`cf.assert`.** Not retained. It writes the default resource, which is how
+it models program termination, so it blocks store-to-load forwarding and
+load CSE across it. It still does not order non-speculatable operations
+without memory effects, such as `arith.remsi`, after it. Its LLVM lowering
+prints with `puts` and calls `abort`, rather than reporting through the
+Fortran runtime. It also has no guarded results.
+
+**A result-less `fir.assert` (the current revision of
+[#223287](https://github.com/llvm/llvm-project/pull/223287)).** The PR
+defines `fir.assert` as a non-speculatable op with a
+`MemWrite<FortranRuntimeResource>` effect, lowers it through `cf.assert`, and
+teaches Flang LICM not to hoist non-speculatable operations across a
+preceding assert in the same block. This design keeps that effect model
+(section 3.2), which gives the right DCE, alias-analysis, and forwarding
+behavior. It is not retained as is, because the ordering of the protected
+operations is not expressed in the IR:
+
+- MLIR has no notion of an operation that may not return. Every pass may
+ assume that all operations in a block execute together, and may move a
+ `fir.load` or an `arith.remsi` above an assert whose effects do not
+ conflict with it. The LICM change protects one pass. Every other pass that
+ moves or recreates operations would need the same fix, and future passes
+ would have to remember it. Examples are sinking, scheduling, and rewrite
+ patterns that build new operations at a different insertion point.
+- Once a protected load or remainder is above the check, LLVM may infer from
+ its undefined behavior that the check cannot fail, and delete it.
+- The guarded results make the ordering a matter of SSA dominance, so no pass
+ has to know about `fir.assert`. The lowering then goes directly to the
+ Fortran runtime entry, so messages match the runtime's format (section
+ 3.5).
+
+The design in this document is therefore the PR's op, with guarded results
+and a direct lowering added.
+
+**A region-based assert, like `shape.assuming`.** The protected operations go
+inside a region attached to the assert, and the values used afterwards are
+yielded out:
+
+```
+%r = fir.assert %ok, conformance, "..." -> (f32) {
+ // inlined DOT_PRODUCT loop nest
+ fir.result %sum : f32
+}
+```
+
+Ordering is structural, so it is as safe as the guarded-result form. It is
+not retained because the region gets in the way of the optimizations the
+checks must not block:
+
+- The protected code is usually a whole loop nest or the rest of a
+ statement, so everything it defines has to be yielded through the region.
+ An allocation-status check would wrap the rest of the block.
+- Passes that work on a single block or match operation sequences see a
+ region boundary instead. Examples are the `OptimizedBufferization` pattern
+ that matches an `hlfir.elemental` and its `hlfir.assign`, store-to-load
+ forwarding, and LICM of loads out of the region.
+- Merging two redundant asserts means merging or inlining regions. Hoisting
+ one means moving its whole region, or splitting it.
+- Several checks on the same statement nest regions inside each other.
+
+The guarded-result form gives the same guarantee at the granularity of a
+value: guarding a loop's extent protects every access in the loop, without
+moving the loop into a region.
+
+## 4. Check categories and default policy
+
+| Kind | Meaning | Default | Option |
+|---|---|---|---|
+| `status` | Allocation status required by the standard (ALLOCATE/DEALLOCATE, assignment to an unallocated or disassociated left-hand side) and allocation failure | Always on, all optimization levels | Engineering option to drop, for staging |
----------------
tblah wrote:
Agreed that this should be on by default. Maybe a flag to turn it off would be useful for diagnosing if the extra if statement regressed anything.
https://github.com/llvm/llvm-project/pull/228502
More information about the flang-commits
mailing list