[flang-commits] [flang] [flang][docs] add design doc for inline runtime checks (PR #228502)
Tom Eccles via flang-commits
flang-commits at lists.llvm.org
Mon Oct 5 03:43:54 PDT 2026
================
@@ -0,0 +1,635 @@
+<!--===- docs/InlineRuntimeChecks.md
+
+ Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+ See https://llvm.org/LICENSE.txt for license information.
+ SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+
+-->
+
+# Inline runtime checks in Flang with `fir.assert`
+
+Status: draft design, for discussion.
+Context: [llvm/llvm-project#223287](https://github.com/llvm/llvm-project/pull/223287).
+Source references are to `llvm-project` at `fe1de64cb3c2` (2026-10-02). Line
+numbers drift, so each reference also names the function or pattern.
+
+Abbreviations used in the tables:
+
+| Short | Path |
+|---|---|
+| `SHI` | `flang/lib/Optimizer/HLFIR/Transforms/SimplifyHLFIRIntrinsics.cpp` |
+| `OB` | `flang/lib/Optimizer/HLFIR/Transforms/OptimizedBufferization.cpp` |
+| `IHA` | `flang/lib/Optimizer/HLFIR/Transforms/InlineHLFIRAssign.cpp` |
+| `SAA` | `flang/lib/Optimizer/HLFIR/Transforms/SeparateAllocatableAssign.cpp` |
+| `C2F` | `flang/lib/Optimizer/HLFIR/Transforms/ConvertToFIR.cpp` |
+| `SI` | `flang/lib/Optimizer/Transforms/SimplifyIntrinsics.cpp` (FIR level) |
+| `IC` | `flang/lib/Optimizer/Builder/IntrinsicCall.cpp` |
+| `X2H` | `flang/lib/Lower/ConvertExprToHLFIR.cpp` |
+| `RT/` | `flang-rt/lib/runtime/` |
+
+## 1. Problem
+
+Flang relies on the Fortran runtime for almost all dynamic checks: shape
+conformance, allocation status, argument consistency of intrinsics. Every
+time an HLFIR or FIR pass replaces a runtime call with inline code, the check
+silently disappears. Three consequences:
+
+1. **Behavior depends on the optimization level.** The inlining passes mostly
+ run at `-O1` and above, so the same non-conforming program stops with a
+ clear message at `-O0` and does something else at `-O2`.
+2. **The inline code is often memory-unsafe on non-conforming input, not just
+ wrong.** Loop bounds are taken from one operand and used to index another
+ (`fir::factory::deduceOptimalExtents`). A size mismatch becomes an
+ out-of-bounds read or write instead of an error.
+3. **Some checks the standard requires are missing on every path.** For
+ example, the inline ALLOCATE/DEALLOCATE of intrinsic-type allocatables
+ does not check the allocation status (F2018 9.7.1.3, 9.7.3.2, 9.7.4).
+
+The few inline checks that exist today (MOD/MODULO `P==0`, NEAREST `S==0`,
+STORAGE_SIZE, the realloc-with-scalar-RHS case) use `fir.if` plus
+`fir.call @_FortranAReportFatalUserError`. That pattern:
+
+- has no data dependence on the operation it protects, so passes may move the
+ protected operation above the check (MLIR models neither "may not return"
+ nor non-local control flow: `mlir/docs/Rationale/SideEffectsAndSpeculation.md`);
+- introduces a call with unknown memory effects, which blocks alias-based
+ optimizations;
+- does not tell LLVM the call never returns (the FIR declaration has no
+ `noreturn`), so LLVM cannot use the checked condition afterwards.
+
+## 2. Goals and non-goals
+
+Goals:
+
+- Keep every user-facing check the runtime performs when the call is inlined,
+ and report it with the runtime's message format:
+ `fatal Fortran runtime error(<file>:<line>): <message>`.
+- Keep the checks cheap and transparent to optimization. They must not block
+ store-to-load forwarding, CSE of loads, LICM of loads, or alias analysis.
+- Let checks be hoisted, merged, folded, and dropped under an option.
+- Provide one mechanism for future opt-in checks (`-fcheck=bounds`,
+ `pointer`, `do`, `bits`).
+
+Non-goals:
+
+- Detecting undefined POINTER association status (only disassociated, that
+ is null, is detectable).
+- Checking explicit-shape dummy size against the actual argument in general
+ (the callee only receives an address).
+- Changing the runtime's own checks, apart from the bugs listed in the
+ appendix.
+
+## 3. The `fir.assert` operation
+
+### 3.1 Semantics and syntax
+
+```
+%g:N = fir.assert %ok, <kind>, "<message>" [values(%v0, ... : T0, ...)] [guard(%x0, ... : U0, ...)]
+```
+
+- `%ok` (`i1`) is the condition that must hold.
+- `<kind>` is a check category (section 4). Options act on categories.
+- `<message>` is the error text, a `printf` format.
+- `values(%v0, ... : T0, ...)` are optional integer or `index` SSA values
+ (`T0`, ... are their types) printed into the message, one per `%jd`
+ directive. They are only read on the failure path. For example, the two
+ extents in `"DOT_PRODUCT: SIZE(VECTOR_A) is %jd but SIZE(VECTOR_B) is %jd"`.
+- `guard(%x0, ... : U0, ...)` are the optional guarded values, of any type.
+ The op has one result per guarded value, with the same type (`%g#i` has
+ type `Ui`).
+- If `%ok` is true, each result equals the corresponding guarded value and
+ nothing else happens. Otherwise the program terminates through the Fortran
+ runtime with the formatted message.
+- Operations that must not run before the check consume the guarded results.
+ The ordering is then enforced by SSA dominance, for every pass, without any
+ pass having to know about `fir.assert`. Asserts without a guard only order
+ themselves with respect to I/O and calls (section 3.2).
+
+What to guard: the value from which the dangerous access is derived. For
+inlined loops this is the **extent used as the loop bound** (or the
+`fir.shape` operands): every access in the loop depends on the induction
+variable, which depends on the guarded extent. Guarding extents rather than
+array entities avoids threading `!hlfir.expr` values through the assert.
+
+### 3.2 Traits
+
+- `MemoryEffects<[MemWrite<FortranRuntimeResource>]>`, where
+ `FortranRuntimeResource` is a non-addressable resource (like the existing
+ `fir::DebuggingResource`).
+ - The op is never trivially dead, so DCE keeps asserts without uses.
+ - It stays ordered with calls and I/O, whose unknown effects cover every
+ resource.
+ - It does not touch user memory. FIR alias analysis skips non-addressable
+ resources (`flang/lib/Optimizer/Analysis/AliasAnalysis.cpp`, `getModRef`),
+ and MLIR CSE ignores writes to resources disjoint from a load's resource.
+ So asserts do not block load CSE, store-to-load forwarding, or FIR LICM of
+ loads. This also means `SimplifyHLFIRIntrinsics` can emit asserts in its
+ first run (`allowNewSideEffects=false`, `SHI:3321-3333`): the concern
+ there is new effects that stop CSE from merging loads, which this op does
+ not have.
+- Not speculatable.
+- **Not transparent to other operations' folds and patterns.** The assert's
+ own canonicalization removes it when `%ok` is a constant true (section 3.4).
+ But no other rewrite may look through an assert to its guarded operand,
+ because that drops the dependence. For example, in
+ `%a = fir.box_addr %b; %a_ok = fir.assert ... guard(%a); %c = fir.embox %a_ok`,
+ a fold that rewrites an `fir.embox` of a `fir.box_addr` into the original
+ box must not replace `%c` with `%b`. In practice:
+ - the op does not implement `ViewLikeOpInterface`;
+ - it is not added to the helpers that skip `fir.declare` or `fir.convert`
+ to find a defining op;
+ - it is not given a folder that returns a guarded operand when `%ok` is
+ unknown.
+
+ Alias analysis may look through the assert, since it only answers
+ aliasing queries and does not rewrite uses.
+
+### 3.3 TableGen sketch
+
+```
+def fir_AssertOp : fir_Op<"assert", [AttrSizedOperandSegments,
+ MemoryEffects<[MemWrite<FortranRuntimeResource>]>]> {
+ let arguments = (ins I1:$condition,
+ fir_RuntimeCheckKindAttr:$kind,
+ StrAttr:$message,
+ Variadic<AnySignlessIntegerOrIndex>:$values,
+ Variadic<AnyType>:$guarded);
+ let results = (outs Variadic<AnyType>:$results); // types == guarded types
+ let hasVerifier = 1;
+ let hasCanonicalizer = 1;
+}
+```
+
+### 3.4 Canonicalization and redundancy elimination
+
+- **Constant true:** erase the assert and replace its results with the guarded
+ operands. This replaces the hand-written `fir::getIntIfConstant` early-outs
+ of today's checks (`IC:6929-6931`). It must be a canonicalization pattern
+ rather than only a `fold`: for an assert with no guard, an empty fold result
+ means an in-place update, not an erasure.
+- **Constant false:** keep it, with no diagnostic. A false condition is common
+ in code that is unreachable but hard to prove unreachable at compile time
+ (for example, after inlining or specialization on constant arguments), so a
+ warning would be noisy.
+
+ A later pass could erase the code strictly dominated by a constant-false
+ assert. This is not proposed for the first version, and it must not run in
+ any pipeline that may drop asserts (for example, device code with checks
+ disabled). There, the assert would disappear together with the pruned
+ code, and a program that should have stopped would silently produce
+ wrong results.
+- **Dominating assert with the same condition:** remove the later one, mapping
+ its results to the earlier one's results. This is valid regardless of what
+ lies between them and regardless of the message, because the earlier check
+ fails first. Only remove it when the earlier assert guards everything the
+ later one guards; otherwise the consumers would lose their dependence.
+ Upstream CSE cannot do this: it skips ops with write effects
+ (`mlir/lib/Transforms/Utils/CSE.cpp`, `simplifyOperation`). So this needs a
+ small dominance-based pass (section 6).
+
+### 3.5 Lowering
+
+Lowering is done in `FIRToLLVMLowering`, where the IR is already a control-flow
+graph:
+
+```
+// fir.assert %ok, conformance, "..." values(%a, %b : index, index) guard(%n : index)
+llvm.cond_br %ok weights([2000, 1]), ^cont, ^fail
+^fail:
+ llvm.call @_FortranAReportFatalUserErrorValues(%msg, %file, %line, %a64, %b64, ...)
+ llvm.unreachable
+^cont: // uses of the result are replaced with %n
+```
+
+- **Runtime entry.** `_FortranAReportFatalUserError(msg, source, line)`
+ (`flang/include/flang/Runtime/stop.h`) is `[[noreturn]]` and device-callable
+ (`RT_API_ATTRS`). It takes no values. Proposed addition:
+ `ReportFatalUserErrorValues(msg, source, line, int64 v0..v3)` with a fixed
+ number of `int64_t` arguments. That keeps the ABI simple on devices, and
+ messages use `%jd` only. The message is passed to `printf`, so a literal
+ `%` in a message must be escaped.
+- **Declaration.** The runtime function declaration must carry `noreturn`,
+ so LLVM can use the checked condition on the continuation path.
+- **Source location.** Take it from the assert's own location. Resolve
+ `FusedLoc` and `CallSiteLoc` to the innermost `FileLineColLoc`: today
+ `fir::factory::locationToFilename` returns null for them, so inlined checks
+ lose file and line.
+- **Device code.** String globals must be created in the nearest symbol table
+ (a `gpu.module` for device code), and the lowering must also exist in any
+ device code generation pipeline. If a device runtime lacks the entry point,
+ fall back to `printf` plus `llvm.trap`. The function is not in the OpenMP
+ offload API group today, unlike `Abort`.
+
+Lowering before `CFGConversion` (to `fir.if` plus call) is also possible, but
+`fir.if` regions cannot end with `fir.unreachable`, and asserts would stop
+being visible to the passes after that point.
+
+### 3.6 Alternatives considered
+
+**`fir.if` plus a runtime call (today's pattern).** Not retained, for the
+reasons in section 1: there is no data dependence on the protected operation,
+the call's unknown memory effects block alias-based optimizations, and LLVM
+is not told that the call never returns.
+
+**`cf.assert`.** Not retained. It writes the default resource, which is how
+it models program termination, so it blocks store-to-load forwarding and
+load CSE across it. It still does not order non-speculatable operations
+without memory effects, such as `arith.remsi`, after it. Its LLVM lowering
+prints with `puts` and calls `abort`, rather than reporting through the
+Fortran runtime. It also has no guarded results.
+
+**A result-less `fir.assert` (the current revision of
+[#223287](https://github.com/llvm/llvm-project/pull/223287)).** The PR
+defines `fir.assert` as a non-speculatable op with a
+`MemWrite<FortranRuntimeResource>` effect, lowers it through `cf.assert`, and
+teaches Flang LICM not to hoist non-speculatable operations across a
+preceding assert in the same block. This design keeps that effect model
+(section 3.2), which gives the right DCE, alias-analysis, and forwarding
+behavior. It is not retained as is, because the ordering of the protected
+operations is not expressed in the IR:
+
+- MLIR has no notion of an operation that may not return. Every pass may
+ assume that all operations in a block execute together, and may move a
+ `fir.load` or an `arith.remsi` above an assert whose effects do not
+ conflict with it. The LICM change protects one pass. Every other pass that
+ moves or recreates operations would need the same fix, and future passes
+ would have to remember it. Examples are sinking, scheduling, and rewrite
+ patterns that build new operations at a different insertion point.
+- Once a protected load or remainder is above the check, LLVM may infer from
+ its undefined behavior that the check cannot fail, and delete it.
+- The guarded results make the ordering a matter of SSA dominance, so no pass
+ has to know about `fir.assert`. The lowering then goes directly to the
+ Fortran runtime entry, so messages match the runtime's format (section
+ 3.5).
+
+The design in this document is therefore the PR's op, with guarded results
+and a direct lowering added.
+
+**A region-based assert, like `shape.assuming`.** The protected operations go
+inside a region attached to the assert, and the values used afterwards are
+yielded out:
+
+```
+%r = fir.assert %ok, conformance, "..." -> (f32) {
+ // inlined DOT_PRODUCT loop nest
+ fir.result %sum : f32
+}
+```
+
+Ordering is structural, so it is as safe as the guarded-result form. It is
+not retained because the region gets in the way of the optimizations the
+checks must not block:
+
+- The protected code is usually a whole loop nest or the rest of a
+ statement, so everything it defines has to be yielded through the region.
+ An allocation-status check would wrap the rest of the block.
+- Passes that work on a single block or match operation sequences see a
+ region boundary instead. Examples are the `OptimizedBufferization` pattern
+ that matches an `hlfir.elemental` and its `hlfir.assign`, store-to-load
+ forwarding, and LICM of loads out of the region.
+- Merging two redundant asserts means merging or inlining regions. Hoisting
+ one means moving its whole region, or splitting it.
+- Several checks on the same statement nest regions inside each other.
+
+The guarded-result form gives the same guarantee at the granularity of a
+value: guarding a loop's extent protects every access in the loop, without
+moving the loop into a region.
+
+## 4. Check categories and default policy
+
+| Kind | Meaning | Default | Option |
+|---|---|---|---|
+| `status` | Allocation status required by the standard (ALLOCATE/DEALLOCATE, assignment to an unallocated or disassociated left-hand side) and allocation failure | Always on, all optimization levels | Engineering option to drop, for staging |
+| `conformance` | Shape or size agreement that the runtime checks when not inlined | On (same behavior as `-O0`) | `-fno-check=conformance` (name to be decided) |
+| `argument` | Intrinsic argument values the runtime checks (DIM range, RESHAPE SHAPE values) | On | same mechanism |
+| `bounds` | Subscript, section, substring, and character-dummy-length checks | Off | `-fcheck=bounds` |
+| `pointer` | Dereference of a disassociated POINTER, an unallocated ALLOCATABLE, or an absent OPTIONAL | Off | `-fcheck=pointer` |
+| `do` | DO loop with zero step | Off | `-fcheck=do` |
+| `bits` | Bit-intrinsic position and shift ranges | Off | `-fcheck=bits` |
+| `divide` | Integer MOD/MODULO with `P==0` | Off | existing `-fcheck-integer-mod-zero-divisor` |
+
+Gating:
+
+- Default-on checks are always emitted. The assert lowering (or an earlier
+ cleanup pass) drops them by category. Dropping replaces the results with the
+ guarded operands.
----------------
tblah wrote:
The opt-in checks are decided by flags so we know everywhere in the pipeline whether the assert will be enabled or not. Why not skip emitting the disabled asserts in the first place instead of adding a cleanup pass?
https://github.com/llvm/llvm-project/pull/228502
More information about the flang-commits
mailing list