[flang-commits] [flang] [llvm] [flang-rt] Prototype: skip copy-out into read-only memory via a process memory-map snapshot (PR #223002)

Eugene Epshteyn via flang-commits flang-commits at lists.llvm.org
Fri Sep 11 10:59:30 PDT 2026


eugeneepshteyn wrote:

## Design points

- **Fail-closed everywhere**: unsupported platform, parse anomaly, allocation
  failure, degenerate descriptor, partial containment, failed confirm — all
  answer "not read-only" and the regular copy-out runs. Initialization is a
  non-blocking atomic state machine (Uninitialized/Building/Ready/Inert);
  allocation failure makes the feature permanently inert rather than
  terminating (no `AllocateMemoryOrCrash` in the module).
- **Host-only**: compiled out of every device path (CUDA/OpenMP offload and
  native GPU); device copy-out behavior is bit-identical to #222101.
- **Documented staleness trade-off** (why this is a *compatibility mode*, not
  a proven-safe optimization): in mode 1 a mapping whose protection changes
  after the snapshot is not seen — a region that *became* read-only is simply
  not recognized (regular copy-out, today's behavior), and a formerly
  read-only region that became writable is still skipped (the copy-out is
  lost). Mode 2 narrows that window with per-hit confirmation but is not
  atomic across VMAs. Programs that `mprotect`/remap regions used as
  copy-out destinations mid-run should not enable this.
- The span check is stride-sign-aware and overflow-checked; the whole
  destination span must be contained in a read-only region.

## Testing

- `check-flang-rt` green on x86-64 and aarch64 Linux (the new unit tests add
  parser fault-injection — no trusted prefix is ever published on any parse
  anomaly — span/containment/classification units, and six
  subprocess-isolated behavioral arms, including both staleness directions
  and a death test proving the feature-off behavior still faults).
- The Windows enumeration/classification is written to the documented API
  contract and unit-tested through a host-portable classifier
  (write-copy protections and guard pages are never treated as read-only),
  but has **not** been run on Windows.

## Measurements (Linux; medians of 9 interleaved reps, pinned core)

- Writable destinations (the common path; the table always misses): **no
  regression** across a size × modified-percentage matrix on either
  architecture (worst cell within noise, ≤ +1.5%).
- Read-only destinations (non-contiguous section of a read-only mapping
  passed to an implicit-interface external in a hot loop, callee never
  modifies): mode 1 replaces the O(n) equal-scan with an O(log R) lookup —
  **19–47% faster** than #222101 alone on x86-64, **44–54%** on aarch64
  (Neoverse-N1). Mode 2 measures at parity with #222101: the confirm system
  call is reached only by executions that would otherwise crash, so
  conforming programs pay zero system calls in every mode.

## Known limitations (deliberate, for discussion)

- The silent skip masks invalid writes into read-only-backed destinations
  (that is its purpose — matching the tolerance of compilers that place
  named-constant arrays in writable static storage). The propagate-and-catch
  behavior of #222101 remains the default and is restored by mode 0.
- Array repacking's copy-back (`fir.unpack_array` →
  `ShallowCopyDirect`) is **not** covered by this prototype; with
  `-frepack-arrays`, a read-only-backed actual associated with an INTENT-less
  assumed-shape dummy still faults on the write-back even in modes 1/2. If
  this design moves forward, that entry point can reuse the same helper.
- Snapshot timing is nondeterministic with respect to `dlopen`: read-only
  segments mapped after the first copy-out are not in the trust table
  (fallback = regular copy-out).


https://github.com/llvm/llvm-project/pull/223002


More information about the flang-commits mailing list