[llvm] [MLGO][Docs] Document EmitC based model embedding (PR #227536)

Bhavesh M via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 19:28:20 PDT 2026


https://github.com/beamandala created https://github.com/llvm/llvm-project/pull/227536

Documents how to embed TOSA models through the MLIR → EmitC → C++ pipeline.

>From 103cff57d0d02b77dc00ade3238fa78c8689e980 Mon Sep 17 00:00:00 2001
From: Bhavesh Mandalapu <mandalapubhavesh at gmail.com>
Date: Tue, 29 Sep 2026 21:26:35 -0500
Subject: [PATCH] [MLGO][Docs] Document EmitC based model embedding

---
 llvm/docs/MLGO.md | 64 ++++++++++++++++++++++++++++++++++++++---------
 1 file changed, 52 insertions(+), 12 deletions(-)

diff --git a/llvm/docs/MLGO.md b/llvm/docs/MLGO.md
index 9e98713fc1ae1..825d3b1cd05e1 100644
--- a/llvm/docs/MLGO.md
+++ b/llvm/docs/MLGO.md
@@ -297,7 +297,7 @@ features.
 
 #### `MLModelRunner` implementations
 
-We currently feature 4 implementations:
+We currently feature 5 implementations:
 
 - `ModelUnderTrainingRunner`. This requires the compiler be built with TFLite
   support. It allows loading a TFLite model dynamically and is primarily
@@ -311,10 +311,11 @@ We currently feature 4 implementations:
   the neural network, together with its weights (essentially, loops performing
   matrix multiplications)
 
-:::{note}
-we are actively working on replacing this with an EmitC implementation
-requiring no out of tree build-time dependencies.
-:::
+
+- `EmitCModelRunner`. This is another inference implementation. At build time,
+  an MLIR pipeline lowers a model expressed in TOSA through EmitC to a C++
+  header, which is compiled into LLVM. It does not require TensorFlow at build
+  time. See {ref}`embed-tosa-models` for configuration and model selection.
 
 - `InteractiveModelRunner`. This is intended for training scenarios where the
   training algorithm drives compilation. This model runner has no special
@@ -655,7 +656,52 @@ For up to date information on custom builds, see the `ml-*`
 
 ### Embed pre-trained models (aka "release" mode)
 
-This supports the `ReleaseModeModelRunner` model runners.
+Release mode supports two ways to embed models at build time: the MLIR-based
+TOSA-to-EmitC path and the existing TensorFlow AOT path. The MLIR path is an
+alternative today and is intended to become the only supported embedding path
+in the future. The two paths cannot be enabled in the same build.
+
+(embed-tosa-models)=
+#### Embed TOSA models with MLIR and EmitC
+
+Supply TOSA models as MLIR files and build `mlir-opt` and `mlir-translate`
+before configuring LLVM. CMake runs the MLIR lowering pipeline, translates the
+resulting EmitC to C++ headers, and compiles those headers into LLVM. The model
+is then available without a TensorFlow runtime dependency.
+
+:::{warning}
+Textual TOSA models produced by `tosa-converter-for-tflite` currently use a
+syntax that this MLIR pipeline does not accept. Until the converter is updated,
+export the model as MLIR bytecode (`.bc`), then run
+`mlir-opt model.bc -o model.mlir` to produce a compatible text MLIR file.
+:::
+
+Set `LLVM_MLGO_MODELS` to a semicolon-separated list of entries in the form
+`<name>,<path-to-model.mlir>,<type>`. The name is the value of the runtime model
+selection flag; the type is `inliner` or `regalloc`. Paths may be absolute or
+relative to the source directory for the corresponding LLVM library
+(`llvm/lib/Analysis` for `inliner`, `llvm/lib/CodeGen` for `regalloc`). For
+example:
+
+```console
+cmake -DLLVM_MLGO_MODELS="size,/absolute/path/to/inliner.mlir,inliner;evict,/absolute/path/to/regalloc.mlir,regalloc" \
+  -DLLVM_MLGO_MLIR_OPT=/absolute/path/to/mlir-opt \
+  -DLLVM_MLGO_MLIR_TRANSLATE=/absolute/path/to/mlir-translate \
+  <...other options...>
+```
+
+`LLVM_MLGO_MLIR_OPT` and `LLVM_MLGO_MLIR_TRANSLATE` default to `mlir-opt` and
+`mlir-translate`, respectively. Set them to the paths of the built tools when
+using a separate MLIR build.
+
+At runtime, select an embedded inliner model with `-mllvm -mlgo-model=<name>`
+and an embedded register allocation eviction model with
+`-mllvm -regalloc-mlgo-model=<name>`. Enable the corresponding release mode
+advisor with `-mllvm -enable-ml-inliner=release` or
+`-mllvm -regalloc-evict-advisor=release`. The `default` model choice uses the
+standard heuristic.
+
+#### Embed TensorFlow Saved Models
 
 You need a tensorflow pip package for the AOT (ahead-of-time) Saved Model compiler
 and a thin wrapper for the native function generated by it. We currently support
@@ -684,12 +730,6 @@ You can also specify a URL for the path, and it is also possible to pre-compile
 the header and object and then just point to the precompiled artifacts. See for
 example `LLVM_OVERRIDE_MODEL_HEADER_INLINERSIZEMODEL`.
 
-:::{note}
-We are transitioning away from the AOT compiler shipping with the
-tensorflow package, and to a EmitC, in-tree solution, so these details will
-change soon.
-:::
-
 ### Using TFLite (aka "development" mode)
 
 This supports the `ModelUnderTrainingRunner` model runners.



More information about the llvm-commits mailing list