[llvm] [MLGO][Docs] Document EmitC based model embedding (PR #227536)
Bhavesh M via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 29 19:28:20 PDT 2026
https://github.com/beamandala created https://github.com/llvm/llvm-project/pull/227536
Documents how to embed TOSA models through the MLIR → EmitC → C++ pipeline.
>From 103cff57d0d02b77dc00ade3238fa78c8689e980 Mon Sep 17 00:00:00 2001
From: Bhavesh Mandalapu <mandalapubhavesh at gmail.com>
Date: Tue, 29 Sep 2026 21:26:35 -0500
Subject: [PATCH] [MLGO][Docs] Document EmitC based model embedding
---
llvm/docs/MLGO.md | 64 ++++++++++++++++++++++++++++++++++++++---------
1 file changed, 52 insertions(+), 12 deletions(-)
diff --git a/llvm/docs/MLGO.md b/llvm/docs/MLGO.md
index 9e98713fc1ae1..825d3b1cd05e1 100644
--- a/llvm/docs/MLGO.md
+++ b/llvm/docs/MLGO.md
@@ -297,7 +297,7 @@ features.
#### `MLModelRunner` implementations
-We currently feature 4 implementations:
+We currently feature 5 implementations:
- `ModelUnderTrainingRunner`. This requires the compiler be built with TFLite
support. It allows loading a TFLite model dynamically and is primarily
@@ -311,10 +311,11 @@ We currently feature 4 implementations:
the neural network, together with its weights (essentially, loops performing
matrix multiplications)
-:::{note}
-we are actively working on replacing this with an EmitC implementation
-requiring no out of tree build-time dependencies.
-:::
+
+- `EmitCModelRunner`. This is another inference implementation. At build time,
+ an MLIR pipeline lowers a model expressed in TOSA through EmitC to a C++
+ header, which is compiled into LLVM. It does not require TensorFlow at build
+ time. See {ref}`embed-tosa-models` for configuration and model selection.
- `InteractiveModelRunner`. This is intended for training scenarios where the
training algorithm drives compilation. This model runner has no special
@@ -655,7 +656,52 @@ For up to date information on custom builds, see the `ml-*`
### Embed pre-trained models (aka "release" mode)
-This supports the `ReleaseModeModelRunner` model runners.
+Release mode supports two ways to embed models at build time: the MLIR-based
+TOSA-to-EmitC path and the existing TensorFlow AOT path. The MLIR path is an
+alternative today and is intended to become the only supported embedding path
+in the future. The two paths cannot be enabled in the same build.
+
+(embed-tosa-models)=
+#### Embed TOSA models with MLIR and EmitC
+
+Supply TOSA models as MLIR files and build `mlir-opt` and `mlir-translate`
+before configuring LLVM. CMake runs the MLIR lowering pipeline, translates the
+resulting EmitC to C++ headers, and compiles those headers into LLVM. The model
+is then available without a TensorFlow runtime dependency.
+
+:::{warning}
+Textual TOSA models produced by `tosa-converter-for-tflite` currently use a
+syntax that this MLIR pipeline does not accept. Until the converter is updated,
+export the model as MLIR bytecode (`.bc`), then run
+`mlir-opt model.bc -o model.mlir` to produce a compatible text MLIR file.
+:::
+
+Set `LLVM_MLGO_MODELS` to a semicolon-separated list of entries in the form
+`<name>,<path-to-model.mlir>,<type>`. The name is the value of the runtime model
+selection flag; the type is `inliner` or `regalloc`. Paths may be absolute or
+relative to the source directory for the corresponding LLVM library
+(`llvm/lib/Analysis` for `inliner`, `llvm/lib/CodeGen` for `regalloc`). For
+example:
+
+```console
+cmake -DLLVM_MLGO_MODELS="size,/absolute/path/to/inliner.mlir,inliner;evict,/absolute/path/to/regalloc.mlir,regalloc" \
+ -DLLVM_MLGO_MLIR_OPT=/absolute/path/to/mlir-opt \
+ -DLLVM_MLGO_MLIR_TRANSLATE=/absolute/path/to/mlir-translate \
+ <...other options...>
+```
+
+`LLVM_MLGO_MLIR_OPT` and `LLVM_MLGO_MLIR_TRANSLATE` default to `mlir-opt` and
+`mlir-translate`, respectively. Set them to the paths of the built tools when
+using a separate MLIR build.
+
+At runtime, select an embedded inliner model with `-mllvm -mlgo-model=<name>`
+and an embedded register allocation eviction model with
+`-mllvm -regalloc-mlgo-model=<name>`. Enable the corresponding release mode
+advisor with `-mllvm -enable-ml-inliner=release` or
+`-mllvm -regalloc-evict-advisor=release`. The `default` model choice uses the
+standard heuristic.
+
+#### Embed TensorFlow Saved Models
You need a tensorflow pip package for the AOT (ahead-of-time) Saved Model compiler
and a thin wrapper for the native function generated by it. We currently support
@@ -684,12 +730,6 @@ You can also specify a URL for the path, and it is also possible to pre-compile
the header and object and then just point to the precompiled artifacts. See for
example `LLVM_OVERRIDE_MODEL_HEADER_INLINERSIZEMODEL`.
-:::{note}
-We are transitioning away from the AOT compiler shipping with the
-tensorflow package, and to a EmitC, in-tree solution, so these details will
-change soon.
-:::
-
### Using TFLite (aka "development" mode)
This supports the `ModelUnderTrainingRunner` model runners.
More information about the llvm-commits
mailing list