[llvm] Add initial version of the opt-llc parameter search scripts (PR #218759)
Giorgi Gvalia via llvm-commits
llvm-commits at lists.llvm.org
Tue Aug 25 12:56:41 PDT 2026
https://github.com/gvalson created https://github.com/llvm/llvm-project/pull/218759
Please see the README.md file for details about the scripts, how to run them and what kind of configuration files they expect.
>From 474afd0d89afb16359771531a49d92da89dd577e Mon Sep 17 00:00:00 2001
From: Giorgi Gvalia <>
Date: Tue, 25 Aug 2026 12:53:13 -0700
Subject: [PATCH] Add initial version of the opt-llc parameter search script
---
.../tools/opt-llc-parameter-search/.gitignore | 3 +
llvm/tools/opt-llc-parameter-search/README.md | 151 +++++
.../tools/opt-llc-parameter-search/autorun.py | 269 +++++++++
.../example-autorun.json | 53 ++
.../example-opt-passes.json | 26 +
.../opt-llc-parameter-search/requirements.txt | 22 +
llvm/tools/opt-llc-parameter-search/run.py | 521 ++++++++++++++++++
7 files changed, 1045 insertions(+)
create mode 100644 llvm/tools/opt-llc-parameter-search/.gitignore
create mode 100644 llvm/tools/opt-llc-parameter-search/README.md
create mode 100755 llvm/tools/opt-llc-parameter-search/autorun.py
create mode 100644 llvm/tools/opt-llc-parameter-search/example-autorun.json
create mode 100644 llvm/tools/opt-llc-parameter-search/example-opt-passes.json
create mode 100644 llvm/tools/opt-llc-parameter-search/requirements.txt
create mode 100755 llvm/tools/opt-llc-parameter-search/run.py
diff --git a/llvm/tools/opt-llc-parameter-search/.gitignore b/llvm/tools/opt-llc-parameter-search/.gitignore
new file mode 100644
index 0000000000000..47ef0caaae4a2
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/.gitignore
@@ -0,0 +1,3 @@
+opt-passes.json
+__pycache__
+.venv
diff --git a/llvm/tools/opt-llc-parameter-search/README.md b/llvm/tools/opt-llc-parameter-search/README.md
new file mode 100644
index 0000000000000..d275a7f3ebb12
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/README.md
@@ -0,0 +1,151 @@
+# LLVM opt and llc parameter search
+
+## Overview
+
+`opt` and `llc` have many hidden flags, shown through the `--help-hidden` option, which can affect performance by adjusting optimization pass parameters. The goal of this repository is to provide scripts that automatically collect data about how different values of a user-defined list of `opt` and `llc` flags affects performance of OpenMP Target GPU kernels. The scripts automatically replay an isolated kernel with a predefined set of `opt`/`llc` flags and their values through AMD's [rocprofv3](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/how-to/using-rocprofv3.html) and collect runtime data into a CSV file.
+
+`autorun.py` is the primary automatic search script, which uses the [hyperopt](https://hyperopt.github.io/) library to automatically choose values for the user-defined flags within the user-defined range. The chosen values are then passed to `llc`, which then modifies the kernel image. The script then runs the modified kernel via rocprofv3 and records runtime data.
+
+`run.py` can be used to run a search through a predefined set of `llc` or `opt` flags and their values. Whereas `autorun.py` automatically selects values, `run.py` requires the user to specify all possible flag values themselves.
+
+## Requirements
+
+- RoCM (rocprofv3)
+- Python 3.10+
+- LLVM
+
+## Setup
+
+### Install Python Dependencies
+
+Make sure that the python version on your system is **at least** 3.10.
+
+Then, install dependencies either globally, using pip:
+
+```bash
+pip install -r requirements.txt
+```
+
+Or into a virtual environment:
+
+```bash
+python -m venv .venv
+source .venv/bin/activate
+pip install -r requirements.txt
+```
+
+### Record your kernel
+
+Firstly, when compiling your application, ensure that the `-fopenmp-target-jit` flag is included in the compilation command. Without a binary compiled with this flag, the scripts won't work. See more information about the flag here: https://openmp.llvm.org/CommandLineArgumentReference.html#fopenmp-target-jit
+
+Assuming that your OpenMP target application is compiled with `-fopenmp-target-jit` and can be run on the GPU, record the kernel by launching your application with the following environment variables:
+
+```bash
+LIBOMPTARGET_RECORD=1 LIBOMPTARGET_RECORD_REPORT=1 LIBOMPTARGET_RECORD_DIR=records ./openmp-app
+```
+
+If all goes well, you should see the output of the kernel recording tool and a new folder called `records`.
+
+> [!IMPORTANT]
+> Make sure that the kernel records directory (which would be `records` if you followed the instruction above) contains a file that has the `.bc` extension. This is the device IR bitcode file, without which the scripts won't work. If you don't see the file, make sure that your binary is compiled with the `-fopenmp-target-jit` option and you're using the latest version of LLVM.
+
+See more about the kernel record and replay process here: https://openmp.llvm.org/design/Runtimes.html#kernel-record-replay
+
+### Create your JSON configuration file
+
+Feel free to copy and modify the example JSON files in this repository. If you're running `autorun.py`, you will need to copy `example-autorun.json`. Else, copy `example-opt-passes.json`. For more details about what each field means, see the [JSON configuration files section](#json-configuration-files).
+
+## Usage
+
+### The autorun script
+
+Assuming that you're on a system with the AMD MI250A GPU, you can run automatic search like this:
+
+```bash
+autorun.py --arch gfx90a records/example.bc example.json
+```
+
+If your system is using a different architecture than `gfx90a`, you need to specify your architecture through the `--arch` argument, otherwise `llc` will return an error. You can find out the correct value for this flag by looking at the output of `rocm-smi --showhw`.
+
+The script will read the JSON configuration file `example.json` and construct a search space, which it will explore for 100 iterations. Feel free to get started by copying and modifying `example-autorun.json`. If a longer run is desired, you can pass the desired number of maximum iteration to the `--max-trials` argument like this:
+
+```bash
+autorun.py --arch gfx90a --max-trials 500 records/example.bc example.json
+```
+
+At the end of the run you should see two new files in your current directory: an `autosearch-results.csv` file that will have the current date prepended to it before the file extension and a `loss_history.png` file with a similar date affix.
+
+You can run `autorun.py --help` to get information about all arguments.
+
+### The run script
+
+The `run.py` script can be invoked like this:
+
+```
+run.py --pipelines-file example-opt-passes.json records/example.bc
+```
+
+By default, the script only transform the device IR bitcode using `opt`. If you also desire to use `llc` to influence the backend, you can use the `--llc` and `--llc-also-use-opt` flags to switch the search to using llc and to pass the bitcode to `opt` first before passing it through `llc`, respectively.
+
+## JSON configuration files
+
+### Autorun script JSON configuration file
+
+The JSON configuration file accepted by `autorun.py` looks like this:
+
+```jsonc
+{
+ "persistent_llc_flags": ["-flag_1", "-flag_2"],
+ "llc_flags": [
+ {
+ "flag": "example",
+ "type": "boolean | value | uint | int | number | string"
+ "range":
+ // For type "boolean"
+ null |
+ {
+ // For types "value" and "string"
+ "choice": ["A", "B", "C"] |
+ // For types "uint", "int", "number"
+ "low": 1,
+ "high": 1e6,
+ "step": 1e2,
+ "distribution": "logarithmic" | "uniform"
+ }
+ }
+ ]
+}
+```
+
+`persistent_llc_flags` is a list of flags that must persist across all trials, such as `-O3`. `llc_flags` is a list of flags with a type and range. This list defines a search space that is automatically explored as the script runs.
+
+There are a few things to take into account when assembling your configuration file:
+
+1. The flag names in the `flag` field are given without hyphens, i.e. `--unroll-threshold` should be represented as `"unroll-threshold"`.
+2. For flags of type `uint`, `int` and `number`, the fields `low`, `high`, `step`, and `distribution` in the `range` object are required.
+3. For flags of type `value` and `string` the field `choice` in the `range` object is required.
+4. If you're using the logarithmic distribution, make sure to not include the value of 0 in the `low` or `high` fields as log(0) = ∞.
+
+### Run script JSON configuration file
+
+The JSON configuration file accepted by `run.py` looks like this:
+
+```
+{
+ "pipeline_1": {
+ "opt_passes": "",
+ "opt_args": ["-O1"],
+ "llc_args": ["-O1"],
+ "replay_args": ["--repetitions", "20"]
+ }
+}
+```
+
+Each pipeline is defined as a `"name": object` pair. In the object:
+
+- `opt_passes` is a string that gets passed to `opt -passes=""`
+- `opt_args` is a list of arguments that gets directly added to the `opt` invocation.
+- `llc_args` is the same but for `llc`
+- `replay_args` is the same but for the `llvm-omp-kernel-replay` tool that is used to launch kernels.
+
+Keep in mind that if you want to specify any flag that takes a value, you should separate the flag name and value like this: `["flag_name", "value"]`. If you don't want to have any arguments for opt, llc, or kernel replay, leave the list empty; do not delete the field as it will result in an error.
diff --git a/llvm/tools/opt-llc-parameter-search/autorun.py b/llvm/tools/opt-llc-parameter-search/autorun.py
new file mode 100755
index 0000000000000..cf2d332d79688
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/autorun.py
@@ -0,0 +1,269 @@
+#!/usr/bin/env python3
+"""
+Run an automatic search to find the fastest paramater values for LLVM
+optimization passes.
+"""
+
+import argparse
+import csv
+import datetime
+import json
+import sys
+import time
+from statistics import median
+from typing import Any, Never
+import matplotlib
+import matplotlib.pyplot as plt
+import numpy as np
+from hyperopt import hp, fmin, Trials, STATUS_OK, STATUS_FAIL
+import run
+
+bitcode_file_name: str = ""
+REPETITIONS: int = 25
+arch:str = ""
+persistent_llc_flags: list[str] = []
+
+def json_to_space(json_path: str) -> dict[str, Any] | Never:
+ """
+ Parses the JSON file at JSON_PATH and transforms it into a space
+ dict used by hyperopt.fmin
+ """
+ space = {}
+ with open(json_path, 'r') as json_file:
+ config = json.load(json_file)
+ # Populate the persistent flags global
+ persistent_llc_flags.extend(config['persistent_llc_flags'])
+ for flag in config['llc_flags']:
+ name = flag['flag']
+ match flag['type']:
+ case "uint" | "int" | "number":
+ value_range = flag['range']
+ distribution = value_range['distribution']
+ if distribution == "logarithmic":
+ low = value_range['low']
+ high = value_range['high']
+ step = value_range['step']
+ if low == 0 or high == 0:
+ print("0 is an invalid number for a log distribution",
+ file=sys.stderr)
+ print(f"Found in {name}", file=sys.stderr)
+ sys.exit(1)
+ space[name] = hp.qloguniform(name,
+ np.log(low),
+ np.log(high),
+ step)
+ elif distribution == "uniform":
+ low = value_range['low']
+ high = value_range['high']
+ step = value_range['step']
+ space[name] = hp.quniform(name, low, high, step)
+ # maybe: support normal distributions too
+ else:
+ print(f"Distribution {distribution} not supported!",
+ file=sys.stderr)
+ sys.exit(1)
+ case "value" | "string":
+ value_range = flag['range']
+ choice = value_range['choice']
+ space[name] = hp.choice(name, choice)
+ case "boolean":
+ space[name] = hp.choice(name, [True, False])
+ return space
+
+def objective(params: dict[str, Any]) -> dict[str, Any]:
+ """
+ The objective function. Returns an object describing the loss
+ (median kernel runtime), status, and flags.
+ """
+ llc_args: list[str] = []
+ for k, v in params.items():
+ if isinstance(v, bool):
+ # Turn pythonic True/False into true/false
+ llc_args.append(f"--{k}={str(v).lower()}")
+ elif isinstance(v, float):
+ # Make sure that floats that don't have decimal parts get
+ # turned into ints first
+ if v - int(v) == 0:
+ llc_args.append(f"--{k}={str(int(v))}")
+ else :
+ llc_args.append(f"--{k}={v}")
+ else:
+ llc_args.append(f"--{k}={v}")
+
+ # Append persistent flags
+ llc_args.extend(persistent_llc_flags)
+
+ config = {"llc_args": llc_args,
+ "replay_args": [f"--repetitions={REPETITIONS}"],
+ "opt_passes": [],
+ "opt_args": ["--O3"]}
+
+ # NOTE: Since we're already invoking `opt --O3` within the
+ # baseline run, we don't need to repeatedly invoke opt when
+ # calling llc.
+ code = run.run_llc(bitcode_file_name,
+ # Dummy argument
+ "AUTORUN",
+ config,
+ False, # dry_run
+ False, # also_opt
+ arch,
+ False) # verbose
+ if code != 0:
+ sys.exit(code)
+ json_path = run.check_record_json_file(bitcode_file_name)
+ output_dir = datetime.datetime.now().strftime("%Y%m%d_%H%M%S%f")
+ runtimes = run.replay_and_measure_kernel(json_path,
+ config,
+ output_dir,
+ False, # run_bitcode
+ False) # dry_run
+ if runtimes:
+ return {
+ 'loss': median(runtimes),
+ 'eval_time': time.time(),
+ 'baseline': "false",
+ 'config': config,
+ 'status': STATUS_OK
+ }
+ return {
+ 'loss': float("NaN"),
+ 'eval_time': time.time(),
+ 'baseline': "false",
+ 'config': config,
+ 'status': STATUS_FAIL
+ }
+
+def write_trials_csv(baseline: dict[str, Any],
+ trial_results: list[dict[str, Any]]) -> None:
+ """
+ Write the trial results out as a CSV file.
+ """
+ now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+ csv_path = f"autosearch-results-{now}.csv"
+ with open(csv_path, 'w', newline='') as csvfile:
+ fieldnames = ['loss', 'eval_time', 'baseline', 'config', 'status']
+ writer = csv.DictWriter(csvfile, fieldnames=fieldnames,
+ quoting=csv.QUOTE_NONNUMERIC)
+ writer.writeheader()
+ # Write baseline results
+ writer.writerow(baseline)
+ # Write trial results
+ for row in trial_results:
+ writer.writerow(row)
+
+def save_loss_history_plot(baseline_loss: float, trials: Trials) -> None:
+ """
+ Save the loss history + the baseline as a scatterplot. Make a
+ horizontal line from the baseline for easy visual comparison.
+ """
+ now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+
+ losses = list(trials.losses())
+ x = range(len(losses) + 1)
+ y = [ baseline_loss ] + losses
+ plt.scatter(x, y)
+ plt.axhline(baseline_loss, c="g")
+ plt.xlabel("Trial iteration")
+ plt.ylabel("Median runtime (ns)")
+ plt.savefig(f"loss_history-{now}.png", dpi=300, bbox_inches="tight")
+ plt.close()
+
+def main():
+ global bitcode_file_name, arch
+ parser = argparse.ArgumentParser(
+ description="Automatically run parameter search for llc or opt flags."
+ )
+ parser.add_argument(
+ "bc_path",
+ help="Path to the recorded device IR bitcode (.bc) file"
+ )
+ parser.add_argument(
+ "json_path",
+ help="Path to the JSON file containing settings for desired flags"
+ )
+ parser.add_argument(
+ "--arch",
+ default="gfx90a",
+ help="GPU Architecture (default: gfx90a)"
+ )
+ parser.add_argument(
+ "--max-trials",
+ default="100",
+ help="Maximum number of trials (defualt: 100)"
+ )
+ parser.add_argument(
+ "--no-plot",
+ default=False,
+ help="Don't save the median runtime history plot as a PNG image",
+ action='store_true'
+ )
+
+ args = parser.parse_args()
+ trials = Trials()
+
+ arch = args.arch
+ bc_path = args.bc_path
+ json_path = args.json_path
+ max_trials = int(args.max_trials)
+
+ # Back up original bc and image files
+ bitcode_file_name = run.get_original_bitcode(bc_path)
+ run.backup_image(run.image_output_file(bc_path))
+
+ # Run the baseline
+ print("=" * 80)
+ print("Replaying the baseline, unmodified, kernel:")
+ print(f"{'=' * 80}")
+ kernel_json_path = run.check_record_json_file(bitcode_file_name)
+ output_dir = datetime.datetime.now().strftime("%Y%m%d_%H%M%S%f")
+ baseline_config = {
+ "opt_passes": [],
+ "opt_args": ["--O3"],
+ "llc_args": ["-O3"],
+ "replay_args": [f"--repetitions={REPETITIONS}"]
+ }
+ status = run.run_llc(bitcode_file_name,
+ "BASELINE",
+ baseline_config,
+ False, # dry_run
+ True, # also_opt
+ arch,
+ False) # verbose
+ if status != 0:
+ sys.exit(status)
+ baseline_runtimes = run.replay_and_measure_kernel(kernel_json_path,
+ baseline_config,
+ output_dir,
+ False, # run_bitcode
+ False) # dry_run
+ if baseline_runtimes:
+ baseline = {
+ "loss": median(baseline_runtimes),
+ "eval_time": time.time(),
+ "baseline": "true",
+ "config": baseline_config,
+ "status": STATUS_OK
+ }
+ else:
+ print("Could not measure baseline runtimes", file=sys.stderr)
+ sys.exit(1)
+
+ # Initialize search space
+ space = json_to_space(json_path)
+
+ # Run search
+ best = fmin(
+ fn=objective, # Objective Function to optimize
+ space=space, # Hyperparameter Search Space
+ max_evals=max_trials, # Number of optimization attempts
+ trials=trials
+ )
+
+ if not args.no_plot:
+ save_loss_history_plot(baseline["loss"], trials)
+
+ write_trials_csv(baseline, trials.results)
+
+if __name__ == "__main__":
+ main()
diff --git a/llvm/tools/opt-llc-parameter-search/example-autorun.json b/llvm/tools/opt-llc-parameter-search/example-autorun.json
new file mode 100644
index 0000000000000..42483fcf4eec9
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/example-autorun.json
@@ -0,0 +1,53 @@
+{
+ "persistent_llc_flags": ["-O3"],
+ "llc_flags": [
+ {
+ "flag": "amdgpu-1",
+ "type": "uint",
+ "range": {
+ "low": 1,
+ "high": 1e6,
+ "step": 1e2,
+ "distribution": "logarithmic"
+ }
+ },
+ {
+ "flag": "amdgpu-2",
+ "type": "value",
+ "range": {
+ "choice": ["A", "B", "C"]
+ }
+ },
+ {
+ "flag": "amdgpu-3",
+ "type": "boolean"
+ },
+ {
+ "flag": "amdgpu-4",
+ "type": "string",
+ "range": {
+ "choice": ["X", "Y", "Z"]
+ }
+ },
+ {
+ "flag": "amdgpu-5",
+ "type": "int",
+ "range": {
+ "low": 0,
+ "high": 100,
+ "step": 5,
+ "distribution": "uniform"
+ }
+ },
+ {
+ "flag": "amdgpu-6",
+ "type": "number",
+ "range": {
+ "low": -1e6,
+ "high": 1e6,
+ "step": 1e4,
+ "distribution": "uniform"
+ }
+ }
+ ]
+}
diff --git a/llvm/tools/opt-llc-parameter-search/example-opt-passes.json b/llvm/tools/opt-llc-parameter-search/example-opt-passes.json
new file mode 100644
index 0000000000000..a54a0011ad827
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/example-opt-passes.json
@@ -0,0 +1,26 @@
+{
+ "O1": {
+ "opt_passes": "",
+ "opt_args": ["-O1"],
+ "llc_args": ["-O1"],
+ "replay_args": ["--repetitions", "20"]
+ },
+ "O2": {
+ "opt_passes": "",
+ "opt_args": ["-O2"],
+ "llc_args": ["-O2"],
+ "replay_args": ["--repetitions", "20"]
+ },
+ "O3": {
+ "opt_passes": "",
+ "opt_args": ["-O3"],
+ "llc_args": ["-O3"],
+ "replay_args": ["--repetitions", "20"]
+ },
+ "opt-custom-pipeline": {
+ "opt_passes": "no-op-module,cgscc(no-op-cgscc,function(no-op-function,loop(no-op-loop))),function(no-op-function,loop(no-op-loop))",
+ "opt_args": ["--debug-pass-manager"],
+ "llc_args": [],
+ "replay_args": ["--repetitions", "20"]
+ }
+}
diff --git a/llvm/tools/opt-llc-parameter-search/requirements.txt b/llvm/tools/opt-llc-parameter-search/requirements.txt
new file mode 100644
index 0000000000000..e7ed157c3a0f3
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/requirements.txt
@@ -0,0 +1,22 @@
+cloudpickle==3.1.2
+contourpy==1.3.3
+cycler==0.12.1
+fonttools==4.63.0
+hyperopt==0.3.0
+kiwisolver==1.5.0
+matplotlib==3.11.1
+mpi4py==4.0.3
+networkx==3.6.1
+numpy==2.5.2
+packaging==26.3
+pillow==12.3.0
+pyparsing==3.3.2
+python-dateutil==2.9.0.post0
+pytorch-triton-rocm==3.3.0
+scipy==1.18.1
+setuptools==70.2.0
+six==1.17.0
+torch==2.7.0+rocm6.3
+torchaudio==2.7.0+rocm6.3
+torchvision==0.22.0+rocm6.3
+tqdm==4.70.0
diff --git a/llvm/tools/opt-llc-parameter-search/run.py b/llvm/tools/opt-llc-parameter-search/run.py
new file mode 100755
index 0000000000000..ba1b487334dbe
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/run.py
@@ -0,0 +1,521 @@
+#!/usr/bin/env python3
+
+"""
+Run opt or llc on a recorded, unoptimized, kernel IR bitcode (bc) file
+to transform it using the pipeline definitions and opt/llc commandline
+flags read from a user-defined JSON file (see example-opt-passes.json
+in the repo). Replay the transformed kernel and measure its
+performance for each transformation using rocprof, output a CSV file
+and a boxplot of the results.
+"""
+
+import argparse
+import datetime
+import json
+import subprocess
+import sys
+import shutil
+import os
+import csv
+
+from collections import defaultdict
+from typing import Generator, Never, TypeAlias
+
+import matplotlib.pyplot as plt
+
+PipelineConfType: TypeAlias = dict[str, list[str]]
+PipelinesJsonType: TypeAlias = dict[str, PipelineConfType]
+# Timeout for running rocprof
+TIMEOUT: int = 90
+
+def load_pipelines(pipelines_file: str) -> PipelinesJsonType | Never:
+ """
+ Load the pipeline definitions from a JSON file path.
+ """
+ if not os.path.exists(pipelines_file):
+ print(f"Error: '{pipelines_file}' not found in current directory.", file=sys.stderr)
+ sys.exit(1)
+
+ try:
+ with open(pipelines_file, 'r') as f:
+ pipelines = json.load(f)
+
+ if not isinstance(pipelines, dict):
+ print(f"Error: '{pipelines_file}' must contain a JSON object.", file=sys.stderr)
+ sys.exit(1)
+
+ if not pipelines:
+ print(f"Error: '{pipelines_file}' is empty.", file=sys.stderr)
+ sys.exit(1)
+
+ return pipelines
+
+ except json.JSONDecodeError as e:
+ print(f"Error: Failed to parse '{pipelines_file}': {e}", file=sys.stderr)
+ sys.exit(1)
+ except Exception as e:
+ print(f"Error: Failed to read '{pipelines_file}': {e}", file=sys.stderr)
+ sys.exit(1)
+
+def backup_bitcode(bc_path: str) -> str:
+ """
+ Make sure that the bitcode file at bc_path is backed up. A backup is a copy
+ of the .bc file with '.original' appended at the end of the file name. If
+ the backup file already exists or the user already invokes the script with
+ the .bc.original file, do nothing.
+ """
+ if bc_path.endswith(".original"):
+ return bc_path
+ bc_backup_path = bc_path + ".original"
+ if not os.path.exists(bc_backup_path):
+ print(f"Backing up the bitcode file to {bc_backup_path}")
+ shutil.copyfile(bc_path, bc_backup_path)
+ return bc_backup_path
+ return bc_backup_path
+
+def backup_image(image_path: str) -> None:
+ """
+ Like backup_bitcode() but for the image file.
+ """
+ image_backup_path = image_path + ".original"
+ if os.path.exists(image_backup_path):
+ return
+ print(f"Backing up the image file to {image_backup_path}")
+ shutil.copyfile(image_path, image_backup_path)
+
+def get_original_bitcode(bc_path: str) -> str | Never:
+ """
+ Always return the path to the original bitcode file, calling
+ backup_bitcode() if it doesn't yet exist. Warning: we at the moment
+ cannot verify that the file ending in '.original' is, in fact, the
+ original.
+ """
+ if not os.path.exists(bc_path):
+ print("Error: the specified bitcode file could not be found", file=sys.stderr)
+ sys.exit(1)
+ if bc_path.endswith(".original"):
+ return bc_path
+ return backup_bitcode(bc_path)
+
+def check_record_json_file(bitcode_file: str) -> str | Never:
+ """
+ Return the path to the JSON file corresponding to the kernel's
+ recording or exit with an error.
+ """
+ json_path = bitcode_file.replace('.bc', '.json').replace('.original', '')
+ if not os.path.exists(json_path):
+ print(f"Error: could not find a JSON file corresponding to the specified bitcode file path: {json_path}", file=sys.stderr)
+ sys.exit(1)
+ return json_path
+
+def opt_output_file(bitcode_file: str) -> str:
+ """
+ Strip the '.original' suffix from the IR bitcode file path if it
+ exists. Necessary due to the fact that the replay tool expects the
+ original '.bc' file extension instead of '.bc.original'.
+ """
+ return bitcode_file.replace(".original", "") if bitcode_file.endswith(".original") else bitcode_file
+
+def llc_output_file(bitcode_file: str) -> str:
+ """
+ Strip the '.original' suffix from the IR bitcode file path if it
+ exists and replace it with '.o'.
+ """
+ output_file = bitcode_file
+ if output_file.endswith(".original"):
+ output_file = output_file.replace(".original", "")
+ return output_file.replace(".bc", ".o")
+
+def image_output_file(bitcode_file: str) -> str:
+ """
+ Strip the '.original' suffix from the IR bitcode file path if it
+ exists and replace it with '.image'.
+ """
+ output_file = bitcode_file
+ if output_file.endswith(".original"):
+ output_file = output_file.replace(".original", "")
+ return output_file.replace(".bc", ".image")
+
+def cleanup_modified_files(bc_path: str) -> None:
+ """
+ Restore the original .image and .bc files. If kernel recording
+ files ending with 'bc.original' and '.image.original' exist,
+ overwrite the
+ """
+ if bc_path.endswith(".bc.original"):
+ bc_original_path = bc_path
+ bc_originalless_path = bc_path.replace(".original", "")
+ elif bc_path.endswith(".bc"):
+ bc_original_path = bc_path + ".original"
+ bc_originalless_path = bc_path
+ else:
+ print(f"Error: path must end with .bc(.original). Path: {bc_path}", file=sys.stderr)
+ sys.exit(1)
+
+ image_originalless_path = image_output_file(bc_path)
+ image_original_path = image_originalless_path + ".original"
+
+ if os.path.exists(bc_original_path):
+ print(f"Cleaning up the file at {bc_originalless_path}, restoring from {bc_original_path}")
+ os.rename(bc_original_path, bc_originalless_path)
+ else:
+ print(f"Warning: couldn't find the file {bc_original_path}. If this is not your first run, make sure that your bitcode file is original.", file=sys.stderr)
+
+ if os.path.exists(image_original_path):
+ os.rename(image_original_path, image_originalless_path)
+ print(f"Cleaning up the file at {image_originalless_path}, restoring from {image_original_path}")
+ else:
+ print(f"Warning: couldn't find the file {bc_original_path}. If this is not your first run, make sure that your image file is original.", file=sys.stderr)
+
+def run_opt(bitcode_file: str, pipeline_name: str, pipeline_config: PipelineConfType, dry_run: bool, verbose: bool) -> int:
+ """
+ Run `opt` on bitcode_file, reading the flags and custom pipeline
+ data from pipeline_config. If dry_run is True, output a stub
+ message and exit. If verbose is True, pass '-debug-pass-manager'
+ to `opt`.
+ """
+ if dry_run:
+ print(f"Simulating run of pipeline {pipeline_name} for bitcode file {bitcode_file}.")
+ return 0
+
+ print(f"{'=' * 80}")
+ print(f"Running pipeline: {pipeline_name}")
+ print(f"{'=' * 80}")
+
+ output_name = opt_output_file(bitcode_file)
+ opt_passes = pipeline_config['opt_passes']
+ opt_args = pipeline_config['opt_args']
+
+ # Using any() here makes sure that arrays with a single empty
+ # string are treated as empty.
+ if not any(opt_passes) and not any(opt_args):
+ print(f"Opt passes and args not specified for pipeline {pipeline_name}!", file=sys.stderr)
+ return 2
+
+ if not bitcode_file.endswith(".original"):
+ print(f"Bitcode file {bitcode_file} does not have the .original extension, exiting to prevent accidental reuse of an already optimized bitcode", file=sys.stderr)
+ return 3
+
+ cmd = ["opt"]
+ if opt_passes:
+ cmd.append(f"-passes={opt_passes}")
+ if verbose:
+ cmd.append("-debug-pass-manager")
+ cmd.append("-o")
+ cmd.append(output_name)
+ cmd = cmd + opt_args
+ cmd.append(bitcode_file)
+
+ print("Running " + ' '.join(cmd))
+
+ try:
+ result = subprocess.run(cmd, capture_output=True, text=True)
+ print(result.stdout)
+ if result.stderr:
+ print(result.stderr, file=sys.stderr)
+ return result.returncode
+ except FileNotFoundError:
+ print("Error: 'opt' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+ return 1
+ except Exception as e:
+ print(f"Error running pipeline: {e}", file=sys.stderr)
+ return 1
+
+def run_llc(bitcode_file: str, pipeline_name: str, pipeline_config: PipelineConfType, dry_run: bool, also_opt: bool, arch: str, verbose: bool) -> int:
+ """
+ Run `llc` on bitcode_file, reading the flags from pipeline_config.
+ If dry_run is True, output a stub message and exit. If also_opt is
+ True, pass the bitcode file through `opt` first by invoking
+ run_opt(). The arch argument is necessary. If verbose is True,
+ pass '-debug-pass-manager' to `opt`.
+ """
+ if dry_run:
+ print(f"Simulating run of backend pipeline {pipeline_name} for bitcode file {bitcode_file}.")
+ return 0
+ if also_opt:
+ print(f"Running backend pipeline: {pipeline_name} with opt first")
+ else:
+ print(f"{'=' * 80}")
+ print(f"Running backend pipeline: {pipeline_name}")
+ print(f"{'=' * 80}")
+ if not bitcode_file.endswith(".original"):
+ print(f"Bitcode file {bitcode_file} does not have the .original extension, use the --llc-also-use-opt flag if you want to pass an already modified bitcode file to llc")
+ return 3
+ if also_opt:
+ status = run_opt(bitcode_file,
+ pipeline_name,
+ pipeline_config,
+ dry_run,
+ verbose)
+ if status > 0:
+ return status
+
+ llc_output_name = llc_output_file(bitcode_file)
+ image_output_name = image_output_file(bitcode_file)
+ llc_args = pipeline_config['llc_args']
+
+ if not any(llc_args):
+ print(f"llc args not specified for pipeline {pipeline_name}!", file=sys.stderr)
+ return 2
+
+ cmd = ["llc", "-mtriple=amdgcn-amd-amdhsa", "-filetype=obj"]
+ if verbose:
+ cmd.append("-debug-pass-manager")
+ cmd.append(f"-mcpu={arch}")
+ cmd.append("-o")
+ cmd.append(llc_output_name)
+ cmd = cmd + llc_args
+ if also_opt:
+ cmd.append(opt_output_file(bitcode_file))
+ else:
+ cmd.append(bitcode_file)
+
+ link_cmd = ["ld.lld", "-flavor", "gnu", "-shared"]
+ link_cmd.append("-o")
+ link_cmd.append(image_output_name)
+ link_cmd.append(llc_output_name)
+
+ print("Running " + ' '.join(cmd))
+
+ try:
+ # Run llc
+ result = subprocess.run(cmd, capture_output=True, text=True)
+ print(result.stdout)
+ if result.stderr:
+ print(result.stderr, file=sys.stderr)
+ if result.returncode > 0:
+ return result.returncode
+ # Run the linker
+ print("...and " + ' '.join(link_cmd))
+ link_result = subprocess.run(link_cmd, capture_output=True, text=True)
+ print(link_result.stdout)
+ if link_result.stderr:
+ print(link_result.stderr, file=sys.stderr)
+ return link_result.returncode
+ except FileNotFoundError as e:
+ print("Error: 'llc' or 'ld.lld' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+ print(e)
+ return 1
+ except Exception as e:
+ print(f"Error running pipeline: {e}", file=sys.stderr)
+ return 1
+
+def run_kernel_replay(json_path: str, pipeline_config: PipelineConfType, rp_output_dir: str, run_bitcode: bool, dry_run: bool) -> int:
+ """
+ Run the LLVM OpenMP kernel replay tool through rocprofv3, using
+ the JSON file at json_path and specifying the profiler output
+ directory rp_output_dir. If run_bitcode is true, pass the
+ '--load-bitcode' flag to the tool to make it JIT-compile the IR
+ bitcode file before replaying (necessary when testing opt). If
+ dry_run is True, output a stub message and exit.
+ """
+ if dry_run:
+ print(f"Simulating dry run for kernel replay for JSON file at {json_path}.")
+ return 0
+
+ cmd = [
+ "rocprofv3",
+ "--stats",
+ "--kernel-trace",
+ "--output-directory", rp_output_dir,
+ "--output-file", "output",
+ "--",
+ "llvm-omp-kernel-replay"
+ ]
+
+ if run_bitcode:
+ cmd.append("--load-bitcode")
+ cmd = cmd + pipeline_config['replay_args']
+ cmd.append(json_path)
+
+ print("Running " + ' '.join(cmd))
+
+ try:
+ result = subprocess.run(cmd, capture_output=True, text=True, timeout=TIMEOUT)
+ print(result.stdout)
+ if result.stderr:
+ print(result.stderr, file=sys.stderr)
+ return result.returncode
+ except FileNotFoundError:
+ print("Error: 'llvm-omp-kernel-replay' or 'rocprofv3' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+ return 1
+ except Exception as e:
+ print(f"Error running pipeline: {e}", file=sys.stderr)
+ return 1
+
+def get_rocprof_output_dir(pipeline_name: str) -> str:
+ """
+ Helper function to format a folder name for a rocprofv3 output
+ directory.
+ """
+ now = datetime.datetime.now().strftime("%Y%m%d_%H%M")
+ return f"{now}-{pipeline_name}"
+
+def profiler_csv_reader(rp_output_dir: str) -> Generator[dict[str, str | float], str, None] | None:
+ """
+ Yield a generator that reads line by line from the kernel trace CSV file
+ located at rp_output_dir. The kernel trace CSV file contains
+ information about one kernel launch per line.
+ """
+ csv_path = f"./{rp_output_dir}/output_kernel_trace.csv"
+ try:
+ with open(csv_path, mode='r', newline='', encoding='utf-8') as kernel_trace_csv:
+ kt_reader = csv.DictReader(kernel_trace_csv, quoting=csv.QUOTE_NONNUMERIC)
+ yield from kt_reader
+ except OSError as e:
+ print(f"Error reading CSV results from {rp_output_dir}: {e}", file=sys.stderr)
+
+def replay_and_measure_kernel(json_path: str, pipeline_config: PipelineConfType, output_dir: str, run_bitcode: bool, dry_run: bool) -> list[float] | None:
+ """
+ Replay the specified kernel via calling run_kernel_replay and
+ return a list that contains kernel runtimes in nanoseconds.
+ """
+ if not dry_run:
+ code = run_kernel_replay(json_path, pipeline_config, output_dir, run_bitcode, dry_run)
+ if code != 0:
+ sys.exit(code)
+ csv_reader = profiler_csv_reader(output_dir)
+ runtimes: list[float] = []
+ if csv_reader:
+ for row in csv_reader:
+ end_ts = row['End_Timestamp']
+ start_ts = row['Start_Timestamp']
+ if isinstance(end_ts, float) and isinstance(start_ts, float):
+ runtime = end_ts - start_ts
+ runtimes.append(runtime)
+ else:
+ print(f"'End_Timestamp' or 'Start_Timestamp' fields in {output_dir} are not floats, exiting")
+ sys.exit(1)
+ return runtimes
+
+def write_results(results: list[dict[str, str | float]]) -> None:
+ """
+ Output results as a CSV file.
+ """
+ now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+ csv_path = f"results-{now}.csv"
+ with open(csv_path, 'w', newline='') as csvfile:
+ fieldnames = ['pipeline', 'runtime_ns']
+ writer = csv.DictWriter(csvfile, fieldnames=fieldnames,
+ quoting=csv.QUOTE_NONNUMERIC)
+ writer.writeheader()
+ for row in results:
+ writer.writerow(row)
+ print(f"Successfully wrote results to {csv_path}")
+
+def write_boxplots(results: list[dict[str, str | float]]) -> None:
+ """
+ Output results as boxplot image.
+ """
+ # Convert results to be associative
+ now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+ png_path = f"results-{now}.png"
+
+ results_assoc = defaultdict(list)
+ for row in results:
+ results_assoc[row["pipeline"]].append(row["runtime_ns"])
+
+ data = list(results_assoc.values())
+ plt.figure(figsize=(14, 6))
+ plt.boxplot(data, tick_labels=list(results_assoc.keys()))
+ plt.xticks(rotation=45, ha='right')
+ plt.title("Kernel runtime")
+ plt.savefig(png_path, dpi=400, bbox_inches='tight')
+
+def main() -> None:
+ parser = argparse.ArgumentParser(
+ description="Run LLVM opt or llc with various pass pipelines on a bitcode file."
+ )
+ parser.add_argument(
+ "bitcode_file",
+ help="Path to the bitcode file (.bc)"
+ )
+ parser.add_argument(
+ "--pipelines-file",
+ default="opt-passes.json",
+ help="Path to a JSON file containing pipeline configurations (default: ./opt-passes.json)"
+ )
+ parser.add_argument(
+ "--pipeline",
+ help="Specific pipeline to run (if not specified, runs all pipelines)"
+ )
+ parser.add_argument(
+ "--dry-run",
+ help="Does not actually invoke opt, replay or other tools, useful for testing",
+ action='store_true'
+ )
+ parser.add_argument(
+ "--arch",
+ help="GPU Architecture (default: gfx90a), only used for --llc",
+ default="gfx90a"
+ )
+ parser.add_argument(
+ "--verbose",
+ help="Use the -debug-pass-manager flag when running opt",
+ action='store_true'
+ )
+ parser.add_argument(
+ "--llc",
+ help="Use the llc-based backend optimization instead of opt",
+ action='store_true'
+ )
+ parser.add_argument(
+ "--llc-also-use-opt",
+ help="When doing backend optimization, also use opt first before running llc",
+ action='store_true'
+ )
+ parser.add_argument(
+ "--no-plot",
+ default=False,
+ help="Don't plot the results as a PNG image",
+ action='store_true'
+ )
+
+ args = parser.parse_args()
+ json_path = check_record_json_file(args.bitcode_file)
+ results = []
+ pipelines = load_pipelines(args.pipelines_file)
+ also_opt = args.llc_also_use_opt
+ arch = args.arch
+
+ original_bitcode_file = get_original_bitcode(args.bitcode_file)
+ backup_image(image_output_file(args.bitcode_file))
+
+ if args.pipeline:
+ pipelines_to_run = [(args.pipeline, pipelines[args.pipeline])]
+ else:
+ pipelines_to_run = list(pipelines.items())
+
+ run_bitcode = not args.llc
+
+ for pipeline_name, pipeline_config in pipelines_to_run:
+ if args.llc:
+ returncode = run_llc(original_bitcode_file, pipeline_name,
+ pipeline_config, args.dry_run,
+ also_opt, arch, args.verbose)
+ else:
+ returncode = run_opt(original_bitcode_file, pipeline_name,
+ pipeline_config, args.dry_run,
+ args.verbose)
+ if returncode != 0:
+ print(f"Warning: Pipeline '{pipeline_name}' exited with code {returncode}",
+ file=sys.stderr)
+ continue
+ rp_output_dir = get_rocprof_output_dir(pipeline_name)
+ runtimes = replay_and_measure_kernel(json_path,
+ pipeline_config,
+ rp_output_dir,
+ run_bitcode,
+ args.dry_run)
+ if runtimes:
+ for runtime in runtimes:
+ runtime_dict = {"pipeline": pipeline_name, "runtime_ns": runtime}
+ results.append(runtime_dict)
+
+ if not args.dry_run and returncode == 0:
+ write_results(results)
+ if not args.no_plot:
+ write_boxplots(results)
+
+if __name__ == "__main__":
+ main()
More information about the llvm-commits
mailing list