[llvm] Add initial version of the opt-llc parameter search scripts (PR #218759)

Giorgi Gvalia via llvm-commits llvm-commits at lists.llvm.org
Tue Aug 25 12:56:41 PDT 2026


https://github.com/gvalson created https://github.com/llvm/llvm-project/pull/218759

Please see the README.md file for details about the scripts, how to run them and what kind of configuration files they expect.

>From 474afd0d89afb16359771531a49d92da89dd577e Mon Sep 17 00:00:00 2001
From: Giorgi Gvalia <>
Date: Tue, 25 Aug 2026 12:53:13 -0700
Subject: [PATCH] Add initial version of the opt-llc parameter search script

---
 .../tools/opt-llc-parameter-search/.gitignore |   3 +
 llvm/tools/opt-llc-parameter-search/README.md | 151 +++++
 .../tools/opt-llc-parameter-search/autorun.py | 269 +++++++++
 .../example-autorun.json                      |  53 ++
 .../example-opt-passes.json                   |  26 +
 .../opt-llc-parameter-search/requirements.txt |  22 +
 llvm/tools/opt-llc-parameter-search/run.py    | 521 ++++++++++++++++++
 7 files changed, 1045 insertions(+)
 create mode 100644 llvm/tools/opt-llc-parameter-search/.gitignore
 create mode 100644 llvm/tools/opt-llc-parameter-search/README.md
 create mode 100755 llvm/tools/opt-llc-parameter-search/autorun.py
 create mode 100644 llvm/tools/opt-llc-parameter-search/example-autorun.json
 create mode 100644 llvm/tools/opt-llc-parameter-search/example-opt-passes.json
 create mode 100644 llvm/tools/opt-llc-parameter-search/requirements.txt
 create mode 100755 llvm/tools/opt-llc-parameter-search/run.py

diff --git a/llvm/tools/opt-llc-parameter-search/.gitignore b/llvm/tools/opt-llc-parameter-search/.gitignore
new file mode 100644
index 0000000000000..47ef0caaae4a2
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/.gitignore
@@ -0,0 +1,3 @@
+opt-passes.json
+__pycache__
+.venv
diff --git a/llvm/tools/opt-llc-parameter-search/README.md b/llvm/tools/opt-llc-parameter-search/README.md
new file mode 100644
index 0000000000000..d275a7f3ebb12
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/README.md
@@ -0,0 +1,151 @@
+# LLVM opt and llc parameter search
+
+## Overview
+
+`opt` and `llc` have many hidden flags, shown through the `--help-hidden` option, which can affect performance by adjusting optimization pass parameters. The goal of this repository is to provide scripts that automatically collect data about how different values of a user-defined list of `opt` and `llc` flags affects performance of OpenMP Target GPU kernels. The scripts automatically replay an isolated kernel with a predefined set of `opt`/`llc` flags and their values through AMD's [rocprofv3](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/how-to/using-rocprofv3.html) and collect runtime data into a CSV file.
+
+`autorun.py` is the primary automatic search script, which uses the [hyperopt](https://hyperopt.github.io/) library to automatically choose values for the user-defined flags within the user-defined range. The chosen values are then passed to `llc`, which then modifies the kernel image. The script then runs the modified kernel via rocprofv3 and records runtime data.
+
+`run.py` can be used to run a search through a predefined set of `llc` or `opt` flags and their values. Whereas `autorun.py` automatically selects values, `run.py` requires the user to specify all possible flag values themselves.
+
+## Requirements
+
+- RoCM (rocprofv3)
+- Python 3.10+
+- LLVM
+
+## Setup
+
+### Install Python Dependencies
+
+Make sure that the python version on your system is **at least** 3.10.
+
+Then, install dependencies either globally, using pip:
+
+```bash
+pip install -r requirements.txt
+```
+
+Or into a virtual environment:
+
+```bash
+python -m venv .venv
+source .venv/bin/activate
+pip install -r requirements.txt
+```
+
+### Record your kernel
+
+Firstly, when compiling your application, ensure that the `-fopenmp-target-jit` flag is included in the compilation command. Without a binary compiled with this flag, the scripts won't work. See more information about the flag here: https://openmp.llvm.org/CommandLineArgumentReference.html#fopenmp-target-jit
+
+Assuming that your OpenMP target application is compiled with `-fopenmp-target-jit` and can be run on the GPU, record the kernel by launching your application with the following environment variables:
+
+```bash
+LIBOMPTARGET_RECORD=1 LIBOMPTARGET_RECORD_REPORT=1 LIBOMPTARGET_RECORD_DIR=records ./openmp-app
+```
+
+If all goes well, you should see the output of the kernel recording tool and a new folder called `records`.
+
+> [!IMPORTANT]
+> Make sure that the kernel records directory (which would be `records` if you followed the instruction above) contains a file that has the `.bc` extension. This is the device IR bitcode file, without which the scripts won't work. If you don't see the file, make sure that your binary is compiled with the `-fopenmp-target-jit` option and you're using the latest version of LLVM.
+
+See more about the kernel record and replay process here: https://openmp.llvm.org/design/Runtimes.html#kernel-record-replay
+
+### Create your JSON configuration file
+
+Feel free to copy and modify the example JSON files in this repository. If you're running `autorun.py`, you will need to copy `example-autorun.json`. Else, copy `example-opt-passes.json`. For more details about what each field means, see the [JSON configuration files section](#json-configuration-files).
+
+## Usage
+
+### The autorun script
+
+Assuming that you're on a system with the AMD MI250A GPU, you can run automatic search like this:
+
+```bash
+autorun.py --arch gfx90a records/example.bc example.json
+```
+
+If your system is using a different architecture than `gfx90a`, you need to specify your architecture through the `--arch` argument, otherwise `llc` will return an error. You can find out the correct value for this flag by looking at the output of `rocm-smi --showhw`.
+
+The script will read the JSON configuration file `example.json` and construct a search space, which it will explore for 100 iterations. Feel free to get started by copying and modifying `example-autorun.json`. If a longer run is desired, you can pass the desired number of maximum iteration to the `--max-trials` argument like this:
+
+```bash
+autorun.py --arch gfx90a --max-trials 500 records/example.bc example.json
+```
+
+At the end of the run you should see two new files in your current directory: an `autosearch-results.csv` file that will have the current date prepended to it before the file extension and a `loss_history.png` file with a similar date affix.
+
+You can run `autorun.py --help` to get information about all arguments.
+
+### The run script
+
+The `run.py` script can be invoked like this:
+
+```
+run.py --pipelines-file example-opt-passes.json records/example.bc
+```
+
+By default, the script only transform the device IR bitcode using `opt`. If you also desire to use `llc` to influence the backend, you can use the `--llc` and `--llc-also-use-opt` flags to switch the search to using llc and to pass the bitcode to `opt` first before passing it through `llc`, respectively.
+
+## JSON configuration files
+
+### Autorun script JSON configuration file
+
+The JSON configuration file accepted by `autorun.py` looks like this:
+
+```jsonc
+{
+  "persistent_llc_flags": ["-flag_1", "-flag_2"],
+  "llc_flags": [
+    {
+      "flag": "example",
+      "type": "boolean | value | uint | int | number | string"
+      "range":
+        // For type "boolean"
+        null |
+        {
+          // For types "value" and "string"
+          "choice": ["A", "B", "C"] |
+          // For types "uint", "int", "number"
+          "low": 1,
+          "high": 1e6,
+          "step": 1e2,
+          "distribution": "logarithmic" | "uniform"
+      }
+    }
+  ]
+}
+```
+
+`persistent_llc_flags` is a list of flags that must persist across all trials, such as `-O3`. `llc_flags` is a list of flags with a type and range. This list defines a search space that is automatically explored as the script runs.
+
+There are a few things to take into account when assembling your configuration file:
+
+1. The flag names in the `flag` field are given without hyphens, i.e. `--unroll-threshold` should be represented as `"unroll-threshold"`.
+2. For flags of type `uint`, `int` and `number`, the fields `low`, `high`, `step`, and `distribution` in the `range` object are required.
+3. For flags of type `value` and `string` the field `choice` in the `range` object is required.
+4. If you're using the logarithmic distribution, make sure to not include the value of 0 in the `low` or `high` fields as log(0) = ∞.
+
+### Run script JSON configuration file
+
+The JSON configuration file accepted by `run.py` looks like this:
+
+```
+{
+  "pipeline_1": {
+    "opt_passes": "",
+    "opt_args": ["-O1"],
+    "llc_args": ["-O1"],
+    "replay_args": ["--repetitions", "20"]
+  }
+}
+```
+
+Each pipeline is defined as a `"name": object` pair. In the object:
+
+- `opt_passes` is a string that gets passed to `opt -passes=""`
+- `opt_args` is a list of arguments that gets directly added to the `opt` invocation.
+- `llc_args` is the same but for `llc`
+- `replay_args` is the same but for the `llvm-omp-kernel-replay` tool that is used to launch kernels.
+
+Keep in mind that if you want to specify any flag that takes a value, you should separate the flag name and value like this: `["flag_name", "value"]`. If you don't want to have any arguments for opt, llc, or kernel replay, leave the list empty; do not delete the field as it will result in an error.
diff --git a/llvm/tools/opt-llc-parameter-search/autorun.py b/llvm/tools/opt-llc-parameter-search/autorun.py
new file mode 100755
index 0000000000000..cf2d332d79688
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/autorun.py
@@ -0,0 +1,269 @@
+#!/usr/bin/env python3
+"""
+Run an automatic search to find the fastest paramater values for LLVM
+optimization passes.
+"""
+
+import argparse
+import csv
+import datetime
+import json
+import sys
+import time
+from statistics import median
+from typing import Any, Never
+import matplotlib
+import matplotlib.pyplot as plt
+import numpy as np
+from hyperopt import hp, fmin, Trials, STATUS_OK, STATUS_FAIL
+import run
+
+bitcode_file_name: str = ""
+REPETITIONS: int = 25
+arch:str = ""
+persistent_llc_flags: list[str] = []
+
+def json_to_space(json_path: str) -> dict[str, Any] | Never:
+    """
+    Parses the JSON file at JSON_PATH and transforms it into a space
+    dict used by hyperopt.fmin
+    """
+    space = {}
+    with open(json_path, 'r') as json_file:
+        config = json.load(json_file)
+        # Populate the persistent flags global
+        persistent_llc_flags.extend(config['persistent_llc_flags'])
+        for flag in config['llc_flags']:
+            name = flag['flag']
+            match flag['type']:
+                case "uint" | "int" | "number":
+                    value_range = flag['range']
+                    distribution = value_range['distribution']
+                    if distribution == "logarithmic":
+                        low = value_range['low']
+                        high = value_range['high']
+                        step = value_range['step']
+                        if low == 0 or high == 0:
+                            print("0 is an invalid number for a log distribution",
+                                  file=sys.stderr)
+                            print(f"Found in {name}", file=sys.stderr)
+                            sys.exit(1)
+                        space[name] = hp.qloguniform(name,
+                                                     np.log(low),
+                                                     np.log(high),
+                                                     step)
+                    elif distribution == "uniform":
+                        low = value_range['low']
+                        high = value_range['high']
+                        step = value_range['step']
+                        space[name] = hp.quniform(name, low, high, step)
+                    # maybe: support normal distributions too
+                    else:
+                        print(f"Distribution {distribution} not supported!",
+                              file=sys.stderr)
+                        sys.exit(1)
+                case "value" | "string":
+                    value_range = flag['range']
+                    choice = value_range['choice']
+                    space[name] = hp.choice(name, choice)
+                case "boolean":
+                    space[name] = hp.choice(name, [True, False])
+    return space
+
+def objective(params: dict[str, Any]) -> dict[str, Any]:
+    """
+    The objective function. Returns an object describing the loss
+    (median kernel runtime), status, and flags.
+    """
+    llc_args: list[str] = []
+    for k, v in params.items():
+        if isinstance(v, bool):
+            # Turn pythonic True/False into true/false
+            llc_args.append(f"--{k}={str(v).lower()}")
+        elif isinstance(v, float):
+            # Make sure that floats that don't have decimal parts get
+            # turned into ints first
+            if v - int(v) == 0:
+                llc_args.append(f"--{k}={str(int(v))}")
+            else :
+                llc_args.append(f"--{k}={v}")
+        else:
+            llc_args.append(f"--{k}={v}")
+
+    # Append persistent flags
+    llc_args.extend(persistent_llc_flags)
+
+    config = {"llc_args": llc_args,
+              "replay_args": [f"--repetitions={REPETITIONS}"],
+              "opt_passes": [],
+              "opt_args": ["--O3"]}
+
+    # NOTE: Since we're already invoking `opt --O3` within the
+    # baseline run, we don't need to repeatedly invoke opt when
+    # calling llc.
+    code = run.run_llc(bitcode_file_name,
+                       # Dummy argument
+                       "AUTORUN",
+                       config,
+                       False,  # dry_run
+                       False,  # also_opt
+                       arch,
+                       False)  # verbose
+    if code != 0:
+        sys.exit(code)
+    json_path = run.check_record_json_file(bitcode_file_name)
+    output_dir = datetime.datetime.now().strftime("%Y%m%d_%H%M%S%f")
+    runtimes = run.replay_and_measure_kernel(json_path,
+                                             config,
+                                             output_dir,
+                                             False,  # run_bitcode
+                                             False)  # dry_run
+    if runtimes:
+        return {
+            'loss': median(runtimes),
+            'eval_time': time.time(),
+            'baseline': "false",
+            'config': config,
+            'status': STATUS_OK
+        }
+    return {
+        'loss': float("NaN"),
+        'eval_time': time.time(),
+        'baseline': "false",
+        'config': config,
+        'status': STATUS_FAIL
+    }
+
+def write_trials_csv(baseline: dict[str, Any],
+                     trial_results: list[dict[str, Any]]) -> None:
+    """
+    Write the trial results out as a CSV file.
+    """
+    now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+    csv_path = f"autosearch-results-{now}.csv"
+    with open(csv_path, 'w', newline='') as csvfile:
+        fieldnames = ['loss', 'eval_time', 'baseline', 'config', 'status']
+        writer = csv.DictWriter(csvfile, fieldnames=fieldnames,
+                                quoting=csv.QUOTE_NONNUMERIC)
+        writer.writeheader()
+        # Write baseline results
+        writer.writerow(baseline)
+        # Write trial results
+        for row in trial_results:
+            writer.writerow(row)
+
+def save_loss_history_plot(baseline_loss: float, trials: Trials) -> None:
+    """
+    Save the loss history + the baseline as a scatterplot. Make a
+    horizontal line from the baseline for easy visual comparison.
+    """
+    now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+
+    losses = list(trials.losses())
+    x = range(len(losses) + 1)
+    y = [ baseline_loss ] + losses
+    plt.scatter(x, y)
+    plt.axhline(baseline_loss, c="g")
+    plt.xlabel("Trial iteration")
+    plt.ylabel("Median runtime (ns)")
+    plt.savefig(f"loss_history-{now}.png", dpi=300, bbox_inches="tight")
+    plt.close()
+
+def main():
+    global bitcode_file_name, arch
+    parser = argparse.ArgumentParser(
+        description="Automatically run parameter search for llc or opt flags."
+    )
+    parser.add_argument(
+        "bc_path",
+        help="Path to the recorded device IR bitcode (.bc) file"
+    )
+    parser.add_argument(
+        "json_path",
+        help="Path to the JSON file containing settings for desired flags"
+    )
+    parser.add_argument(
+        "--arch",
+        default="gfx90a",
+        help="GPU Architecture (default: gfx90a)"
+    )
+    parser.add_argument(
+        "--max-trials",
+        default="100",
+        help="Maximum number of trials (defualt: 100)"
+    )
+    parser.add_argument(
+        "--no-plot",
+        default=False,
+        help="Don't save the median runtime history plot as a PNG image",
+        action='store_true'
+    )
+
+    args = parser.parse_args()
+    trials = Trials()
+
+    arch = args.arch
+    bc_path = args.bc_path
+    json_path = args.json_path
+    max_trials = int(args.max_trials)
+
+    # Back up original bc and image files
+    bitcode_file_name = run.get_original_bitcode(bc_path)
+    run.backup_image(run.image_output_file(bc_path))
+
+    # Run the baseline
+    print("=" * 80)
+    print("Replaying the baseline, unmodified, kernel:")
+    print(f"{'=' * 80}")
+    kernel_json_path = run.check_record_json_file(bitcode_file_name)
+    output_dir = datetime.datetime.now().strftime("%Y%m%d_%H%M%S%f")
+    baseline_config = {
+        "opt_passes": [],
+        "opt_args": ["--O3"],
+        "llc_args": ["-O3"],
+        "replay_args": [f"--repetitions={REPETITIONS}"]
+    }
+    status = run.run_llc(bitcode_file_name,
+                         "BASELINE",
+                         baseline_config,
+                         False, # dry_run
+                         True,  # also_opt
+                         arch,
+                         False) # verbose
+    if status != 0:
+        sys.exit(status)
+    baseline_runtimes = run.replay_and_measure_kernel(kernel_json_path,
+                                                      baseline_config,
+                                                      output_dir,
+                                                      False, # run_bitcode
+                                                      False) # dry_run
+    if baseline_runtimes:
+        baseline = {
+            "loss": median(baseline_runtimes),
+            "eval_time": time.time(),
+            "baseline": "true",
+            "config": baseline_config,
+            "status": STATUS_OK
+        }
+    else:
+        print("Could not measure baseline runtimes", file=sys.stderr)
+        sys.exit(1)
+
+    # Initialize search space
+    space = json_to_space(json_path)
+
+    # Run search
+    best = fmin(
+        fn=objective,         # Objective Function to optimize
+        space=space,          # Hyperparameter Search Space
+        max_evals=max_trials, # Number of optimization attempts
+        trials=trials
+    )
+
+    if not args.no_plot:
+        save_loss_history_plot(baseline["loss"], trials)
+
+    write_trials_csv(baseline, trials.results)
+
+if __name__ == "__main__":
+    main()
diff --git a/llvm/tools/opt-llc-parameter-search/example-autorun.json b/llvm/tools/opt-llc-parameter-search/example-autorun.json
new file mode 100644
index 0000000000000..42483fcf4eec9
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/example-autorun.json
@@ -0,0 +1,53 @@
+{
+  "persistent_llc_flags": ["-O3"],
+  "llc_flags": [
+    {
+      "flag": "amdgpu-1",
+      "type": "uint",
+      "range": {
+        "low": 1,
+        "high": 1e6,
+        "step": 1e2,
+        "distribution": "logarithmic"
+      }
+    },
+    {
+      "flag": "amdgpu-2",
+      "type": "value",
+      "range": {
+        "choice": ["A", "B", "C"]
+      }
+    },
+    {
+      "flag": "amdgpu-3",
+      "type": "boolean"
+    },
+    {
+      "flag": "amdgpu-4",
+      "type": "string",
+      "range": {
+        "choice": ["X", "Y", "Z"]
+      }
+    },
+    {
+      "flag": "amdgpu-5",
+      "type": "int",
+      "range": {
+        "low": 0,
+        "high": 100,
+        "step": 5,
+        "distribution": "uniform"
+      }
+    },
+    {
+      "flag": "amdgpu-6",
+      "type": "number",
+      "range": {
+        "low": -1e6,
+        "high": 1e6,
+        "step": 1e4,
+        "distribution": "uniform"
+      }
+    }
+  ]
+}
diff --git a/llvm/tools/opt-llc-parameter-search/example-opt-passes.json b/llvm/tools/opt-llc-parameter-search/example-opt-passes.json
new file mode 100644
index 0000000000000..a54a0011ad827
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/example-opt-passes.json
@@ -0,0 +1,26 @@
+{
+  "O1": {
+    "opt_passes": "",
+    "opt_args": ["-O1"],
+    "llc_args": ["-O1"],
+    "replay_args": ["--repetitions", "20"]
+  },
+  "O2": {
+    "opt_passes": "",
+    "opt_args": ["-O2"],
+    "llc_args": ["-O2"],
+    "replay_args": ["--repetitions", "20"]
+  },
+  "O3": {
+    "opt_passes": "",
+    "opt_args": ["-O3"],
+    "llc_args": ["-O3"],
+    "replay_args": ["--repetitions", "20"]
+  },
+  "opt-custom-pipeline": {
+    "opt_passes": "no-op-module,cgscc(no-op-cgscc,function(no-op-function,loop(no-op-loop))),function(no-op-function,loop(no-op-loop))",
+    "opt_args": ["--debug-pass-manager"],
+    "llc_args": [],
+    "replay_args": ["--repetitions", "20"]
+  }
+}
diff --git a/llvm/tools/opt-llc-parameter-search/requirements.txt b/llvm/tools/opt-llc-parameter-search/requirements.txt
new file mode 100644
index 0000000000000..e7ed157c3a0f3
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/requirements.txt
@@ -0,0 +1,22 @@
+cloudpickle==3.1.2
+contourpy==1.3.3
+cycler==0.12.1
+fonttools==4.63.0
+hyperopt==0.3.0
+kiwisolver==1.5.0
+matplotlib==3.11.1
+mpi4py==4.0.3
+networkx==3.6.1
+numpy==2.5.2
+packaging==26.3
+pillow==12.3.0
+pyparsing==3.3.2
+python-dateutil==2.9.0.post0
+pytorch-triton-rocm==3.3.0
+scipy==1.18.1
+setuptools==70.2.0
+six==1.17.0
+torch==2.7.0+rocm6.3
+torchaudio==2.7.0+rocm6.3
+torchvision==0.22.0+rocm6.3
+tqdm==4.70.0
diff --git a/llvm/tools/opt-llc-parameter-search/run.py b/llvm/tools/opt-llc-parameter-search/run.py
new file mode 100755
index 0000000000000..ba1b487334dbe
--- /dev/null
+++ b/llvm/tools/opt-llc-parameter-search/run.py
@@ -0,0 +1,521 @@
+#!/usr/bin/env python3
+
+"""
+Run opt or llc on a recorded, unoptimized, kernel IR bitcode (bc) file
+to transform it using the pipeline definitions and opt/llc commandline
+flags read from a user-defined JSON file (see example-opt-passes.json
+in the repo). Replay the transformed kernel and measure its
+performance for each transformation using rocprof, output a CSV file
+and a boxplot of the results.
+"""
+
+import argparse
+import datetime
+import json
+import subprocess
+import sys
+import shutil
+import os
+import csv
+
+from collections import defaultdict
+from typing import Generator, Never, TypeAlias
+
+import matplotlib.pyplot as plt
+
+PipelineConfType: TypeAlias = dict[str, list[str]]
+PipelinesJsonType: TypeAlias = dict[str, PipelineConfType]
+# Timeout for running rocprof
+TIMEOUT: int = 90
+
+def load_pipelines(pipelines_file: str) -> PipelinesJsonType | Never:
+    """
+    Load the pipeline definitions from a JSON file path.
+    """
+    if not os.path.exists(pipelines_file):
+        print(f"Error: '{pipelines_file}' not found in current directory.", file=sys.stderr)
+        sys.exit(1)
+
+    try:
+        with open(pipelines_file, 'r') as f:
+            pipelines = json.load(f)
+
+        if not isinstance(pipelines, dict):
+            print(f"Error: '{pipelines_file}' must contain a JSON object.", file=sys.stderr)
+            sys.exit(1)
+
+        if not pipelines:
+            print(f"Error: '{pipelines_file}' is empty.", file=sys.stderr)
+            sys.exit(1)
+
+        return pipelines
+
+    except json.JSONDecodeError as e:
+        print(f"Error: Failed to parse '{pipelines_file}': {e}", file=sys.stderr)
+        sys.exit(1)
+    except Exception as e:
+        print(f"Error: Failed to read '{pipelines_file}': {e}", file=sys.stderr)
+        sys.exit(1)
+
+def backup_bitcode(bc_path: str) -> str:
+    """
+    Make sure that the bitcode file at bc_path is backed up. A backup is a copy
+    of the .bc file with '.original' appended at the end of the file name. If
+    the backup file already exists or the user already invokes the script with
+    the .bc.original file, do nothing.
+    """
+    if bc_path.endswith(".original"):
+        return bc_path
+    bc_backup_path = bc_path + ".original"
+    if not os.path.exists(bc_backup_path):
+        print(f"Backing up the bitcode file to {bc_backup_path}")
+        shutil.copyfile(bc_path, bc_backup_path)
+        return bc_backup_path
+    return bc_backup_path
+
+def backup_image(image_path: str) -> None:
+    """
+    Like backup_bitcode() but for the image file.
+    """
+    image_backup_path = image_path + ".original"
+    if os.path.exists(image_backup_path):
+        return
+    print(f"Backing up the image file to {image_backup_path}")
+    shutil.copyfile(image_path, image_backup_path)
+
+def get_original_bitcode(bc_path: str) -> str | Never:
+    """
+    Always return the path to the original bitcode file, calling
+    backup_bitcode() if it doesn't yet exist. Warning: we at the moment
+    cannot verify that the file ending in '.original' is, in fact, the
+    original.
+    """
+    if not os.path.exists(bc_path):
+        print("Error: the specified bitcode file could not be found", file=sys.stderr)
+        sys.exit(1)
+    if bc_path.endswith(".original"):
+        return bc_path
+    return backup_bitcode(bc_path)
+
+def check_record_json_file(bitcode_file: str) -> str | Never:
+    """
+    Return the path to the JSON file corresponding to the kernel's
+    recording or exit with an error.
+    """
+    json_path = bitcode_file.replace('.bc', '.json').replace('.original', '')
+    if not os.path.exists(json_path):
+        print(f"Error: could not find a JSON file corresponding to the specified bitcode file path: {json_path}", file=sys.stderr)
+        sys.exit(1)
+    return json_path
+
+def opt_output_file(bitcode_file: str) -> str:
+    """
+    Strip the '.original' suffix from the IR bitcode file path if it
+    exists. Necessary due to the fact that the replay tool expects the
+    original '.bc' file extension instead of '.bc.original'.
+    """
+    return bitcode_file.replace(".original", "") if bitcode_file.endswith(".original") else bitcode_file
+
+def llc_output_file(bitcode_file: str) -> str:
+    """
+    Strip the '.original' suffix from the IR bitcode file path if it
+    exists and replace it with '.o'.
+    """
+    output_file = bitcode_file
+    if output_file.endswith(".original"):
+        output_file = output_file.replace(".original", "")
+    return output_file.replace(".bc", ".o")
+
+def image_output_file(bitcode_file: str) -> str:
+    """
+    Strip the '.original' suffix from the IR bitcode file path if it
+    exists and replace it with '.image'.
+    """
+    output_file = bitcode_file
+    if output_file.endswith(".original"):
+        output_file = output_file.replace(".original", "")
+    return output_file.replace(".bc", ".image")
+
+def cleanup_modified_files(bc_path: str) -> None:
+    """
+    Restore the original .image and .bc files. If kernel recording
+    files ending with 'bc.original' and '.image.original' exist,
+    overwrite the
+    """
+    if bc_path.endswith(".bc.original"):
+        bc_original_path = bc_path
+        bc_originalless_path = bc_path.replace(".original", "")
+    elif bc_path.endswith(".bc"):
+        bc_original_path = bc_path + ".original"
+        bc_originalless_path = bc_path
+    else:
+        print(f"Error: path must end with .bc(.original). Path: {bc_path}", file=sys.stderr)
+        sys.exit(1)
+
+    image_originalless_path = image_output_file(bc_path)
+    image_original_path = image_originalless_path + ".original"
+
+    if os.path.exists(bc_original_path):
+        print(f"Cleaning up the file at {bc_originalless_path}, restoring from {bc_original_path}")
+        os.rename(bc_original_path, bc_originalless_path)
+    else:
+        print(f"Warning: couldn't find the file {bc_original_path}. If this is not your first run, make sure that your bitcode file is original.", file=sys.stderr)
+
+    if os.path.exists(image_original_path):
+        os.rename(image_original_path, image_originalless_path)
+        print(f"Cleaning up the file at {image_originalless_path}, restoring from {image_original_path}")
+    else:
+        print(f"Warning: couldn't find the file {bc_original_path}. If this is not your first run, make sure that your image file is original.", file=sys.stderr)
+
+def run_opt(bitcode_file: str, pipeline_name: str, pipeline_config: PipelineConfType, dry_run: bool, verbose: bool) -> int:
+    """
+    Run `opt` on bitcode_file, reading the flags and custom pipeline
+    data from pipeline_config. If dry_run is True, output a stub
+    message and exit. If verbose is True, pass '-debug-pass-manager'
+    to `opt`.
+    """
+    if dry_run:
+        print(f"Simulating run of pipeline {pipeline_name} for bitcode file {bitcode_file}.")
+        return 0
+
+    print(f"{'=' * 80}")
+    print(f"Running pipeline: {pipeline_name}")
+    print(f"{'=' * 80}")
+
+    output_name = opt_output_file(bitcode_file)
+    opt_passes = pipeline_config['opt_passes']
+    opt_args = pipeline_config['opt_args']
+
+    # Using any() here makes sure that arrays with a single empty
+    # string are treated as empty.
+    if not any(opt_passes) and not any(opt_args):
+        print(f"Opt passes and args not specified for pipeline {pipeline_name}!", file=sys.stderr)
+        return 2
+
+    if not bitcode_file.endswith(".original"):
+        print(f"Bitcode file {bitcode_file} does not have the .original extension, exiting to prevent accidental reuse of an already optimized bitcode", file=sys.stderr)
+        return 3
+
+    cmd = ["opt"]
+    if opt_passes:
+        cmd.append(f"-passes={opt_passes}")
+    if verbose:
+        cmd.append("-debug-pass-manager")
+    cmd.append("-o")
+    cmd.append(output_name)
+    cmd = cmd + opt_args
+    cmd.append(bitcode_file)
+
+    print("Running " + ' '.join(cmd))
+
+    try:
+        result = subprocess.run(cmd, capture_output=True, text=True)
+        print(result.stdout)
+        if result.stderr:
+            print(result.stderr, file=sys.stderr)
+        return result.returncode
+    except FileNotFoundError:
+        print("Error: 'opt' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+        return 1
+    except Exception as e:
+        print(f"Error running pipeline: {e}", file=sys.stderr)
+        return 1
+
+def run_llc(bitcode_file: str, pipeline_name: str, pipeline_config: PipelineConfType, dry_run: bool, also_opt: bool, arch: str, verbose: bool) -> int:
+    """
+    Run `llc` on bitcode_file, reading the flags from pipeline_config.
+    If dry_run is True, output a stub message and exit. If also_opt is
+    True, pass the bitcode file through `opt` first by invoking
+    run_opt(). The arch argument is necessary. If verbose is True,
+    pass '-debug-pass-manager' to `opt`.
+    """
+    if dry_run:
+        print(f"Simulating run of backend pipeline {pipeline_name} for bitcode file {bitcode_file}.")
+        return 0
+    if also_opt:
+        print(f"Running backend pipeline: {pipeline_name} with opt first")
+    else:
+        print(f"{'=' * 80}")
+        print(f"Running backend pipeline: {pipeline_name}")
+        print(f"{'=' * 80}")
+    if not bitcode_file.endswith(".original"):
+        print(f"Bitcode file {bitcode_file} does not have the .original extension, use the --llc-also-use-opt flag if you want to pass an already modified bitcode file to llc")
+        return 3
+    if also_opt:
+        status = run_opt(bitcode_file,
+                            pipeline_name,
+                            pipeline_config,
+                            dry_run,
+                            verbose)
+        if status > 0:
+            return status
+
+    llc_output_name = llc_output_file(bitcode_file)
+    image_output_name = image_output_file(bitcode_file)
+    llc_args = pipeline_config['llc_args']
+
+    if not any(llc_args):
+        print(f"llc args not specified for pipeline {pipeline_name}!", file=sys.stderr)
+        return 2
+
+    cmd = ["llc", "-mtriple=amdgcn-amd-amdhsa", "-filetype=obj"]
+    if verbose:
+        cmd.append("-debug-pass-manager")
+    cmd.append(f"-mcpu={arch}")
+    cmd.append("-o")
+    cmd.append(llc_output_name)
+    cmd = cmd + llc_args
+    if also_opt:
+        cmd.append(opt_output_file(bitcode_file))
+    else:
+        cmd.append(bitcode_file)
+
+    link_cmd = ["ld.lld", "-flavor", "gnu", "-shared"]
+    link_cmd.append("-o")
+    link_cmd.append(image_output_name)
+    link_cmd.append(llc_output_name)
+
+    print("Running " + ' '.join(cmd))
+
+    try:
+        # Run llc
+        result = subprocess.run(cmd, capture_output=True, text=True)
+        print(result.stdout)
+        if result.stderr:
+            print(result.stderr, file=sys.stderr)
+        if result.returncode > 0:
+            return result.returncode
+        # Run the linker
+        print("...and " + ' '.join(link_cmd))
+        link_result = subprocess.run(link_cmd, capture_output=True, text=True)
+        print(link_result.stdout)
+        if link_result.stderr:
+            print(link_result.stderr, file=sys.stderr)
+        return link_result.returncode
+    except FileNotFoundError as e:
+        print("Error: 'llc' or 'ld.lld' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+        print(e)
+        return 1
+    except Exception as e:
+        print(f"Error running pipeline: {e}", file=sys.stderr)
+        return 1
+
+def run_kernel_replay(json_path: str, pipeline_config: PipelineConfType, rp_output_dir: str, run_bitcode: bool, dry_run: bool) -> int:
+    """
+    Run the LLVM OpenMP kernel replay tool through rocprofv3, using
+    the JSON file at json_path and specifying the profiler output
+    directory rp_output_dir. If run_bitcode is true, pass the
+    '--load-bitcode' flag to the tool to make it JIT-compile the IR
+    bitcode file before replaying (necessary when testing opt). If
+    dry_run is True, output a stub message and exit.
+    """
+    if dry_run:
+        print(f"Simulating dry run for kernel replay for JSON file at {json_path}.")
+        return 0
+
+    cmd = [
+        "rocprofv3",
+        "--stats",
+        "--kernel-trace",
+        "--output-directory", rp_output_dir,
+        "--output-file", "output",
+        "--",
+        "llvm-omp-kernel-replay"
+    ]
+
+    if run_bitcode:
+        cmd.append("--load-bitcode")
+    cmd = cmd + pipeline_config['replay_args']
+    cmd.append(json_path)
+
+    print("Running " + ' '.join(cmd))
+
+    try:
+        result = subprocess.run(cmd, capture_output=True, text=True, timeout=TIMEOUT)
+        print(result.stdout)
+        if result.stderr:
+            print(result.stderr, file=sys.stderr)
+        return result.returncode
+    except FileNotFoundError:
+        print("Error: 'llvm-omp-kernel-replay' or 'rocprofv3' command not found. Make sure LLVM is installed and in PATH.", file=sys.stderr)
+        return 1
+    except Exception as e:
+        print(f"Error running pipeline: {e}", file=sys.stderr)
+        return 1
+
+def get_rocprof_output_dir(pipeline_name: str) -> str:
+    """
+    Helper function to format a folder name for a rocprofv3 output
+    directory.
+    """
+    now = datetime.datetime.now().strftime("%Y%m%d_%H%M")
+    return f"{now}-{pipeline_name}"
+
+def profiler_csv_reader(rp_output_dir: str) -> Generator[dict[str, str | float], str, None] | None:
+    """
+    Yield a generator that reads line by line from the kernel trace CSV file
+    located at rp_output_dir. The kernel trace CSV file contains
+    information about one kernel launch per line.
+    """
+    csv_path = f"./{rp_output_dir}/output_kernel_trace.csv"
+    try:
+        with open(csv_path, mode='r', newline='', encoding='utf-8') as kernel_trace_csv:
+            kt_reader = csv.DictReader(kernel_trace_csv, quoting=csv.QUOTE_NONNUMERIC)
+            yield from kt_reader
+    except OSError as e:
+        print(f"Error reading CSV results from {rp_output_dir}: {e}", file=sys.stderr)
+
+def replay_and_measure_kernel(json_path: str, pipeline_config: PipelineConfType, output_dir: str, run_bitcode: bool, dry_run: bool) -> list[float] | None:
+    """
+    Replay the specified kernel via calling run_kernel_replay and
+    return a list that contains kernel runtimes in nanoseconds.
+    """
+    if not dry_run:
+        code = run_kernel_replay(json_path, pipeline_config, output_dir, run_bitcode, dry_run)
+        if code != 0:
+            sys.exit(code)
+        csv_reader = profiler_csv_reader(output_dir)
+        runtimes: list[float] = []
+        if csv_reader:
+            for row in csv_reader:
+                end_ts = row['End_Timestamp']
+                start_ts = row['Start_Timestamp']
+                if isinstance(end_ts, float) and isinstance(start_ts, float):
+                    runtime = end_ts - start_ts
+                    runtimes.append(runtime)
+                else:
+                    print(f"'End_Timestamp' or 'Start_Timestamp' fields in {output_dir} are not floats, exiting")
+                    sys.exit(1)
+        return runtimes
+
+def write_results(results: list[dict[str, str | float]]) -> None:
+    """
+    Output results as a CSV file.
+    """
+    now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+    csv_path = f"results-{now}.csv"
+    with open(csv_path, 'w', newline='') as csvfile:
+        fieldnames = ['pipeline', 'runtime_ns']
+        writer = csv.DictWriter(csvfile, fieldnames=fieldnames,
+                                quoting=csv.QUOTE_NONNUMERIC)
+        writer.writeheader()
+        for row in results:
+            writer.writerow(row)
+    print(f"Successfully wrote results to {csv_path}")
+
+def write_boxplots(results: list[dict[str, str | float]]) -> None:
+    """
+    Output results as boxplot image.
+    """
+    # Convert results to be associative
+    now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
+    png_path = f"results-{now}.png"
+
+    results_assoc = defaultdict(list)
+    for row in results:
+        results_assoc[row["pipeline"]].append(row["runtime_ns"])
+
+    data = list(results_assoc.values())
+    plt.figure(figsize=(14, 6))
+    plt.boxplot(data, tick_labels=list(results_assoc.keys()))
+    plt.xticks(rotation=45, ha='right')
+    plt.title("Kernel runtime")
+    plt.savefig(png_path, dpi=400, bbox_inches='tight')
+
+def main() -> None:
+    parser = argparse.ArgumentParser(
+        description="Run LLVM opt or llc with various pass pipelines on a bitcode file."
+    )
+    parser.add_argument(
+        "bitcode_file",
+        help="Path to the bitcode file (.bc)"
+    )
+    parser.add_argument(
+        "--pipelines-file",
+        default="opt-passes.json",
+        help="Path to a JSON file containing pipeline configurations (default: ./opt-passes.json)"
+    )
+    parser.add_argument(
+        "--pipeline",
+        help="Specific pipeline to run (if not specified, runs all pipelines)"
+    )
+    parser.add_argument(
+        "--dry-run",
+        help="Does not actually invoke opt, replay or other tools, useful for testing",
+        action='store_true'
+    )
+    parser.add_argument(
+        "--arch",
+        help="GPU Architecture (default: gfx90a), only used for --llc",
+        default="gfx90a"
+    )
+    parser.add_argument(
+        "--verbose",
+        help="Use the -debug-pass-manager flag when running opt",
+        action='store_true'
+    )
+    parser.add_argument(
+        "--llc",
+        help="Use the llc-based backend optimization instead of opt",
+        action='store_true'
+    )
+    parser.add_argument(
+        "--llc-also-use-opt",
+        help="When doing backend optimization, also use opt first before running llc",
+        action='store_true'
+    )
+    parser.add_argument(
+        "--no-plot",
+        default=False,
+        help="Don't plot the results as a PNG image",
+        action='store_true'
+    )
+
+    args = parser.parse_args()
+    json_path = check_record_json_file(args.bitcode_file)
+    results = []
+    pipelines = load_pipelines(args.pipelines_file)
+    also_opt = args.llc_also_use_opt
+    arch = args.arch
+
+    original_bitcode_file = get_original_bitcode(args.bitcode_file)
+    backup_image(image_output_file(args.bitcode_file))
+
+    if args.pipeline:
+        pipelines_to_run = [(args.pipeline, pipelines[args.pipeline])]
+    else:
+        pipelines_to_run = list(pipelines.items())
+
+    run_bitcode = not args.llc
+
+    for pipeline_name, pipeline_config in pipelines_to_run:
+        if args.llc:
+            returncode = run_llc(original_bitcode_file, pipeline_name,
+                                 pipeline_config, args.dry_run,
+                                 also_opt, arch, args.verbose)
+        else:
+            returncode = run_opt(original_bitcode_file, pipeline_name,
+                                 pipeline_config, args.dry_run,
+                                 args.verbose)
+        if returncode != 0:
+            print(f"Warning: Pipeline '{pipeline_name}' exited with code {returncode}",
+                  file=sys.stderr)
+            continue
+        rp_output_dir = get_rocprof_output_dir(pipeline_name)
+        runtimes = replay_and_measure_kernel(json_path,
+                                             pipeline_config,
+                                             rp_output_dir,
+                                             run_bitcode,
+                                             args.dry_run)
+        if runtimes:
+            for runtime in runtimes:
+                runtime_dict = {"pipeline": pipeline_name, "runtime_ns": runtime}
+                results.append(runtime_dict)
+
+    if not args.dry_run and returncode == 0:
+        write_results(results)
+        if not args.no_plot:
+            write_boxplots(results)
+
+if __name__ == "__main__":
+    main()



More information about the llvm-commits mailing list