[llvm] [llvm-calc-occupancy] Add an AMDGPU occupancy calculator tool (PR #208727)

via llvm-commits llvm-commits at lists.llvm.org
Mon Jul 13 02:32:04 PDT 2026


https://github.com/michaelselehov updated https://github.com/llvm/llvm-project/pull/208727

>From 51447e07817c0a2c9653769610a51ea3230bbf64 Mon Sep 17 00:00:00 2001
From: mselehov <mselehov at amd.com>
Date: Fri, 10 Jul 2026 08:42:15 -0500
Subject: [PATCH 1/4] [llvm-calc-occupancy] Add an AMDGPU occupancy calculator
 tool

Add a command-line utility that reports the occupancy (waves/EU) an AMDGPU
kernel would reach for a given workgroup size, VGPR/SGPR usage and LDS. It
reuses the backend's own occupancy math, so the numbers match the compiler.
Includes a lit test and a CommandGuide man page.

Assisted-by: Claude Opus
---
 llvm/docs/CommandGuide/index.md               |   2 +
 .../docs/CommandGuide/llvm-calc-occupancy.rst | 107 +++++
 llvm/test/lit.cfg.py                          |   1 +
 .../test/tools/llvm-calc-occupancy/basic.test |  58 +++
 llvm/tools/llvm-calc-occupancy/CMakeLists.txt |  30 ++
 .../llvm-calc-occupancy.cpp                   | 410 ++++++++++++++++++
 6 files changed, 608 insertions(+)
 create mode 100644 llvm/docs/CommandGuide/llvm-calc-occupancy.rst
 create mode 100644 llvm/test/tools/llvm-calc-occupancy/basic.test
 create mode 100644 llvm/tools/llvm-calc-occupancy/CMakeLists.txt
 create mode 100644 llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp

diff --git a/llvm/docs/CommandGuide/index.md b/llvm/docs/CommandGuide/index.md
index ac90d9468b848..b9c14f30c3010 100644
--- a/llvm/docs/CommandGuide/index.md
+++ b/llvm/docs/CommandGuide/index.md
@@ -15,6 +15,7 @@ llvm-addr2line
 llvm-ar
 llvm-as
 llvm-bcanalyzer
+llvm-calc-occupancy
 llvm-cgdata
 llvm-config
 llvm-cov
@@ -143,6 +144,7 @@ interested in.
 * {doc}`llvm-profgen`
 * {doc}`llvm-tli-checker`
 * {doc}`llvm-offload-binary`
+* {doc}`llvm-calc-occupancy`
 
 ## Remarks Tools
 
diff --git a/llvm/docs/CommandGuide/llvm-calc-occupancy.rst b/llvm/docs/CommandGuide/llvm-calc-occupancy.rst
new file mode 100644
index 0000000000000..61cebf52b3223
--- /dev/null
+++ b/llvm/docs/CommandGuide/llvm-calc-occupancy.rst
@@ -0,0 +1,107 @@
+llvm-calc-occupancy - AMDGPU occupancy calculator
+==================================================
+
+.. program:: llvm-calc-occupancy
+
+SYNOPSIS
+--------
+
+:program:`llvm-calc-occupancy` -mcpu=<gfxNNN> [*options*]
+
+DESCRIPTION
+-----------
+
+:program:`llvm-calc-occupancy` reports the occupancy (waves per execution unit)
+that an AMDGPU kernel would achieve for a given combination of workgroup size,
+VGPR usage, SGPR usage and LDS allocation. It is a thin front-end over the same
+occupancy math the AMDGPU backend uses (``GCNSubtarget``), so the numbers match
+what the compiler computes for an equivalent kernel.
+
+Any resource that is not specified is treated as unconstrained. When the
+workgroup size is left unspecified (or given as a range), the occupancy is
+reported as a range as well.
+
+The tool only supports the ``amdgcn`` target.
+
+EXAMPLE
+-------
+
+.. code-block:: console
+
+  $ llvm-calc-occupancy -mcpu=gfx90a --wg-size=512 --vgprs=50 --sgprs=30
+  llvm-calc-occupancy - AMDGPU occupancy calculator
+
+  Target
+    Triple:              amdgcn-amd-amdhsa
+    GPU (-mcpu):         gfx90a
+    Wavefront size:      64
+    Max waves/EU:        8 (waves per SIMD, hardware limit)
+    EUs (SIMDs) per CU:  4
+    ...
+
+  Per-constraint occupancy (waves/EU)
+    Workgroup + LDS:     8
+    VGPRs:               8
+    SGPRs:               8
+
+  Result
+    Occupancy:           8 waves/EU (32 waves/CU)
+    Limited by:          workgroup size / LDS, VGPRs, SGPRs
+    Next step:           already at the hardware maximum
+
+OPTIONS
+-------
+
+.. option:: -mcpu=<gfxNNN>
+
+  Target GPU, for example ``gfx90a``. Required.
+
+.. option:: -mtriple=<triple>
+
+  Target triple. Defaults to ``amdgcn-amd-amdhsa``. Must be an ``amdgcn``
+  triple.
+
+.. option:: -mattr=<features>
+
+  Comma-separated list of subtarget features, for example
+  ``+wavefrontsize32``.
+
+.. option:: --wg-size=<N>, --flat-workgroup-size=<N>
+
+  Flat workgroup size. Accepts a single value ``N`` or a range ``MIN:MAX``.
+  When omitted, the full legal range is assumed and the occupancy is reported
+  as a range.
+
+.. option:: --vgprs=<N>
+
+  Number of VGPRs used per lane. When omitted, VGPRs do not constrain the
+  occupancy.
+
+.. option:: --sgprs=<N>
+
+  Number of SGPRs used per wave. When omitted, SGPRs do not constrain the
+  occupancy.
+
+.. option:: --lds=<size>
+
+  LDS bytes allocated per workgroup. Accepts an optional ``k``/``kb`` or
+  ``m``/``mb`` suffix (base 1024).
+
+.. option:: --dynamic-vgpr-block-size=<N>
+
+  Dynamic VGPR block size. ``0`` (the default) disables dynamic VGPR mode.
+
+.. option:: --limits
+
+  Print, for each occupancy level supported by the GPU, the maximum number of
+  VGPRs and SGPRs that still reaches that level.
+
+.. option:: --help
+
+  Print a summary of command line options and exit.
+
+EXIT STATUS
+-----------
+
+:program:`llvm-calc-occupancy` returns 0 on success and a non-zero exit code if
+the arguments are invalid (for example a missing or non-``amdgcn`` target).
diff --git a/llvm/test/lit.cfg.py b/llvm/test/lit.cfg.py
index cd028963dd59e..4f6f33c544ba7 100644
--- a/llvm/test/lit.cfg.py
+++ b/llvm/test/lit.cfg.py
@@ -347,6 +347,7 @@ def get_asan_rtlib():
         ToolSubst("OrcV2CBindingsLazy", unresolved="ignore"),
         ToolSubst("OrcV2CBindingsVeryLazy", unresolved="ignore"),
         ToolSubst("dxil-dis", unresolved="ignore"),
+        ToolSubst("llvm-calc-occupancy", unresolved="ignore"),
     ]
 )
 
diff --git a/llvm/test/tools/llvm-calc-occupancy/basic.test b/llvm/test/tools/llvm-calc-occupancy/basic.test
new file mode 100644
index 0000000000000..9c1656070973d
--- /dev/null
+++ b/llvm/test/tools/llvm-calc-occupancy/basic.test
@@ -0,0 +1,58 @@
+## Basic checks for the AMDGPU occupancy calculator.
+REQUIRES: amdgpu-registered-target
+
+## VGPR-limited kernel on gfx1030 (wave32, 16 waves/EU HW limit).
+RUN: llvm-calc-occupancy -mcpu=gfx1030 --vgprs=128 | FileCheck %s --check-prefix=VGPR
+VGPR: GPU (-mcpu):         gfx1030
+VGPR: Max waves/EU:        16
+VGPR: VGPRs:               8
+VGPR: Occupancy:           8 waves/EU (32 waves/CU)
+VGPR: Limited by:          VGPRs
+
+## VGPR-limited kernel reports how far VGPRs must drop for the next level.
+RUN: llvm-calc-occupancy -mcpu=gfx950 --vgprs=250 --wg-size=64 | FileCheck %s --check-prefix=NEXT
+NEXT: Limited by:          VGPRs
+NEXT: Next step:           reach 3 waves/EU:
+NEXT: VGPRs <= 168 (currently 250)
+
+## When several factors tie, all of them must be relaxed.
+RUN: llvm-calc-occupancy -mcpu=gfx90a --vgprs=80 --sgprs=102 | FileCheck %s --check-prefix=BOTH
+BOTH: Limited by:          VGPRs, SGPRs
+BOTH: Next step:           reach 7 waves/EU (requires all of):
+BOTH: VGPRs <= 72 (currently 80)
+BOTH: SGPRs <= 96 (currently 102)
+
+## An LDS-limited kernel is told how much LDS to shed.
+RUN: llvm-calc-occupancy -mcpu=gfx90a --wg-size=256 --lds=40k | FileCheck %s --check-prefix=NEXTLDS
+NEXTLDS: Limited by:          workgroup size / LDS
+NEXTLDS: Next step:           reach 2 waves/EU:
+NEXTLDS: LDS <= 32768 bytes (currently 40960)
+
+## No hint when already at the hardware maximum.
+RUN: llvm-calc-occupancy -mcpu=gfx90a --vgprs=40 --sgprs=20 --wg-size=64 | FileCheck %s --check-prefix=MAXED
+MAXED: Occupancy:           8 waves/EU
+MAXED: Next step:           already at the hardware maximum
+
+## No workgroup size given: the workgroup dimension is reported as a range.
+RUN: llvm-calc-occupancy -mcpu=gfx90a | FileCheck %s --check-prefix=FUZZY
+FUZZY: Workgroup size:      1 .. 1024 (unspecified -> full range)
+FUZZY: Occupancy:           8 waves/EU (32 waves/CU)
+
+## Requesting more LDS than the GPU has warns and pins occupancy to 1.
+RUN: llvm-calc-occupancy -mcpu=gfx90a --lds=103kb 2>&1 | FileCheck %s --check-prefix=BIGLDS
+BIGLDS: warning:{{.*}}LDS request (105472 bytes) exceeds addressable LDS
+BIGLDS: Occupancy:           1 waves/EU
+
+## The register-limit table can be dumped for a GPU.
+RUN: llvm-calc-occupancy -mcpu=gfx90a --limits | FileCheck %s --check-prefix=LIMITS
+LIMITS: Occupancy      Max VGPRs      Max SGPRs
+LIMITS: 8              64             80
+LIMITS: 1              512            102
+
+## Missing -mcpu is an error.
+RUN: not llvm-calc-occupancy 2>&1 | FileCheck %s --check-prefix=NOCPU
+NOCPU: error:{{.*}}no GPU specified
+
+## A non-AMDGPU triple is rejected.
+RUN: not llvm-calc-occupancy -mtriple=x86_64-- -mcpu=gfx90a 2>&1 | FileCheck %s --check-prefix=BADTRIPLE
+BADTRIPLE: error:{{.*}}only supports the amdgcn target
diff --git a/llvm/tools/llvm-calc-occupancy/CMakeLists.txt b/llvm/tools/llvm-calc-occupancy/CMakeLists.txt
new file mode 100644
index 0000000000000..9580d855f06a3
--- /dev/null
+++ b/llvm/tools/llvm-calc-occupancy/CMakeLists.txt
@@ -0,0 +1,30 @@
+# This tool depends on AMDGPU backend internals (GCNSubtarget). Only build it
+# when the AMDGPU target is enabled.
+if(NOT "AMDGPU" IN_LIST LLVM_TARGETS_TO_BUILD)
+  return()
+endif()
+
+include_directories(
+  ${LLVM_MAIN_SRC_DIR}/lib/Target/AMDGPU
+  ${LLVM_BINARY_DIR}/lib/Target/AMDGPU
+  )
+
+set(LLVM_LINK_COMPONENTS
+  AMDGPUCodeGen
+  AMDGPUDesc
+  AMDGPUInfo
+  AMDGPUUtils
+  CodeGen
+  Core
+  MC
+  Support
+  Target
+  TargetParser
+  )
+
+add_llvm_tool(llvm-calc-occupancy
+  llvm-calc-occupancy.cpp
+
+  DEPENDS
+  AMDGPUCommonTableGen
+  )
diff --git a/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
new file mode 100644
index 0000000000000..422161e7fa7ea
--- /dev/null
+++ b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
@@ -0,0 +1,410 @@
+//===-- llvm-calc-occupancy.cpp - AMDGPU occupancy calculator -------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+//
+// A small standalone utility that answers "what occupancy do I get?" for an
+// AMDGPU kernel, given some subset of its resource usage: workgroup size,
+// VGPRs, SGPRs and LDS. Fields that are left unspecified are treated as
+// unconstrained, and the result is reported as a range (waves per EU).
+//
+// It reuses the compiler's own occupancy math (GCNSubtarget) so the numbers
+// match what the backend would compute for the same inputs.
+//
+// Example:
+//   llvm-calc-occupancy -mcpu=gfx90a --wg-size=512 --vgprs=50 --sgprs=30 \
+//                       --lds=103kb
+//
+//===----------------------------------------------------------------------===//
+
+#include "AMDGPUTargetMachine.h"
+#include "GCNSubtarget.h"
+#include "Utils/AMDGPUBaseInfo.h"
+#include "llvm/ADT/SmallVector.h"
+#include "llvm/ADT/StringExtras.h"
+#include "llvm/ADT/StringRef.h"
+#include "llvm/MC/TargetRegistry.h"
+#include "llvm/Support/CommandLine.h"
+#include "llvm/Support/Format.h"
+#include "llvm/Support/InitLLVM.h"
+#include "llvm/Support/TargetSelect.h"
+#include "llvm/Support/WithColor.h"
+#include "llvm/Support/raw_ostream.h"
+#include "llvm/Target/TargetMachine.h"
+#include "llvm/TargetParser/Triple.h"
+#include <optional>
+
+using namespace llvm;
+
+namespace {
+cl::OptionCategory OccCategory("llvm-calc-occupancy options");
+
+cl::opt<std::string> TripleName("mtriple", cl::desc("Target triple"),
+                                cl::init("amdgcn-amd-amdhsa"),
+                                cl::cat(OccCategory));
+
+cl::opt<std::string> MCPU("mcpu", cl::desc("Target GPU (e.g. gfx90a)"),
+                          cl::init(""), cl::cat(OccCategory));
+
+cl::opt<std::string> MAttr("mattr",
+                           cl::desc("Comma-separated subtarget features "
+                                    "(e.g. +wavefrontsize32)"),
+                           cl::init(""), cl::cat(OccCategory));
+
+cl::opt<std::string>
+    WGSizeStr("wg-size",
+              cl::desc("Flat workgroup size: single value 'N' or range "
+                       "'MIN:MAX' (default: 1:1024)"),
+              cl::init(""), cl::cat(OccCategory));
+cl::alias WGSizeAlias("flat-workgroup-size", cl::aliasopt(WGSizeStr));
+
+cl::opt<int> NumVGPRs("vgprs", cl::desc("VGPRs used per lane (default: none)"),
+                      cl::init(-1), cl::cat(OccCategory));
+
+cl::opt<int> NumSGPRs("sgprs", cl::desc("SGPRs used per wave (default: none)"),
+                      cl::init(-1), cl::cat(OccCategory));
+
+cl::opt<std::string> LDSStr("lds",
+                            cl::desc("LDS bytes per workgroup, accepts k/kb/m "
+                                     "suffixes (default: 0)"),
+                            cl::init(""), cl::cat(OccCategory));
+
+cl::opt<unsigned>
+    DynVGPRBlockSize("dynamic-vgpr-block-size",
+                     cl::desc("Dynamic VGPR block size (0 = disabled)"),
+                     cl::init(0), cl::cat(OccCategory));
+
+cl::opt<bool> ShowLimits("limits",
+                         cl::desc("Print the per-occupancy VGPR/SGPR limit "
+                                  "table for this GPU"),
+                         cl::init(false), cl::cat(OccCategory));
+} // namespace
+
+// Parse a byte size with an optional binary suffix (k/kb/kib/m/mb/mib), all
+// base 1024. A bare number is interpreted as bytes.
+static bool parseSize(StringRef S, uint64_t &Out) {
+  S = S.trim();
+  if (S.empty())
+    return false;
+  uint64_t Mult = 1;
+  StringRef Lower = S.lower();
+  static const std::pair<StringRef, uint64_t> Suffixes[] = {
+      {"kib", 1024},        {"kb", 1024},        {"k", 1024},
+      {"mib", 1024 * 1024}, {"mb", 1024 * 1024}, {"m", 1024 * 1024}};
+  for (const auto &[Suf, M] : Suffixes) {
+    if (Lower.ends_with(Suf)) {
+      Mult = M;
+      S = S.drop_back(Suf.size()).rtrim();
+      break;
+    }
+  }
+  uint64_t Value;
+  if (S.getAsInteger(10, Value))
+    return false;
+  Out = Value * Mult;
+  return true;
+}
+
+// Parse "N" or "MIN:MAX" (also accepts "MIN-MAX") into a flat workgroup range.
+static bool parseWGRange(StringRef S, unsigned &Min, unsigned &Max) {
+  S = S.trim();
+  StringRef LHS, RHS;
+  if (S.contains(':'))
+    std::tie(LHS, RHS) = S.split(':');
+  else if (S.contains('-'))
+    std::tie(LHS, RHS) = S.split('-');
+  else
+    LHS = RHS = S;
+
+  unsigned Lo, Hi;
+  if (LHS.trim().getAsInteger(10, Lo) || RHS.trim().getAsInteger(10, Hi))
+    return false;
+  if (Lo == 0 || Hi == 0 || Lo > Hi)
+    return false;
+  Min = Lo;
+  Max = Hi;
+  return true;
+}
+
+static std::string formatBytes(uint64_t Bytes) {
+  if (Bytes && Bytes % 1024 == 0)
+    return (Twine(Bytes) + " bytes (" + Twine(Bytes / 1024) + " KiB)").str();
+  return (Twine(Bytes) + " bytes").str();
+}
+
+int main(int argc, char **argv) {
+  InitLLVM X(argc, argv);
+
+  cl::HideUnrelatedOptions(OccCategory);
+  cl::ParseCommandLineOptions(
+      argc, argv,
+      "AMDGPU occupancy calculator\n\n"
+      "  Prints the occupancy (waves per EU) implied by a given workgroup "
+      "size,\n"
+      "  VGPR/SGPR usage and LDS allocation. Unspecified fields are reported "
+      "as\n"
+      "  a range.\n");
+
+  LLVMInitializeAMDGPUTargetInfo();
+  LLVMInitializeAMDGPUTarget();
+  LLVMInitializeAMDGPUTargetMC();
+
+  if (MCPU.empty()) {
+    WithColor::error(errs(), "llvm-calc-occupancy")
+        << "no GPU specified; pass -mcpu=<gfxNNN> (e.g. -mcpu=gfx90a)\n";
+    return 1;
+  }
+
+  Triple TT(Triple::normalize(TripleName));
+  if (!TT.isAMDGCN()) {
+    WithColor::error(errs(), "llvm-calc-occupancy")
+        << "this tool only supports the amdgcn target; got triple '" << TT.str()
+        << "'\n";
+    return 1;
+  }
+
+  std::string Error;
+  const Target *T = TargetRegistry::lookupTarget(TT, Error);
+  if (!T) {
+    WithColor::error(errs(), "llvm-calc-occupancy") << Error << "\n";
+    return 1;
+  }
+
+  TargetOptions Options;
+  std::unique_ptr<TargetMachine> TM(T->createTargetMachine(
+      TT, MCPU, MAttr, Options, std::nullopt, std::nullopt));
+  if (!TM) {
+    WithColor::error(errs(), "llvm-calc-occupancy")
+        << "failed to create target machine for '" << MCPU << "'\n";
+    return 1;
+  }
+
+  GCNSubtarget ST(TM->getTargetTriple(), std::string(TM->getTargetCPU()),
+                  std::string(TM->getTargetFeatureString()),
+                  *static_cast<GCNTargetMachine *>(TM.get()));
+
+  const MCSubtargetInfo &STI = ST;
+
+  // Parse inputs.
+  unsigned WGMin = 1, WGMax = AMDGPU::IsaInfo::getMaxFlatWorkGroupSize();
+  bool WGSpecified = !WGSizeStr.empty();
+  if (WGSpecified && !parseWGRange(WGSizeStr, WGMin, WGMax)) {
+    WithColor::error(errs(), "llvm-calc-occupancy")
+        << "invalid --wg-size '" << WGSizeStr << "'\n";
+    return 1;
+  }
+
+  uint64_t LDSBytes = 0;
+  bool LDSSpecified = !LDSStr.empty();
+  if (LDSSpecified && !parseSize(LDSStr, LDSBytes)) {
+    WithColor::error(errs(), "llvm-calc-occupancy")
+        << "invalid --lds '" << LDSStr << "'\n";
+    return 1;
+  }
+
+  bool VGPRSpecified = NumVGPRs >= 0;
+  bool SGPRSpecified = NumSGPRs >= 0;
+
+  // Hardware characteristics.
+  unsigned WaveSize = AMDGPU::IsaInfo::getWavefrontSize(STI);
+  unsigned MaxWaves = AMDGPU::IsaInfo::getMaxWavesPerEU(STI);
+  unsigned EUsPerCU = AMDGPU::IsaInfo::getEUsPerCU(STI);
+  unsigned LocalMemSize = AMDGPU::IsaInfo::getLocalMemorySize(STI);
+  unsigned AddrLocalMem = AMDGPU::IsaInfo::getAddressableLocalMemorySize(STI);
+  unsigned AddrVGPRs =
+      AMDGPU::IsaInfo::getAddressableNumVGPRs(STI, DynVGPRBlockSize);
+  unsigned AddrSGPRs = AMDGPU::IsaInfo::getAddressableNumSGPRs(STI);
+  unsigned MaxWGSize = AMDGPU::IsaInfo::getMaxFlatWorkGroupSize();
+
+  // Warn about inputs that exceed the hardware's physical capacity: such a
+  // kernel could not actually launch, so the reported occupancy is only the
+  // math extrapolated past the limit.
+  auto Warn = [](const Twine &Msg) {
+    WithColor::warning(errs(), "llvm-calc-occupancy") << Msg << "\n";
+  };
+  if (LDSSpecified && LDSBytes > AddrLocalMem)
+    Warn("LDS request (" + Twine(LDSBytes) +
+         " bytes) exceeds addressable LDS "
+         "per workgroup (" +
+         Twine(AddrLocalMem) + " bytes)");
+  if (VGPRSpecified && static_cast<unsigned>(NumVGPRs) > AddrVGPRs)
+    Warn("VGPR request (" + Twine(static_cast<int>(NumVGPRs)) +
+         ") exceeds addressable "
+         "VGPRs (" +
+         Twine(AddrVGPRs) + ")");
+  if (SGPRSpecified && static_cast<unsigned>(NumSGPRs) > AddrSGPRs)
+    Warn("SGPR request (" + Twine(static_cast<int>(NumSGPRs)) +
+         ") exceeds addressable "
+         "SGPRs (" +
+         Twine(AddrSGPRs) + ")");
+  if (WGMax > MaxWGSize)
+    Warn("workgroup size (" + Twine(WGMax) +
+         ") exceeds the maximum flat "
+         "workgroup size (" +
+         Twine(MaxWGSize) + ")");
+
+  outs() << "llvm-calc-occupancy - AMDGPU occupancy calculator\n\n";
+  outs() << "Target\n";
+  outs() << format("  %-20s %s\n", "Triple:", TT.str().c_str());
+  outs() << format("  %-20s %s\n",
+                   "GPU (-mcpu):", std::string(TM->getTargetCPU()).c_str());
+  outs() << format("  %-20s %u\n", "Wavefront size:", WaveSize);
+  outs() << format("  %-20s %u (waves per SIMD, hardware limit)\n",
+                   "Max waves/EU:", MaxWaves);
+  outs() << format("  %-20s %u\n", "EUs (SIMDs) per CU:", EUsPerCU);
+  outs() << format("  %-20s %s\n",
+                   "LDS per CU:", formatBytes(LocalMemSize).c_str());
+  outs() << format("  %-20s %s\n", "Addressable LDS:",
+                   (formatBytes(AddrLocalMem) + " per workgroup").c_str());
+
+  outs() << "\nInputs\n";
+  if (WGSpecified) {
+    if (WGMin == WGMax)
+      outs() << format("  %-20s %u\n", "Workgroup size:", WGMin);
+    else
+      outs() << format("  %-20s %u .. %u\n", "Workgroup size:", WGMin, WGMax);
+  } else {
+    outs() << format("  %-20s %u .. %u (unspecified -> full range)\n",
+                     "Workgroup size:", WGMin, WGMax);
+  }
+  if (VGPRSpecified)
+    outs() << format("  %-20s %d\n", "VGPRs per lane:", (int)NumVGPRs);
+  else
+    outs() << format("  %-20s %s\n", "VGPRs per lane:", "unspecified");
+  if (SGPRSpecified)
+    outs() << format("  %-20s %d\n", "SGPRs per wave:", (int)NumSGPRs);
+  else
+    outs() << format("  %-20s %s\n", "SGPRs per wave:", "unspecified");
+  if (LDSSpecified)
+    outs() << format("  %-20s %s\n",
+                     "LDS per workgroup:", formatBytes(LDSBytes).c_str());
+  else
+    outs() << format("  %-20s %s\n", "LDS per workgroup:", "unspecified (0)");
+
+  // Per-constraint occupancy (all in waves/EU).
+  std::pair<unsigned, unsigned> WGOcc = ST.getOccupancyWithWorkGroupSizes(
+      static_cast<uint32_t>(LDSBytes), {WGMin, WGMax});
+  unsigned VGPROcc =
+      VGPRSpecified ? ST.getOccupancyWithNumVGPRs(NumVGPRs, DynVGPRBlockSize)
+                    : MaxWaves;
+  unsigned SGPROcc =
+      SGPRSpecified ? ST.getOccupancyWithNumSGPRs(NumSGPRs) : MaxWaves;
+
+  outs() << "\nPer-constraint occupancy (waves/EU)\n";
+  if (WGOcc.first == WGOcc.second)
+    outs() << format("  %-20s %u\n", "Workgroup + LDS:", WGOcc.second);
+  else
+    outs() << format("  %-20s %u .. %u\n", "Workgroup + LDS:", WGOcc.first,
+                     WGOcc.second);
+  if (VGPRSpecified)
+    outs() << format("  %-20s %u\n", "VGPRs:", VGPROcc);
+  if (SGPRSpecified)
+    outs() << format("  %-20s %u\n", "SGPRs:", SGPROcc);
+
+  // Combine like GCNSubtarget::computeOccupancy.
+  unsigned MaxOcc = std::min({WGOcc.second, VGPROcc, SGPROcc});
+  unsigned MinOcc = std::min(WGOcc.first, MaxOcc);
+
+  // Identify what pins the maximum occupancy.
+  SmallVector<StringRef, 3> LimitedBy;
+  if (WGOcc.second == MaxOcc)
+    LimitedBy.push_back("workgroup size / LDS");
+  if (VGPRSpecified && VGPROcc == MaxOcc)
+    LimitedBy.push_back("VGPRs");
+  if (SGPRSpecified && SGPROcc == MaxOcc)
+    LimitedBy.push_back("SGPRs");
+
+  outs() << "\nResult\n";
+  if (MinOcc == MaxOcc)
+    outs() << format("  %-20s %u waves/EU (%u waves/CU)\n",
+                     "Occupancy:", MaxOcc, MaxOcc * EUsPerCU);
+  else
+    outs() << format("  %-20s %u .. %u waves/EU (%u .. %u waves/CU)\n",
+                     "Occupancy:", MinOcc, MaxOcc, MinOcc * EUsPerCU,
+                     MaxOcc * EUsPerCU);
+  outs() << format("  %-20s %s\n",
+                   "Limited by:", join(LimitedBy, ", ").c_str());
+
+  // Hint: what would it take to gain one more wave/EU. Every factor that
+  // currently pins the occupancy has to be relaxed, so list each one.
+  if (MaxOcc >= MaxWaves) {
+    outs() << format("  %-20s already at the hardware maximum\n", "Next step:");
+  } else {
+    unsigned TargetOcc = MaxOcc + 1;
+    outs() << format("  %-20s reach %u waves/EU%s\n", "Next step:", TargetOcc,
+                     LimitedBy.size() > 1 ? " (requires all of):" : ":");
+
+    if (VGPRSpecified && VGPROcc == MaxOcc) {
+      unsigned MaxV =
+          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, TargetOcc, DynVGPRBlockSize);
+      outs() << format("      VGPRs <= %u (currently %d)\n", MaxV,
+                       static_cast<int>(NumVGPRs));
+    }
+    if (SGPRSpecified && SGPROcc == MaxOcc) {
+      unsigned MaxS =
+          AMDGPU::IsaInfo::getMaxNumSGPRs(STI, TargetOcc, /*Addressable=*/true);
+      outs() << format("      SGPRs <= %u (currently %d)\n", MaxS,
+                       static_cast<int>(NumSGPRs));
+    }
+    if (WGOcc.second == MaxOcc) {
+      auto WGLDSOcc = [&](uint32_t LDS, unsigned Lo, unsigned Hi) {
+        return ST.getOccupancyWithWorkGroupSizes(LDS, {Lo, Hi}).second;
+      };
+      uint32_t LDS32 = static_cast<uint32_t>(LDSBytes);
+      bool Suggested = false;
+      // LDS lever: largest LDS that still reaches TargetOcc at current WG.
+      if (LDSSpecified && LDSBytes > 0 &&
+          WGLDSOcc(0, WGMin, WGMax) >= TargetOcc) {
+        uint64_t Lo = 0, Hi = LDSBytes;
+        while (Lo < Hi) {
+          uint64_t Mid = Lo + (Hi - Lo + 1) / 2;
+          if (WGLDSOcc(static_cast<uint32_t>(Mid), WGMin, WGMax) >= TargetOcc)
+            Lo = Mid;
+          else
+            Hi = Mid - 1;
+        }
+        outs() << format("      LDS <= %llu bytes (currently %llu)\n",
+                         static_cast<unsigned long long>(Lo),
+                         static_cast<unsigned long long>(LDSBytes));
+        Suggested = true;
+      }
+      // Workgroup lever: largest flat workgroup size that reaches TargetOcc.
+      if (WGMax > 1 && WGLDSOcc(LDS32, 1, 1) >= TargetOcc) {
+        unsigned Lo = 1, Hi = WGMax;
+        while (Lo < Hi) {
+          unsigned Mid = Lo + (Hi - Lo + 1) / 2;
+          if (WGLDSOcc(LDS32, Mid, Mid) >= TargetOcc)
+            Lo = Mid;
+          else
+            Hi = Mid - 1;
+        }
+        if (Lo < WGMax) {
+          outs() << format("      workgroup size <= %u (currently %u)\n", Lo,
+                           WGMax);
+          Suggested = true;
+        }
+      }
+      if (!Suggested)
+        outs() << "      reduce workgroup size and/or LDS\n";
+    }
+  }
+
+  if (ShowLimits) {
+    outs() << "\nPer-occupancy register limits (max regs to still reach "
+              "each level)\n";
+    outs() << format("  %-14s %-14s %-14s\n", "Occupancy", "Max VGPRs",
+                     "Max SGPRs");
+    for (unsigned Occ = MaxWaves; Occ >= 1; --Occ) {
+      unsigned MaxV =
+          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, Occ, DynVGPRBlockSize);
+      unsigned MaxS =
+          AMDGPU::IsaInfo::getMaxNumSGPRs(STI, Occ, /*Addressable=*/true);
+      outs() << format("  %-14u %-14u %-14u\n", Occ, MaxV, MaxS);
+    }
+  }
+
+  return 0;
+}

>From e9e0468a3b64847502cacd9f03812c6037a1b5e7 Mon Sep 17 00:00:00 2001
From: mselehov <mselehov at amd.com>
Date: Fri, 10 Jul 2026 10:03:22 -0500
Subject: [PATCH 2/4] [llvm-calc-occupancy] Address review feedback

- Fix dynamic VGPR block size: fall back to the subtarget default when 0
  and dynamic VGPRs are enabled, matching GCNSubtarget::computeOccupancy.
- Default to the amdgpu triple and advise it in the docs (legacy amdgcn
  still accepted).
- Add a TODO to drop the codegen dependency once the occupancy math is
  available through TargetParser.
- Use a structured binding for the workgroup occupancy range.
- Use argv[0] instead of a hardcoded name in diagnostics.

Assisted-by: Cursor
---
 .../docs/CommandGuide/llvm-calc-occupancy.rst | 12 ++--
 .../test/tools/llvm-calc-occupancy/basic.test |  2 +-
 .../llvm-calc-occupancy.cpp                   | 65 +++++++++++--------
 3 files changed, 47 insertions(+), 32 deletions(-)

diff --git a/llvm/docs/CommandGuide/llvm-calc-occupancy.rst b/llvm/docs/CommandGuide/llvm-calc-occupancy.rst
index 61cebf52b3223..72f972ccae097 100644
--- a/llvm/docs/CommandGuide/llvm-calc-occupancy.rst
+++ b/llvm/docs/CommandGuide/llvm-calc-occupancy.rst
@@ -21,7 +21,9 @@ Any resource that is not specified is treated as unconstrained. When the
 workgroup size is left unspecified (or given as a range), the occupancy is
 reported as a range as well.
 
-The tool only supports the ``amdgcn`` target.
+The tool only supports the AMDGPU target. New callers should use the
+``amdgpu`` triple (for example ``amdgpu-amd-amdhsa``); the legacy ``amdgcn``
+spelling is still accepted.
 
 EXAMPLE
 -------
@@ -32,7 +34,7 @@ EXAMPLE
   llvm-calc-occupancy - AMDGPU occupancy calculator
 
   Target
-    Triple:              amdgcn-amd-amdhsa
+    Triple:              amdgpu-amd-amdhsa
     GPU (-mcpu):         gfx90a
     Wavefront size:      64
     Max waves/EU:        8 (waves per SIMD, hardware limit)
@@ -58,8 +60,8 @@ OPTIONS
 
 .. option:: -mtriple=<triple>
 
-  Target triple. Defaults to ``amdgcn-amd-amdhsa``. Must be an ``amdgcn``
-  triple.
+  Target triple. Defaults to ``amdgpu-amd-amdhsa``. Must be an AMDGPU triple
+  (the ``amdgpu`` or legacy ``amdgcn`` arch).
 
 .. option:: -mattr=<features>
 
@@ -104,4 +106,4 @@ EXIT STATUS
 -----------
 
 :program:`llvm-calc-occupancy` returns 0 on success and a non-zero exit code if
-the arguments are invalid (for example a missing or non-``amdgcn`` target).
+the arguments are invalid (for example a missing or non-AMDGPU target).
diff --git a/llvm/test/tools/llvm-calc-occupancy/basic.test b/llvm/test/tools/llvm-calc-occupancy/basic.test
index 9c1656070973d..3c2f63229df59 100644
--- a/llvm/test/tools/llvm-calc-occupancy/basic.test
+++ b/llvm/test/tools/llvm-calc-occupancy/basic.test
@@ -55,4 +55,4 @@ NOCPU: error:{{.*}}no GPU specified
 
 ## A non-AMDGPU triple is rejected.
 RUN: not llvm-calc-occupancy -mtriple=x86_64-- -mcpu=gfx90a 2>&1 | FileCheck %s --check-prefix=BADTRIPLE
-BADTRIPLE: error:{{.*}}only supports the amdgcn target
+BADTRIPLE: error:{{.*}}only supports the AMDGPU target
diff --git a/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
index 422161e7fa7ea..764089e63ad73 100644
--- a/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
+++ b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
@@ -14,6 +14,11 @@
 // It reuses the compiler's own occupancy math (GCNSubtarget) so the numbers
 // match what the backend would compute for the same inputs.
 //
+// TODO: This links the AMDGPU codegen libraries only because the occupancy
+// math currently lives in GCNSubtarget. Once that subtarget information is
+// exposed through TargetParser, this tool should depend on TargetParser alone
+// and drop the codegen dependency.
+//
 // Example:
 //   llvm-calc-occupancy -mcpu=gfx90a --wg-size=512 --vgprs=50 --sgprs=30 \
 //                       --lds=103kb
@@ -43,7 +48,7 @@ namespace {
 cl::OptionCategory OccCategory("llvm-calc-occupancy options");
 
 cl::opt<std::string> TripleName("mtriple", cl::desc("Target triple"),
-                                cl::init("amdgcn-amd-amdhsa"),
+                                cl::init("amdgpu-amd-amdhsa"),
                                 cl::cat(OccCategory));
 
 cl::opt<std::string> MCPU("mcpu", cl::desc("Target GPU (e.g. gfx90a)"),
@@ -137,6 +142,7 @@ static std::string formatBytes(uint64_t Bytes) {
 
 int main(int argc, char **argv) {
   InitLLVM X(argc, argv);
+  const char *ToolName = argv[0];
 
   cl::HideUnrelatedOptions(OccCategory);
   cl::ParseCommandLineOptions(
@@ -153,15 +159,15 @@ int main(int argc, char **argv) {
   LLVMInitializeAMDGPUTargetMC();
 
   if (MCPU.empty()) {
-    WithColor::error(errs(), "llvm-calc-occupancy")
+    WithColor::error(errs(), ToolName)
         << "no GPU specified; pass -mcpu=<gfxNNN> (e.g. -mcpu=gfx90a)\n";
     return 1;
   }
 
   Triple TT(Triple::normalize(TripleName));
   if (!TT.isAMDGCN()) {
-    WithColor::error(errs(), "llvm-calc-occupancy")
-        << "this tool only supports the amdgcn target; got triple '" << TT.str()
+    WithColor::error(errs(), ToolName)
+        << "this tool only supports the AMDGPU target; got triple '" << TT.str()
         << "'\n";
     return 1;
   }
@@ -169,7 +175,7 @@ int main(int argc, char **argv) {
   std::string Error;
   const Target *T = TargetRegistry::lookupTarget(TT, Error);
   if (!T) {
-    WithColor::error(errs(), "llvm-calc-occupancy") << Error << "\n";
+    WithColor::error(errs(), ToolName) << Error << "\n";
     return 1;
   }
 
@@ -177,7 +183,7 @@ int main(int argc, char **argv) {
   std::unique_ptr<TargetMachine> TM(T->createTargetMachine(
       TT, MCPU, MAttr, Options, std::nullopt, std::nullopt));
   if (!TM) {
-    WithColor::error(errs(), "llvm-calc-occupancy")
+    WithColor::error(errs(), ToolName)
         << "failed to create target machine for '" << MCPU << "'\n";
     return 1;
   }
@@ -188,11 +194,17 @@ int main(int argc, char **argv) {
 
   const MCSubtargetInfo &STI = ST;
 
+  // Mirror GCNSubtarget::computeOccupancy: when the block size is not given
+  // explicitly, fall back to the subtarget's default if dynamic VGPRs are on.
+  unsigned DynVGPRBlockSizeEff = DynVGPRBlockSize;
+  if (DynVGPRBlockSizeEff == 0 && ST.isDynamicVGPREnabled())
+    DynVGPRBlockSizeEff = ST.getDynamicVGPRBlockSize();
+
   // Parse inputs.
   unsigned WGMin = 1, WGMax = AMDGPU::IsaInfo::getMaxFlatWorkGroupSize();
   bool WGSpecified = !WGSizeStr.empty();
   if (WGSpecified && !parseWGRange(WGSizeStr, WGMin, WGMax)) {
-    WithColor::error(errs(), "llvm-calc-occupancy")
+    WithColor::error(errs(), ToolName)
         << "invalid --wg-size '" << WGSizeStr << "'\n";
     return 1;
   }
@@ -200,8 +212,7 @@ int main(int argc, char **argv) {
   uint64_t LDSBytes = 0;
   bool LDSSpecified = !LDSStr.empty();
   if (LDSSpecified && !parseSize(LDSStr, LDSBytes)) {
-    WithColor::error(errs(), "llvm-calc-occupancy")
-        << "invalid --lds '" << LDSStr << "'\n";
+    WithColor::error(errs(), ToolName) << "invalid --lds '" << LDSStr << "'\n";
     return 1;
   }
 
@@ -215,15 +226,15 @@ int main(int argc, char **argv) {
   unsigned LocalMemSize = AMDGPU::IsaInfo::getLocalMemorySize(STI);
   unsigned AddrLocalMem = AMDGPU::IsaInfo::getAddressableLocalMemorySize(STI);
   unsigned AddrVGPRs =
-      AMDGPU::IsaInfo::getAddressableNumVGPRs(STI, DynVGPRBlockSize);
+      AMDGPU::IsaInfo::getAddressableNumVGPRs(STI, DynVGPRBlockSizeEff);
   unsigned AddrSGPRs = AMDGPU::IsaInfo::getAddressableNumSGPRs(STI);
   unsigned MaxWGSize = AMDGPU::IsaInfo::getMaxFlatWorkGroupSize();
 
   // Warn about inputs that exceed the hardware's physical capacity: such a
   // kernel could not actually launch, so the reported occupancy is only the
   // math extrapolated past the limit.
-  auto Warn = [](const Twine &Msg) {
-    WithColor::warning(errs(), "llvm-calc-occupancy") << Msg << "\n";
+  auto Warn = [ToolName](const Twine &Msg) {
+    WithColor::warning(errs(), ToolName) << Msg << "\n";
   };
   if (LDSSpecified && LDSBytes > AddrLocalMem)
     Warn("LDS request (" + Twine(LDSBytes) +
@@ -271,11 +282,13 @@ int main(int argc, char **argv) {
                      "Workgroup size:", WGMin, WGMax);
   }
   if (VGPRSpecified)
-    outs() << format("  %-20s %d\n", "VGPRs per lane:", (int)NumVGPRs);
+    outs() << format("  %-20s %d\n",
+                     "VGPRs per lane:", static_cast<int>(NumVGPRs));
   else
     outs() << format("  %-20s %s\n", "VGPRs per lane:", "unspecified");
   if (SGPRSpecified)
-    outs() << format("  %-20s %d\n", "SGPRs per wave:", (int)NumSGPRs);
+    outs() << format("  %-20s %d\n",
+                     "SGPRs per wave:", static_cast<int>(NumSGPRs));
   else
     outs() << format("  %-20s %s\n", "SGPRs per wave:", "unspecified");
   if (LDSSpecified)
@@ -285,32 +298,32 @@ int main(int argc, char **argv) {
     outs() << format("  %-20s %s\n", "LDS per workgroup:", "unspecified (0)");
 
   // Per-constraint occupancy (all in waves/EU).
-  std::pair<unsigned, unsigned> WGOcc = ST.getOccupancyWithWorkGroupSizes(
+  auto [WGMinOcc, WGMaxOcc] = ST.getOccupancyWithWorkGroupSizes(
       static_cast<uint32_t>(LDSBytes), {WGMin, WGMax});
   unsigned VGPROcc =
-      VGPRSpecified ? ST.getOccupancyWithNumVGPRs(NumVGPRs, DynVGPRBlockSize)
+      VGPRSpecified ? ST.getOccupancyWithNumVGPRs(NumVGPRs, DynVGPRBlockSizeEff)
                     : MaxWaves;
   unsigned SGPROcc =
       SGPRSpecified ? ST.getOccupancyWithNumSGPRs(NumSGPRs) : MaxWaves;
 
   outs() << "\nPer-constraint occupancy (waves/EU)\n";
-  if (WGOcc.first == WGOcc.second)
-    outs() << format("  %-20s %u\n", "Workgroup + LDS:", WGOcc.second);
+  if (WGMinOcc == WGMaxOcc)
+    outs() << format("  %-20s %u\n", "Workgroup + LDS:", WGMaxOcc);
   else
-    outs() << format("  %-20s %u .. %u\n", "Workgroup + LDS:", WGOcc.first,
-                     WGOcc.second);
+    outs() << format("  %-20s %u .. %u\n", "Workgroup + LDS:", WGMinOcc,
+                     WGMaxOcc);
   if (VGPRSpecified)
     outs() << format("  %-20s %u\n", "VGPRs:", VGPROcc);
   if (SGPRSpecified)
     outs() << format("  %-20s %u\n", "SGPRs:", SGPROcc);
 
   // Combine like GCNSubtarget::computeOccupancy.
-  unsigned MaxOcc = std::min({WGOcc.second, VGPROcc, SGPROcc});
-  unsigned MinOcc = std::min(WGOcc.first, MaxOcc);
+  unsigned MaxOcc = std::min({WGMaxOcc, VGPROcc, SGPROcc});
+  unsigned MinOcc = std::min(WGMinOcc, MaxOcc);
 
   // Identify what pins the maximum occupancy.
   SmallVector<StringRef, 3> LimitedBy;
-  if (WGOcc.second == MaxOcc)
+  if (WGMaxOcc == MaxOcc)
     LimitedBy.push_back("workgroup size / LDS");
   if (VGPRSpecified && VGPROcc == MaxOcc)
     LimitedBy.push_back("VGPRs");
@@ -339,7 +352,7 @@ int main(int argc, char **argv) {
 
     if (VGPRSpecified && VGPROcc == MaxOcc) {
       unsigned MaxV =
-          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, TargetOcc, DynVGPRBlockSize);
+          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, TargetOcc, DynVGPRBlockSizeEff);
       outs() << format("      VGPRs <= %u (currently %d)\n", MaxV,
                        static_cast<int>(NumVGPRs));
     }
@@ -349,7 +362,7 @@ int main(int argc, char **argv) {
       outs() << format("      SGPRs <= %u (currently %d)\n", MaxS,
                        static_cast<int>(NumSGPRs));
     }
-    if (WGOcc.second == MaxOcc) {
+    if (WGMaxOcc == MaxOcc) {
       auto WGLDSOcc = [&](uint32_t LDS, unsigned Lo, unsigned Hi) {
         return ST.getOccupancyWithWorkGroupSizes(LDS, {Lo, Hi}).second;
       };
@@ -399,7 +412,7 @@ int main(int argc, char **argv) {
                      "Max SGPRs");
     for (unsigned Occ = MaxWaves; Occ >= 1; --Occ) {
       unsigned MaxV =
-          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, Occ, DynVGPRBlockSize);
+          AMDGPU::IsaInfo::getMaxNumVGPRs(STI, Occ, DynVGPRBlockSizeEff);
       unsigned MaxS =
           AMDGPU::IsaInfo::getMaxNumSGPRs(STI, Occ, /*Addressable=*/true);
       outs() << format("  %-14u %-14u %-14u\n", Occ, MaxV, MaxS);

>From 4270aa02e60e1c6fef62d305f001cb3d8a237862 Mon Sep 17 00:00:00 2001
From: mselehov <mselehov at amd.com>
Date: Mon, 13 Jul 2026 02:54:50 -0500
Subject: [PATCH 3/4] [llvm-calc-occupancy] Add tool to check-llvm test
 dependencies

The lit test runs under REQUIRES: amdgpu-registered-target and the tool
substitution is registered as unresolved="ignore", so on a clean build
where llvm-calc-occupancy was not built the RUN line executed the literal
command and failed with exit code 127. Add the tool to LLVM_TEST_DEPENDS
(guarded on the target, since it is only built with the AMDGPU target) so
check-llvm builds it before running the tests.

Assisted-by: Cursor
---
 llvm/test/CMakeLists.txt | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/llvm/test/CMakeLists.txt b/llvm/test/CMakeLists.txt
index 201f0ee86d7ff..3f1ee7c18bde7 100644
--- a/llvm/test/CMakeLists.txt
+++ b/llvm/test/CMakeLists.txt
@@ -187,6 +187,11 @@ if (TARGET llvm-mt)
   list(APPEND LLVM_TEST_DEPENDS llvm-mt)
 endif ()
 
+# Only built when the AMDGPU target is enabled.
+if (TARGET llvm-calc-occupancy)
+  list(APPEND LLVM_TEST_DEPENDS llvm-calc-occupancy)
+endif ()
+
 if(LLVM_INCLUDE_EXAMPLES)
   list(APPEND LLVM_TEST_DEPENDS
     Kaleidoscope-Ch3

>From 536512d466324ff253e5af0983802edc82fbe739 Mon Sep 17 00:00:00 2001
From: mselehov <mselehov at amd.com>
Date: Mon, 13 Jul 2026 04:06:11 -0500
Subject: [PATCH 4/4] [llvm-calc-occupancy] Fix dangling StringRef in parseSize

'StringRef Lower = S.lower();' bound a StringRef to the temporary
std::string returned by StringRef::lower(), which clang diagnoses as
-Wdangling-gsl (an error under -Werror on the premerge Linux builders).
The dangling was previously latent because the tool was not compiled in
CI until it was wired into the test dependencies. Drop the temporary and
compare suffixes case-insensitively in place via equals_insensitive.

Assisted-by: Cursor
---
 llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
index 764089e63ad73..916c9fb941626 100644
--- a/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
+++ b/llvm/tools/llvm-calc-occupancy/llvm-calc-occupancy.cpp
@@ -95,12 +95,11 @@ static bool parseSize(StringRef S, uint64_t &Out) {
   if (S.empty())
     return false;
   uint64_t Mult = 1;
-  StringRef Lower = S.lower();
   static const std::pair<StringRef, uint64_t> Suffixes[] = {
       {"kib", 1024},        {"kb", 1024},        {"k", 1024},
       {"mib", 1024 * 1024}, {"mb", 1024 * 1024}, {"m", 1024 * 1024}};
   for (const auto &[Suf, M] : Suffixes) {
-    if (Lower.ends_with(Suf)) {
+    if (S.take_back(Suf.size()).equals_insensitive(Suf)) {
       Mult = M;
       S = S.drop_back(Suf.size()).rtrim();
       break;



More information about the llvm-commits mailing list