[llvm] [AMDGPU][docs] Describe cluster broadcast using async loads (PR #225310)
Sameer Sahasrabuddhe via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 23:58:19 PDT 2026
https://github.com/ssahasra created https://github.com/llvm/llvm-project/pull/225310
Assisted-By: AI Coding Assistant
>From 851a76f16ed29918633a19654c85dc8e76e84bd3 Mon Sep 17 00:00:00 2001
From: Sameer Sahasrabuddhe <sameer.sahasrabuddhe at amd.com>
Date: Thu, 6 Aug 2026 18:27:46 +0530
Subject: [PATCH] [AMDGPU][docs] Describe cluster broadcast using async loads
Assisted-By: AI Coding Assistant
---
llvm/docs/AMDGPUDMAOperations.md | 49 ++++++++++++++++++++++++++++++--
llvm/docs/AMDGPUMemoryModel.md | 2 +-
llvm/docs/AMDGPUUsage.rst | 46 ++++++++++++++++++++++++++++++
3 files changed, 93 insertions(+), 4 deletions(-)
diff --git a/llvm/docs/AMDGPUDMAOperations.md b/llvm/docs/AMDGPUDMAOperations.md
index eff09848e4fed..00e08c05f16ea 100644
--- a/llvm/docs/AMDGPUDMAOperations.md
+++ b/llvm/docs/AMDGPUDMAOperations.md
@@ -99,7 +99,7 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
i32 immarg %offset, ; offset (immediate) applied to both global and LDS address
i32 immarg %cpol, ; cache policy (immediate)
- [i32 %m0]) ; workgroup broadcast mask, cluster variants only (in M0)
+ [i32 %mask]) ; workgroup broadcast mask, cluster variants only (wave-uniform, in M0)
```
The bit-size encoded in the name can be 8, 32, 64 or 128.
@@ -107,8 +107,51 @@ The bit-size encoded in the name can be 8, 32, 64 or 128.
Loads data from global memory to LDS. The `%offset` is applied to both the
global and LDS addresses.
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-broadcast-dma)=
+
+**Cluster Broadcast**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup broadcast mask*: bit `i` corresponds to the
+ workgroup with cluster index `i`. Each workgroup that wants to participate in
+ the combined fetch must set its own bit and all participating workgroups must
+ supply an identical mask (and the same address). If the mask is all-zero, the
+ operation behaves like an ordinary, non-cluster load: it returns only to the
+ requesting workgroup.
+- Bit `[16]` selects the timeout behavior. When clear, a target-defined timeout
+ is used: requests are combined if they arrive within that window. When set,
+ an *early timeout* is used: as soon as the L2 cache supplies the data, it is
+ returned to whichever waves have already issued their requests.
+
+Each participating workgroup issues its own request, with one wave per
+workgroup. For requests that arrive within the timeout, the L2 cache is accessed
+only once, and a copy of the loaded data is written into the LDS of each of
+those workgroups. Workgroups that issue their request later trigger a new,
+separate combined fetch. The data is written at the same LDS location
+(`%lds_base` plus `%offset`) in each participating workgroup.
+
+Because each wave's request only ever writes to its own workgroup's LDS, this is
+an ordinary `addrspace(3)` access with scope "workgroup", exactly like a
+non-cluster variant of this load. Completion is tracked independently for each
+request using the requesting wave's own
+{ref}`asyncmarks<amdgpu-async-operations>`; there is no signal that tells a wave
+when the *other* participating workgroups have completed their own requests.
+Applications that need every participating workgroup to observe the shared data
+must synchronize separately across those workgroups.
+
+The `cluster` variants are only available on targets that have the
+`mcast-load-insts` subtarget feature (e.g., GFX1250). Using them on any other
+target is not supported. Note that this feature is distinct from workgroup
+cluster support itself: a target may support clusters without providing these
+broadcast DMA instructions.
+
+On the `gfx1250-strict` subtarget, the mask is always forced to zero and no
+broadcast occurs, regardless of the value supplied.
```llvm
void @llvm.amdgcn.global.store.async.from.lds.b<N>(
diff --git a/llvm/docs/AMDGPUMemoryModel.md b/llvm/docs/AMDGPUMemoryModel.md
index 0a7f023bf19ac..938daabe3fc01 100644
--- a/llvm/docs/AMDGPUMemoryModel.md
+++ b/llvm/docs/AMDGPUMemoryModel.md
@@ -74,7 +74,7 @@ target-defined scopes and constraints:
- *system scope* (same as LLVM)
- "agent" scope
-- "cluster" scope
+- "cluster" scope (see {ref}`amdgpu-clusters`)
- "workgroup" scope
- "wavefront" scope
- "singlethread" scope (same as LLVM)
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 2d6263402164f..fd25f5cd3ad69 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1451,6 +1451,52 @@ upper 32 bits of the generic pointer.
As the LDS aperture is defined by its 16 most significant bits, we can theoretically
support up to ``(1 << 16) - 1`` synthetic apertures safely.
+.. _amdgpu-clusters:
+
+Clusters
+--------
+
+On targets that support workgroup clusters (for example ``gfx1250``), the
+workgroups of a grid can be grouped into *clusters*. A cluster is a fixed-size
+block of workgroups that are co-scheduled on the same shader engine so that
+they can cooperate more closely than workgroups in different clusters.
+
+On each subtarget, the maximum size of a cluster is limited by the number of
+WGPs on a shader engine; each workgroup in a cluster runs on a separate WGP. A
+WGP may simultaneously host workgroups from different clusters. Clusters may be
+1D, 2D, or 3D. The cluster dimensions are specified with the
+``"amdgpu-cluster-dims"`` function attribute (see
+:ref:`amdgpu-llvm-ir-attributes-table`). A value of ``0,0,0`` disables
+clustering for the function.
+
+Every workgroup in a cluster is the same size, and every cluster in a dispatch
+is the same size.
+
+Within a cluster, several workgroups can combine matching load requests via
+:ref:`cluster broadcast DMA operations <amdgpu-cluster-broadcast-dma>` so that
+each populates its own LDS from a single shared fetch of global memory. The
+``cluster`` memory scope (see
+:ref:`amdgpu-memory-scopes`) synchronizes operations performed by threads in
+workgroups of the same cluster. On targets that do not support clusters,
+``cluster`` scope behaves like ``agent`` scope.
+
+A workgroup can query its position within the grid and its cluster using the
+following intrinsics:
+
+* ``llvm.amdgcn.cluster.id.{x,y,z}`` -- the coordinates of this workgroup's
+ cluster within the grid.
+* ``llvm.amdgcn.cluster.workgroup.id.{x,y,z}`` -- the coordinates of this
+ workgroup within its cluster.
+* ``llvm.amdgcn.cluster.workgroup.flat.id`` -- the flattened index of this
+ workgroup within its cluster. The linearization treats X as the innermost
+ dimension and Z as the outermost:
+
+ ``flat_id = x + y * cluster_dim_x + z * cluster_dim_x * cluster_dim_y``.
+* ``llvm.amdgcn.cluster.workgroup.max.id.{x,y,z}`` -- the largest workgroup
+ index within the cluster in each dimension.
+* ``llvm.amdgcn.cluster.workgroup.max.flat.id`` -- the largest flattened
+ workgroup within the cluster.
+
.. _amdgpu-memory-scopes:
Memory Scopes
More information about the llvm-commits
mailing list