[llvm] [AMDGPU][docs] clusters and cluster broadcast (PR #225310)

Sameer Sahasrabuddhe via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 30 02:17:03 PDT 2026


https://github.com/ssahasra updated https://github.com/llvm/llvm-project/pull/225310

>From 851a76f16ed29918633a19654c85dc8e76e84bd3 Mon Sep 17 00:00:00 2001
From: Sameer Sahasrabuddhe <sameer.sahasrabuddhe at amd.com>
Date: Thu, 6 Aug 2026 18:27:46 +0530
Subject: [PATCH 1/2] [AMDGPU][docs] Describe cluster broadcast using async
 loads

Assisted-By: AI Coding Assistant
---
 llvm/docs/AMDGPUDMAOperations.md | 49 ++++++++++++++++++++++++++++++--
 llvm/docs/AMDGPUMemoryModel.md   |  2 +-
 llvm/docs/AMDGPUUsage.rst        | 46 ++++++++++++++++++++++++++++++
 3 files changed, 93 insertions(+), 4 deletions(-)

diff --git a/llvm/docs/AMDGPUDMAOperations.md b/llvm/docs/AMDGPUDMAOperations.md
index eff09848e4fed..00e08c05f16ea 100644
--- a/llvm/docs/AMDGPUDMAOperations.md
+++ b/llvm/docs/AMDGPUDMAOperations.md
@@ -99,7 +99,7 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
     ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
     i32 immarg %offset,         ; offset (immediate) applied to both global and LDS address
     i32 immarg %cpol,           ; cache policy (immediate)
-    [i32 %m0])                  ; workgroup broadcast mask, cluster variants only (in M0)
+    [i32 %mask])                ; workgroup broadcast mask, cluster variants only (wave-uniform, in M0)
 ```
 
 The bit-size encoded in the name can be 8, 32, 64 or 128.
@@ -107,8 +107,51 @@ The bit-size encoded in the name can be 8, 32, 64 or 128.
 Loads data from global memory to LDS. The `%offset` is applied to both the
 global and LDS addresses.
 
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-broadcast-dma)=
+
+**Cluster Broadcast**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup broadcast mask*: bit `i` corresponds to the
+  workgroup with cluster index `i`. Each workgroup that wants to participate in
+  the combined fetch must set its own bit and all participating workgroups must
+  supply an identical mask (and the same address). If the mask is all-zero, the
+  operation behaves like an ordinary, non-cluster load: it returns only to the
+  requesting workgroup.
+- Bit `[16]` selects the timeout behavior. When clear, a target-defined timeout
+  is used: requests are combined if they arrive within that window. When set,
+  an *early timeout* is used: as soon as the L2 cache supplies the data, it is
+  returned to whichever waves have already issued their requests.
+
+Each participating workgroup issues its own request, with one wave per
+workgroup. For requests that arrive within the timeout, the L2 cache is accessed
+only once, and a copy of the loaded data is written into the LDS of each of
+those workgroups. Workgroups that issue their request later trigger a new,
+separate combined fetch. The data is written at the same LDS location
+(`%lds_base` plus `%offset`) in each participating workgroup.
+
+Because each wave's request only ever writes to its own workgroup's LDS, this is
+an ordinary `addrspace(3)` access with scope "workgroup", exactly like a
+non-cluster variant of this load. Completion is tracked independently for each
+request using the requesting wave's own
+{ref}`asyncmarks<amdgpu-async-operations>`; there is no signal that tells a wave
+when the *other* participating workgroups have completed their own requests.
+Applications that need every participating workgroup to observe the shared data
+must synchronize separately across those workgroups.
+
+The `cluster` variants are only available on targets that have the
+`mcast-load-insts` subtarget feature (e.g., GFX1250). Using them on any other
+target is not supported. Note that this feature is distinct from workgroup
+cluster support itself: a target may support clusters without providing these
+broadcast DMA instructions.
+
+On the `gfx1250-strict` subtarget, the mask is always forced to zero and no
+broadcast occurs, regardless of the value supplied.
 
 ```llvm
 void @llvm.amdgcn.global.store.async.from.lds.b<N>(
diff --git a/llvm/docs/AMDGPUMemoryModel.md b/llvm/docs/AMDGPUMemoryModel.md
index 0a7f023bf19ac..938daabe3fc01 100644
--- a/llvm/docs/AMDGPUMemoryModel.md
+++ b/llvm/docs/AMDGPUMemoryModel.md
@@ -74,7 +74,7 @@ target-defined scopes and constraints:
 
 - *system scope* (same as LLVM)
 - "agent" scope
-- "cluster" scope
+- "cluster" scope (see {ref}`amdgpu-clusters`)
 - "workgroup" scope
 - "wavefront" scope
 - "singlethread" scope (same as LLVM)
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 2d6263402164f..fd25f5cd3ad69 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1451,6 +1451,52 @@ upper 32 bits of the generic pointer.
 As the LDS aperture is defined by its 16 most significant bits, we can theoretically
 support up to ``(1 << 16) - 1`` synthetic apertures safely.
 
+.. _amdgpu-clusters:
+
+Clusters
+--------
+
+On targets that support workgroup clusters (for example ``gfx1250``), the
+workgroups of a grid can be grouped into *clusters*. A cluster is a fixed-size
+block of workgroups that are co-scheduled on the same shader engine so that
+they can cooperate more closely than workgroups in different clusters.
+
+On each subtarget, the maximum size of a cluster is limited by the number of
+WGPs on a shader engine; each workgroup in a cluster runs on a separate WGP. A
+WGP may simultaneously host workgroups from different clusters. Clusters may be
+1D, 2D, or 3D. The cluster dimensions are specified with the
+``"amdgpu-cluster-dims"`` function attribute (see
+:ref:`amdgpu-llvm-ir-attributes-table`). A value of ``0,0,0`` disables
+clustering for the function.
+
+Every workgroup in a cluster is the same size, and every cluster in a dispatch
+is the same size.
+
+Within a cluster, several workgroups can combine matching load requests via
+:ref:`cluster broadcast DMA operations <amdgpu-cluster-broadcast-dma>` so that
+each populates its own LDS from a single shared fetch of global memory. The
+``cluster`` memory scope (see
+:ref:`amdgpu-memory-scopes`) synchronizes operations performed by threads in
+workgroups of the same cluster. On targets that do not support clusters,
+``cluster`` scope behaves like ``agent`` scope.
+
+A workgroup can query its position within the grid and its cluster using the
+following intrinsics:
+
+* ``llvm.amdgcn.cluster.id.{x,y,z}`` -- the coordinates of this workgroup's
+  cluster within the grid.
+* ``llvm.amdgcn.cluster.workgroup.id.{x,y,z}`` -- the coordinates of this
+  workgroup within its cluster.
+* ``llvm.amdgcn.cluster.workgroup.flat.id`` -- the flattened index of this
+  workgroup within its cluster. The linearization treats X as the innermost
+  dimension and Z as the outermost:
+
+  ``flat_id = x + y * cluster_dim_x + z * cluster_dim_x * cluster_dim_y``.
+* ``llvm.amdgcn.cluster.workgroup.max.id.{x,y,z}`` -- the largest workgroup
+  index within the cluster in each dimension.
+* ``llvm.amdgcn.cluster.workgroup.max.flat.id`` -- the largest flattened
+  workgroup within the cluster.
+
 .. _amdgpu-memory-scopes:
 
 Memory Scopes

>From 32d482224d1367c9f5ecf537bdffaaf58d9b9b22 Mon Sep 17 00:00:00 2001
From: Sameer Sahasrabuddhe <sameer.sahasrabuddhe at amd.com>
Date: Wed, 30 Sep 2026 14:45:44 +0530
Subject: [PATCH 2/2] it's multicast, not broadcast; also don't assume a
 formula for flat ID

---
 llvm/docs/AMDGPUDMAOperations.md | 22 +++++++++++-----------
 llvm/docs/AMDGPUUsage.rst        | 13 +++++++------
 2 files changed, 18 insertions(+), 17 deletions(-)

diff --git a/llvm/docs/AMDGPUDMAOperations.md b/llvm/docs/AMDGPUDMAOperations.md
index 00e08c05f16ea..e8f5dd029387e 100644
--- a/llvm/docs/AMDGPUDMAOperations.md
+++ b/llvm/docs/AMDGPUDMAOperations.md
@@ -99,7 +99,7 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
     ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
     i32 immarg %offset,         ; offset (immediate) applied to both global and LDS address
     i32 immarg %cpol,           ; cache policy (immediate)
-    [i32 %mask])                ; workgroup broadcast mask, cluster variants only (wave-uniform, in M0)
+    [i32 %mask])                ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
 ```
 
 The bit-size encoded in the name can be 8, 32, 64 or 128.
@@ -107,9 +107,9 @@ The bit-size encoded in the name can be 8, 32, 64 or 128.
 Loads data from global memory to LDS. The `%offset` is applied to both the
 global and LDS addresses.
 
-(amdgpu-cluster-broadcast-dma)=
+(amdgpu-cluster-multicast-dma)=
 
-**Cluster Broadcast**
+**Cluster Multicast Loads**
 
 The `cluster` variants let several workgroups in a {ref}`workgroup cluster
 <amdgpu-clusters>` share a single fetch from global memory when they are all
@@ -117,22 +117,22 @@ loading the same data to their respective LDS. The `%mask` argument identifies
 the group of workgroups whose requests may be combined, and controls the request
 timeout. The mask is wave-uniform and passed in the `M0` register:
 
-- Bits `[15:0]` are the *workgroup broadcast mask*: bit `i` corresponds to the
+- Bits `[15:0]` are the *workgroup multicast mask*: bit `i` corresponds to the
   workgroup with cluster index `i`. Each workgroup that wants to participate in
   the combined fetch must set its own bit and all participating workgroups must
   supply an identical mask (and the same address). If the mask is all-zero, the
   operation behaves like an ordinary, non-cluster load: it returns only to the
   requesting workgroup.
 - Bit `[16]` selects the timeout behavior. When clear, a target-defined timeout
-  is used: requests are combined if they arrive within that window. When set,
+  is used: requests may be combined if they arrive within that window. When set,
   an *early timeout* is used: as soon as the L2 cache supplies the data, it is
   returned to whichever waves have already issued their requests.
 
 Each participating workgroup issues its own request, with one wave per
-workgroup. For requests that arrive within the timeout, the L2 cache is accessed
-only once, and a copy of the loaded data is written into the LDS of each of
-those workgroups. Workgroups that issue their request later trigger a new,
-separate combined fetch. The data is written at the same LDS location
+workgroup. When requests are combined, they share a single L2 fetch, and a copy
+of the loaded data is written into the LDS of each requesting workgroup.
+Workgroups that issue their request later make a separate request, which may
+combine with other later requests. The data is written at the same LDS location
 (`%lds_base` plus `%offset`) in each participating workgroup.
 
 Because each wave's request only ever writes to its own workgroup's LDS, this is
@@ -148,10 +148,10 @@ The `cluster` variants are only available on targets that have the
 `mcast-load-insts` subtarget feature (e.g., GFX1250). Using them on any other
 target is not supported. Note that this feature is distinct from workgroup
 cluster support itself: a target may support clusters without providing these
-broadcast DMA instructions.
+multicast load instructions.
 
 On the `gfx1250-strict` subtarget, the mask is always forced to zero and no
-broadcast occurs, regardless of the value supplied.
+multicast occurs, regardless of the value supplied.
 
 ```llvm
 void @llvm.amdgcn.global.store.async.from.lds.b<N>(
diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index fd25f5cd3ad69..2026008dae689 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1473,7 +1473,7 @@ Every workgroup in a cluster is the same size, and every cluster in a dispatch
 is the same size.
 
 Within a cluster, several workgroups can combine matching load requests via
-:ref:`cluster broadcast DMA operations <amdgpu-cluster-broadcast-dma>` so that
+:ref:`cluster multicast DMA operations <amdgpu-cluster-multicast-dma>` so that
 each populates its own LDS from a single shared fetch of global memory. The
 ``cluster`` memory scope (see
 :ref:`amdgpu-memory-scopes`) synchronizes operations performed by threads in
@@ -1487,11 +1487,12 @@ following intrinsics:
   cluster within the grid.
 * ``llvm.amdgcn.cluster.workgroup.id.{x,y,z}`` -- the coordinates of this
   workgroup within its cluster.
-* ``llvm.amdgcn.cluster.workgroup.flat.id`` -- the flattened index of this
-  workgroup within its cluster. The linearization treats X as the innermost
-  dimension and Z as the outermost:
-
-  ``flat_id = x + y * cluster_dim_x + z * cluster_dim_x * cluster_dim_y``.
+* ``llvm.amdgcn.cluster.workgroup.flat.id`` -- this workgroup's position within
+  its cluster: a value between ``0`` and the number of workgroups in the
+  cluster minus one, assigned to each workgroup in the cluster. Every wave in
+  the workgroup is initialized with its workgroup's value. This is the same
+  index used to select destination workgroups in the mask of a
+  :ref:`cluster multicast DMA <amdgpu-cluster-multicast-dma>`.
 * ``llvm.amdgcn.cluster.workgroup.max.id.{x,y,z}`` -- the largest workgroup
   index within the cluster in each dimension.
 * ``llvm.amdgcn.cluster.workgroup.max.flat.id`` -- the largest flattened



More information about the llvm-commits mailing list