[llvm] [AMDGPU][docs] define cluster scope and multicast loads (gfx1250) (PR #225310)
Pierre van Houtryve via llvm-commits
llvm-commits at lists.llvm.org
Fri Oct 2 04:49:17 PDT 2026
================
@@ -99,16 +99,59 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
i32 immarg %offset, ; offset (immediate) applied to both global and LDS address
i32 immarg %cpol, ; cache policy (immediate)
- [i32 %m0]) ; workgroup broadcast mask, cluster variants only (in M0)
+ [i32 %mask]) ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
```
The bit-size encoded in the name can be 8, 32, 64 or 128.
Loads data from global memory to LDS. The `%offset` is applied to both the
global and LDS addresses.
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-multicast-dma)=
+
+**Cluster Multicast Loads**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup multicast mask*: bit `i` corresponds to the
+ workgroup with cluster index `i`. Each workgroup that wants to participate in
+ the combined fetch must set its own bit and all participating workgroups must
+ supply an identical mask (and the same address). If the mask is all-zero, the
+ operation behaves like an ordinary, non-cluster load: it returns only to the
+ requesting workgroup.
+- Bit `[16]` selects the timeout behavior. When clear, a target-defined timeout
+ is used: requests may be combined if they arrive within that window. When set,
+ an *early timeout* is used: as soon as the L2 cache supplies the data, it is
+ returned to whichever waves have already issued their requests.
+
+Each participating workgroup issues its own request, with one wave per
+workgroup. When requests are combined, they share a single L2 fetch, and a copy
+of the loaded data is written into the LDS of each requesting workgroup.
+Workgroups that issue their request later make a separate request, which may
+combine with other later requests. The data is written at the same LDS location
+(`%lds_base` plus `%offset`) in each participating workgroup.
+
+Because each wave's request only ever writes to its own workgroup's LDS, this is
+an ordinary `addrspace(3)` access with scope "workgroup", exactly like a
----------------
Pierre-vh wrote:
```suggestion
an ordinary `addrspace(3)` access with scope `workgroup`, exactly like a
```
https://github.com/llvm/llvm-project/pull/225310
More information about the llvm-commits
mailing list