[llvm] [AMDGPU][docs] define cluster scope and multicast loads (gfx1250) (PR #225310)

Sameer Sahasrabuddhe via llvm-commits llvm-commits at lists.llvm.org
Sun Oct 4 23:24:40 PDT 2026


================
@@ -99,16 +99,59 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
     ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
     i32 immarg %offset,         ; offset (immediate) applied to both global and LDS address
     i32 immarg %cpol,           ; cache policy (immediate)
-    [i32 %m0])                  ; workgroup broadcast mask, cluster variants only (in M0)
+    [i32 %mask])                ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
 ```
 
 The bit-size encoded in the name can be 8, 32, 64 or 128.
 
 Loads data from global memory to LDS. The `%offset` is applied to both the
 global and LDS addresses.
 
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-multicast-dma)=
+
+**Cluster Multicast Loads**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup multicast mask*: bit `i` corresponds to the
+  workgroup with cluster index `i`. Each workgroup that wants to participate in
+  the combined fetch must set its own bit and all participating workgroups must
+  supply an identical mask (and the same address). If the mask is all-zero, the
+  operation behaves like an ordinary, non-cluster load: it returns only to the
+  requesting workgroup.
+- Bit `[16]` selects the timeout behavior. When clear, a target-defined timeout
+  is used: requests may be combined if they arrive within that window. When set,
+  an *early timeout* is used: as soon as the L2 cache supplies the data, it is
+  returned to whichever waves have already issued their requests.
+
+Each participating workgroup issues its own request, with one wave per
----------------
ssahasra wrote:

> But also, it's UB for multiple waves to make the same request, yeah?

I hadn't thought of that! It's not obvious to me that it should be UB. If two waves from the same workgroup make multicast requests which are not in happens before, and even if the requests are identical, there is no reason why each of them should not be completed independently by the implementation. Will need to go and check the details.

https://github.com/llvm/llvm-project/pull/225310


More information about the llvm-commits mailing list