[llvm] [AMDGPU][docs] define cluster scope and multicast loads (gfx1250) (PR #225310)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 04:49:17 PDT 2026


================
@@ -99,16 +99,59 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
     ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
     i32 immarg %offset,         ; offset (immediate) applied to both global and LDS address
     i32 immarg %cpol,           ; cache policy (immediate)
-    [i32 %m0])                  ; workgroup broadcast mask, cluster variants only (in M0)
+    [i32 %mask])                ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
 ```
 
 The bit-size encoded in the name can be 8, 32, 64 or 128.
 
 Loads data from global memory to LDS. The `%offset` is applied to both the
 global and LDS addresses.
 
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-multicast-dma)=
+
+**Cluster Multicast Loads**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup multicast mask*: bit `i` corresponds to the
+  workgroup with cluster index `i`. Each workgroup that wants to participate in
+  the combined fetch must set its own bit and all participating workgroups must
----------------
Pierre-vh wrote:

ditto with fetch

https://github.com/llvm/llvm-project/pull/225310


More information about the llvm-commits mailing list