[llvm] [AMDGPU][docs] define cluster scope and multicast loads (gfx1250) (PR #225310)

Pierre van Houtryve via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 04:49:16 PDT 2026


================
@@ -99,16 +99,59 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
     ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
     i32 immarg %offset,         ; offset (immediate) applied to both global and LDS address
     i32 immarg %cpol,           ; cache policy (immediate)
-    [i32 %m0])                  ; workgroup broadcast mask, cluster variants only (in M0)
+    [i32 %mask])                ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
 ```
 
 The bit-size encoded in the name can be 8, 32, 64 or 128.
 
 Loads data from global memory to LDS. The `%offset` is applied to both the
 global and LDS addresses.
 
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-multicast-dma)=
+
+**Cluster Multicast Loads**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
----------------
Pierre-vh wrote:

```suggestion
<amdgpu-clusters>` share a single load from global memory when they are all
```
For consistency, I don't think we use fetch often, we generally use read/access/load?

https://github.com/llvm/llvm-project/pull/225310


More information about the llvm-commits mailing list