[llvm] [AMDGPU][docs] define cluster scope and multicast loads (gfx1250) (PR #225310)
Pierre van Houtryve via llvm-commits
llvm-commits at lists.llvm.org
Fri Oct 2 04:49:17 PDT 2026
================
@@ -99,16 +99,59 @@ void @llvm.amdgcn.{global|cluster}.load.async.to.lds.b<N>(
ptr addrspace(3) %lds_base, ; LDS base pointer (per-lane)
i32 immarg %offset, ; offset (immediate) applied to both global and LDS address
i32 immarg %cpol, ; cache policy (immediate)
- [i32 %m0]) ; workgroup broadcast mask, cluster variants only (in M0)
+ [i32 %mask]) ; workgroup multicast mask, cluster variants only (wave-uniform, in M0)
```
The bit-size encoded in the name can be 8, 32, 64 or 128.
Loads data from global memory to LDS. The `%offset` is applied to both the
global and LDS addresses.
-The `cluster` variants add a `%m0` argument for workgroup broadcast. The
-broadcast mask selects which workgroups within a cluster participate in the load.
+(amdgpu-cluster-multicast-dma)=
+
+**Cluster Multicast Loads**
+
+The `cluster` variants let several workgroups in a {ref}`workgroup cluster
+<amdgpu-clusters>` share a single fetch from global memory when they are all
+loading the same data to their respective LDS. The `%mask` argument identifies
+the group of workgroups whose requests may be combined, and controls the request
+timeout. The mask is wave-uniform and passed in the `M0` register:
+
+- Bits `[15:0]` are the *workgroup multicast mask*: bit `i` corresponds to the
+ workgroup with cluster index `i`. Each workgroup that wants to participate in
+ the combined fetch must set its own bit and all participating workgroups must
----------------
Pierre-vh wrote:
ditto with fetch
https://github.com/llvm/llvm-project/pull/225310
More information about the llvm-commits
mailing list