[llvm] [docs][AMDGPU] move DMA operations to a separate file (PR #206917)
Sameer Sahasrabuddhe via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 2 01:56:53 PDT 2026
================
@@ -0,0 +1,145 @@
+(amdgpu-dma-operations)=
+
+# AMDGPU DMA Operations
+
+```{contents}
+:local:
+```
+
+## Introduction
+
+DMA operations transfer data between different kinds of memory directly without
+occupying registers in the invoking wave. They are usually
+{ref}`asynchronous<amdgpu-async-operations>` asynchronous, and require the user
+to explicitly track completion using {ref}`asyncmarks<amdgpu-async-operations>`.
+
+All DMA operations support the same cache modifiers as ordinary load/store
+operations from registers. They cannot be performed atomically.
+
+### GFX9 DMA
+
+Each GFX9 DMA instruction has a synchronous counterpart (e.g.,
+``@llvm.amdgcn.load.to.lds`` for ``@llvm.amdgcn.load.async.to.lds``). The
+synchronous variants perform the same operation, but the compiler automatically
+ensures completion before their side-effects are used.
+
+GFX9 DMA instructions implement volatile (via ``aux/cpol`` bit 31) and
+nontemporal (via metadata) as if they were loads from the global address space.
+
+**Flat/Global Addressing**
+
+```llvm
+void @llvm.amdgcn.load[.async].to.lds.pN(
+ ptr addrspace(N) %src, ; base pointer to load from (per-lane)
+ ptr addrspace(3) %lds_base, ; LDS base pointer (wave-uniform)
+ i32 immarg %size, ; data byte size: 1/2/4 (12/16 for gfx950)
+ i32 immarg %offset, ; offset applied to both src and LDS address
+ i32 immarg %cpol) ; cache policy
+```
+
+Loads data from global memory to LDS. The data size can be 1, 2, or 4 bytes
+(gfx950 also allows 12 or 16 bytes). The LDS address is implicitly offset by
+``4 * lane_id`` bytes for sizes up to 4 bytes, and by ``16 * lane_id`` bytes
+for larger sizes.
+
+The ``%lds_base`` pointer must be wave-uniform.
+
+The source pointer is overloaded on address space. Supported address spaces are
+flat (0), global (1), and buffer fat pointer (7).
+
+``@llvm.amdgcn.load[.async].to.lds.p7`` (buffer pointer) is lowered to
+``@llvm.amdgcn.raw.ptr.buffer.load[.async].lds`` before instruction selection.
+
+```llvm
+void @llvm.amdgcn.global.load[.async].lds(
+ ptr addrspace(1) %src, ; global base pointer to load from (per-lane)
+ ptr addrspace(3) %lds_base, ; LDS base pointer (wave-uniform)
+ i32 immarg %size, ; data byte size: 1/2/4 (12/16 for gfx950)
+ i32 immarg %offset, ; offset applied to both global and LDS address
+ i32 immarg %cpol) ; cache policy
+```
+
+This is identical to ``@llvm.amdgcn.load[.async].to.lds.p1``.
+
+**Buffer Addressing**
+
+```llvm
+void @llvm.amdgcn.{raw|struct}[.ptr].buffer.load[.async].lds(
+ %rsrc, ; buffer resource descriptor (SGPR):
+ ; <4 x i32> or ptr addrspace(8)
+ ptr addrspace(3) %lds_base, ; LDS base pointer (wave-uniform)
+ i32 immarg %size, ; data byte size: 1/2/4 (12/16 for gfx950)
+ [i32 %vindex,] ; VGPR buffer index (struct variants only)
+ i32 %voffset, ; VGPR offset (included in bounds checking)
+ i32 %soffset, ; SGPR/imm offset (excluded from bounds checking)
+ i32 immarg %offset, ; imm offset (included in bounds checking)
+ i32 immarg %cpol) ; cache policy
+```
+
+Loads data from a buffer resource to LDS.
+
+The ``%lds_base`` pointer must be wave-uniform.
+
+The intrinsics differ in two orthogonal ways:
----------------
ssahasra wrote:
"Overload" typically implies that the base name of the function remains the same but the arguments are different, and their types get included in the automatic mangling. But these don't really fit that pattern very well.
https://github.com/llvm/llvm-project/pull/206917
More information about the llvm-commits
mailing list