[clang] [llvm] [AMDGPU] Implement AsyncMark Stages (PR #220442)

Sameer Sahasrabuddhe via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 22:53:44 PDT 2026


================
@@ -16,31 +16,81 @@ internally by the compiler. A thread that initiates one or more async operations
 An *asyncmark* created by a thread can be used to track async operations
 initiated by that thread.
 
+### Stages
+
+A *stage* names a kind of async operation. Each async operation *belongs to* the
+one stage determined by the instruction that initiates it.
+
+The stages are:
+
+| Bit | Stage | Async operations |
+|---|---|---|
+| 0 | `TENSOR` | tensor loads and stores |
+| 1 | `GLOBAL_LOAD_ASYNC_TO_LDS` | global loads async to LDS |
+| 2 | `GLOBAL_LOAD_ASYNC_TO_LDS_MCAST` | multicast (cluster) global loads async to LDS |
+| 3 | `GLOBAL_STORE_ASYNC_FROM_LDS` | async global stores from LDS |
+| 5 | `BUFFER_GLOBAL_LOAD` | buffer loads to LDS and pre-gfx1250 global loads to LDS |
+
+Bits 4 and 6 through 10 are reserved for future async operations, and no
+operation belongs to them yet.
+
+Which async operations a given subtarget actually has is described in
+{ref}`AMDGPU DMA Operations <amdgpu-dma-operations>`. A stage exists on every
+subtarget that supports asyncmarks, whether or not that subtarget has any
+operation belonging to it.
+
+### Stage Masks
+
+Both intrinsics take a *stage mask*: an 11-bit value in which a set bit names a
----------------
ssahasra wrote:

```suggestion
`asyncmark` intrinsics take a *stage mask*: an 11-bit value in which a set bit names a
```

https://github.com/llvm/llvm-project/pull/220442


More information about the llvm-commits mailing list