[clang] [llvm] [AMDGPU] Implement AsyncMark Stages (PR #220442)

Sameer Sahasrabuddhe via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 29 22:53:49 PDT 2026


================
@@ -233,10 +357,13 @@ After inlining, both `B` and `C` are *completed-at* `Y`.
 ### Optimization
 
 The implementation may eliminate asyncmark/wait intrinsics in the following
-cases. These are just examples and not meant to be an exhaustive list.
+cases. These are just examples and not meant to be an exhaustive list. Each
+applies per stage: the sequences are independent, so an asyncmark in one stage
+is unaffected by the waits of another.
 
-1. An `asyncmark` operation which remains in the current sequence along every
-   path that reaches the function exit.
+1. An `asyncmark` in a stage `S` which remains in the current sequence of `S`
+   along every path that reaches the function exit. An `asyncmark` operation is
+   eliminated entirely only once this holds for every stage it covers.
----------------
ssahasra wrote:

Not really true for independent stages, right? Some stages can be removed while keeping the asyncmark operation in place?

https://github.com/llvm/llvm-project/pull/220442


More information about the llvm-commits mailing list