[clang] [llvm] [AMDGPU] Implement AsyncMark Stages (PR #220442)
Sameer Sahasrabuddhe via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 29 22:53:49 PDT 2026
================
@@ -233,10 +357,13 @@ After inlining, both `B` and `C` are *completed-at* `Y`.
### Optimization
The implementation may eliminate asyncmark/wait intrinsics in the following
-cases. These are just examples and not meant to be an exhaustive list.
+cases. These are just examples and not meant to be an exhaustive list. Each
+applies per stage: the sequences are independent, so an asyncmark in one stage
+is unaffected by the waits of another.
-1. An `asyncmark` operation which remains in the current sequence along every
- path that reaches the function exit.
+1. An `asyncmark` in a stage `S` which remains in the current sequence of `S`
+ along every path that reaches the function exit. An `asyncmark` operation is
+ eliminated entirely only once this holds for every stage it covers.
----------------
ssahasra wrote:
Not really true for independent stages, right? Some stages can be removed while keeping the asyncmark operation in place?
https://github.com/llvm/llvm-project/pull/220442
More information about the llvm-commits
mailing list