[clang] [llvm] [AMDGPU] Implement AsyncMark Stages (PR #220442)

Sameer Sahasrabuddhe via cfe-commits cfe-commits at lists.llvm.org
Tue Sep 29 22:53:45 PDT 2026


================
@@ -1107,20 +1143,32 @@ void WaitcntBrackets::updateByEvent(HWEvents E, MachineInstr &Inst) {
   }
 }
 
-void WaitcntBrackets::recordAsyncMark(MachineInstr &Inst) {
+void WaitcntBrackets::recordAsyncMark(MachineInstr &Inst, uint32_t StageMask) {
   // In the absence of loops, AsyncMarks can grow linearly with the program
   // until we encounter an ASYNCMARK_WAIT. We could drop the oldest mark above a
   // limit every time we push a new mark, but that seems like unnecessary work
   // in practical cases. We do separately truncate the array when processing a
   // loop, which should be sufficient.
-  AsyncMarks.push_back(AsyncScore);
-  LLVM_DEBUG({
-    dbgs() << "recordAsyncMark:\n" << Inst;
-    for (const auto &Mark : AsyncMarks) {
-      llvm::interleaveComma(Mark, dbgs());
-      dbgs() << '\n';
-    }
-  });
+  //
+  // The mark joins every stage it names. Each stage gets its own copy of the
----------------
ssahasra wrote:

```suggestion
  // A mark is inserted in every stage named. Each stage gets a copy of its corresponding score.
```

https://github.com/llvm/llvm-project/pull/220442


More information about the cfe-commits mailing list