[llvm] [AMDGPU] Track async events for vm_vsrc simplifications (PR #211333)

Jeffrey Byrnes via llvm-commits llvm-commits at lists.llvm.org
Fri Jul 24 15:51:32 PDT 2026


================
@@ -1322,8 +1314,8 @@ void WaitcntBrackets::simplifyVmVsrc(const AMDGPU::Waitcnt &CheckWait,
   // have read its VGPR sources, but only if there are no other outstanding VMEM
   // operations that use a different counter (like SAMPLE_CNT).
   static constexpr AMDGPU::InstCounterType VmemCounters[] = {
-      AMDGPU::LOAD_CNT, AMDGPU::STORE_CNT, AMDGPU::SAMPLE_CNT, AMDGPU::BVH_CNT,
-      AMDGPU::DS_CNT};
+      AMDGPU::LOAD_CNT, AMDGPU::STORE_CNT, AMDGPU::SAMPLE_CNT,
+      AMDGPU::BVH_CNT,  AMDGPU::DS_CNT,    AMDGPU::ASYNC_CNT};
----------------
jrbyrnes wrote:

> Presumably we would want to treat the outstanding ASYNCMARK as real memory state, and thus this simplification should honor it, right?

The latest just disallows wait based simplifications in the presence of outstanding asyncmarks. The optimization relies on being able to do meaningful comparisons between the wait for the vm_vsrc and the wait for various memory types. Since the outstanding asyncmark are not contained in this structure, we lose the meaningful comparison. This is a conservative approach, and likely can be cleaned up a bit.

https://github.com/llvm/llvm-project/pull/211333


More information about the llvm-commits mailing list