[llvm] [AMDGPU] Pre-commit tests for loop-carried memory waits (PR #220356)

Krzysztof Drewniak via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 1 14:50:18 PDT 2026


https://github.com/krzysz00 updated https://github.com/llvm/llvm-project/pull/220356

>From 4004b0b89a5f3160548ac46d3c33a16bb67af57d Mon Sep 17 00:00:00 2001
From: Krzysztof Drewniak <Krzysztof.Drewniak at amd.com>
Date: Tue, 1 Sep 2026 19:12:49 +0000
Subject: [PATCH 1/2] [AMDGPU] Pre-commit tests for loop-carried memory waits

SIInsertWaitcnts is currently dropping waits in a
(fence; read; write; branch) loop in some cases. Pre-commit tests to
show the problem.

AI disclosure: test generated by AI, but I poked them into not being terrible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply at anthropic.com>
---
 ...waitcnt-loop-carried-fence-drain-gfx12.mir | 57 +++++++++++++++++++
 .../waitcnt-loop-carried-fence-drain.mir      | 53 +++++++++++++++++
 2 files changed, 110 insertions(+)
 create mode 100644 llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
 create mode 100644 llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir

diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
new file mode 100644
index 0000000000000..28fefe5704c5d
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
@@ -0,0 +1,57 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+
+# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
+
+---
+name: block_sync_lds_loop
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: block_sync_lds_loop
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT:   liveins: $vgpr1, $vgpr2, $sgpr0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_WAIT_LOADCNT_DSCNT .Loadcnt_0_Dscnt_0
+  ; CHECK-NEXT:   S_WAIT_EXPCNT 0
+  ; CHECK-NEXT:   S_WAIT_SAMPLECNT 0
+  ; CHECK-NEXT:   S_WAIT_BVHCNT 0
+  ; CHECK-NEXT:   S_WAIT_KMCNT 0
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.1(0x40000000), %bb.2(0x40000000)
+  ; CHECK-NEXT:   liveins: $vgpr1, $vgpr2, $sgpr0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_BARRIER
+  ; CHECK-NEXT:   renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+  ; CHECK-NEXT:   S_WAIT_DSCNT 0
+  ; CHECK-NEXT:   DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+  ; CHECK-NEXT:   S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.1, implicit $scc
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    liveins: $vgpr1, $vgpr2, $sgpr0
+
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.1, %bb.2
+    liveins: $vgpr1, $vgpr2, $sgpr0
+
+    S_WAIT_DSCNT_soft 0
+    S_BARRIER
+    S_WAIT_DSCNT_soft 0
+    renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+    DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+    S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
new file mode 100644
index 0000000000000..48d0296ba2920
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
@@ -0,0 +1,53 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+
+# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
+
+---
+name: block_sync_lds_loop
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: block_sync_lds_loop
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT:   liveins: $vgpr1, $vgpr2, $sgpr0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_WAITCNT .Vmcnt_0_Expcnt_0_Lgkmcnt_0
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.1(0x40000000), %bb.2(0x40000000)
+  ; CHECK-NEXT:   liveins: $vgpr1, $vgpr2, $sgpr0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_BARRIER
+  ; CHECK-NEXT:   renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+  ; CHECK-NEXT:   S_WAITCNT .Lgkmcnt_0
+  ; CHECK-NEXT:   DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+  ; CHECK-NEXT:   S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.1, implicit $scc
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    liveins: $vgpr1, $vgpr2, $sgpr0
+
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.1, %bb.2
+    liveins: $vgpr1, $vgpr2, $sgpr0
+
+    S_WAITCNT_soft 49279
+    S_BARRIER
+    S_WAITCNT_soft 49279
+    renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+    DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+    S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.1, implicit $scc
+    S_BRANCH %bb.2
+
+  bb.2:
+    S_ENDPGM 0
+...

>From c3bff081aadeb9bc8ca86f66a2f0d1e93139c553 Mon Sep 17 00:00:00 2001
From: Krzysztof Drewniak <Krzysztof.Drewniak at amd.com>
Date: Tue, 1 Sep 2026 21:49:44 +0000
Subject: [PATCH 2/2] Fix RUN line style

---
 .../CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir   | 2 +-
 llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir   | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
index 28fefe5704c5d..70ac188505e4a 100644
--- a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
@@ -1,5 +1,5 @@
 # NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
-# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass=si-insert-waitcnts -o - %s | FileCheck %s
 
 # Check that writes from one loop iteration are waited on at the top of the next loop iteration.
 
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
index 48d0296ba2920..220782473b863 100644
--- a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
@@ -1,5 +1,5 @@
 # NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
-# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass=si-insert-waitcnts -o - %s | FileCheck %s
 
 # Check that writes from one loop iteration are waited on at the top of the next loop iteration.
 



More information about the llvm-commits mailing list