[llvm] [AMDGPU] Pre-commit tests for loop-carried memory waits (PR #220356)
Krzysztof Drewniak via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 1 14:50:18 PDT 2026
https://github.com/krzysz00 updated https://github.com/llvm/llvm-project/pull/220356
>From 4004b0b89a5f3160548ac46d3c33a16bb67af57d Mon Sep 17 00:00:00 2001
From: Krzysztof Drewniak <Krzysztof.Drewniak at amd.com>
Date: Tue, 1 Sep 2026 19:12:49 +0000
Subject: [PATCH 1/2] [AMDGPU] Pre-commit tests for loop-carried memory waits
SIInsertWaitcnts is currently dropping waits in a
(fence; read; write; branch) loop in some cases. Pre-commit tests to
show the problem.
AI disclosure: test generated by AI, but I poked them into not being terrible.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply at anthropic.com>
---
...waitcnt-loop-carried-fence-drain-gfx12.mir | 57 +++++++++++++++++++
.../waitcnt-loop-carried-fence-drain.mir | 53 +++++++++++++++++
2 files changed, 110 insertions(+)
create mode 100644 llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
create mode 100644 llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
new file mode 100644
index 0000000000000..28fefe5704c5d
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
@@ -0,0 +1,57 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+
+# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
+
+---
+name: block_sync_lds_loop
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: block_sync_lds_loop
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr1, $vgpr2, $sgpr0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_WAIT_LOADCNT_DSCNT .Loadcnt_0_Dscnt_0
+ ; CHECK-NEXT: S_WAIT_EXPCNT 0
+ ; CHECK-NEXT: S_WAIT_SAMPLECNT 0
+ ; CHECK-NEXT: S_WAIT_BVHCNT 0
+ ; CHECK-NEXT: S_WAIT_KMCNT 0
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.1(0x40000000), %bb.2(0x40000000)
+ ; CHECK-NEXT: liveins: $vgpr1, $vgpr2, $sgpr0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_BARRIER
+ ; CHECK-NEXT: renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+ ; CHECK-NEXT: S_WAIT_DSCNT 0
+ ; CHECK-NEXT: DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+ ; CHECK-NEXT: S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.1, implicit $scc
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ liveins: $vgpr1, $vgpr2, $sgpr0
+
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.1, %bb.2
+ liveins: $vgpr1, $vgpr2, $sgpr0
+
+ S_WAIT_DSCNT_soft 0
+ S_BARRIER
+ S_WAIT_DSCNT_soft 0
+ renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+ DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+ S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.1, implicit $scc
+ S_BRANCH %bb.2
+
+ bb.2:
+ S_ENDPGM 0
+...
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
new file mode 100644
index 0000000000000..48d0296ba2920
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
@@ -0,0 +1,53 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+
+# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
+
+---
+name: block_sync_lds_loop
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: block_sync_lds_loop
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr1, $vgpr2, $sgpr0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_WAITCNT .Vmcnt_0_Expcnt_0_Lgkmcnt_0
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.1(0x40000000), %bb.2(0x40000000)
+ ; CHECK-NEXT: liveins: $vgpr1, $vgpr2, $sgpr0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_BARRIER
+ ; CHECK-NEXT: renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+ ; CHECK-NEXT: S_WAITCNT .Lgkmcnt_0
+ ; CHECK-NEXT: DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+ ; CHECK-NEXT: S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.1, implicit $scc
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ liveins: $vgpr1, $vgpr2, $sgpr0
+
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.1, %bb.2
+ liveins: $vgpr1, $vgpr2, $sgpr0
+
+ S_WAITCNT_soft 49279
+ S_BARRIER
+ S_WAITCNT_soft 49279
+ renamable $vgpr3 = DS_READ_B32_gfx9 $vgpr2, 0, 0, implicit $exec :: (load (s32) from `ptr addrspace(3) poison`, addrspace 3)
+ DS_WRITE_B32_gfx9 $vgpr1, $vgpr3, 0, 0, implicit $exec :: (store (s32) into `ptr addrspace(3) poison`, addrspace 3)
+ S_CMP_EQ_U32 $sgpr0, 0, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.1, implicit $scc
+ S_BRANCH %bb.2
+
+ bb.2:
+ S_ENDPGM 0
+...
>From c3bff081aadeb9bc8ca86f66a2f0d1e93139c553 Mon Sep 17 00:00:00 2001
From: Krzysztof Drewniak <Krzysztof.Drewniak at amd.com>
Date: Tue, 1 Sep 2026 21:49:44 +0000
Subject: [PATCH 2/2] Fix RUN line style
---
.../CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir | 2 +-
llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
index 28fefe5704c5d..70ac188505e4a 100644
--- a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain-gfx12.mir
@@ -1,5 +1,5 @@
# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
-# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgpu12.00-amd-amdhsa -run-pass=si-insert-waitcnts -o - %s | FileCheck %s
# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
diff --git a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
index 48d0296ba2920..220782473b863 100644
--- a/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
+++ b/llvm/test/CodeGen/AMDGPU/waitcnt-loop-carried-fence-drain.mir
@@ -1,5 +1,5 @@
# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
-# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass si-insert-waitcnts -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgpu9.0a-amd-amdhsa -run-pass=si-insert-waitcnts -o - %s | FileCheck %s
# Check that writes from one loop iteration are waited on at the top of the next loop iteration.
More information about the llvm-commits
mailing list