[llvm] 565fdb7 - [AMDGPU] SIInsertWaitcnts: rebase async marks into the merged frame at CFG joins (#211688)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 30 09:01:22 PDT 2026
Author: Yoonseo Choi
Date: 2026-07-30T09:01:16-07:00
New Revision: 565fdb70e6063e27230f6397d2cbccf1003beaed
URL: https://github.com/llvm/llvm-project/commit/565fdb70e6063e27230f6397d2cbccf1003beaed
DIFF: https://github.com/llvm/llvm-project/commit/565fdb70e6063e27230f6397d2cbccf1003beaed.diff
LOG: [AMDGPU] SIInsertWaitcnts: rebase async marks into the merged frame at CFG joins (#211688)
This PR fixes `WaitcntBrackets::mergeAsyncMarks` to call `mergeScore` at
CFG join block even when one of its predecessors has no AsyncMark
At a CFG join block, the upper bounds of counts from predecessors are
merged. Then,
`mergeScore` rebase each predecessor’s `Score` by the merged upper
bound. Even if a predecessor has no AsyncMark, scores of other
predecessors with AsyncMarks should be rebased by `mergeScore`.
Suppose a predecessor, bb B, without AsyncMark (Score 0) visited later
than another predecessor with AsyncMarks (positive Score), bb A. If
merging upper bounds of bb B into that of bb A increases the new global
upper bounds to UB’ from UB, bb A’s Score should be rebased by the new
UB’. Previously in that case, only UB was merged but Score of bb A was
not updated as `mergeScore` was not called due to bb B’s having zero
score.
Detailed example can be seen in added test,
`llvm/test/CodeGen/AMDGPU/asyncmark-merge-rebase-pregfx12.mir`. Here are
the descriptions on the test.
The join (bb.3) has two predecessors:
- (bb.2) an AsyncMark predecessor: one async LDS DMA + ASYNCMARK, so one
AsyncMark recorded at LOAD_CNT score 1 (small frame, UB = 1).
- (bb.1) a load predecessor: several async LDS DMAs with no ASYNCMARK.
These stay
outstanding (they write LDS, no VGPR result) and extend the frame to UB
= 4,
but contribute no async marks.
si-insert-waitcnts visits the branch target bb.2 before the fall-through
bb.1. So, bb.2 seeds bb.3's incoming state and bb.1 with an empty
AsyncMarks list is merged in.
When UB of LOAD_CNT is still merged to 4 = max(1,4).
If bb.2's Score 1 is not rebased to new UB 4, the vmcnt value wrongly
becomes 3 (= 4 - 1), which was the previous behavior.
Score 1 must be rebased to 4 (= 1 + (4 - 1)) to make correct vmcnt, 0 (=
4 - 4). That rebase logic is already within `mergeScore`. `mergeScore`
should be called even when one of the predecessors has zero Score.
Notice that the bug doesn’t show up if bb.1 was visited before bb.2.
Added:
llvm/test/CodeGen/AMDGPU/asyncmark-merge-rebase-pregfx12.mir
Modified:
llvm/lib/Target/AMDGPU/SIInsertWaitcnts.cpp
Removed:
################################################################################
diff --git a/llvm/lib/Target/AMDGPU/SIInsertWaitcnts.cpp b/llvm/lib/Target/AMDGPU/SIInsertWaitcnts.cpp
index d7a994a45757d..b3390a8f5fac0 100644
--- a/llvm/lib/Target/AMDGPU/SIInsertWaitcnts.cpp
+++ b/llvm/lib/Target/AMDGPU/SIInsertWaitcnts.cpp
@@ -2796,21 +2796,28 @@ bool WaitcntBrackets::mergeAsyncMarks(ArrayRef<MergeInfo> MergeInfos,
});
// Merge element-wise using the existing mergeScore function and the
- // appropriate MergeInfo for each counter type. Iterate only while we have
- // elements in both vectors.
- unsigned OtherSize = OtherMarks.size();
- unsigned OurSize = AsyncMarks.size();
- unsigned MergeCount = std::min(OtherSize, OurSize);
- // OtherMarks is empty -> OtherSize == 0 -> MergeCount == 0.
- // Our existing marks are the conservative result; return early to avoid
- // passing MergeCount == 0 to seq_inclusive which asserts Begin <= End.
- if (MergeCount == 0)
- return StrictDom;
- for (auto Idx : seq_inclusive<unsigned>(1, MergeCount)) {
- for (auto T : inst_counter_types(Context->MaxCounter)) {
- StrictDom |= mergeScore(MergeInfos[T], AsyncMarks[OurSize - Idx][T],
- OtherMarks[OtherSize - Idx][T]);
- }
+ // appropriate MergeInfo for each counter type. Rebase and merge EVERY
+ // surviving mark into the new (widened) frame, aligned by recency from the
+ // back. A mark with no counterpart on the other side merges against the
+ // zero/identity mark: mergeScore still applies MyShift, so the mark is
+ // re-expressed in the new frame instead of being left with a stale
+ // (pre-merge) score.
+ const unsigned OtherSize = OtherMarks.size();
+ const unsigned OurSize = AsyncMarks.size();
+ // After the both-empty early-return above, max(AsyncMarks, OtherMarks) >= 1,
+ // and the erase/pad steps normalize AsyncMarks to exactly MaxSize. Hence
+ // OurSize == MaxSize >= 1 (as long as MaxAsyncMarks != 0), so the
+ // seq_inclusive(1, OurSize) below never trips its "Begin <= End" assertion
+ // the way seq_inclusive(1, MergeCount) could when OtherSize == 0.
+ assert(OurSize >= 1 &&
+ "AsyncMarks padded to MaxSize >= 1 (needs MaxAsyncMarks != 0)");
+
+ for (auto Idx : seq_inclusive<unsigned>(1, OurSize)) {
+ const CounterValueArray &OtherMark =
+ Idx <= OtherSize ? OtherMarks[OtherSize - Idx] : ZeroMark;
+ for (auto T : inst_counter_types(Context->MaxCounter))
+ StrictDom |=
+ mergeScore(MergeInfos[T], AsyncMarks[OurSize - Idx][T], OtherMark[T]);
}
LLVM_DEBUG({
diff --git a/llvm/test/CodeGen/AMDGPU/asyncmark-merge-rebase-pregfx12.mir b/llvm/test/CodeGen/AMDGPU/asyncmark-merge-rebase-pregfx12.mir
new file mode 100644
index 0000000000000..dd41980f42b14
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/asyncmark-merge-rebase-pregfx12.mir
@@ -0,0 +1,166 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgpu9.50-amd-amdhsa -run-pass=si-insert-waitcnts -o - %s | FileCheck %s
+
+# Regression test for WaitcntBrackets::mergeAsyncMarks rebasing surviving async
+# marks into the merged (widened) counter frame at a CFG join.
+#
+# WAIT_ASYNCMARK at bb.3, CFG join of bb.1 and bb.2, should be zero.
+# bb.1 has 4 LDS loads but no ASYNCMARK.
+
+---
+name: async_mark_pred_merged_first
+tracksRegLiveness: true
+machineFunctionInfo:
+ occupancy: 8
+body: |
+ ; CHECK-LABEL: name: async_mark_pred_merged_first
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x40000000), %bb.2(0x40000000)
+ ; CHECK-NEXT: liveins: $sgpr0, $sgpr1, $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_WAITCNT .Vmcnt_0_Expcnt_0_Lgkmcnt_0
+ ; CHECK-NEXT: S_CMP_LG_U32 $sgpr0, $sgpr1, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.2, implicit killed $scc
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: $m0 = S_MOV_B32 0
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: $m0 = S_MOV_B32 0
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: ASYNCMARK
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.3:
+ ; CHECK-NEXT: liveins: $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: WAIT_ASYNCMARK 0
+ ; CHECK-NEXT: S_WAITCNT .Vmcnt_0
+ ; CHECK-NEXT: renamable $vgpr0 = DS_READ_B32_gfx9 killed $vgpr2, 0, 0, implicit $exec :: (load (s32), addrspace 3)
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $sgpr0, $sgpr1, $vgpr0_vgpr1, $vgpr2
+
+ S_CMP_LG_U32 $sgpr0, $sgpr1, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.2, implicit killed $scc
+
+ ; fall-through: load predecessor, merged second with empty OtherMarks
+ bb.1:
+ successors: %bb.3
+ liveins: $vgpr0_vgpr1, $vgpr2
+
+ $m0 = S_MOV_B32 0
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ S_BRANCH %bb.3
+
+ ; branch target: async-mark predecessor, merged first (seeds incoming state)
+ bb.2:
+ successors: %bb.3
+ liveins: $vgpr0_vgpr1, $vgpr2
+
+ $m0 = S_MOV_B32 0
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ASYNCMARK
+ S_BRANCH %bb.3
+
+ bb.3:
+ liveins: $vgpr2
+
+ WAIT_ASYNCMARK 0
+ renamable $vgpr0 = DS_READ_B32_gfx9 killed $vgpr2, 0, 0, implicit $exec :: (load (s32), addrspace 3)
+ S_ENDPGM 0
+...
+
+---
+name: load_pred_merged_first
+tracksRegLiveness: true
+machineFunctionInfo:
+ occupancy: 8
+body: |
+ ; CHECK-LABEL: name: load_pred_merged_first
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x40000000), %bb.2(0x40000000)
+ ; CHECK-NEXT: liveins: $sgpr0, $sgpr1, $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_WAITCNT .Vmcnt_0_Expcnt_0_Lgkmcnt_0
+ ; CHECK-NEXT: S_CMP_LG_U32 $sgpr0, $sgpr1, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.2, implicit killed $scc
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: $m0 = S_MOV_B32 0
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: ASYNCMARK
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr0_vgpr1, $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: $m0 = S_MOV_B32 0
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.3:
+ ; CHECK-NEXT: liveins: $vgpr2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: WAIT_ASYNCMARK 0
+ ; CHECK-NEXT: S_WAITCNT .Vmcnt_0
+ ; CHECK-NEXT: renamable $vgpr0 = DS_READ_B32_gfx9 killed $vgpr2, 0, 0, implicit $exec :: (load (s32), addrspace 3)
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1, %bb.2
+ liveins: $sgpr0, $sgpr1, $vgpr0_vgpr1, $vgpr2
+
+ S_CMP_LG_U32 $sgpr0, $sgpr1, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.2, implicit killed $scc
+
+ ; fall-through: async-mark predecessor, merged second (non-empty OtherMarks)
+ bb.1:
+ successors: %bb.3
+ liveins: $vgpr0_vgpr1, $vgpr2
+
+ $m0 = S_MOV_B32 0
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ ASYNCMARK
+ S_BRANCH %bb.3
+
+ ; branch target: load predecessor, merged first (seeds incoming state)
+ bb.2:
+ successors: %bb.3
+ liveins: $vgpr0_vgpr1, $vgpr2
+
+ $m0 = S_MOV_B32 0
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ GLOBAL_LOAD_LDS_DWORD $vgpr0_vgpr1, 0, 0, 1, implicit $m0, implicit $exec :: (load (s32), addrspace 1), (store (s32), addrspace 3)
+ S_BRANCH %bb.3
+
+ bb.3:
+ liveins: $vgpr2
+
+ WAIT_ASYNCMARK 0
+ renamable $vgpr0 = DS_READ_B32_gfx9 killed $vgpr2, 0, 0, implicit $exec :: (load (s32), addrspace 3)
+ S_ENDPGM 0
+...
More information about the llvm-commits
mailing list