[llvm] [AMDGPU] Correct bidirectional pending-candidate arbitration (PR #226156)
Alan Li via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 24 06:13:15 PDT 2026
https://github.com/lialan created https://github.com/llvm/llvm-project/pull/226156
## Summary
- Pass the actual opposing candidate (`TryCand`) to `tryPendingCandidate`. The current call compares the top candidate with itself when the bottom candidate is pending, so the bottom candidate cannot win the cross-boundary comparison.
- Refresh queue membership for reused candidates and set `PickedPending` from the selected node, so pending-cycle and hazard handling follow the winner.
- Keep this scheduler correctness fix separate from the occupancy-policy changes in #220975.
## Tests
- Add two MIR regressions: a pending bottom PhysReg COPY must beat the top candidate, and `-verify-misched` must not change the schedule through stale cached pending state. Both fail on the original unmodified base (`03fb86920dc0`) and pass with this fix.
- Update five scheduling-sensitive fixtures for the changed instruction order while retaining their semantic assertions.
- On rebased `upstream/main` (`0d7414aec9d6`), `check-llvm-codegen-amdgpu` passed: 5,081 passed, 6 expected failures, 1 unsupported.
>From 3412623590cedc4bb681c0890a6dad09e518c45c Mon Sep 17 00:00:00 2001
From: Alan Li <me at alanli.org>
Date: Wed, 23 Sep 2026 15:23:50 -0700
Subject: [PATCH] fix(amdgpu): correct pending arbitration
Bidirectional scheduling first chooses BotCand and TopCand within their
own queues, then compares one against the other. When the bottom
candidate is pending, Cand is TopCand and TryCand is BotCand. The
pending path instead passed TopCand to tryPendingCandidate(), comparing
the top candidate with itself. The pending bottom candidate could not
win; in the regression, a physical-register COPY lost to an unrelated
top instruction despite the PhysReg bias.
Reused candidates expose a related bookkeeping error. BotPending and
TopPending were reset on every pick but updated only when their queues
were rescanned. A cached candidate can move between Available and
Pending, leaving those flags stale. Read queue membership immediately
before cross-boundary arbitration, compare Cand with TryCand in both
paths, and derive PickedPending from the final winner. The cycle and
hazard advancement in pickNode then follows the node actually selected.
Add two MIR regressions. On the original unmodified base, 03fb86920dc0,
the first picks top SU(0) instead of the pending physical-register COPY
SU(8). The second emits different MIR with and without -verify-misched
because of a stale cached pending flag. Both pass with the fix. These
are scheduler behavior tests, not workload performance measurements.
Correct arbitration changes instruction order and, in some fixtures,
spill layout. Pin the WMMA local-reassignment stress test to a top-down
schedule so it still exercises reassignment. Disable machine scheduling
in the two joint-dominance tests so their intentionally non-dominating
spill layout remains under test. Replace full generated output checks in
the MFMA spill-cost and single-exit-copy tests with their semantic
invariants: all 15 MFMAs stay in VGPR form with no AGPR-form opcode, and
one four-register AGPR-to-VGPR copy feeds both exit MFMAs. This removes
incidental register-number and instruction-order checks while retaining
the regression assertions.
Validation on the original base: check-llvm-codegen-amdgpu passed (5,077
passed, 6 XFAIL, 1 unsupported). Both new MIR tests fail on the
unmodified original base, 03fb86920dc0, and pass with this fix.
Refs #220975
---
llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp | 8 +-
.../AMDGPU/regalloc-spill-wmma-scale.ll | 8 +-
.../rewrite-mfma-form-spill-cost-reset.ll | 482 +---------------
.../AMDGPU/rewrite-mfma-single-exit-copy.ll | 542 +-----------------
...-vgpr-mfma-to-agpr-spill-joint-dom-mir.mir | 2 +
...write-vgpr-mfma-to-agpr-spill-joint-dom.ll | 2 +
.../schedule-pending-bidirectional-cache.mir | 32 ++
...schedule-pending-bidirectional-physreg.mir | 33 ++
8 files changed, 101 insertions(+), 1008 deletions(-)
create mode 100644 llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-cache.mir
create mode 100644 llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-physreg.mir
diff --git a/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp b/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
index 602489ab3a5c14..bfe49cac7d200d 100644
--- a/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+++ b/llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
@@ -608,13 +608,16 @@ SUnit *GCNSchedStrategy::pickNodeBidirectional(bool &IsTopNode,
// Pick best from BotCand and TopCand.
LLVM_DEBUG(dbgs() << "Top Cand: "; traceCandidate(TopCand);
dbgs() << "Bot Cand: "; traceCandidate(BotCand););
+ // Cached candidates may have moved between the available and pending queues
+ // since they were last picked.
+ BotPending = Bot.Pending.isInQueue(BotCand.SU);
+ TopPending = Top.Pending.isInQueue(TopCand.SU);
SchedCandidate Cand = BotPending ? TopCand : BotCand;
SchedCandidate TryCand = BotPending ? BotCand : TopCand;
- PickedPending = BotPending && TopPending;
TryCand.Reason = NoCand;
if (BotPending || TopPending) {
- PickedPending |= tryPendingCandidate(Cand, TopCand, nullptr);
+ tryPendingCandidate(Cand, TryCand, nullptr);
} else {
tryCandidate(Cand, TryCand, nullptr);
}
@@ -622,6 +625,7 @@ SUnit *GCNSchedStrategy::pickNodeBidirectional(bool &IsTopNode,
if (TryCand.Reason != NoCand) {
Cand.setBest(TryCand);
}
+ PickedPending = Cand.AtTop ? TopPending : BotPending;
LLVM_DEBUG(dbgs() << "Picking: "; traceCandidate(Cand););
diff --git a/llvm/test/CodeGen/AMDGPU/regalloc-spill-wmma-scale.ll b/llvm/test/CodeGen/AMDGPU/regalloc-spill-wmma-scale.ll
index b30f463d29d728..bcd235c88654ae 100644
--- a/llvm/test/CodeGen/AMDGPU/regalloc-spill-wmma-scale.ll
+++ b/llvm/test/CodeGen/AMDGPU/regalloc-spill-wmma-scale.ll
@@ -1,11 +1,13 @@
; RUN: llc -mtriple=amdgpu12.50 < %s | FileCheck %s
-; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -stress-regalloc=16 -verify-machineinstrs -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=REASSIGN %}
-; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -stress-regalloc=16 -verify-machineinstrs -enable-local-reassign=true -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=REASSIGN %}
-; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -stress-regalloc=16 -verify-machineinstrs -enable-local-reassign=false -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=NO-REASSIGN --implicit-check-not="can reassign:" %}
+; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -misched-prera-direction=topdown -stress-regalloc=16 -verify-machineinstrs -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=REASSIGN %}
+; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -misched-prera-direction=topdown -stress-regalloc=16 -verify-machineinstrs -enable-local-reassign=true -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=REASSIGN %}
+; RUN: %if asserts %{ llc -mtriple=amdgpu12.50 -mcpu=gfx1250 -misched-prera-direction=topdown -stress-regalloc=16 -verify-machineinstrs -enable-local-reassign=false -debug-only=regalloc -o /dev/null %s 2>&1 | FileCheck %s --check-prefix=NO-REASSIGN --implicit-check-not="can reassign:" %}
; REASSIGN: can reassign:
; NO-REASSIGN: GREEDY REGISTER ALLOCATION
+; Use a fixed scheduling direction so the local-reassign checks see the same
+; register-pressure pattern regardless of bidirectional arbitration.
; Scale operands of WMMA are limited to low 256 VGPRs
; Make sure we do not spill scale operands because of the low 256 restriction.
; CHECK: ; ScratchSize: 0
diff --git a/llvm/test/CodeGen/AMDGPU/rewrite-mfma-form-spill-cost-reset.ll b/llvm/test/CodeGen/AMDGPU/rewrite-mfma-form-spill-cost-reset.ll
index d115f9bd028f27..770067104858dc 100644
--- a/llvm/test/CodeGen/AMDGPU/rewrite-mfma-form-spill-cost-reset.ll
+++ b/llvm/test/CodeGen/AMDGPU/rewrite-mfma-form-spill-cost-reset.ll
@@ -1,5 +1,4 @@
-; NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
-; RUN: llc -mtriple=amdgpu9.50-amd-amdhsa -amdgpu-disable-rewrite-mfma-form-sched-stage=false -verify-machineinstrs -stop-after=machine-scheduler < %s | FileCheck %s
+; RUN: llc -mtriple=amdgpu9.50-amd-amdhsa -amdgpu-disable-rewrite-mfma-form-sched-stage=false -verify-machineinstrs -stop-after=machine-scheduler < %s | FileCheck %s --implicit-check-not=V_MFMA_F32_16X16X32_F16_e64
;
; Regression test for resetRewriteCandsToVGPR() called on the SpillCost > 0
; early-return path inside getRewriteCost().
@@ -31,6 +30,8 @@
; resetRewriteCandsToVGPR() is called before returning, restoring all
; candidate MFMA register classes and opcodes to VGPR form. rewrite() is
; never called. MachineVerifier passes cleanly.
+; Check that all 15 MFMAs remain in VGPR form without depending on their
+; scheduling order or virtual-register numbers.
;
; Expected behavior WITHOUT the fix
; ----------------------------------
@@ -63,480 +64,9 @@ define amdgpu_kernel void @test_spill_cost_reset(
; leaving AGPR-form opcodes (V_MFMA_F32_16X16X32_F16_e64) in the MIR.
; With the fix, all MFMAs must use the VGPR-form opcode (_vgprcd_e64).
; CHECK-LABEL: name: test_spill_cost_reset
- ; CHECK: bb.0.entry:
- ; CHECK-NEXT: successors: %bb.1(0x80000000)
- ; CHECK-NEXT: liveins: $sgpr4_sgpr5
- ; CHECK-NEXT: {{ $}}
- ; CHECK-NEXT: [[COPY:%[0-9]+]]:sgpr_64(p4) = COPY $sgpr4_sgpr5
- ; CHECK-NEXT: early-clobber %1891:sgpr_256 = S_LOAD_DWORDX8_IMM_ec [[COPY]](p4), 16, 0 :: (dereferenceable invariant load (s256) from %ir.a0.kernarg.offset, align 16, addrspace 4)
- ; CHECK-NEXT: early-clobber %1892:sgpr_256 = S_LOAD_DWORDX8_IMM_ec [[COPY]](p4), 48, 0 :: (dereferenceable invariant load (s256) from %ir.b0.kernarg.offset, align 16, addrspace 4)
- ; CHECK-NEXT: [[S_LOAD_DWORD_IMM:%[0-9]+]]:sreg_32_xm0_xexec = S_LOAD_DWORD_IMM [[COPY]](p4), 80, 0 :: (dereferenceable invariant load (s32) from %ir.n.kernarg.offset, align 16, addrspace 4)
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_:%[0-9]+]].sub0:vreg_64_align2 = V_MOV_B32_e32 0, implicit $exec
- ; CHECK-NEXT: [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_1:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_2:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_2:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_2:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_2:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_3:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_3:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_3:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_3:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_4:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_4:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_4:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_4:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_5:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_5:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_5:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_5:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_6:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_6:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_6:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_6:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_7:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_7:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_7:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_7:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_8:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_8:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_8:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_8:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_9:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_9:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_9:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_9:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_10:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_10:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_10:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_10:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_11:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_11:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_11:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_11:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: undef [[V_MOV_B32_e32_12:%[0-9]+]].sub0:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_12:%[0-9]+]].sub1:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_12:%[0-9]+]].sub2:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_12:%[0-9]+]].sub3:vreg_128_align2 = V_MOV_B32_e32 0, implicit $exec, implicit $exec
- ; CHECK-NEXT: [[COPY1:%[0-9]+]]:av_128_align2 = COPY %1891.sub0_sub1_sub2_sub3
- ; CHECK-NEXT: [[COPY2:%[0-9]+]]:av_128_align2 = COPY %1892.sub0_sub1_sub2_sub3
- ; CHECK-NEXT: [[COPY3:%[0-9]+]]:av_128_align2 = COPY %1891.sub4_sub5_sub6_sub7
- ; CHECK-NEXT: [[COPY4:%[0-9]+]]:av_128_align2 = COPY %1892.sub4_sub5_sub6_sub7
- ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY5:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY5:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY6:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY6:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY7:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY7:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY8:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY8:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY9:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY9:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY10:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY10:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY11:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY11:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY12:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY12:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY13:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY13:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY14:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY14:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY15:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY15:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY16:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY16:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY17:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY17:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY18:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY18:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY19:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY19:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY20:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY20:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY21:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY21:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY22:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY22:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY23:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY23:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY24:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY24:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY25:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY25:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY26:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY26:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY27:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY27:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY28:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY28:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY29:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY29:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY30:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY30:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY31:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY31:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY32:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY32:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY33:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY33:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY34:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY34:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY35:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY35:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY36:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY36:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY37:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY37:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY38:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY38:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY39:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY39:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY40:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY40:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY41:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY41:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY42:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY42:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY43:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY43:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY44:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY44:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY45:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY45:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY46:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY46:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY47:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY47:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY48:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY48:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY49:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY49:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY50:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY50:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY51:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY51:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY52:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY52:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY53:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY53:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY54:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY54:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY55:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY55:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY56:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY56:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY57:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY57:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY58:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY58:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY59:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY59:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY60:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY60:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY61:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY61:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY62:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY62:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY63:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY63:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY64:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY64:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY65:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY65:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY66:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY66:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY67:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY67:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY68:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY68:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY69:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY69:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY70:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY70:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY71:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY71:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY72:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY72:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY73:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY73:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY74:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY74:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY75:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY75:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY76:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY76:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY77:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY77:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY78:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY78:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY79:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY79:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY80:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY80:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY81:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY81:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY82:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY82:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY83:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY83:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY84:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY84:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY85:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY85:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY86:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY86:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY87:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY87:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY88:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY88:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY89:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY89:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY90:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY90:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY91:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY91:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY92:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY92:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY93:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY93:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY94:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY94:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY95:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY95:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY96:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY96:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY97:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY97:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY98:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY98:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY99:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY99:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY100:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY100:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY101:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY101:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY102:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY102:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY103:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY103:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY104:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY104:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY105:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY105:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY106:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY106:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY107:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY107:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY108:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY108:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY109:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY109:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY110:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY110:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY111:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY111:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY112:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY112:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY113:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY113:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY114:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY114:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY115:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY115:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY116:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY116:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY117:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY117:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY118:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY118:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY119:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY119:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY120:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY120:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY121:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY121:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY122:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY122:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY123:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY123:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY124:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY124:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY125:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY125:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY126:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY126:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY127:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY127:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY128:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY128:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY129:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY129:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY130:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY130:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: undef [[COPY131:%[0-9]+]].sub0:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: [[COPY131:%[0-9]+]].sub1:vreg_64_align2 = COPY [[V_MOV_B32_e32_]].sub0
- ; CHECK-NEXT: {{ $}}
- ; CHECK-NEXT: bb.1.loop:
- ; CHECK-NEXT: successors: %bb.2(0x04000000), %bb.1(0x7c000000)
- ; CHECK-NEXT: {{ $}}
- ; CHECK-NEXT: [[V_MOV_B32_e32_12:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_12]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY131:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY131]], 8, [[COPY131]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY130:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY130]], 8, [[COPY130]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY129:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY129]], 8, [[COPY129]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY128:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY128]], 8, [[COPY128]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY127:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY127]], 8, [[COPY127]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_11:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_11]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY126:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY126]], 8, [[COPY126]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY125:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY125]], 8, [[COPY125]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY124:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY124]], 8, [[COPY124]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_10:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_10]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY123:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY123]], 8, [[COPY123]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY122:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY122]], 8, [[COPY122]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY121:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY121]], 8, [[COPY121]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_9:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_9]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY120:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY120]], 8, [[COPY120]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY119:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY119]], 8, [[COPY119]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY118:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY118]], 8, [[COPY118]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_8:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_8]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY117:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY117]], 8, [[COPY117]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY116:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY116]], 8, [[COPY116]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY115:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY115]], 8, [[COPY115]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_7:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_7]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY114:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY114]], 8, [[COPY114]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY113:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY113]], 8, [[COPY113]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY112:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY112]], 8, [[COPY112]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_6:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_6]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY111:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY111]], 8, [[COPY111]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY110:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY110]], 8, [[COPY110]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY109:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY109]], 8, [[COPY109]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_5:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_5]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY108:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY108]], 8, [[COPY108]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY107:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY107]], 8, [[COPY107]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY106:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY106]], 8, [[COPY106]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_4:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_4]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY105:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY105]], 8, [[COPY105]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY104:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY104]], 8, [[COPY104]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY103:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY103]], 8, [[COPY103]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_3:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_3]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY102:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY102]], 8, [[COPY102]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY101:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY101]], 8, [[COPY101]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY100:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY100]], 8, [[COPY100]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_2:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_2]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY99:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY99]], 8, [[COPY99]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY98:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY98]], 8, [[COPY98]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY97:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY97]], 8, [[COPY97]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MOV_B32_e32_1]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY96:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY96]], 8, [[COPY96]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY95:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY95]], 8, [[COPY95]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY94:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY94]], 8, [[COPY94]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY93:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY93]], 8, [[COPY93]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY92:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY92]], 8, [[COPY92]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY91:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY91]], 8, [[COPY91]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY3]], [[COPY4]], [[V_MOV_B32_e32_12]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_1:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY1]], [[COPY2]], [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY90:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY90]], 8, [[COPY90]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY89:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY89]], 8, [[COPY89]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY88:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY88]], 8, [[COPY88]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_2:%[0-9]+]]:vreg_128_align2 = V_MFMA_F32_16X16X32_F16_vgprcd_e64 [[COPY3]], [[COPY4]], [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_1]], 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY87:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY87]], 8, [[COPY87]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY86:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY86]], 8, [[COPY86]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY85:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY85]], 8, [[COPY85]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY84:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY84]], 8, [[COPY84]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY83:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY83]], 8, [[COPY83]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY82:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY82]], 8, [[COPY82]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY81:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY81]], 8, [[COPY81]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY80:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY80]], 8, [[COPY80]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY79:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY79]], 8, [[COPY79]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY78:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY78]], 8, [[COPY78]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY77:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY77]], 8, [[COPY77]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY76:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY76]], 8, [[COPY76]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY75:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY75]], 8, [[COPY75]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY74:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY74]], 8, [[COPY74]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY73:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY73]], 8, [[COPY73]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY72:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY72]], 8, [[COPY72]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY71:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY71]], 8, [[COPY71]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY70:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY70]], 8, [[COPY70]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY69:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY69]], 8, [[COPY69]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY68:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY68]], 8, [[COPY68]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY67:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY67]], 8, [[COPY67]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY66:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY66]], 8, [[COPY66]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY65:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY65]], 8, [[COPY65]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY64:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY64]], 8, [[COPY64]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY63:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY63]], 8, [[COPY63]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY62:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY62]], 8, [[COPY62]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY61:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY61]], 8, [[COPY61]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY60:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY60]], 8, [[COPY60]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY59:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY59]], 8, [[COPY59]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY58:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY58]], 8, [[COPY58]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY57:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY57]], 8, [[COPY57]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY56:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY56]], 8, [[COPY56]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY55:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY55]], 8, [[COPY55]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY54:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY54]], 8, [[COPY54]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY53:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY53]], 8, [[COPY53]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY52:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY52]], 8, [[COPY52]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY51:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY51]], 8, [[COPY51]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY50:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY50]], 8, [[COPY50]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY49:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY49]], 8, [[COPY49]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY48:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY48]], 8, [[COPY48]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY47:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY47]], 8, [[COPY47]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY46:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY46]], 8, [[COPY46]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY45:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY45]], 8, [[COPY45]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY44:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY44]], 8, [[COPY44]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY43:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY43]], 8, [[COPY43]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY42:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY42]], 8, [[COPY42]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY41:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY41]], 8, [[COPY41]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY40:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY40]], 8, [[COPY40]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY39:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY39]], 8, [[COPY39]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY38:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY38]], 8, [[COPY38]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY37:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY37]], 8, [[COPY37]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY36:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY36]], 8, [[COPY36]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY35:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY35]], 8, [[COPY35]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY34:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY34]], 8, [[COPY34]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY33:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY33]], 8, [[COPY33]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY32:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY32]], 8, [[COPY32]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY31:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY31]], 8, [[COPY31]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY30:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY30]], 8, [[COPY30]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY29:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY29]], 8, [[COPY29]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY28:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY28]], 8, [[COPY28]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY27:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY27]], 8, [[COPY27]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY26:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY26]], 8, [[COPY26]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY25:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY25]], 8, [[COPY25]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY24:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY24]], 8, [[COPY24]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY23:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY23]], 8, [[COPY23]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY22:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY22]], 8, [[COPY22]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY21:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY21]], 8, [[COPY21]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY20:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY20]], 8, [[COPY20]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY19:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY19]], 8, [[COPY19]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY18:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY18]], 8, [[COPY18]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY17:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY17]], 8, [[COPY17]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY16:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY16]], 8, [[COPY16]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY15:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY15]], 8, [[COPY15]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY14:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY14]], 8, [[COPY14]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY13:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY13]], 8, [[COPY13]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY12:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY12]], 8, [[COPY12]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY11:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY11]], 8, [[COPY11]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY10:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY10]], 8, [[COPY10]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY9:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY9]], 8, [[COPY9]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY8:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY8]], 8, [[COPY8]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY7:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY7]], 8, [[COPY7]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY6:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY6]], 8, [[COPY6]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[COPY5:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[COPY5]], 8, [[COPY5]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_ADD_I32 [[S_MOV_B32_]], 1, implicit-def dead $scc
- ; CHECK-NEXT: S_CMP_LT_I32 [[S_MOV_B32_]], [[S_LOAD_DWORD_IMM]], implicit-def $scc
- ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vreg_64_align2 = nofpexcept V_PK_ADD_F32 8, [[V_MOV_B32_e32_]], 8, [[V_MOV_B32_e32_]], 0, 0, 0, 0, 0, implicit $mode, implicit $exec
- ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.1, implicit killed $scc
- ; CHECK-NEXT: S_BRANCH %bb.2
- ; CHECK-NEXT: {{ $}}
- ; CHECK-NEXT: bb.2.epilogue:
- ; CHECK-NEXT: [[S_LOAD_DWORDX2_IMM:%[0-9]+]]:sreg_64_xexec_xnull = S_LOAD_DWORDX2_IMM [[COPY]](p4), 0, 0 :: (dereferenceable invariant load (s64) from %ir.out.kernarg.offset550, align 16, addrspace 4)
- ; CHECK-NEXT: [[V_MOV_B32_e32_13:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 0, implicit $exec
- ; CHECK-NEXT: GLOBAL_STORE_DWORD_SADDR [[V_MOV_B32_e32_13]], [[V_MFMA_F32_16X16X32_F16_vgprcd_e64_2]].sub0, [[S_LOAD_DWORDX2_IMM]], 0, 0, implicit $exec :: (store (s32) into %ir.out.load, addrspace 1)
- ; CHECK-NEXT: S_ENDPGM 0
+ ; CHECK: bb.1.loop:
+ ; CHECK-COUNT-15: V_MFMA_F32_16X16X32_F16_vgprcd_e64
+ ; CHECK: bb.2.epilogue:
ptr addrspace(1) %out,
<8 x half> %a0, <8 x half> %a1,
<8 x half> %b0, <8 x half> %b1,
diff --git a/llvm/test/CodeGen/AMDGPU/rewrite-mfma-single-exit-copy.ll b/llvm/test/CodeGen/AMDGPU/rewrite-mfma-single-exit-copy.ll
index cfbb80c7696a09..174542b27460e4 100644
--- a/llvm/test/CodeGen/AMDGPU/rewrite-mfma-single-exit-copy.ll
+++ b/llvm/test/CodeGen/AMDGPU/rewrite-mfma-single-exit-copy.ll
@@ -1,4 +1,3 @@
-; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
; RUN: llc -mtriple=amdgpu9.50-amd-amdhsa -amdgpu-disable-rewrite-mfma-form-sched-stage=false < %s | FileCheck %s
; Test that the MFMA rewrite stage generates only one AGPR->VGPR copy
@@ -15,532 +14,21 @@ declare <4 x float> @llvm.amdgcn.mfma.f32.16x16x32.f16(<8 x half>, <8 x half>, <
define amdgpu_kernel void @single_exit_copy(
; CHECK-LABEL: single_exit_copy:
-; CHECK: .Lfunc_begin0:
-; CHECK-NEXT: .file 1 "/tmp" "test"
-; CHECK-NEXT: .loc 1 1 0 ; test:1:0
-; CHECK-NEXT: .cfi_sections .debug_frame
-; CHECK-NEXT: .cfi_startproc
-; CHECK-NEXT: ; %bb.0: ; %entry
-; CHECK-NEXT: .cfi_escape 0x0f, 0x04, 0x30, 0x36, 0xe9, 0x02 ; CFA is 0 in private_wave aspace
-; CHECK-NEXT: .cfi_undefined 16
-; CHECK-NEXT: s_load_dwordx8 s[8:15], s[4:5], 0x0
-; CHECK-NEXT: s_load_dword s0, s[4:5], 0x20
-; CHECK-NEXT: s_mov_b32 s1, 0
-; CHECK-NEXT: v_mov_b32_e32 v31, 0
-; CHECK-NEXT: v_mov_b32_e32 v30, 0
-; CHECK-NEXT: v_mov_b32_e32 v29, 0
-; CHECK-NEXT: v_mov_b32_e32 v28, 0
-; CHECK-NEXT: v_mov_b32_e32 v27, 0
-; CHECK-NEXT: v_mov_b32_e32 v26, 0
-; CHECK-NEXT: v_mov_b32_e32 v25, 0
-; CHECK-NEXT: v_mov_b32_e32 v24, 0
-; CHECK-NEXT: v_mov_b32_e32 v23, 0
-; CHECK-NEXT: v_mov_b32_e32 v22, 0
-; CHECK-NEXT: v_mov_b32_e32 v21, 0
-; CHECK-NEXT: v_mov_b32_e32 v20, 0
-; CHECK-NEXT: v_mov_b32_e32 v19, 0
-; CHECK-NEXT: v_mov_b32_e32 v18, 0
-; CHECK-NEXT: v_mov_b32_e32 v17, 0
-; CHECK-NEXT: v_mov_b32_e32 v16, 0
-; CHECK-NEXT: v_mov_b32_e32 v15, 0
-; CHECK-NEXT: v_mov_b32_e32 v14, 0
-; CHECK-NEXT: v_mov_b32_e32 v13, 0
-; CHECK-NEXT: v_mov_b32_e32 v12, 0
-; CHECK-NEXT: v_mov_b32_e32 v11, 0
-; CHECK-NEXT: v_mov_b32_e32 v10, 0
-; CHECK-NEXT: v_mov_b32_e32 v9, 0
-; CHECK-NEXT: v_mov_b32_e32 v8, 0
-; CHECK-NEXT: v_mov_b32_e32 v7, 0
-; CHECK-NEXT: v_mov_b32_e32 v6, 0
-; CHECK-NEXT: v_mov_b32_e32 v5, 0
-; CHECK-NEXT: v_mov_b32_e32 v4, 0
-; CHECK-NEXT: v_mov_b32_e32 v3, 0
-; CHECK-NEXT: v_mov_b32_e32 v2, 0
-; CHECK-NEXT: v_mov_b32_e32 v1, 0
-; CHECK-NEXT: v_mov_b32_e32 v0, 0
-; CHECK-NEXT: v_mov_b32_e32 v63, 0
-; CHECK-NEXT: v_mov_b32_e32 v62, 0
-; CHECK-NEXT: v_mov_b32_e32 v61, 0
-; CHECK-NEXT: v_mov_b32_e32 v60, 0
-; CHECK-NEXT: v_mov_b32_e32 v59, 0
-; CHECK-NEXT: v_mov_b32_e32 v58, 0
-; CHECK-NEXT: v_mov_b32_e32 v57, 0
-; CHECK-NEXT: v_mov_b32_e32 v56, 0
-; CHECK-NEXT: v_mov_b32_e32 v55, 0
-; CHECK-NEXT: v_mov_b32_e32 v54, 0
-; CHECK-NEXT: v_mov_b32_e32 v53, 0
-; CHECK-NEXT: v_mov_b32_e32 v52, 0
-; CHECK-NEXT: v_mov_b32_e32 v51, 0
-; CHECK-NEXT: v_mov_b32_e32 v50, 0
-; CHECK-NEXT: v_mov_b32_e32 v49, 0
-; CHECK-NEXT: v_mov_b32_e32 v48, 0
-; CHECK-NEXT: v_mov_b32_e32 v47, 0
-; CHECK-NEXT: v_mov_b32_e32 v46, 0
-; CHECK-NEXT: v_mov_b32_e32 v45, 0
-; CHECK-NEXT: v_mov_b32_e32 v44, 0
-; CHECK-NEXT: v_mov_b32_e32 v43, 0
-; CHECK-NEXT: v_mov_b32_e32 v42, 0
-; CHECK-NEXT: v_mov_b32_e32 v41, 0
-; CHECK-NEXT: v_mov_b32_e32 v40, 0
-; CHECK-NEXT: v_mov_b32_e32 v39, 0
-; CHECK-NEXT: v_mov_b32_e32 v38, 0
-; CHECK-NEXT: v_mov_b32_e32 v37, 0
-; CHECK-NEXT: v_mov_b32_e32 v36, 0
-; CHECK-NEXT: v_mov_b32_e32 v35, 0
-; CHECK-NEXT: v_mov_b32_e32 v34, 0
-; CHECK-NEXT: v_mov_b32_e32 v33, 0
-; CHECK-NEXT: v_mov_b32_e32 v32, 0
-; CHECK-NEXT: v_mov_b32_e32 v95, 0
-; CHECK-NEXT: v_mov_b32_e32 v94, 0
-; CHECK-NEXT: v_mov_b32_e32 v93, 0
-; CHECK-NEXT: v_mov_b32_e32 v92, 0
-; CHECK-NEXT: v_mov_b32_e32 v91, 0
-; CHECK-NEXT: v_mov_b32_e32 v90, 0
-; CHECK-NEXT: v_mov_b32_e32 v89, 0
-; CHECK-NEXT: v_mov_b32_e32 v88, 0
-; CHECK-NEXT: v_mov_b32_e32 v87, 0
-; CHECK-NEXT: v_mov_b32_e32 v86, 0
-; CHECK-NEXT: v_mov_b32_e32 v85, 0
-; CHECK-NEXT: v_mov_b32_e32 v84, 0
-; CHECK-NEXT: v_mov_b32_e32 v83, 0
-; CHECK-NEXT: v_mov_b32_e32 v82, 0
-; CHECK-NEXT: v_mov_b32_e32 v81, 0
-; CHECK-NEXT: v_mov_b32_e32 v80, 0
-; CHECK-NEXT: v_mov_b32_e32 v79, 0
-; CHECK-NEXT: v_mov_b32_e32 v78, 0
-; CHECK-NEXT: v_mov_b32_e32 v77, 0
-; CHECK-NEXT: v_mov_b32_e32 v76, 0
-; CHECK-NEXT: v_mov_b32_e32 v75, 0
-; CHECK-NEXT: v_mov_b32_e32 v74, 0
-; CHECK-NEXT: v_mov_b32_e32 v73, 0
-; CHECK-NEXT: v_mov_b32_e32 v72, 0
-; CHECK-NEXT: v_mov_b32_e32 v71, 0
-; CHECK-NEXT: v_mov_b32_e32 v70, 0
-; CHECK-NEXT: v_mov_b32_e32 v69, 0
-; CHECK-NEXT: v_mov_b32_e32 v68, 0
-; CHECK-NEXT: v_mov_b32_e32 v67, 0
-; CHECK-NEXT: v_mov_b32_e32 v66, 0
-; CHECK-NEXT: v_mov_b32_e32 v65, 0
-; CHECK-NEXT: v_mov_b32_e32 v64, 0
-; CHECK-NEXT: v_mov_b32_e32 v127, 0
-; CHECK-NEXT: v_mov_b32_e32 v126, 0
-; CHECK-NEXT: v_mov_b32_e32 v125, 0
-; CHECK-NEXT: v_mov_b32_e32 v124, 0
-; CHECK-NEXT: v_mov_b32_e32 v123, 0
-; CHECK-NEXT: v_mov_b32_e32 v122, 0
-; CHECK-NEXT: v_mov_b32_e32 v121, 0
-; CHECK-NEXT: v_mov_b32_e32 v120, 0
-; CHECK-NEXT: v_mov_b32_e32 v119, 0
-; CHECK-NEXT: v_mov_b32_e32 v118, 0
-; CHECK-NEXT: v_mov_b32_e32 v117, 0
-; CHECK-NEXT: v_mov_b32_e32 v116, 0
-; CHECK-NEXT: v_mov_b32_e32 v115, 0
-; CHECK-NEXT: v_mov_b32_e32 v114, 0
-; CHECK-NEXT: v_mov_b32_e32 v113, 0
-; CHECK-NEXT: v_mov_b32_e32 v112, 0
-; CHECK-NEXT: v_mov_b32_e32 v111, 0
-; CHECK-NEXT: v_mov_b32_e32 v110, 0
-; CHECK-NEXT: v_mov_b32_e32 v109, 0
-; CHECK-NEXT: v_mov_b32_e32 v108, 0
-; CHECK-NEXT: v_mov_b32_e32 v107, 0
-; CHECK-NEXT: v_mov_b32_e32 v106, 0
-; CHECK-NEXT: v_mov_b32_e32 v105, 0
-; CHECK-NEXT: v_mov_b32_e32 v104, 0
-; CHECK-NEXT: v_mov_b32_e32 v103, 0
-; CHECK-NEXT: v_mov_b32_e32 v102, 0
-; CHECK-NEXT: v_mov_b32_e32 v101, 0
-; CHECK-NEXT: v_mov_b32_e32 v100, 0
-; CHECK-NEXT: v_mov_b32_e32 v99, 0
-; CHECK-NEXT: v_mov_b32_e32 v98, 0
-; CHECK-NEXT: v_mov_b32_e32 v97, 0
-; CHECK-NEXT: v_mov_b32_e32 v96, 0
-; CHECK-NEXT: v_mov_b32_e32 v159, 0
-; CHECK-NEXT: v_mov_b32_e32 v158, 0
-; CHECK-NEXT: v_mov_b32_e32 v157, 0
-; CHECK-NEXT: v_mov_b32_e32 v156, 0
-; CHECK-NEXT: v_mov_b32_e32 v155, 0
-; CHECK-NEXT: v_mov_b32_e32 v154, 0
-; CHECK-NEXT: v_mov_b32_e32 v153, 0
-; CHECK-NEXT: v_mov_b32_e32 v152, 0
-; CHECK-NEXT: v_mov_b32_e32 v151, 0
-; CHECK-NEXT: v_mov_b32_e32 v150, 0
-; CHECK-NEXT: v_mov_b32_e32 v149, 0
-; CHECK-NEXT: v_mov_b32_e32 v148, 0
-; CHECK-NEXT: v_mov_b32_e32 v147, 0
-; CHECK-NEXT: v_mov_b32_e32 v146, 0
-; CHECK-NEXT: v_mov_b32_e32 v145, 0
-; CHECK-NEXT: v_mov_b32_e32 v144, 0
-; CHECK-NEXT: v_mov_b32_e32 v143, 0
-; CHECK-NEXT: v_mov_b32_e32 v142, 0
-; CHECK-NEXT: v_mov_b32_e32 v141, 0
-; CHECK-NEXT: v_mov_b32_e32 v140, 0
-; CHECK-NEXT: v_mov_b32_e32 v139, 0
-; CHECK-NEXT: v_mov_b32_e32 v138, 0
-; CHECK-NEXT: v_mov_b32_e32 v137, 0
-; CHECK-NEXT: v_mov_b32_e32 v136, 0
-; CHECK-NEXT: v_mov_b32_e32 v135, 0
-; CHECK-NEXT: v_mov_b32_e32 v134, 0
-; CHECK-NEXT: v_mov_b32_e32 v133, 0
-; CHECK-NEXT: v_mov_b32_e32 v132, 0
-; CHECK-NEXT: v_mov_b32_e32 v131, 0
-; CHECK-NEXT: v_mov_b32_e32 v130, 0
-; CHECK-NEXT: v_mov_b32_e32 v129, 0
-; CHECK-NEXT: v_mov_b32_e32 v128, 0
-; CHECK-NEXT: v_mov_b32_e32 v191, 0
-; CHECK-NEXT: v_mov_b32_e32 v190, 0
-; CHECK-NEXT: v_mov_b32_e32 v189, 0
-; CHECK-NEXT: v_mov_b32_e32 v188, 0
-; CHECK-NEXT: v_mov_b32_e32 v187, 0
-; CHECK-NEXT: v_mov_b32_e32 v186, 0
-; CHECK-NEXT: v_mov_b32_e32 v185, 0
-; CHECK-NEXT: v_mov_b32_e32 v184, 0
-; CHECK-NEXT: v_mov_b32_e32 v183, 0
-; CHECK-NEXT: v_mov_b32_e32 v182, 0
-; CHECK-NEXT: v_mov_b32_e32 v181, 0
-; CHECK-NEXT: v_mov_b32_e32 v180, 0
-; CHECK-NEXT: v_mov_b32_e32 v179, 0
-; CHECK-NEXT: v_mov_b32_e32 v178, 0
-; CHECK-NEXT: v_mov_b32_e32 v177, 0
-; CHECK-NEXT: v_mov_b32_e32 v176, 0
-; CHECK-NEXT: v_mov_b32_e32 v175, 0
-; CHECK-NEXT: v_mov_b32_e32 v174, 0
-; CHECK-NEXT: v_mov_b32_e32 v173, 0
-; CHECK-NEXT: v_mov_b32_e32 v172, 0
-; CHECK-NEXT: v_mov_b32_e32 v171, 0
-; CHECK-NEXT: v_mov_b32_e32 v170, 0
-; CHECK-NEXT: v_mov_b32_e32 v169, 0
-; CHECK-NEXT: v_mov_b32_e32 v168, 0
-; CHECK-NEXT: v_mov_b32_e32 v167, 0
-; CHECK-NEXT: v_mov_b32_e32 v166, 0
-; CHECK-NEXT: v_mov_b32_e32 v165, 0
-; CHECK-NEXT: v_mov_b32_e32 v164, 0
-; CHECK-NEXT: v_mov_b32_e32 v163, 0
-; CHECK-NEXT: v_mov_b32_e32 v162, 0
-; CHECK-NEXT: v_mov_b32_e32 v161, 0
-; CHECK-NEXT: v_mov_b32_e32 v160, 0
-; CHECK-NEXT: v_mov_b32_e32 v223, 0
-; CHECK-NEXT: v_mov_b32_e32 v222, 0
-; CHECK-NEXT: v_mov_b32_e32 v221, 0
-; CHECK-NEXT: v_mov_b32_e32 v220, 0
-; CHECK-NEXT: v_mov_b32_e32 v219, 0
-; CHECK-NEXT: v_mov_b32_e32 v218, 0
-; CHECK-NEXT: v_mov_b32_e32 v217, 0
-; CHECK-NEXT: v_mov_b32_e32 v216, 0
-; CHECK-NEXT: v_mov_b32_e32 v215, 0
-; CHECK-NEXT: v_mov_b32_e32 v214, 0
-; CHECK-NEXT: v_mov_b32_e32 v213, 0
-; CHECK-NEXT: v_mov_b32_e32 v212, 0
-; CHECK-NEXT: v_mov_b32_e32 v211, 0
-; CHECK-NEXT: v_mov_b32_e32 v210, 0
-; CHECK-NEXT: v_mov_b32_e32 v209, 0
-; CHECK-NEXT: v_mov_b32_e32 v208, 0
-; CHECK-NEXT: v_mov_b32_e32 v207, 0
-; CHECK-NEXT: v_mov_b32_e32 v206, 0
-; CHECK-NEXT: v_mov_b32_e32 v205, 0
-; CHECK-NEXT: v_mov_b32_e32 v204, 0
-; CHECK-NEXT: v_mov_b32_e32 v203, 0
-; CHECK-NEXT: v_mov_b32_e32 v202, 0
-; CHECK-NEXT: v_mov_b32_e32 v201, 0
-; CHECK-NEXT: v_mov_b32_e32 v200, 0
-; CHECK-NEXT: v_mov_b32_e32 v199, 0
-; CHECK-NEXT: v_mov_b32_e32 v198, 0
-; CHECK-NEXT: v_mov_b32_e32 v197, 0
-; CHECK-NEXT: v_mov_b32_e32 v196, 0
-; CHECK-NEXT: v_mov_b32_e32 v195, 0
-; CHECK-NEXT: v_mov_b32_e32 v194, 0
-; CHECK-NEXT: v_mov_b32_e32 v193, 0
-; CHECK-NEXT: v_mov_b32_e32 v192, 0
-; CHECK-NEXT: v_mov_b32_e32 v243, 0
-; CHECK-NEXT: v_mov_b32_e32 v242, 0
-; CHECK-NEXT: v_mov_b32_e32 v241, 0
-; CHECK-NEXT: v_mov_b32_e32 v240, 0
-; CHECK-NEXT: v_mov_b32_e32 v247, 0
-; CHECK-NEXT: v_mov_b32_e32 v246, 0
-; CHECK-NEXT: v_mov_b32_e32 v245, 0
-; CHECK-NEXT: v_mov_b32_e32 v244, 0
-; CHECK-NEXT: v_mov_b32_e32 v251, 0
-; CHECK-NEXT: v_mov_b32_e32 v250, 0
-; CHECK-NEXT: v_mov_b32_e32 v249, 0
-; CHECK-NEXT: v_mov_b32_e32 v248, 0
-; CHECK-NEXT: v_mov_b32_e32 v255, 0
-; CHECK-NEXT: v_mov_b32_e32 v254, 0
-; CHECK-NEXT: v_mov_b32_e32 v253, 0
-; CHECK-NEXT: v_mov_b32_e32 v252, 0
-; CHECK-NEXT: v_mov_b32_e32 v239, 0
-; CHECK-NEXT: v_mov_b32_e32 v238, 0
-; CHECK-NEXT: v_mov_b32_e32 v237, 0
-; CHECK-NEXT: v_mov_b32_e32 v236, 0
-; CHECK-NEXT: v_mov_b32_e32 v235, 0
-; CHECK-NEXT: v_mov_b32_e32 v234, 0
-; CHECK-NEXT: v_mov_b32_e32 v233, 0
-; CHECK-NEXT: v_mov_b32_e32 v232, 0
-; CHECK-NEXT: v_mov_b32_e32 v231, 0
-; CHECK-NEXT: v_mov_b32_e32 v230, 0
-; CHECK-NEXT: v_mov_b32_e32 v229, 0
-; CHECK-NEXT: v_mov_b32_e32 v228, 0
-; CHECK-NEXT: v_mov_b32_e32 v227, 0
-; CHECK-NEXT: v_mov_b32_e32 v226, 0
-; CHECK-NEXT: v_mov_b32_e32 v225, 0
-; CHECK-NEXT: v_mov_b32_e32 v224, 0
-; CHECK-NEXT: v_accvgpr_write_b32 a0, 0
-; CHECK-NEXT: v_accvgpr_write_b32 a1, 0
-; CHECK-NEXT: v_accvgpr_write_b32 a2, 0
-; CHECK-NEXT: v_accvgpr_write_b32 a3, 0
-; CHECK-NEXT: .loc 1 1 0 prologue_end ; test:1:0
-; CHECK-NEXT: s_waitcnt lgkmcnt(0)
-; CHECK-NEXT: v_accvgpr_write_b32 a4, s8
-; CHECK-NEXT: v_accvgpr_write_b32 a5, s9
-; CHECK-NEXT: v_accvgpr_write_b32 a6, s10
-; CHECK-NEXT: v_accvgpr_write_b32 a7, s11
-; CHECK-NEXT: v_accvgpr_write_b32 a8, s12
-; CHECK-NEXT: v_accvgpr_write_b32 a9, s13
-; CHECK-NEXT: v_accvgpr_write_b32 a10, s14
-; CHECK-NEXT: v_accvgpr_write_b32 a11, s15
-; CHECK-NEXT: .loc 1 0 0 is_stmt 0 ; :0:0
-; CHECK-NEXT: .Ltmp0:
-; CHECK-NEXT: .p2align 5, , 4
-; CHECK-NEXT: .LBB0_1: ; %loop
-; CHECK-NEXT: ; =>This Inner Loop Header: Depth=1
-; CHECK-NEXT: s_nop 1
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 a[0:3], a[4:7], a[8:11], a[0:3]
-; CHECK-NEXT: s_add_i32 s1, s1, 1
-; CHECK-NEXT: v_pk_add_f32 v[242:243], v[242:243], v[242:243]
-; CHECK-NEXT: v_pk_add_f32 v[240:241], v[240:241], v[240:241]
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 a[0:3], a[4:7], a[8:11], a[0:3]
-; CHECK-NEXT: v_add_f32_e64 v246, v246, v246
-; CHECK-NEXT: v_add_f32_e64 v247, v247, v247
-; CHECK-NEXT: v_pk_add_f32 v[244:245], v[244:245], v[244:245]
-; CHECK-NEXT: v_pk_add_f32 v[250:251], v[250:251], v[250:251]
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 a[0:3], a[4:7], a[8:11], a[0:3]
-; CHECK-NEXT: v_add_f32_e64 v248, v248, v248
-; CHECK-NEXT: v_add_f32_e64 v249, v249, v249
-; CHECK-NEXT: v_pk_add_f32 v[254:255], v[254:255], v[254:255]
-; CHECK-NEXT: v_pk_add_f32 v[252:253], v[252:253], v[252:253]
-; CHECK-NEXT: v_pk_add_f32 v[238:239], v[238:239], v[238:239]
-; CHECK-NEXT: v_pk_add_f32 v[236:237], v[236:237], v[236:237]
-; CHECK-NEXT: v_pk_add_f32 v[234:235], v[234:235], v[234:235]
-; CHECK-NEXT: v_pk_add_f32 v[232:233], v[232:233], v[232:233]
-; CHECK-NEXT: v_pk_add_f32 v[230:231], v[230:231], v[230:231]
-; CHECK-NEXT: v_pk_add_f32 v[228:229], v[228:229], v[228:229]
-; CHECK-NEXT: v_pk_add_f32 v[226:227], v[226:227], v[226:227]
-; CHECK-NEXT: v_pk_add_f32 v[224:225], v[224:225], v[224:225]
-; CHECK-NEXT: v_pk_add_f32 v[222:223], v[222:223], v[222:223]
-; CHECK-NEXT: v_pk_add_f32 v[220:221], v[220:221], v[220:221]
-; CHECK-NEXT: v_pk_add_f32 v[218:219], v[218:219], v[218:219]
-; CHECK-NEXT: v_pk_add_f32 v[216:217], v[216:217], v[216:217]
-; CHECK-NEXT: v_pk_add_f32 v[214:215], v[214:215], v[214:215]
-; CHECK-NEXT: v_pk_add_f32 v[212:213], v[212:213], v[212:213]
-; CHECK-NEXT: v_pk_add_f32 v[210:211], v[210:211], v[210:211]
-; CHECK-NEXT: v_pk_add_f32 v[208:209], v[208:209], v[208:209]
-; CHECK-NEXT: v_pk_add_f32 v[206:207], v[206:207], v[206:207]
-; CHECK-NEXT: v_pk_add_f32 v[204:205], v[204:205], v[204:205]
-; CHECK-NEXT: v_pk_add_f32 v[202:203], v[202:203], v[202:203]
-; CHECK-NEXT: v_pk_add_f32 v[200:201], v[200:201], v[200:201]
-; CHECK-NEXT: v_pk_add_f32 v[198:199], v[198:199], v[198:199]
-; CHECK-NEXT: v_pk_add_f32 v[196:197], v[196:197], v[196:197]
-; CHECK-NEXT: v_pk_add_f32 v[194:195], v[194:195], v[194:195]
-; CHECK-NEXT: v_pk_add_f32 v[192:193], v[192:193], v[192:193]
-; CHECK-NEXT: v_pk_add_f32 v[190:191], v[190:191], v[190:191]
-; CHECK-NEXT: v_pk_add_f32 v[188:189], v[188:189], v[188:189]
-; CHECK-NEXT: v_pk_add_f32 v[186:187], v[186:187], v[186:187]
-; CHECK-NEXT: v_pk_add_f32 v[184:185], v[184:185], v[184:185]
-; CHECK-NEXT: v_pk_add_f32 v[182:183], v[182:183], v[182:183]
-; CHECK-NEXT: v_pk_add_f32 v[180:181], v[180:181], v[180:181]
-; CHECK-NEXT: v_pk_add_f32 v[178:179], v[178:179], v[178:179]
-; CHECK-NEXT: v_pk_add_f32 v[176:177], v[176:177], v[176:177]
-; CHECK-NEXT: v_pk_add_f32 v[174:175], v[174:175], v[174:175]
-; CHECK-NEXT: v_pk_add_f32 v[172:173], v[172:173], v[172:173]
-; CHECK-NEXT: v_pk_add_f32 v[170:171], v[170:171], v[170:171]
-; CHECK-NEXT: v_pk_add_f32 v[168:169], v[168:169], v[168:169]
-; CHECK-NEXT: v_pk_add_f32 v[166:167], v[166:167], v[166:167]
-; CHECK-NEXT: v_pk_add_f32 v[164:165], v[164:165], v[164:165]
-; CHECK-NEXT: v_pk_add_f32 v[162:163], v[162:163], v[162:163]
-; CHECK-NEXT: v_pk_add_f32 v[160:161], v[160:161], v[160:161]
-; CHECK-NEXT: v_pk_add_f32 v[158:159], v[158:159], v[158:159]
-; CHECK-NEXT: v_pk_add_f32 v[156:157], v[156:157], v[156:157]
-; CHECK-NEXT: v_pk_add_f32 v[154:155], v[154:155], v[154:155]
-; CHECK-NEXT: v_pk_add_f32 v[152:153], v[152:153], v[152:153]
-; CHECK-NEXT: v_pk_add_f32 v[150:151], v[150:151], v[150:151]
-; CHECK-NEXT: v_pk_add_f32 v[148:149], v[148:149], v[148:149]
-; CHECK-NEXT: v_pk_add_f32 v[146:147], v[146:147], v[146:147]
-; CHECK-NEXT: v_pk_add_f32 v[144:145], v[144:145], v[144:145]
-; CHECK-NEXT: v_pk_add_f32 v[142:143], v[142:143], v[142:143]
-; CHECK-NEXT: v_pk_add_f32 v[140:141], v[140:141], v[140:141]
-; CHECK-NEXT: v_pk_add_f32 v[138:139], v[138:139], v[138:139]
-; CHECK-NEXT: v_pk_add_f32 v[136:137], v[136:137], v[136:137]
-; CHECK-NEXT: v_pk_add_f32 v[134:135], v[134:135], v[134:135]
-; CHECK-NEXT: v_pk_add_f32 v[132:133], v[132:133], v[132:133]
-; CHECK-NEXT: v_pk_add_f32 v[130:131], v[130:131], v[130:131]
-; CHECK-NEXT: v_pk_add_f32 v[128:129], v[128:129], v[128:129]
-; CHECK-NEXT: v_pk_add_f32 v[126:127], v[126:127], v[126:127]
-; CHECK-NEXT: v_pk_add_f32 v[124:125], v[124:125], v[124:125]
-; CHECK-NEXT: v_pk_add_f32 v[122:123], v[122:123], v[122:123]
-; CHECK-NEXT: v_pk_add_f32 v[120:121], v[120:121], v[120:121]
-; CHECK-NEXT: v_pk_add_f32 v[118:119], v[118:119], v[118:119]
-; CHECK-NEXT: v_pk_add_f32 v[116:117], v[116:117], v[116:117]
-; CHECK-NEXT: v_pk_add_f32 v[114:115], v[114:115], v[114:115]
-; CHECK-NEXT: v_pk_add_f32 v[112:113], v[112:113], v[112:113]
-; CHECK-NEXT: v_pk_add_f32 v[110:111], v[110:111], v[110:111]
-; CHECK-NEXT: v_pk_add_f32 v[108:109], v[108:109], v[108:109]
-; CHECK-NEXT: v_pk_add_f32 v[106:107], v[106:107], v[106:107]
-; CHECK-NEXT: v_pk_add_f32 v[104:105], v[104:105], v[104:105]
-; CHECK-NEXT: v_pk_add_f32 v[102:103], v[102:103], v[102:103]
-; CHECK-NEXT: v_pk_add_f32 v[100:101], v[100:101], v[100:101]
-; CHECK-NEXT: v_pk_add_f32 v[98:99], v[98:99], v[98:99]
-; CHECK-NEXT: v_pk_add_f32 v[96:97], v[96:97], v[96:97]
-; CHECK-NEXT: v_pk_add_f32 v[94:95], v[94:95], v[94:95]
-; CHECK-NEXT: v_pk_add_f32 v[92:93], v[92:93], v[92:93]
-; CHECK-NEXT: v_pk_add_f32 v[90:91], v[90:91], v[90:91]
-; CHECK-NEXT: v_pk_add_f32 v[88:89], v[88:89], v[88:89]
-; CHECK-NEXT: v_pk_add_f32 v[86:87], v[86:87], v[86:87]
-; CHECK-NEXT: v_pk_add_f32 v[84:85], v[84:85], v[84:85]
-; CHECK-NEXT: v_pk_add_f32 v[82:83], v[82:83], v[82:83]
-; CHECK-NEXT: v_pk_add_f32 v[80:81], v[80:81], v[80:81]
-; CHECK-NEXT: v_pk_add_f32 v[78:79], v[78:79], v[78:79]
-; CHECK-NEXT: v_pk_add_f32 v[76:77], v[76:77], v[76:77]
-; CHECK-NEXT: v_pk_add_f32 v[74:75], v[74:75], v[74:75]
-; CHECK-NEXT: v_pk_add_f32 v[72:73], v[72:73], v[72:73]
-; CHECK-NEXT: v_pk_add_f32 v[70:71], v[70:71], v[70:71]
-; CHECK-NEXT: v_pk_add_f32 v[68:69], v[68:69], v[68:69]
-; CHECK-NEXT: v_pk_add_f32 v[66:67], v[66:67], v[66:67]
-; CHECK-NEXT: v_pk_add_f32 v[64:65], v[64:65], v[64:65]
-; CHECK-NEXT: v_pk_add_f32 v[62:63], v[62:63], v[62:63]
-; CHECK-NEXT: v_pk_add_f32 v[60:61], v[60:61], v[60:61]
-; CHECK-NEXT: v_pk_add_f32 v[58:59], v[58:59], v[58:59]
-; CHECK-NEXT: v_pk_add_f32 v[56:57], v[56:57], v[56:57]
-; CHECK-NEXT: v_pk_add_f32 v[54:55], v[54:55], v[54:55]
-; CHECK-NEXT: v_pk_add_f32 v[52:53], v[52:53], v[52:53]
-; CHECK-NEXT: v_pk_add_f32 v[50:51], v[50:51], v[50:51]
-; CHECK-NEXT: v_pk_add_f32 v[48:49], v[48:49], v[48:49]
-; CHECK-NEXT: v_pk_add_f32 v[46:47], v[46:47], v[46:47]
-; CHECK-NEXT: v_pk_add_f32 v[44:45], v[44:45], v[44:45]
-; CHECK-NEXT: v_pk_add_f32 v[42:43], v[42:43], v[42:43]
-; CHECK-NEXT: v_pk_add_f32 v[40:41], v[40:41], v[40:41]
-; CHECK-NEXT: v_pk_add_f32 v[38:39], v[38:39], v[38:39]
-; CHECK-NEXT: v_pk_add_f32 v[36:37], v[36:37], v[36:37]
-; CHECK-NEXT: v_pk_add_f32 v[34:35], v[34:35], v[34:35]
-; CHECK-NEXT: v_pk_add_f32 v[32:33], v[32:33], v[32:33]
-; CHECK-NEXT: v_pk_add_f32 v[30:31], v[30:31], v[30:31]
-; CHECK-NEXT: v_pk_add_f32 v[28:29], v[28:29], v[28:29]
-; CHECK-NEXT: v_pk_add_f32 v[26:27], v[26:27], v[26:27]
-; CHECK-NEXT: v_pk_add_f32 v[24:25], v[24:25], v[24:25]
-; CHECK-NEXT: v_pk_add_f32 v[22:23], v[22:23], v[22:23]
-; CHECK-NEXT: v_pk_add_f32 v[20:21], v[20:21], v[20:21]
-; CHECK-NEXT: v_pk_add_f32 v[18:19], v[18:19], v[18:19]
-; CHECK-NEXT: v_pk_add_f32 v[16:17], v[16:17], v[16:17]
-; CHECK-NEXT: v_pk_add_f32 v[14:15], v[14:15], v[14:15]
-; CHECK-NEXT: v_pk_add_f32 v[12:13], v[12:13], v[12:13]
-; CHECK-NEXT: v_pk_add_f32 v[10:11], v[10:11], v[10:11]
-; CHECK-NEXT: v_pk_add_f32 v[8:9], v[8:9], v[8:9]
-; CHECK-NEXT: v_pk_add_f32 v[6:7], v[6:7], v[6:7]
-; CHECK-NEXT: v_pk_add_f32 v[4:5], v[4:5], v[4:5]
-; CHECK-NEXT: v_pk_add_f32 v[2:3], v[2:3], v[2:3]
-; CHECK-NEXT: s_cmp_lt_i32 s1, s0
-; CHECK-NEXT: v_pk_add_f32 v[0:1], v[0:1], v[0:1]
-; CHECK-NEXT: s_cbranch_scc1 .LBB0_1
-; CHECK-NEXT: ; %bb.2: ; %exit
-; CHECK-NEXT: .Ltmp1:
-; CHECK-NEXT: .loc 1 10 1 is_stmt 1 ; test:10:1
-; CHECK-NEXT: v_accvgpr_write_b32 a4, s8
-; CHECK-NEXT: v_accvgpr_write_b32 a5, s9
-; CHECK-NEXT: v_accvgpr_write_b32 a6, s10
-; CHECK-NEXT: v_accvgpr_write_b32 a7, s11
-; CHECK-NEXT: s_load_dwordx2 s[0:1], s[4:5], 0x28
-; CHECK-NEXT: v_accvgpr_write_b32 a8, s12
-; CHECK-NEXT: v_accvgpr_write_b32 a9, s13
-; CHECK-NEXT: v_accvgpr_write_b32 a10, s14
-; CHECK-NEXT: v_accvgpr_write_b32 a11, s15
-; CHECK-NEXT: v_accvgpr_write_b32 a15, v3
-; CHECK-NEXT: v_accvgpr_write_b32 a14, v2
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 a[0:3], a[4:7], a[8:11], a[0:3]
-; CHECK-NEXT: v_accvgpr_write_b32 a13, v1
-; CHECK-NEXT: v_accvgpr_write_b32 a12, v0
-; CHECK-NEXT: v_mov_b32_e32 v0, 0
-; CHECK-NEXT: s_waitcnt lgkmcnt(0)
-; CHECK-NEXT: global_store_dwordx4 v0, v[240:243], s[0:1] offset:2032
-; CHECK-NEXT: v_mov_b32_e32 v0, 0
-; CHECK-NEXT: global_store_dwordx4 v0, v[244:247], s[0:1] offset:2016
-; CHECK-NEXT: global_store_dwordx4 v0, v[248:251], s[0:1] offset:2000
-; CHECK-NEXT: v_mov_b32_e32 v0, 0
-; CHECK-NEXT: v_accvgpr_read_b32 v243, a3
-; CHECK-NEXT: v_accvgpr_read_b32 v242, a2
-; CHECK-NEXT: v_accvgpr_read_b32 v241, a1
-; CHECK-NEXT: v_accvgpr_read_b32 v240, a0
-; CHECK-NEXT: global_store_dwordx4 v0, v[252:255], s[0:1] offset:1984
-; CHECK-NEXT: .loc 1 20 1 ; test:20:1
-; CHECK-NEXT: s_nop 0
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 v[244:247], a[4:7], a[8:11], v[240:243]
-; CHECK-NEXT: s_nop 7
-; CHECK-NEXT: v_mov_b32_e32 v245, 0
-; CHECK-NEXT: global_store_dwordx4 v245, v[236:239], s[0:1] offset:1968
-; CHECK-NEXT: .loc 1 30 1 ; test:30:1
-; CHECK-NEXT: s_nop 1
-; CHECK-NEXT: v_mfma_f32_16x16x32_f16 v[236:239], a[8:11], a[4:7], v[240:243]
-; CHECK-NEXT: s_nop 7
-; CHECK-NEXT: v_add_f32_e32 v238, 1.0, v244
-; CHECK-NEXT: v_add_f32_e32 v239, 2.0, v236
-; CHECK-NEXT: global_store_dwordx2 v245, v[238:239], s[0:1]
-; CHECK-NEXT: global_store_dwordx4 v245, v[232:235], s[0:1] offset:1952
-; CHECK-NEXT: global_store_dwordx4 v245, v[228:231], s[0:1] offset:1936
-; CHECK-NEXT: global_store_dwordx4 v245, v[224:227], s[0:1] offset:1920
-; CHECK-NEXT: global_store_dwordx4 v245, v[220:223], s[0:1] offset:368
-; CHECK-NEXT: global_store_dwordx4 v245, v[216:219], s[0:1] offset:352
-; CHECK-NEXT: global_store_dwordx4 v245, v[212:215], s[0:1] offset:336
-; CHECK-NEXT: global_store_dwordx4 v245, v[208:211], s[0:1] offset:320
-; CHECK-NEXT: global_store_dwordx4 v245, v[204:207], s[0:1] offset:304
-; CHECK-NEXT: global_store_dwordx4 v245, v[200:203], s[0:1] offset:288
-; CHECK-NEXT: global_store_dwordx4 v245, v[196:199], s[0:1] offset:272
-; CHECK-NEXT: global_store_dwordx4 v245, v[192:195], s[0:1] offset:256
-; CHECK-NEXT: global_store_dwordx4 v245, v[188:191], s[0:1] offset:496
-; CHECK-NEXT: global_store_dwordx4 v245, v[184:187], s[0:1] offset:480
-; CHECK-NEXT: global_store_dwordx4 v245, v[180:183], s[0:1] offset:464
-; CHECK-NEXT: global_store_dwordx4 v245, v[176:179], s[0:1] offset:448
-; CHECK-NEXT: global_store_dwordx4 v245, v[172:175], s[0:1] offset:432
-; CHECK-NEXT: global_store_dwordx4 v245, v[168:171], s[0:1] offset:416
-; CHECK-NEXT: global_store_dwordx4 v245, v[164:167], s[0:1] offset:400
-; CHECK-NEXT: global_store_dwordx4 v245, v[160:163], s[0:1] offset:384
-; CHECK-NEXT: global_store_dwordx4 v245, v[156:159], s[0:1] offset:624
-; CHECK-NEXT: global_store_dwordx4 v245, v[152:155], s[0:1] offset:608
-; CHECK-NEXT: global_store_dwordx4 v245, v[148:151], s[0:1] offset:592
-; CHECK-NEXT: global_store_dwordx4 v245, v[144:147], s[0:1] offset:576
-; CHECK-NEXT: global_store_dwordx4 v245, v[140:143], s[0:1] offset:560
-; CHECK-NEXT: global_store_dwordx4 v245, v[136:139], s[0:1] offset:544
-; CHECK-NEXT: global_store_dwordx4 v245, v[132:135], s[0:1] offset:528
-; CHECK-NEXT: global_store_dwordx4 v245, v[128:131], s[0:1] offset:512
-; CHECK-NEXT: global_store_dwordx4 v245, v[124:127], s[0:1] offset:752
-; CHECK-NEXT: global_store_dwordx4 v245, v[120:123], s[0:1] offset:736
-; CHECK-NEXT: global_store_dwordx4 v245, v[116:119], s[0:1] offset:720
-; CHECK-NEXT: global_store_dwordx4 v245, v[112:115], s[0:1] offset:704
-; CHECK-NEXT: global_store_dwordx4 v245, v[108:111], s[0:1] offset:688
-; CHECK-NEXT: global_store_dwordx4 v245, v[104:107], s[0:1] offset:672
-; CHECK-NEXT: global_store_dwordx4 v245, v[100:103], s[0:1] offset:656
-; CHECK-NEXT: global_store_dwordx4 v245, v[96:99], s[0:1] offset:640
-; CHECK-NEXT: global_store_dwordx4 v245, v[92:95], s[0:1] offset:880
-; CHECK-NEXT: global_store_dwordx4 v245, v[88:91], s[0:1] offset:864
-; CHECK-NEXT: global_store_dwordx4 v245, v[84:87], s[0:1] offset:848
-; CHECK-NEXT: global_store_dwordx4 v245, v[80:83], s[0:1] offset:832
-; CHECK-NEXT: global_store_dwordx4 v245, v[76:79], s[0:1] offset:816
-; CHECK-NEXT: global_store_dwordx4 v245, v[72:75], s[0:1] offset:800
-; CHECK-NEXT: global_store_dwordx4 v245, v[68:71], s[0:1] offset:784
-; CHECK-NEXT: global_store_dwordx4 v245, v[64:67], s[0:1] offset:768
-; CHECK-NEXT: global_store_dwordx4 v245, v[60:63], s[0:1] offset:1008
-; CHECK-NEXT: global_store_dwordx4 v245, v[56:59], s[0:1] offset:992
-; CHECK-NEXT: global_store_dwordx4 v245, v[52:55], s[0:1] offset:976
-; CHECK-NEXT: global_store_dwordx4 v245, v[48:51], s[0:1] offset:960
-; CHECK-NEXT: global_store_dwordx4 v245, v[44:47], s[0:1] offset:944
-; CHECK-NEXT: global_store_dwordx4 v245, v[40:43], s[0:1] offset:928
-; CHECK-NEXT: global_store_dwordx4 v245, v[36:39], s[0:1] offset:912
-; CHECK-NEXT: global_store_dwordx4 v245, v[32:35], s[0:1] offset:896
-; CHECK-NEXT: global_store_dwordx4 v245, v[28:31], s[0:1] offset:1136
-; CHECK-NEXT: global_store_dwordx4 v245, v[24:27], s[0:1] offset:1120
-; CHECK-NEXT: global_store_dwordx4 v245, v[20:23], s[0:1] offset:1104
-; CHECK-NEXT: global_store_dwordx4 v245, v[16:19], s[0:1] offset:1088
-; CHECK-NEXT: global_store_dwordx4 v245, v[12:15], s[0:1] offset:1072
-; CHECK-NEXT: global_store_dwordx4 v245, v[8:11], s[0:1] offset:1056
-; CHECK-NEXT: global_store_dwordx4 v245, v[4:7], s[0:1] offset:1040
-; CHECK-NEXT: global_store_dwordx4 v245, a[12:15], s[0:1] offset:1024
-; CHECK-NEXT: s_endpgm
-; CHECK-NEXT: .Ltmp2:
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: ; %bb.2: ; %exit
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: v_mfma_f32_16x16x32_f16 a[0:3], a[4:7], a[8:11], a[0:3]
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: v_accvgpr_read_b32 v[[#HI:]], a3
+; CHECK-NEXT: v_accvgpr_read_b32 v[[#HI - 1]], a2
+; CHECK-NEXT: v_accvgpr_read_b32 v[[#HI - 2]], a1
+; CHECK-NEXT: v_accvgpr_read_b32 v[[#HI - 3]], a0
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: v_mfma_f32_16x16x32_f16 v{{\[[0-9]+:[0-9]+\]}}, a[4:7], a[8:11], v{{\[}}[[#HI - 3]]:[[#HI]]{{\]}}
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: v_mfma_f32_16x16x32_f16 v{{\[[0-9]+:[0-9]+\]}}, a[8:11], a[4:7], v{{\[}}[[#HI - 3]]:[[#HI]]{{\]}}
+; CHECK-NOT: v_accvgpr_read_b32
+; CHECK: s_endpgm
; Loop body: MFMAs in AGPR form.
; Exit block: one MFMA consuming the loop result (AGPR form), then
; exactly 4 v_accvgpr_read (one 128-bit copy), shared by both vgprcd MFMAs.
diff --git a/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom-mir.mir b/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom-mir.mir
index d916eac55e68c1..8c7304d9059d14 100644
--- a/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom-mir.mir
+++ b/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom-mir.mir
@@ -1,5 +1,6 @@
# REQUIRES: asserts
# RUN: llc -mtriple=amdgpu9.50-amd-amdhsa \
+# RUN: -enable-misched=false \
# RUN: -start-before=register-coalescer \
# RUN: -stop-after=amdgpu-rewrite-agpr-copy-mfma \
# RUN: -debug-only=amdgpu-rewrite-agpr-copy-mfma -filetype=null %s 2>&1 \
@@ -8,6 +9,7 @@
# It is legal for a spill reload to not be jointly dominated by the slot's
# spill stores. The AGPR rewrite pass must not unspill such a slot into a
# vreg, otherwise the compiler will crash.
+# Disable machine scheduling to preserve the spill layout exercising this case.
# CHECK: Skipping ${{[a-zA-Z0-9_]+}}: some reachable load not jointly dominated by stores
diff --git a/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom.ll b/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom.ll
index c233012fd709b8..9f024b28b3d8f6 100644
--- a/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom.ll
+++ b/llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-joint-dom.ll
@@ -1,5 +1,6 @@
; REQUIRES: asserts
; RUN: llc -O3 -mtriple=amdgpu9.50-amd-amdhsa \
+; RUN: -enable-misched=false \
; RUN: -stop-after=amdgpu-rewrite-agpr-copy-mfma \
; RUN: -debug-only=amdgpu-rewrite-agpr-copy-mfma -filetype=null %s 2>&1 \
; RUN: | FileCheck %s
@@ -9,6 +10,7 @@
; It is legal for a spill reload to not be jointly dominated by the slot's
; spill stores. The AGPR rewrite pass must not unspill such a slot into a
; vreg, otherwise the compiler will crash.
+; Disable machine scheduling to preserve the spill layout exercising this case.
; CHECK: Skipping ${{[a-zA-Z0-9_]+}}: some reachable load not jointly dominated by stores
diff --git a/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-cache.mir b/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-cache.mir
new file mode 100644
index 00000000000000..e43d4577dd10cb
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-cache.mir
@@ -0,0 +1,32 @@
+# RUN: llc -mtriple=amdgpu9.08 -run-pass=machine-scheduler -misched-prera-direction=bidirectional -misched-limit=1 -verify-machineinstrs %s -o %t.normal
+# RUN: llc -mtriple=amdgpu9.08 -run-pass=machine-scheduler -misched-prera-direction=bidirectional -misched-limit=1 -verify-machineinstrs -verify-misched %s -o %t.verified
+# RUN: cmp %t.normal %t.verified
+# RUN: FileCheck %s < %t.normal
+# REQUIRES: asserts
+
+# A cached candidate must retain its pending status when scheduling switches
+# directions. Enabling scheduler verification must not change the schedule.
+# CHECK-LABEL: name: pending_queue_ready_cycle
+# CHECK: %0:sgpr_128 = IMPLICIT_DEF
+# CHECK-NEXT: %4:vreg_128 = BUFFER_LOAD_DWORDX4_OFFSET
+
+---
+name: pending_queue_ready_cycle
+tracksRegLiveness: true
+body: |
+ bb.0:
+ liveins: $sgpr4_sgpr5
+
+ %2:sgpr_128 = IMPLICIT_DEF
+ %14:vgpr_32 = IMPLICIT_DEF
+ %15:vgpr_32 = IMPLICIT_DEF
+ %18:areg_512 = IMPLICIT_DEF
+ %18:areg_512 = V_MFMA_F32_16X16X1F32_mac_e64 %15, %14, %18, 0, 0, 0, implicit $mode, implicit $exec
+ %5:vreg_128 = BUFFER_LOAD_DWORDX4_OFFSET %2, 0, 0, 0, 0, implicit $exec
+ %18:areg_512 = V_MFMA_F32_16X16X1F32_mac_e64 %15, %14, %18, 0, 0, 0, implicit $mode, implicit $exec
+ undef %84.sub0:vreg_128_align2 = V_ADD_U32_e32 %5.sub0, %14, implicit $exec
+ %7:vreg_512 = COPY %18
+ SCHED_BARRIER 0
+ S_NOP 0, implicit %18, implicit %7, implicit %84
+ S_ENDPGM 0
+...
diff --git a/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-physreg.mir b/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-physreg.mir
new file mode 100644
index 00000000000000..4d001a492ac46f
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/schedule-pending-bidirectional-physreg.mir
@@ -0,0 +1,33 @@
+# RUN: llc -mtriple=amdgpu9.08 -run-pass=machine-scheduler -misched-prera-direction=bidirectional -misched-limit=1 -verify-machineinstrs -debug-only=machine-scheduler -o /dev/null %s 2>&1 | FileCheck %s
+# REQUIRES: asserts
+
+# The physical-register COPY is bottom-ready but pending because the one-slot
+# available queue is full. It must win the cross-boundary PhysReg comparison.
+# CHECK-LABEL: pending_physreg_cross_boundary:%bb.0
+# CHECK: Queue BotQ.P: 8
+# CHECK: Top Cand: Cand SU(0) FIRST
+# CHECK-NEXT: Bot Cand: Cand SU(8) PHYS-REG
+# CHECK-NEXT: Picking: Cand SU(8) PHYS-REG
+
+---
+name: pending_physreg_cross_boundary
+tracksRegLiveness: true
+registers:
+ - { id: 99, class: vgpr_32 }
+body: |
+ bb.0:
+ liveins: $sgpr4_sgpr5
+
+ %2:sgpr_128 = IMPLICIT_DEF
+ %14:vgpr_32 = IMPLICIT_DEF
+ %15:vgpr_32 = IMPLICIT_DEF
+ %18:areg_512 = IMPLICIT_DEF
+ %18:areg_512 = V_MFMA_F32_16X16X1F32_mac_e64 %15, %14, %18, 0, 0, 0, implicit $mode, implicit $exec
+ %5:vreg_128 = BUFFER_LOAD_DWORDX4_OFFSET %2, 0, 0, 0, 0, implicit $exec
+ %18:areg_512 = V_MFMA_F32_16X16X1F32_mac_e64 %15, %14, %18, 0, 0, 0, implicit $mode, implicit $exec
+ undef %84.sub0:vreg_128_align2 = V_ADD_U32_e32 %5.sub0, %14, implicit $exec
+ $vgpr0_vgpr1_vgpr2_vgpr3_vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15 = COPY %18, implicit-def %99
+ SCHED_BARRIER 0
+ S_NOP 0, implicit %18, implicit %84, implicit %99, implicit $vgpr0_vgpr1_vgpr2_vgpr3_vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+ S_ENDPGM 0
+...
More information about the llvm-commits
mailing list