[llvm] [InlineSpiller][AMDGPU] Skip undef lanes when spilling a partial-def super-register (PR #226689)
via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 26 05:56:49 PDT 2026
https://github.com/xgxanq created https://github.com/llvm/llvm-project/pull/226689
When the greedy allocator live-range-splits a wide register (e.g. an sgpr_128 buffer descriptor) on a single sub-register, it produces sibling vregs each defined on one sub-lane and undef elsewhere. InlineSpiller shares one stack slot among all descendants of the Original value, so a full-width store of such a partially-defined sibling writes its undef lanes over a live sibling value in the same slot and corrupts it. On gfx950 this was observed (via asm data-flow analysis) as an in-loop buffer-descriptor lane being clobbered by an unrelated in-loop spill sharing the slot.
The hazard is target-independent (InlineSpiller's shared-slot logic is generic; SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs). Fix it in two places:
1. hoistSpillInsideBB: bail out unless every sub-range of the hoisted value is live at the hoist point, mirroring the cross-BB guard in isSpillCandBB (#177703) that the in-BB path was missing.
2. insertSpill: for a real (partial) spill, compute the defined-lane mask and hand it to the target via a new TargetInstrInfo::setSpillDefinedLaneMask hook (default no-op). AMDGPU records it in the SI_SPILL_S*_SAVE $lanemask operand; spillSGPR then writelanes only the defined dwords, leaving the slot's other lanes (a sibling value) intact. A -1 mask preserves the original full-width behavior.
Existing SGPR-spill tests are regenerated to carry the new -1 $lanemask operand. Adds inline-spiller-partial-subreg-shared-slot.ll as a gfx950 end-to-end regression check.
>From c3e394a80253ba9e3c4bcff430c216af900aa12f Mon Sep 17 00:00:00 2001
From: anqfu <anqfu at amd.com>
Date: Sat, 26 Sep 2026 12:36:43 +0000
Subject: [PATCH] [InlineSpiller][AMDGPU] Skip undef lanes when spilling a
partial-def super-register
When the greedy allocator live-range-splits a wide register (e.g. an sgpr_128
buffer descriptor) on a single sub-register, it produces sibling vregs each
defined on one sub-lane and undef elsewhere. InlineSpiller shares one stack slot
among all descendants of the Original value, so a full-width store of such a
partially-defined sibling writes its undef lanes over a live sibling value in
the same slot and corrupts it. On gfx950 this was observed (via asm data-flow
analysis) as an in-loop buffer-descriptor lane being clobbered by an unrelated
in-loop spill sharing the slot.
The hazard is target-independent (InlineSpiller's shared-slot logic is generic;
SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs). Fix it in two
places:
1. hoistSpillInsideBB: bail out unless every sub-range of the hoisted value is
live at the hoist point, mirroring the cross-BB guard in isSpillCandBB
(#177703) that the in-BB path was missing.
2. insertSpill: for a real (partial) spill, compute the defined-lane mask and
hand it to the target via a new TargetInstrInfo::setSpillDefinedLaneMask
hook (default no-op). AMDGPU records it in the SI_SPILL_S*_SAVE $lanemask
operand; spillSGPR then writelanes only the defined dwords, leaving the
slot's other lanes (a sibling value) intact. A -1 mask preserves the
original full-width behavior.
Existing SGPR-spill tests are regenerated to carry the new -1 $lanemask
operand. Adds inline-spiller-partial-subreg-shared-slot.ll as a gfx950
end-to-end regression check.
---
llvm/include/llvm/CodeGen/TargetInstrInfo.h | 6 +
llvm/lib/CodeGen/InlineSpiller.cpp | 34 ++-
llvm/lib/Target/AMDGPU/SIInstrInfo.cpp | 37 ++-
llvm/lib/Target/AMDGPU/SIInstrInfo.h | 3 +
llvm/lib/Target/AMDGPU/SIInstructions.td | 10 +-
llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp | 51 +++-
.../CodeGen/AMDGPU/bug-undef-spilled-agpr.mir | 2 +-
.../greedy-alloc-fail-sgpr1024-spill.mir | 2 +-
...nfloop-subrange-spill-inspect-subrange.mir | 2 +-
...line-spiller-partial-subreg-shared-slot.ll | 257 ++++++++++++++++++
.../sgpr-spill-dead-frame-in-dbg-value.mir | 4 +-
...ip-processing-stack-arg-dbg-value-list.mir | 2 +-
...fi-skip-processing-stack-arg-dbg-value.mir | 2 +-
.../AMDGPU/sgpr-spill-overlap-wwm-reserve.mir | 2 +-
.../AMDGPU/sgpr-spill-partially-undef.mir | 18 +-
...pr-spill-to-vmem-scc-clobber-unhandled.mir | 2 +-
.../AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir | 12 +-
.../AMDGPU/sgpr-spill-vmem-large-frame.mir | 4 +-
.../AMDGPU/sgpr-spill-wrong-stack-id.mir | 10 +-
llvm/test/CodeGen/AMDGPU/sgpr-spill.mir | 22 +-
.../si-lower-sgpr-spills-cycle-header.mir | 4 +-
...wer-sgpr-spills-initial-insert-in-body.mir | 2 +-
...er-sgpr-spills-initial-insert-in-latch.mir | 2 +-
...si-lower-sgpr-spills-multi-entry-cycle.mir | 4 +-
.../AMDGPU/spill-reg-tuple-super-reg-use.mir | 4 +-
.../CodeGen/AMDGPU/spill-scavenge-offset.ll | 14 +-
.../AMDGPU/spill-sgpr-to-virtual-vgpr.mir | 12 +-
.../CodeGen/AMDGPU/spill-special-sgpr.mir | 4 +-
llvm/test/CodeGen/AMDGPU/spill192.mir | 2 +-
llvm/test/CodeGen/AMDGPU/spill224.mir | 2 +-
llvm/test/CodeGen/AMDGPU/spill288.mir | 2 +-
llvm/test/CodeGen/AMDGPU/spill320.mir | 2 +-
llvm/test/CodeGen/AMDGPU/spill352.mir | 2 +-
llvm/test/CodeGen/AMDGPU/spill384.mir | 2 +-
.../stack-slot-color-sgpr-vgpr-spills.mir | 2 +-
.../AMDGPU/undefined-physreg-sgpr-spill.mir | 6 +-
.../AMDGPU/wwm-regalloc-partial-pool.mir | 6 +-
.../wwm-regalloc-preallocation-guard.mir | 18 +-
.../AMDGPU/wwm-spill-superclass-pseudo.mir | 2 +-
.../CodeGen/Hexagon/regalloc-bad-undef.mir | 2 +-
40 files changed, 470 insertions(+), 106 deletions(-)
create mode 100644 llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
diff --git a/llvm/include/llvm/CodeGen/TargetInstrInfo.h b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
index b013511d33313..022c39d0ca92d 100644
--- a/llvm/include/llvm/CodeGen/TargetInstrInfo.h
+++ b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
@@ -1226,6 +1226,12 @@ class LLVM_ABI TargetInstrInfo : public MCInstrInfo {
"TargetInstrInfo::storeRegToStackSlot!");
}
+ /// Tell the target which lanes of spill \p SpillMI hold a real value
+ /// (\p DefinedLanes); it can skip the undef ones when lowering, so a shared
+ /// slot's neighbor isn't stomped. Default: no-op.
+ virtual void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const {}
+
/// Load the specified register of the given register class from the specified
/// stack frame index. The load instruction is to be added to the given
/// machine basic block before the specified machine instruction. If \p
diff --git a/llvm/lib/CodeGen/InlineSpiller.cpp b/llvm/lib/CodeGen/InlineSpiller.cpp
index f3682a2e24808..a7a960a0cf441 100644
--- a/llvm/lib/CodeGen/InlineSpiller.cpp
+++ b/llvm/lib/CodeGen/InlineSpiller.cpp
@@ -454,6 +454,17 @@ bool InlineSpiller::hoistSpillInsideBB(LiveInterval &SpillLI,
if (DefMBB != CopyMI.getParent() || !SrcQ.isKill())
return false;
+ // Every sub-range of the hoisted value must be live at the hoist point. The
+ // store below is full-width, so a dead sub-range would write a stale sibling
+ // value into the shared spill slot and clobber it. This mirrors the sub-range
+ // liveness guard on the cross-BB path in isSpillCandBB (#177703), which the
+ // in-BB path here was missing.
+ if (SrcLI.hasSubRanges() &&
+ !all_of(SrcLI.subranges(), [&](const LiveInterval::SubRange &SR) {
+ return SR.getVNInfoAt(Idx) != nullptr;
+ }))
+ return false;
+
MachineBasicBlock *MBB = DefMBB;
MachineBasicBlock::iterator MII;
if (SrcVNI->isPHIDef())
@@ -1285,10 +1296,29 @@ void InlineSpiller::insertSpill(Register NewVReg, bool isKill,
MachineBasicBlock::iterator SpillBefore = std::next(MI);
bool IsRealSpill = isRealSpill(*MI);
- if (IsRealSpill)
+ if (IsRealSpill) {
TII.storeRegToStackSlot(MBB, SpillBefore, NewVReg, isKill, StackSlot,
MRI.getRegClass(NewVReg), Register());
- else
+
+ MachineInstr &SpillMI = *std::next(MI);
+
+ // Compute the lanes of NewVReg defined by MI. For a partial def (undef
+ // lanes remain in the full-width store above), pass the defined-lane mask
+ // so the target skips them and avoids clobbering a sibling in the shared
+ // stack slot. A fully-undef def never reaches here (filtered by
+ // isRealSpill).
+ LaneBitmask FullMask = MRI.getMaxLaneMaskForVReg(NewVReg);
+ LaneBitmask DefinedLanes;
+ for (const MachineOperand &MO : MI->all_defs()) {
+ if (MO.getReg() != NewVReg)
+ continue;
+ DefinedLanes |= MO.getSubReg()
+ ? TRI.getSubRegIndexLaneMask(MO.getSubReg())
+ : FullMask;
+ }
+ if (DefinedLanes != FullMask)
+ TII.setSpillDefinedLaneMask(SpillMI, DefinedLanes);
+ } else
// Don't spill undef value.
// Anything works for undef, in particular keeping the memory
// uninitialized is a viable option and it saves code size and
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
index 0346ebc254361..28a06d6921aa4 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
@@ -1800,10 +1800,11 @@ void SIInstrInfo::storeRegToStackSlotImpl(
}
BuildMI(MBB, MI, DL, OpDesc)
- .addReg(SrcReg, getKillRegState(isKill)) // data
- .addFrameIndex(FrameIndex) // addr
- .addMemOperand(MMO)
- .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
+ .addReg(SrcReg, getKillRegState(isKill)) // data
+ .addFrameIndex(FrameIndex) // addr
+ .addImm(-1) // lanemask (all lanes; refined by spiller)
+ .addMemOperand(MMO)
+ .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
return;
}
@@ -1837,6 +1838,34 @@ void SIInstrInfo::storeRegToStackSlotCFI(MachineBasicBlock &MBB,
MachineInstr::NoFlags, true);
}
+void SIInstrInfo::setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const {
+ // Only SGPR spill saves carry a $lanemask operand (see SI_SPILL_SGPR).
+ int Idx =
+ AMDGPU::getNamedOperandIdx(SpillMI.getOpcode(), AMDGPU::OpName::lanemask);
+ if (Idx == -1)
+ return;
+
+ // Build the per-dword mask consumed by spillSGPR: bit i is set when dword i
+ // of the spilled super-register intersects a defined lane. spillSGPR lowers
+ // the spill dword by dword and consults this mask to skip the undef ones.
+ const MachineRegisterInfo &MRI = SpillMI.getMF()->getRegInfo();
+ Register Data = SpillMI.getOperand(0).getReg();
+ const TargetRegisterClass *RC = MRI.getRegClass(Data);
+ unsigned NumDwords = RI.getRegSizeInBits(*RC) / 32;
+ // DwordMask is a u32; the widest SGPR tuple (SReg_1024) has 32 dwords, so
+ // this holds and keeps the 1u << i shift below well-defined.
+ assert(NumDwords <= 32 && "SGPR spill wider than DwordMask can represent");
+ unsigned DwordMask = 0;
+ for (unsigned i = 0; i != NumDwords; ++i) {
+ LaneBitmask DwordLanes =
+ RI.getSubRegIndexLaneMask(RI.getSubRegFromChannel(i));
+ if ((DefinedLanes & DwordLanes).any())
+ DwordMask |= 1u << i;
+ }
+ SpillMI.getOperand(Idx).setImm(DwordMask);
+}
+
static unsigned getSGPRSpillRestoreOpcode(unsigned Size) {
switch (Size) {
case 4:
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.h b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
index 48eba6b2a567d..42afa3cd8c60e 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.h
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
@@ -356,6 +356,9 @@ class SIInstrInfo final : public AMDGPUGenInstrInfo {
bool isKill, int FrameIndex, const TargetRegisterClass *RC, Register VReg,
MachineInstr::MIFlag Flags = MachineInstr::NoFlags) const override;
+ void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const override;
+
void loadRegFromStackSlot(
MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, Register DestReg,
int FrameIndex, const TargetRegisterClass *RC, Register VReg,
diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 3cf20ffb1fcf4..3faae31302689 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -1178,14 +1178,20 @@ def V_INDIRECT_REG_READ_GPR_IDX_B32_V32 : V_INDIRECT_REG_READ_GPR_IDX_pseudo<VRe
multiclass SI_SPILL_SGPR <RegisterClass sgpr_class> {
let UseNamedOperandTable = 1, Spill = 1, SALU = 1, Uses = [EXEC] in {
+ // $lanemask: per-dword bitmask of which dwords of the spilled super-register
+ // are defined. Undef dwords must not be written into the (possibly shared)
+ // stack slot. -1 (all dwords defined) is the ordinary full-width spill; a
+ // partial mask is set by the spiller for partially-defined values. i32imm's
+ // 32 bits cover the widest SGPR tuple SReg_1024 (32 dwords).
def _SAVE : PseudoInstSI <
(outs),
- (ins sgpr_class:$data, i32imm:$addr)> {
+ (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
let mayStore = 1;
let mayLoad = 0;
}
- def _CFI_SAVE : PseudoInstSI<(outs), (ins sgpr_class:$data, i32imm:$addr)> {
+ def _CFI_SAVE : PseudoInstSI<(outs),
+ (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
let mayStore = 1;
let mayLoad = 0;
}
diff --git a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
index 1aa423e1da647..5a8dd05c239ab 100644
--- a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
@@ -2180,6 +2180,18 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
if (OnlyToVGPR && !SpillToVGPR)
return false;
+ // A partial-def spill (see InlineSpiller / setSpillDefinedLaneMask) records
+ // which dwords are real values in its $lanemask operand; a cleared bit means
+ // that dword is undef and must NOT be written to the (possibly shared) stack
+ // slot, so it does not clobber a sibling value living in the same slot. A
+ // mask of -1 (default) means "all lanes defined" = original behavior.
+ const MachineOperand *LaneMask =
+ SB.TII.getNamedOperand(*MI, AMDGPU::OpName::lanemask);
+ int64_t DefinedDwordMask = LaneMask ? LaneMask->getImm() : -1;
+ auto DwordDefined = [&](unsigned i) {
+ return (DefinedDwordMask & (int64_t(1) << i)) != 0;
+ };
+
const SIFrameLowering *TFL = ST.getFrameLowering();
assert(SpillToVGPR || (SB.SuperReg != SB.MFI.getStackPtrOffsetReg() &&
@@ -2195,15 +2207,29 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
"Num of SGPRs spilled should be less than or equal to num of "
"the VGPR lanes.");
+ // With a partial-def mask, find the first/last dword that is actually
+ // defined, so the ImplicitDefine of the super-register and the kill flag
+ // can be re-anchored onto emitted writelanes (undef dwords are skipped).
+ unsigned FirstDefined = 0, LastDefined = SB.NumSubRegs - 1;
+ while (FirstDefined < SB.NumSubRegs && !DwordDefined(FirstDefined))
+ ++FirstDefined;
+ while (LastDefined > 0 && !DwordDefined(LastDefined))
+ --LastDefined;
+
for (unsigned i = 0, e = SB.NumSubRegs; i < e; ++i) {
+ // Skip undef dwords: not writing them to a (possibly shared) stack slot
+ // preserves a sibling value living in the same slot.
+ if (!DwordDefined(i))
+ continue;
+
Register SubReg =
SB.NumSubRegs == 1
? SB.SuperReg
: Register(getSubReg(SB.SuperReg, SB.SplitParts[i]));
SpilledReg Spill = VGPRSpills[i];
- bool IsFirstSubreg = i == 0;
- bool IsLastSubreg = i == SB.NumSubRegs - 1;
+ bool IsFirstSubreg = i == FirstDefined;
+ bool IsLastSubreg = i == LastDefined;
bool UseKill = SB.IsKill && IsLastSubreg;
@@ -2259,13 +2285,30 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
// Per VGPR helper data
auto PVD = SB.getPerVGPRData();
+ // For a partial-def spill (undef dwords, see $lanemask), the whole TmpVGPR
+ // is stored at once, so first load the slot's current contents; the skipped
+ // dwords then keep the sibling value already there instead of writing
+ // undef.
+ bool IsPartialDef = SB.NumSubRegs > 1 && DefinedDwordMask != -1;
+
for (unsigned Offset = 0; Offset < PVD.NumVGPRs; ++Offset) {
RegState TmpVGPRFlags = RegState::Undef;
+ if (IsPartialDef) {
+ // Seed TmpVGPR with the slot's current contents, so the dwords we skip
+ // (undef in this partial def) keep the sibling value already living in
+ // the slot when the whole TmpVGPR is stored back below.
+ SB.readWriteTmpVGPR(Offset, /*IsLoad*/ true);
+ TmpVGPRFlags = {};
+ }
+
// Write sub registers into the VGPR
for (unsigned i = Offset * PVD.PerVGPR,
e = std::min((Offset + 1) * PVD.PerVGPR, SB.NumSubRegs);
i < e; ++i) {
+ if (IsPartialDef && !DwordDefined(i))
+ continue; // Undef dword: keep the sibling value seeded from the slot.
+
Register SubReg =
SB.NumSubRegs == 1
? SB.SuperReg
@@ -2286,8 +2329,8 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
Indexes->insertMachineInstrInMaps(*WriteLane);
}
- // There could be undef components of a spilled super register.
- // TODO: Can we detect this and skip the spill?
+ // Undef components of the spilled super-register are detected via the
+ // $lanemask operand and skipped above (see IsPartialDef).
if (SB.NumSubRegs > 1) {
// The last implicit use of the SB.SuperReg carries the "Kill" flag.
RegState SuperKillState = {};
diff --git a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
index 5c564decd1e7d..a026d13e0cce5 100644
--- a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
+++ b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
@@ -75,7 +75,7 @@ body: |
liveins: $sgpr12, $sgpr13, $sgpr14, $sgpr15
%45:vgpr_32 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE $sgpr15, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr15, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
%16:vgpr_32 = V_AND_B32_e32 1, %45, implicit $exec
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
index 94c22b1aa8664..5c186b9a0470b 100644
--- a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
+++ b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
@@ -57,7 +57,7 @@ body: |
; CHECK-NEXT: renamable $sgpr86 = S_MOV_B32 0
; CHECK-NEXT: renamable $sgpr87 = S_MOV_B32 0
; CHECK-NEXT: renamable $sgpr88 = S_MOV_B32 0
- ; CHECK-NEXT: SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
+ ; CHECK-NEXT: SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, 65535, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
; CHECK-NEXT: renamable $sgpr52_sgpr53 = IMPLICIT_DEF
; CHECK-NEXT: ADJCALLSTACKUP 0, 0, implicit-def dead $scc, implicit-def $sgpr32, implicit $sgpr32
; CHECK-NEXT: dead $sgpr30_sgpr31 = SI_CALL renamable $sgpr52_sgpr53, 0, csr_amdgpu, implicit $sgpr0_sgpr1_sgpr2_sgpr3
diff --git a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
index 63ee4473a56e5..43669338dd308 100644
--- a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
+++ b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
@@ -55,7 +55,7 @@ body: |
; CHECK-NEXT: liveins: $sgpr24_sgpr25_sgpr26_sgpr27:0x000000000000000F, $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19:0x000000000000FFFF, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51:0x000000000000FFFF
; CHECK-NEXT: {{ $}}
; CHECK-NEXT: renamable $sgpr12 = IMPLICIT_DEF
- ; CHECK-NEXT: SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
+ ; CHECK-NEXT: SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, 255, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
; CHECK-NEXT: renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = IMPLICIT_DEF
; CHECK-NEXT: dead undef [[IMAGE_SAMPLE_LZ_V1_V2_2:%[0-9]+]].sub0:vreg_96 = IMAGE_SAMPLE_LZ_V1_V2 undef [[DEF2]], killed renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43, renamable $sgpr12_sgpr13_sgpr14_sgpr15, 1, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s32), addrspace 8)
; CHECK-NEXT: renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = SI_SPILL_S512_RESTORE %stack.0, implicit $exec, implicit $sgpr32 :: (load (s512) from %stack.0, align 4, addrspace 5)
diff --git a/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
new file mode 100644
index 0000000000000..bbff66888c652
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
@@ -0,0 +1,257 @@
+; RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -stop-after=greedy %s -o - | FileCheck %s
+;
+; Regression test for the InlineSpiller shared-stack-slot partial-subregister
+; hazard.
+;
+; The hazard is target-independent: InlineSpiller's shared-slot logic is generic,
+; and SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs. When the
+; greedy allocator live-range-splits an sgpr_128 descriptor on one sub-register,
+; it produces sibling vregs each defined on a single sub-lane
+; (`undef %V.sub0:sgpr_128 = ...`), undef elsewhere. Since InlineSpiller shares
+; one stack slot among descendants of the Original value, a full-width store of
+; such a sibling writes its undef lanes over a live sibling in the SAME slot. This
+; input is a gfx950 reproducer (there the clobber corrupts an in-loop buffer
+; descriptor -- the kind of corruption that leads to an HSA memory fault), but the
+; fix and this check apply to all AMDGPU.
+;
+; The fix records in the save pseudo's $lanemask operand which dwords the store
+; defines; spillSGPR writelanes only those, leaving the slot's other lanes intact.
+; This pins that the two partial siblings sharing one slot each store with a
+; partial mask (matched as a non-negative immediate: the buggy full-width store
+; used -1, which the {{[0-9]+}} pattern cannot match).
+;
+; Driven from IR through -stop-after=greedy: a .mir + -run-pass=greedy form does
+; not reproduce, since the trigger needs pre-greedy PreRARemat pressure state that
+; MIR serialization drops. The CHECK block is vreg/slot-number independent.
+;
+; First partial sibling (only .sub0 defined), stored with a partial mask:
+; CHECK: undef [[DESC:%[0-9]+]].sub0:sgpr_128 = S_MOV_B32 0
+; CHECK-NEXT: SI_SPILL_S128_SAVE [[DESC]], [[SLOT:%stack\.[0-9]+]], {{[0-9]+}},
+; Second partial sibling into the SAME slot, also with a partial mask:
+; CHECK: SI_SPILL_S128_SAVE %{{[0-9]+}}, [[SLOT]], {{[0-9]+}},
+;
+target datalayout = "e-m:e-p:64:64-p1:64:64-p2:32:32-p3:32:32-p4:64:64-p5:32:32-p6:32:32-p7:160:256:256:32-p8:128:128:128:48-p9:192:256:256:32-i64:64-v16:16-v24:32-v32:32-v48:64-v96:128-v192:256-v256:256-v512:512-v1024:1024-v2048:2048-n32:64-S32-A5-G1-ni:7:8:9-p10:32:32-p11:32:32-p12:32:32-p13:32:32-p14:32:32-p15:32:32"
+target triple = "amdgcn-amd-amdhsa"
+
+define amdgpu_kernel void @_attn_fwd_IS_CAUSAL_1_NUM_Q_HEADS_8_NUM_K_HEADS_1_BLOCK_M_256_BLOCK_N_64_BLOCK_DMODEL_64_RETURN_SCORES_1_ENABLE_DROPOUT_1_IS_FP8_0_VARLEN_1_NUM_XCD_8_USE_INT64_STRIDES_1_ENABLE_SINK_0_SLIDING_WINDOW_0(ptr addrspace(1) inreg %0, ptr addrspace(1) inreg %1, i32 %2, i32 %3, i32 %4, i32 %5, i32 %6, <1 x i32> %7, i32 %8, i32 %9, i32 %10, i32 %11, i32 %12, i32 %13, i32 %14, i32 %15, i32 %16, i32 %17, i32 %18, i32 %19, i32 %20, i32 %21, i32 %22, i32 %23, i32 %24, i32 %25, i32 %26, i64 %27, i64 %28, i64 %29, i64 %30, i64 %31, i64 %32, i64 %33, i64 %34, i64 %35, i64 %36, i64 %37, i1 %38, i1 %39, i1 %40, i1 %41, i1 %42, i1 %43, i1 %44, i64 %sext452, i64 %.pn241.in795, i64 %sext496, i64 %.pn233.in799, i64 %sext500, i64 %.pn231.in800, i64 %sext501, i64 %.pn229.in801, i64 %sext503, i64 %.pn223.in804, i64 %sext505, i64 %.pn221.in805, i64 %sext506) #0 {
+..loopexit_crit_edge:
+ %45 = icmp slt i32 %8, 0
+ %46 = icmp slt i32 %6, 0
+ %47 = icmp slt i32 %9, 0
+ %48 = icmp slt i32 %10, 0
+ %49 = icmp slt i32 %12, 0
+ %50 = icmp slt i32 %2, 0
+ br label %51
+
+51: ; preds = %51, %..loopexit_crit_edge
+ %.pn265.in835 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext503, %51 ]
+ %.pn179.in826 = phi i64 [ 0, %..loopexit_crit_edge ], [ %27, %51 ]
+ %.pn181.in8251 = phi i64 [ 0, %..loopexit_crit_edge ], [ %137, %51 ]
+ %.pn183.in824 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext506, %51 ]
+ %.pn185.in8232 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext501, %51 ]
+ %.pn195.in818 = phi i64 [ 0, %..loopexit_crit_edge ], [ %135, %51 ]
+ %.pn197.in8177 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext452, %51 ]
+ %.pn199.in8168 = phi i64 [ 0, %..loopexit_crit_edge ], [ %.pn229.in801, %51 ]
+ %.pn201.in8159 = phi i64 [ 0, %..loopexit_crit_edge ], [ %30, %51 ]
+ %.pn203.in81410 = phi i64 [ 0, %..loopexit_crit_edge ], [ %31, %51 ]
+ %.pn205.in813 = phi i64 [ 0, %..loopexit_crit_edge ], [ 2, %51 ]
+ %.pn209.in811 = phi i64 [ 0, %..loopexit_crit_edge ], [ 1, %51 ]
+ %.pn215.in80811 = phi i64 [ 0, %..loopexit_crit_edge ], [ %.pn231.in800, %51 ]
+ %.pn217.in80712 = phi i64 [ 0, %..loopexit_crit_edge ], [ %32, %51 ]
+ %.pn219.in80613 = phi i64 [ 0, %..loopexit_crit_edge ], [ %33, %51 ]
+ %.pn223.in80415 = phi i64 [ 0, %..loopexit_crit_edge ], [ %136, %51 ]
+ %.pn227.in80216 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext505, %51 ]
+ %.pn229.in80117 = phi i64 [ 0, %..loopexit_crit_edge ], [ %sext496, %51 ]
+ %.pn231.in80018 = phi i64 [ 0, %..loopexit_crit_edge ], [ %35, %51 ]
+ %.pn233.in79919 = phi i64 [ 0, %..loopexit_crit_edge ], [ %36, %51 ]
+ %.pn241.in79520 = phi i64 [ 0, %..loopexit_crit_edge ], [ %37, %51 ]
+ %52 = phi i32 [ 0, %..loopexit_crit_edge ], [ 2, %51 ]
+ %53 = phi <2 x float> [ zeroinitializer, %..loopexit_crit_edge ], [ %138, %51 ]
+ %54 = or i32 %52, %20
+ %55 = tail call ptr addrspace(8) @llvm.amdgcn.make.buffer.rsrc.p8.p1.i64(ptr addrspace(1) %1, i16 0, i64 1, i32 1)
+ %.pn265 = trunc i64 %.pn265.in835 to i32
+ %56 = shl i32 %.pn265, 1
+ %57 = tail call <2 x i32> @llvm.amdgcn.raw.ptr.buffer.load.v2i32(ptr addrspace(8) %55, i32 %56, i32 0, i32 0)
+ %58 = icmp slt i32 %52, %5
+ %59 = select i1 %58, i32 0, i32 1
+ %60 = tail call <4 x i32> @llvm.amdgcn.raw.ptr.buffer.load.v4i32(ptr addrspace(8) null, i32 %59, i32 0, i32 0)
+ %61 = icmp slt i32 %54, 0
+ %62 = and i1 %46, %61
+ %63 = and i1 %47, %61
+ %64 = and i1 %48, %61
+ %65 = icmp slt i32 %3, 0
+ %66 = and i1 %65, %61
+ %67 = icmp slt i32 %13, 0
+ %68 = and i1 %67, %61
+ %69 = icmp slt i32 %14, 0
+ %70 = and i1 %69, %61
+ %71 = and i1 %38, %61
+ %72 = and i1 %50, %61
+ %73 = and i1 %39, %61
+ %74 = icmp slt i32 %23, 0
+ %75 = and i1 %74, %61
+ %76 = and i1 %49, %61
+ %77 = icmp slt i32 %25, 0
+ %78 = and i1 %77, %61
+ %79 = icmp slt i32 %18, 0
+ %80 = and i1 %79, %61
+ %81 = and i1 %40, %61
+ %82 = icmp slt i32 %24, 0
+ %83 = and i1 %82, %61
+ store <2 x i32> %57, ptr addrspace(3) null, align 8
+ tail call void @llvm.amdgcn.s.barrier()
+ %84 = icmp slt i32 %52, 1
+ %85 = and i1 %45, %84
+ %.pn241 = trunc i64 %.pn241.in79520 to i32
+ %86 = select i1 %62, i32 %.pn241, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %86, i32 0, i32 0)
+ %87 = select i1 %63, i32 %3, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %87, i32 0, i32 0)
+ %88 = select i1 %64, i32 1, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %88, i32 0, i32 0)
+ %89 = and i1 %44, %61
+ %.pn239 = trunc i64 %.pn209.in811 to i32
+ %90 = select i1 %89, i32 %.pn239, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %90, i32 0, i32 0)
+ %91 = and i1 %39, %38
+ %.pn233 = trunc i64 %.pn233.in79919 to i32
+ %92 = select i1 %91, i32 %.pn233, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %92, i32 0, i32 0)
+ %.pn231 = trunc i64 %.pn231.in80018 to i32
+ %93 = select i1 %39, i32 %.pn231, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %93, i32 0, i32 0)
+ %.pn229 = trunc i64 %.pn229.in80117 to i32
+ %94 = select i1 %38, i32 %.pn229, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %94, i32 0, i32 0)
+ %.pn227 = trunc i64 %.pn227.in80216 to i32
+ %95 = select i1 %39, i32 %.pn227, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %95, i32 0, i32 0)
+ %96 = select i1 %66, i32 0, i32 1
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %96, i32 0, i32 0)
+ %.pn223 = trunc i64 %.pn223.in80415 to i32
+ %97 = shl i32 %.pn223, 1
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %97, i32 0, i32 0)
+ %98 = select i1 %68, i32 %14, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %98, i32 0, i32 0)
+ %.pn219 = trunc i64 %.pn219.in80613 to i32
+ %99 = select i1 %70, i32 %.pn219, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %99, i32 0, i32 0)
+ %100 = and i1 %43, %38
+ %.pn217 = trunc i64 %.pn217.in80712 to i32
+ %101 = select i1 %100, i32 %.pn217, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %101, i32 0, i32 0)
+ %.pn215 = trunc i64 %.pn215.in80811 to i32
+ %102 = select i1 %39, i32 %.pn215, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %102, i32 0, i32 0)
+ %.pn213 = trunc i64 %.pn205.in813 to i32
+ %103 = select i1 %38, i32 %.pn213, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %103, i32 0, i32 0)
+ %104 = select i1 %71, i32 1, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %104, i32 0, i32 0)
+ %105 = select i1 %72, i32 1, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %105, i32 0, i32 0)
+ %.pn207 = trunc i64 %28 to i32
+ %106 = select i1 %73, i32 %.pn207, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %106, i32 0, i32 0)
+ %107 = select i1 %75, i32 %16, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %107, i32 0, i32 0)
+ %.pn203 = trunc i64 %.pn203.in81410 to i32
+ %108 = select i1 %76, i32 %.pn203, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %108, i32 0, i32 0)
+ %109 = and i1 %41, %38
+ %.pn201 = trunc i64 %.pn201.in8159 to i32
+ %110 = select i1 %109, i32 %.pn201, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %110, i32 0, i32 0)
+ %.pn199 = trunc i64 %.pn199.in8168 to i32
+ %111 = shl i32 %.pn199, 0
+ %112 = select i1 %38, i32 %111, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %112, i32 0, i32 0)
+ %.pn197 = trunc i64 %.pn197.in8177 to i32
+ %113 = select i1 %78, i32 %.pn197, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %113, i32 0, i32 0)
+ %114 = and i1 %38, %40
+ %.pn195 = trunc i64 %.pn195.in818 to i32
+ %115 = select i1 %114, i32 %.pn195, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %115, i32 0, i32 0)
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %54, i32 0, i32 0)
+ %116 = select i1 %80, i32 %21, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %116, i32 0, i32 0)
+ %.pn189 = trunc i64 %.pn179.in826 to i32
+ %117 = select i1 %42, i32 %.pn189, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %117, i32 0, i32 0)
+ %118 = select i1 %81, i32 %.pn265, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %118, i32 0, i32 0)
+ %.pn185 = trunc i64 %.pn185.in8232 to i32
+ %119 = select i1 %50, i32 %.pn185, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %119, i32 0, i32 0)
+ %.pn183 = trunc i64 %.pn183.in824 to i32
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %.pn183, i32 0, i32 0)
+ %.pn181 = trunc i64 %.pn181.in8251 to i32
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %.pn181, i32 0, i32 0)
+ %120 = select i1 %83, i32 %3, i32 0
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %120, i32 0, i32 0)
+ %121 = select i1 %85, i32 0, i32 -2147483648
+ tail call void @llvm.amdgcn.raw.ptr.buffer.store.f32(float 0.000000e+00, ptr addrspace(8) null, i32 %121, i32 0, i32 0)
+ %122 = bitcast <4 x i32> %60 to <8 x bfloat>
+ %123 = shufflevector <8 x bfloat> %122, <8 x bfloat> zeroinitializer, <4 x i32> <i32 0, i32 1, i32 2, i32 3>
+ store <4 x bfloat> %123, ptr addrspace(3) null, align 8
+ %124 = shufflevector <2 x float> zeroinitializer, <2 x float> %53, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison>
+ %125 = shufflevector <16 x float> %124, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 16, i32 17, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison>
+ %126 = shufflevector <16 x float> %125, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 16, i32 17, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison>
+ %127 = shufflevector <16 x float> %126, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7, i32 16, i32 17, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison, i32 poison>
+ %128 = shufflevector <16 x float> %127, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 16, i32 17, i32 poison, i32 poison, i32 poison, i32 poison>
+ %129 = shufflevector <16 x float> %128, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 16, i32 17, i32 poison, i32 poison>
+ %130 = shufflevector <16 x float> %129, <16 x float> zeroinitializer, <16 x i32> <i32 0, i32 1, i32 2, i32 3, i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 16, i32 17>
+ %131 = tail call <16 x float> @llvm.amdgcn.mfma.f32.32x32x16.bf16(<8 x bfloat> splat (bfloat 1.000000e+00), <8 x bfloat> splat (bfloat 1.000000e+00), <16 x float> %130, i32 0, i32 0, i32 0)
+ %132 = tail call <16 x float> @llvm.amdgcn.mfma.f32.32x32x16.bf16(<8 x bfloat> splat (bfloat 1.000000e+00), <8 x bfloat> splat (bfloat 1.000000e+00), <16 x float> %131, i32 0, i32 0, i32 0)
+ %133 = tail call <16 x float> @llvm.amdgcn.mfma.f32.32x32x16.bf16(<8 x bfloat> splat (bfloat +qnan), <8 x bfloat> zeroinitializer, <16 x float> %132, i32 0, i32 0, i32 0)
+ %134 = tail call <16 x float> @llvm.amdgcn.mfma.f32.32x32x16.bf16(<8 x bfloat> splat (bfloat +qnan), <8 x bfloat> splat (bfloat 1.000000e+00), <16 x float> %133, i32 0, i32 0, i32 0)
+ %135 = ashr i64 %27, 1
+ %136 = or i64 %.pn233.in799, 1
+ %137 = or i64 %sext500, 1
+ %138 = shufflevector <16 x float> %134, <16 x float> zeroinitializer, <2 x i32> <i32 2, i32 3>
+ br label %51
+}
+
+; Function Attrs: nocallback nofree nosync nounwind speculatable willreturn memory(none)
+declare noundef range(i32 0, 1024) i32 @llvm.amdgcn.workitem.id.x() #1
+
+; Function Attrs: nocallback nofree nosync nounwind speculatable willreturn memory(none)
+declare noundef i32 @llvm.amdgcn.workgroup.id.x() #1
+
+; Function Attrs: nocallback nofree nosync nounwind willreturn memory(argmem: read)
+declare <4 x i32> @llvm.amdgcn.raw.ptr.buffer.load.v4i32(ptr addrspace(8) readonly captures(none), i32, i32, i32 immarg) #2
+
+; Function Attrs: convergent nocallback nofree nounwind willreturn
+declare void @llvm.amdgcn.s.barrier() #3
+
+; Function Attrs: nocallback nofree nosync nounwind willreturn memory(argmem: write)
+declare void @llvm.amdgcn.raw.ptr.buffer.store.f32(float, ptr addrspace(8) writeonly captures(none), i32, i32, i32 immarg) #4
+
+; Function Attrs: nocallback nofree nosync nounwind willreturn memory(argmem: read)
+declare <2 x i32> @llvm.amdgcn.raw.ptr.buffer.load.v2i32(ptr addrspace(8) readonly captures(none), i32, i32, i32 immarg) #2
+
+; Function Attrs: convergent nocallback nocreateundeforpoison nofree nosync nounwind willreturn memory(none)
+declare <16 x float> @llvm.amdgcn.mfma.f32.32x32x16.bf16(<8 x bfloat>, <8 x bfloat>, <16 x float>, i32 immarg, i32 immarg, i32 immarg) #5
+
+; Function Attrs: nocallback nocreateundeforpoison nofree nosync nounwind speculatable willreturn memory(none)
+declare float @llvm.amdgcn.exp2.f32(float) #6
+
+; Function Attrs: convergent nocallback nofree nounwind willreturn memory(argmem: read)
+declare <4 x bfloat> @llvm.amdgcn.ds.read.tr16.b64.v4bf16(ptr addrspace(3) captures(none)) #7
+
+; Function Attrs: nocallback nocreateundeforpoison nofree nosync nounwind speculatable willreturn memory(none)
+declare ptr addrspace(8) @llvm.amdgcn.make.buffer.rsrc.p8.p1.i64(ptr addrspace(1) readnone, i16, i64, i32) #6
+
+; uselistorder directives
+uselistorder ptr @llvm.amdgcn.raw.ptr.buffer.store.f32, { 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0 }
+uselistorder ptr @llvm.amdgcn.mfma.f32.32x32x16.bf16, { 3, 2, 1, 0 }
+
+attributes #0 = { "amdgpu-agpr-alloc"="0" "amdgpu-no-dispatch-id" "amdgpu-no-dispatch-ptr" "amdgpu-no-queue-ptr" "amdgpu-waves-per-eu"="8,8" }
+attributes #1 = { nocallback nofree nosync nounwind speculatable willreturn memory(none) }
+attributes #2 = { nocallback nofree nosync nounwind willreturn memory(argmem: read) }
+attributes #3 = { convergent nocallback nofree nounwind willreturn }
+attributes #4 = { nocallback nofree nosync nounwind willreturn memory(argmem: write) }
+attributes #5 = { convergent nocallback nocreateundeforpoison nofree nosync nounwind willreturn memory(none) }
+attributes #6 = { nocallback nocreateundeforpoison nofree nosync nounwind speculatable willreturn memory(none) }
+attributes #7 = { convergent nocallback nofree nounwind willreturn memory(argmem: read) }
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir
index 22dfc0ebd4f31..8fe51ec578fe0 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir
@@ -49,7 +49,7 @@ body: |
; SGPR_SPILL-NEXT: renamable $sgpr10 = IMPLICIT_DEF
; SGPR_SPILL-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = IMPLICIT_DEF
; SGPR_SPILL-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = SI_SPILL_S32_TO_VGPR killed $sgpr10, 0, [[DEF]]
- ; SGPR_SPILL-NEXT: DBG_VALUE $noreg, 0
+ ; SGPR_SPILL-NEXT: DBG_VALUE $noreg, 0, !0, !DIExpression(), debug-location !DILocation(line: 10, column: 9, scope: !1)
; SGPR_SPILL-NEXT: {{ $}}
; SGPR_SPILL-NEXT: bb.1:
; SGPR_SPILL-NEXT: $sgpr10 = SI_RESTORE_S32_FROM_VGPR [[DEF]], 0
@@ -70,7 +70,7 @@ body: |
; PEI-NEXT: S_ENDPGM 0
bb.0:
renamable $sgpr10 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
DBG_VALUE %stack.0, 0, !1, !8, debug-location !9
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir
index 04ae8f11f3143..55f8b577e8350 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir
@@ -48,7 +48,7 @@ body: |
bb.0:
renamable $sgpr10 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
DBG_VALUE_LIST !1, !8, %stack.0, 0, debug-location !9
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir
index 67008b760a135..df434d84be4b9 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir
@@ -49,7 +49,7 @@ body: |
; CHECK: DBG_VALUE
bb.0:
renamable $sgpr10 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
DBG_VALUE %fixed-stack.0, 0, !1, !8, debug-location !9
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir
index 55ec919041e00..3ee3de66c23a1 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir
@@ -279,7 +279,7 @@ body: |
liveins: $vgpr0
$sgpr22 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE $sgpr22, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr22, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
$sgpr0 = ENTER_STRICT_WWM -1, implicit-def $exec, implicit-def $scc, implicit $exec
%0:vgpr_32 = V_SET_INACTIVE_B32 0, $vgpr0, 0, 0, $sgpr_null, implicit $exec, implicit-def $scc
$exec_lo = EXIT_STRICT_WWM $sgpr0
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir
index ce0105863826a..33454fd51eaa7 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir
@@ -25,7 +25,7 @@ body: |
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = IMPLICIT_DEF
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = SI_SPILL_S32_TO_VGPR $sgpr4, 0, [[DEF]], implicit-def $sgpr4_sgpr5, implicit $sgpr4_sgpr5
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = SI_SPILL_S32_TO_VGPR $sgpr5, 1, [[DEF]], implicit $sgpr4_sgpr5
- SI_SPILL_S64_SAVE renamable $sgpr4_sgpr5, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE renamable $sgpr4_sgpr5, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
...
@@ -51,7 +51,7 @@ body: |
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = IMPLICIT_DEF
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = SI_SPILL_S32_TO_VGPR $sgpr4, 0, [[DEF]], implicit-def $sgpr4_sgpr5, implicit $sgpr4_sgpr5
; CHECK-NEXT: [[DEF:%[0-9]+]]:vgpr_32 = SI_SPILL_S32_TO_VGPR $sgpr5, 1, [[DEF]], implicit $sgpr4_sgpr5
- SI_SPILL_S64_SAVE renamable $sgpr4_sgpr5, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE renamable $sgpr4_sgpr5, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
...
@@ -67,12 +67,7 @@ stack:
- { id: 0, type: spill-slot, size: 4, alignment: 4, stack-id: sgpr-spill }
body: |
bb.0:
- ; CHECK-LABEL: name: sgpr_spill_s32_undef
- ; CHECK: body:
- ; CHECK-NEXT: bb.0:
- ; CHECK-NOT: {{.+}}
- ; CHECK: ...
- SI_SPILL_S32_SAVE undef $sgpr8, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s32) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE undef $sgpr8, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s32) into %stack.0, align 4, addrspace 5)
...
@@ -88,11 +83,6 @@ stack:
- { id: 0, type: spill-slot, size: 8, alignment: 4, stack-id: sgpr-spill }
body: |
bb.0:
- ; CHECK-LABEL: name: sgpr_spill_s64_undef
- ; CHECK: body:
- ; CHECK-NEXT: bb.0:
- ; CHECK-NOT: {{.+}}
- ; CHECK: ...
- SI_SPILL_S64_SAVE undef $sgpr8_sgpr9, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE undef $sgpr8_sgpr9, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32 :: (store (s64) into %stack.0, align 4, addrspace 5)
...
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir
index c8cc6fe170598..a53d36914e9e4 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir
@@ -28,7 +28,7 @@ body: |
liveins: $sgpr0_sgpr1_sgpr2_sgpr3_sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15, $sgpr16_sgpr17_sgpr18_sgpr19_sgpr20_sgpr21_sgpr22_sgpr23_sgpr24_sgpr25_sgpr26_sgpr27_sgpr28_sgpr29_sgpr30_sgpr31, $sgpr32_sgpr33_sgpr34_sgpr35_sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47, $sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63, $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79, $sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95, $sgpr96_sgpr97_sgpr98_sgpr99, $sgpr100_sgpr101_sgpr102_sgpr103, $sgpr104_128
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE $sgpr8, %stack.1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr8, %stack.1, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.1, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir
index f021b958bf2d5..dda4fefc55f80 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir
@@ -47,7 +47,7 @@ body: |
liveins: $sgpr8
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE $sgpr8, %stack.1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr8, %stack.1, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.1, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
@@ -99,7 +99,7 @@ body: |
bb.0:
liveins: $sgpr8_sgpr9
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S64_SAVE $sgpr8_sgpr9, %stack.1, implicit $exec, implicit $sgpr32:: (store (s64) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE $sgpr8_sgpr9, %stack.1, -1, implicit $exec, implicit $sgpr32:: (store (s64) into %stack.1, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
@@ -319,7 +319,7 @@ body: |
bb.0:
liveins: $vgpr0_vgpr1_vgpr2_vgpr3_vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15, $vgpr16_vgpr17_vgpr18_vgpr19_vgpr20_vgpr21_vgpr22_vgpr23_vgpr24_vgpr25_vgpr26_vgpr27_vgpr28_vgpr29_vgpr30_vgpr31, $vgpr32_vgpr33_vgpr34_vgpr35_vgpr36_vgpr37_vgpr38_vgpr39_vgpr40_vgpr41_vgpr42_vgpr43_vgpr44_vgpr45_vgpr46_vgpr47, $vgpr48_vgpr49_vgpr50_vgpr51_vgpr52_vgpr53_vgpr54_vgpr55_vgpr56_vgpr57_vgpr58_vgpr59_vgpr60_vgpr61_vgpr62_vgpr63, $vgpr64_vgpr65_vgpr66_vgpr67_vgpr68_vgpr69_vgpr70_vgpr71_vgpr72_vgpr73_vgpr74_vgpr75_vgpr76_vgpr77_vgpr78_vgpr79, $vgpr80_vgpr81_vgpr82_vgpr83_vgpr84_vgpr85_vgpr86_vgpr87_vgpr88_vgpr89_vgpr90_vgpr91_vgpr92_vgpr93_vgpr94_vgpr95, $vgpr96_vgpr97_vgpr98_vgpr99_vgpr100_vgpr101_vgpr102_vgpr103_vgpr104_vgpr105_vgpr106_vgpr107_vgpr108_vgpr109_vgpr110_vgpr111, $vgpr112_vgpr113_vgpr114_vgpr115_vgpr116_vgpr117_vgpr118_vgpr119_vgpr120_vgpr121_vgpr122_vgpr123_vgpr124_vgpr125_vgpr126_vgpr127, $vgpr128_vgpr129_vgpr130_vgpr131_vgpr132_vgpr133_vgpr134_vgpr135_vgpr136_vgpr137_vgpr138_vgpr139_vgpr140_vgpr141_vgpr142_vgpr143, $vgpr144_vgpr145_vgpr146_vgpr147_vgpr148_vgpr149_vgpr150_vgpr151_vgpr152_vgpr153_vgpr154_vgpr155_vgpr156_vgpr157_vgpr158_vgpr159, $vgpr160_vgpr161_vgpr162_vgpr163_vgpr164_vgpr165_vgpr166_vgpr167_vgpr168_vgpr169_vgpr170_vgpr171_vgpr172_vgpr173_vgpr174_vgpr175, $vgpr176_vgpr177_vgpr178_vgpr179_vgpr180_vgpr181_vgpr182_vgpr183_vgpr184_vgpr185_vgpr186_vgpr187_vgpr188_vgpr189_vgpr190_vgpr191, $vgpr192_vgpr193_vgpr194_vgpr195_vgpr196_vgpr197_vgpr198_vgpr199_vgpr200_vgpr201_vgpr202_vgpr203_vgpr204_vgpr205_vgpr206_vgpr207, $vgpr208_vgpr209_vgpr210_vgpr211_vgpr212_vgpr213_vgpr214_vgpr215_vgpr216_vgpr217_vgpr218_vgpr219_vgpr220_vgpr221_vgpr222_vgpr223, $vgpr224_vgpr225_vgpr226_vgpr227_vgpr228_vgpr229_vgpr230_vgpr231_vgpr232_vgpr233_vgpr234_vgpr235_vgpr236_vgpr237_vgpr238_vgpr239, $vgpr240_vgpr241_vgpr242_vgpr243_vgpr244_vgpr245_vgpr246_vgpr247, $vgpr248_vgpr249_vgpr250_vgpr251, $vgpr252_vgpr253_vgpr254_vgpr255, $sgpr8
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE $sgpr8, %stack.1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr8, %stack.1, -1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.1, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
@@ -557,7 +557,7 @@ body: |
bb.0:
liveins: $vgpr0_vgpr1_vgpr2_vgpr3_vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15, $vgpr16_vgpr17_vgpr18_vgpr19_vgpr20_vgpr21_vgpr22_vgpr23_vgpr24_vgpr25_vgpr26_vgpr27_vgpr28_vgpr29_vgpr30_vgpr31, $vgpr32_vgpr33_vgpr34_vgpr35_vgpr36_vgpr37_vgpr38_vgpr39_vgpr40_vgpr41_vgpr42_vgpr43_vgpr44_vgpr45_vgpr46_vgpr47, $vgpr48_vgpr49_vgpr50_vgpr51_vgpr52_vgpr53_vgpr54_vgpr55_vgpr56_vgpr57_vgpr58_vgpr59_vgpr60_vgpr61_vgpr62_vgpr63, $vgpr64_vgpr65_vgpr66_vgpr67_vgpr68_vgpr69_vgpr70_vgpr71_vgpr72_vgpr73_vgpr74_vgpr75_vgpr76_vgpr77_vgpr78_vgpr79, $vgpr80_vgpr81_vgpr82_vgpr83_vgpr84_vgpr85_vgpr86_vgpr87_vgpr88_vgpr89_vgpr90_vgpr91_vgpr92_vgpr93_vgpr94_vgpr95, $vgpr96_vgpr97_vgpr98_vgpr99_vgpr100_vgpr101_vgpr102_vgpr103_vgpr104_vgpr105_vgpr106_vgpr107_vgpr108_vgpr109_vgpr110_vgpr111, $vgpr112_vgpr113_vgpr114_vgpr115_vgpr116_vgpr117_vgpr118_vgpr119_vgpr120_vgpr121_vgpr122_vgpr123_vgpr124_vgpr125_vgpr126_vgpr127, $vgpr128_vgpr129_vgpr130_vgpr131_vgpr132_vgpr133_vgpr134_vgpr135_vgpr136_vgpr137_vgpr138_vgpr139_vgpr140_vgpr141_vgpr142_vgpr143, $vgpr144_vgpr145_vgpr146_vgpr147_vgpr148_vgpr149_vgpr150_vgpr151_vgpr152_vgpr153_vgpr154_vgpr155_vgpr156_vgpr157_vgpr158_vgpr159, $vgpr160_vgpr161_vgpr162_vgpr163_vgpr164_vgpr165_vgpr166_vgpr167_vgpr168_vgpr169_vgpr170_vgpr171_vgpr172_vgpr173_vgpr174_vgpr175, $vgpr176_vgpr177_vgpr178_vgpr179_vgpr180_vgpr181_vgpr182_vgpr183_vgpr184_vgpr185_vgpr186_vgpr187_vgpr188_vgpr189_vgpr190_vgpr191, $vgpr192_vgpr193_vgpr194_vgpr195_vgpr196_vgpr197_vgpr198_vgpr199_vgpr200_vgpr201_vgpr202_vgpr203_vgpr204_vgpr205_vgpr206_vgpr207, $vgpr208_vgpr209_vgpr210_vgpr211_vgpr212_vgpr213_vgpr214_vgpr215_vgpr216_vgpr217_vgpr218_vgpr219_vgpr220_vgpr221_vgpr222_vgpr223, $vgpr224_vgpr225_vgpr226_vgpr227_vgpr228_vgpr229_vgpr230_vgpr231_vgpr232_vgpr233_vgpr234_vgpr235_vgpr236_vgpr237_vgpr238_vgpr239, $vgpr240_vgpr241_vgpr242_vgpr243_vgpr244_vgpr245_vgpr246_vgpr247, $vgpr248_vgpr249_vgpr250_vgpr251, $vgpr252_vgpr253_vgpr254_vgpr255, $sgpr8_sgpr9
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S64_SAVE $sgpr8_sgpr9, %stack.1, implicit $exec, implicit $sgpr32:: (store (s64) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE $sgpr8_sgpr9, %stack.1, -1, implicit $exec, implicit $sgpr32:: (store (s64) into %stack.1, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
@@ -808,8 +808,8 @@ body: |
bb.0:
liveins: $vgpr0_vgpr1_vgpr2_vgpr3_vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15, $vgpr16_vgpr17_vgpr18_vgpr19_vgpr20_vgpr21_vgpr22_vgpr23_vgpr24_vgpr25_vgpr26_vgpr27_vgpr28_vgpr29_vgpr30_vgpr31, $vgpr32_vgpr33_vgpr34_vgpr35_vgpr36_vgpr37_vgpr38_vgpr39_vgpr40_vgpr41_vgpr42_vgpr43_vgpr44_vgpr45_vgpr46_vgpr47, $vgpr48_vgpr49_vgpr50_vgpr51_vgpr52_vgpr53_vgpr54_vgpr55_vgpr56_vgpr57_vgpr58_vgpr59_vgpr60_vgpr61_vgpr62_vgpr63, $vgpr64_vgpr65_vgpr66_vgpr67_vgpr68_vgpr69_vgpr70_vgpr71_vgpr72_vgpr73_vgpr74_vgpr75_vgpr76_vgpr77_vgpr78_vgpr79, $vgpr80_vgpr81_vgpr82_vgpr83_vgpr84_vgpr85_vgpr86_vgpr87_vgpr88_vgpr89_vgpr90_vgpr91_vgpr92_vgpr93_vgpr94_vgpr95, $vgpr96_vgpr97_vgpr98_vgpr99_vgpr100_vgpr101_vgpr102_vgpr103_vgpr104_vgpr105_vgpr106_vgpr107_vgpr108_vgpr109_vgpr110_vgpr111, $vgpr112_vgpr113_vgpr114_vgpr115_vgpr116_vgpr117_vgpr118_vgpr119_vgpr120_vgpr121_vgpr122_vgpr123_vgpr124_vgpr125_vgpr126_vgpr127, $vgpr128_vgpr129_vgpr130_vgpr131_vgpr132_vgpr133_vgpr134_vgpr135_vgpr136_vgpr137_vgpr138_vgpr139_vgpr140_vgpr141_vgpr142_vgpr143, $vgpr144_vgpr145_vgpr146_vgpr147_vgpr148_vgpr149_vgpr150_vgpr151_vgpr152_vgpr153_vgpr154_vgpr155_vgpr156_vgpr157_vgpr158_vgpr159, $vgpr160_vgpr161_vgpr162_vgpr163_vgpr164_vgpr165_vgpr166_vgpr167_vgpr168_vgpr169_vgpr170_vgpr171_vgpr172_vgpr173_vgpr174_vgpr175, $vgpr176_vgpr177_vgpr178_vgpr179_vgpr180_vgpr181_vgpr182_vgpr183_vgpr184_vgpr185_vgpr186_vgpr187_vgpr188_vgpr189_vgpr190_vgpr191, $vgpr192_vgpr193_vgpr194_vgpr195_vgpr196_vgpr197_vgpr198_vgpr199_vgpr200_vgpr201_vgpr202_vgpr203_vgpr204_vgpr205_vgpr206_vgpr207, $vgpr208_vgpr209_vgpr210_vgpr211_vgpr212_vgpr213_vgpr214_vgpr215_vgpr216_vgpr217_vgpr218_vgpr219_vgpr220_vgpr221_vgpr222_vgpr223, $vgpr224_vgpr225_vgpr226_vgpr227_vgpr228_vgpr229_vgpr230_vgpr231_vgpr232_vgpr233_vgpr234_vgpr235_vgpr236_vgpr237_vgpr238_vgpr239, $vgpr240_vgpr241_vgpr242_vgpr243_vgpr244_vgpr245_vgpr246_vgpr247, $vgpr248_vgpr249_vgpr250_vgpr251, $vgpr252_vgpr253_vgpr254_vgpr255, $sgpr8, $sgpr9
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE $sgpr8, %stack.1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.1, align 4, addrspace 5)
- SI_SPILL_S32_SAVE $sgpr9, %stack.1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.2, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr8, %stack.1, -1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.1, align 4, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr9, %stack.1, -1, implicit $exec, implicit $sgpr32:: (store (s32) into %stack.2, align 4, addrspace 5)
S_CBRANCH_SCC1 %bb.2, implicit $scc
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir
index 8b64e4fb09308..c86539eaf95cd 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir
@@ -49,7 +49,7 @@ body: |
; CHECK-NEXT: $exec = S_MOV_B64 killed $sgpr4_sgpr5, implicit killed $vgpr1
; CHECK-NEXT: S_SETPC_B64 $sgpr30_sgpr31, implicit $scc
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_SETPC_B64 $sgpr30_sgpr31, implicit $scc
...
@@ -93,7 +93,7 @@ body: |
; CHECK-NEXT: $exec = S_MOV_B64 killed $sgpr4_sgpr5, implicit killed $vgpr1
; CHECK-NEXT: S_SETPC_B64 $sgpr30_sgpr31, implicit $scc
S_CMP_EQ_U32 0, 0, implicit-def $scc
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_SETPC_B64 $sgpr30_sgpr31, implicit $scc
...
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir
index 2b795daf812e7..4898e38349d38 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir
@@ -36,9 +36,9 @@
# SHARE: stack-id: sgpr-spill, callee-saved-register: '', callee-saved-restored: true,
# SHARE: debug-info-variable: '', debug-info-expression: '', debug-info-location: '' }
-# SHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.2, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.2, addrspace 5)
+# SHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.2, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.2, addrspace 5)
# SHARE: SI_SPILL_V32_SAVE killed $vgpr0, %stack.0, $sgpr32, 0, implicit $exec :: (store (s32) into %stack.0, addrspace 5)
-# SHARE: SI_SPILL_S64_SAVE killed renamable $sgpr4_sgpr5, %stack.1, implicit $exec, implicit $sgpr32 :: (store (s64) into %stack.1, align 4, addrspace 5)
+# SHARE: SI_SPILL_S64_SAVE killed renamable $sgpr4_sgpr5, %stack.1, -1, implicit $exec, implicit $sgpr32 :: (store (s64) into %stack.1, align 4, addrspace 5)
# SHARE: renamable $sgpr4_sgpr5 = SI_SPILL_S64_RESTORE %stack.1, implicit $exec, implicit $sgpr32 :: (load (s64) from %stack.1, align 4, addrspace 5)
# SHARE: dead $sgpr30_sgpr31 = SI_CALL killed renamable $sgpr4_sgpr5, @func, csr_amdgpu, implicit undef $vgpr0
# SHARE: $sgpr32 = SI_SPILL_S32_RESTORE %stack.2, implicit $exec, implicit $sgpr32 :: (load (s32) from %stack.2, addrspace 5)
@@ -61,13 +61,13 @@
# NOSHARE: stack-id: sgpr-spill, callee-saved-register: '', callee-saved-restored: true,
# NOSHARE: debug-info-variable: '', debug-info-expression: '', debug-info-location: '' }
-# NOSHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.2, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.2, addrspace 5)
+# NOSHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.2, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.2, addrspace 5)
# NOSHARE: SI_SPILL_V32_SAVE killed $vgpr0, %stack.0, $sgpr32, 0, implicit $exec :: (store (s32) into %stack.0, addrspace 5)
-# NOSHARE: SI_SPILL_S64_SAVE killed renamable $sgpr4_sgpr5, %stack.1, implicit $exec, implicit $sgpr32 :: (store (s64) into %stack.1, align 4, addrspace 5)
+# NOSHARE: SI_SPILL_S64_SAVE killed renamable $sgpr4_sgpr5, %stack.1, -1, implicit $exec, implicit $sgpr32 :: (store (s64) into %stack.1, align 4, addrspace 5)
# NOSHARE: renamable $sgpr4_sgpr5 = SI_SPILL_S64_RESTORE %stack.1, implicit $exec, implicit $sgpr32 :: (load (s64) from %stack.1, align 4, addrspace 5)
# NOSHARE: dead $sgpr30_sgpr31 = SI_CALL killed renamable $sgpr4_sgpr5, @func, csr_amdgpu, implicit undef $vgpr0
# NOSHARE: $sgpr32 = SI_SPILL_S32_RESTORE %stack.2, implicit $exec, implicit $sgpr32 :: (load (s32) from %stack.2, addrspace 5)
-# NOSHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.3, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.3, addrspace 5)
+# NOSHARE: SI_SPILL_S32_SAVE $sgpr32, %stack.3, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.3, addrspace 5)
# NOSHARE: $vgpr0 = SI_SPILL_V32_RESTORE %stack.0, $sgpr32, 0, implicit $exec :: (load (s32) from %stack.0, addrspace 5)
# NOSHARE: renamable $sgpr4_sgpr5 = SI_SPILL_S64_RESTORE %stack.1, implicit $exec, implicit $sgpr32 :: (load (s64) from %stack.1, align 4, addrspace 5)
# NOSHARE: dead $sgpr30_sgpr31 = SI_CALL killed renamable $sgpr4_sgpr5, @func, csr_amdgpu, implicit $vgpr0
diff --git a/llvm/test/CodeGen/AMDGPU/sgpr-spill.mir b/llvm/test/CodeGen/AMDGPU/sgpr-spill.mir
index 045fd3ce7f58b..ff94ce8df210e 100644
--- a/llvm/test/CodeGen/AMDGPU/sgpr-spill.mir
+++ b/llvm/test/CodeGen/AMDGPU/sgpr-spill.mir
@@ -543,37 +543,37 @@ body: |
; GCN64-FLATSCR-NEXT: $vgpr0 = SCRATCH_LOAD_DWORD_SADDR $sgpr33, 0, 0, implicit $exec, implicit $flat_scr :: ("amdgpu-thread-private" load (s32) from %stack.9, addrspace 5)
; GCN64-FLATSCR-NEXT: $exec = S_MOV_B64 killed $sgpr0_sgpr1, implicit killed $vgpr0
renamable $sgpr12 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE killed $sgpr12, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr12, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE $sgpr12, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE $sgpr12, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13 = IMPLICIT_DEF
- SI_SPILL_S64_SAVE killed $sgpr12_sgpr13, %stack.1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S64_SAVE killed $sgpr12_sgpr13, %stack.1, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13 = IMPLICIT_DEF
- SI_SPILL_S64_SAVE $sgpr12_sgpr13, %stack.1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S64_SAVE $sgpr12_sgpr13, %stack.1, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13_sgpr14 = IMPLICIT_DEF
- SI_SPILL_S96_SAVE killed $sgpr12_sgpr13_sgpr14, %stack.2, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S96_SAVE killed $sgpr12_sgpr13_sgpr14, %stack.2, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13_sgpr14_sgpr15 = IMPLICIT_DEF
- SI_SPILL_S128_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15, %stack.3, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S128_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15, %stack.3, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16 = IMPLICIT_DEF
- SI_SPILL_S160_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16, %stack.4, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S160_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16, %stack.4, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19 = IMPLICIT_DEF
- SI_SPILL_S256_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19, %stack.5, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S256_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19, %stack.5, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19_sgpr20_sgpr21_sgpr22_sgpr23_sgpr24_sgpr25_sgpr26_sgpr27 = IMPLICIT_DEF
- SI_SPILL_S512_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19_sgpr20_sgpr21_sgpr22_sgpr23_sgpr24_sgpr25_sgpr26_sgpr27, %stack.6, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S512_SAVE killed $sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19_sgpr20_sgpr21_sgpr22_sgpr23_sgpr24_sgpr25_sgpr26_sgpr27, %stack.6, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95 = IMPLICIT_DEF
- SI_SPILL_S1024_SAVE killed $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95, %stack.7, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95, %stack.7, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
renamable $sgpr12 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE $sgpr12, %stack.8, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S32_SAVE $sgpr12, %stack.8, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
...
---
diff --git a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir
index 469a10d47fc1e..615500996906b 100644
--- a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir
+++ b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir
@@ -99,14 +99,14 @@ body: |
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_CMP_EQ_U32 $sgpr11, 0, implicit-def $scc
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_CBRANCH_SCC1 %bb.2, implicit killed $scc
S_BRANCH %bb.3
bb.2:
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
$sgpr10 = S_MOV_B32 1
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.1
bb.3:
liveins: $sgpr30_sgpr31
diff --git a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir
index 735899506e6ac..acf263a7ba2ef 100644
--- a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir
+++ b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir
@@ -62,7 +62,7 @@ body: |
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
$sgpr10 = S_MOV_B32 1
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.3
bb.3:
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
diff --git a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir
index 10b7a670e6384..3867abe0756a3 100644
--- a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir
+++ b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir
@@ -56,7 +56,7 @@ body: |
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
$sgpr10 = S_MOV_B32 1
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.1
bb.3:
liveins: $sgpr30_sgpr31
diff --git a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir
index 63939ff5d69ae..1cc9bea0dfb0f 100644
--- a/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir
+++ b/llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir
@@ -134,14 +134,14 @@ body: |
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_CMP_EQ_U32 $sgpr11, 0, implicit-def $scc
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_CBRANCH_SCC1 %bb.3, implicit killed $scc
S_BRANCH %bb.5
bb.3:
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
$sgpr10 = S_MOV_B32 1
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.2
bb.5:
liveins: $sgpr30_sgpr31
diff --git a/llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir b/llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir
index 0e74535ecaf61..49f25ffb70313 100644
--- a/llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir
@@ -46,7 +46,7 @@ body: |
; GCN-NEXT: $exec = S_MOV_B64 killed $sgpr0_sgpr1
; GCN-NEXT: S_ENDPGM 0, implicit $sgpr8
renamable $sgpr1 = COPY $sgpr2
- SI_SPILL_S128_SAVE renamable $sgpr0_sgpr1_sgpr2_sgpr3, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s128) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S128_SAVE renamable $sgpr0_sgpr1_sgpr2_sgpr3, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s128) into %stack.0, align 4, addrspace 5)
renamable $sgpr8 = COPY killed renamable $sgpr1
S_ENDPGM 0, implicit $sgpr8
...
@@ -91,7 +91,7 @@ body: |
; GCN-NEXT: $exec = S_MOV_B64 killed $sgpr0_sgpr1
; GCN-NEXT: S_ENDPGM 0
renamable $sgpr1 = COPY $sgpr2
- SI_SPILL_S128_SAVE renamable killed $sgpr0_sgpr1_sgpr2_sgpr3, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s128) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S128_SAVE renamable killed $sgpr0_sgpr1_sgpr2_sgpr3, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s128) into %stack.0, align 4, addrspace 5)
S_ENDPGM 0
...
diff --git a/llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll b/llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll
index d45833b1fc262..31f6237e38ce9 100644
--- a/llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll
+++ b/llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll
@@ -9756,12 +9756,12 @@ define amdgpu_kernel void @test_limited_sgpr(ptr addrspace(1) %out, ptr addrspac
; GFX6-NEXT: s_mov_b32 s7, 0xf000
; GFX6-NEXT: s_mov_b64 exec, 15
; GFX6-NEXT: buffer_store_dword v1, off, s[40:43], 0
-; GFX6-NEXT: s_waitcnt expcnt(0) lgkmcnt(0)
+; GFX6-NEXT: s_mov_b32 s8, 0x80400
+; GFX6-NEXT: s_waitcnt expcnt(0)
+; GFX6-NEXT: buffer_load_dword v1, off, s[40:43], s8 ; 4-byte Folded Reload
+; GFX6-NEXT: s_waitcnt vmcnt(0) lgkmcnt(0)
; GFX6-NEXT: v_writelane_b32 v1, s0, 0
; GFX6-NEXT: v_writelane_b32 v1, s1, 1
-; GFX6-NEXT: v_writelane_b32 v1, s2, 2
-; GFX6-NEXT: v_writelane_b32 v1, s3, 3
-; GFX6-NEXT: s_mov_b32 s8, 0x80400
; GFX6-NEXT: buffer_store_dword v1, off, s[40:43], s8 ; 4-byte Folded Spill
; GFX6-NEXT: s_waitcnt expcnt(0)
; GFX6-NEXT: buffer_load_dword v1, off, s[40:43], 0
@@ -9885,12 +9885,12 @@ define amdgpu_kernel void @test_limited_sgpr(ptr addrspace(1) %out, ptr addrspac
; GFX6-NEXT: s_mov_b64 s[2:3], s[6:7]
; GFX6-NEXT: s_mov_b64 exec, 15
; GFX6-NEXT: buffer_store_dword v4, off, s[40:43], 0
+; GFX6-NEXT: s_mov_b32 s10, 0x80800
; GFX6-NEXT: s_waitcnt expcnt(0)
-; GFX6-NEXT: v_writelane_b32 v4, s0, 0
-; GFX6-NEXT: v_writelane_b32 v4, s1, 1
+; GFX6-NEXT: buffer_load_dword v4, off, s[40:43], s10 ; 4-byte Folded Reload
+; GFX6-NEXT: s_waitcnt vmcnt(0)
; GFX6-NEXT: v_writelane_b32 v4, s2, 2
; GFX6-NEXT: v_writelane_b32 v4, s3, 3
-; GFX6-NEXT: s_mov_b32 s10, 0x80800
; GFX6-NEXT: buffer_store_dword v4, off, s[40:43], s10 ; 4-byte Folded Spill
; GFX6-NEXT: s_waitcnt expcnt(0)
; GFX6-NEXT: buffer_load_dword v4, off, s[40:43], 0
diff --git a/llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir b/llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir
index f7fafa0248835..6ef3cc50b2a18 100644
--- a/llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir
@@ -27,7 +27,7 @@ body: |
; GCN-NEXT: $sgpr10 = SI_RESTORE_S32_FROM_VGPR [[DEF]], 0
; GCN-NEXT: S_SETPC_B64 $sgpr30_sgpr31
S_NOP 0
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_SETPC_B64 $sgpr30_sgpr31
...
@@ -158,8 +158,8 @@ body: |
; GCN-NEXT: $sgpr10 = SI_RESTORE_S32_FROM_VGPR [[DEF]], 0
; GCN-NEXT: S_SETPC_B64 $sgpr30_sgpr31
S_NOP 0
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S1024_SAVE killed $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95, %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95, %stack.1, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_NOP 0
renamable $sgpr64_sgpr65_sgpr66_sgpr67_sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95 = SI_SPILL_S1024_RESTORE %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr10 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
@@ -221,12 +221,12 @@ body: |
bb.1:
liveins: $sgpr10, $sgpr30_sgpr31
$sgpr10 = S_MOV_B32 10
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.3
bb.2:
liveins: $sgpr10, $sgpr30_sgpr31
$sgpr10 = S_MOV_B32 20
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_BRANCH %bb.3
bb.3:
liveins: $sgpr10, $sgpr30_sgpr31
@@ -308,7 +308,7 @@ body: |
bb.3:
liveins: $sgpr10, $sgpr11, $sgpr30_sgpr31
$sgpr10 = S_MOV_B32 10
- SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_CMP_EQ_U32 $sgpr11, 0, implicit-def $scc
S_CBRANCH_SCC1 %bb.2, implicit killed $scc
S_BRANCH %bb.1
diff --git a/llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir b/llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir
index 9e15d46feaf26..6c982ed3c5ce7 100644
--- a/llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir
@@ -142,10 +142,10 @@ body: |
; GFX11-NEXT: $vgpr0 = SCRATCH_LOAD_DWORD_SADDR $sgpr33, 8, 0, implicit $exec, implicit $flat_scr :: ("amdgpu-thread-private" load (s32) from %stack.1, addrspace 5)
; GFX11-NEXT: $exec = S_MOV_B64 killed $sgpr0_sgpr1, implicit killed $vgpr0
$vcc = IMPLICIT_DEF
- SI_SPILL_S64_SAVE $vcc, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S64_SAVE $vcc, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
$vcc = IMPLICIT_DEF
- SI_SPILL_S64_SAVE killed $vcc, %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
+ SI_SPILL_S64_SAVE killed $vcc, %stack.0, -1, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
$vcc = SI_SPILL_S64_RESTORE %stack.0, implicit $exec, implicit $sgpr96_sgpr97_sgpr98_sgpr99, implicit $sgpr32
...
diff --git a/llvm/test/CodeGen/AMDGPU/spill192.mir b/llvm/test/CodeGen/AMDGPU/spill192.mir
index 2f8daa0c5a721..fa6c423f1685c 100644
--- a/llvm/test/CodeGen/AMDGPU/spill192.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill192.mir
@@ -21,7 +21,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9
- ; SPILLED-NEXT: SI_SPILL_S192_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s192) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S192_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s192) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/spill224.mir b/llvm/test/CodeGen/AMDGPU/spill224.mir
index 9fe13c0541064..f40645d7598b3 100644
--- a/llvm/test/CodeGen/AMDGPU/spill224.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill224.mir
@@ -17,7 +17,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10
- ; SPILLED-NEXT: SI_SPILL_S224_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s224) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S224_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s224) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/spill288.mir b/llvm/test/CodeGen/AMDGPU/spill288.mir
index e370f9b1c3d50..53bcab7c429e0 100644
--- a/llvm/test/CodeGen/AMDGPU/spill288.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill288.mir
@@ -17,7 +17,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12
- ; SPILLED-NEXT: SI_SPILL_S288_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s288) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S288_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s288) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/spill320.mir b/llvm/test/CodeGen/AMDGPU/spill320.mir
index b942f925417dd..38013e98567d6 100644
--- a/llvm/test/CodeGen/AMDGPU/spill320.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill320.mir
@@ -17,7 +17,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13
- ; SPILLED-NEXT: SI_SPILL_S320_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s320) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S320_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s320) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/spill352.mir b/llvm/test/CodeGen/AMDGPU/spill352.mir
index e7090e69965e1..40cd772f57228 100644
--- a/llvm/test/CodeGen/AMDGPU/spill352.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill352.mir
@@ -17,7 +17,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14
- ; SPILLED-NEXT: SI_SPILL_S352_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s352) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S352_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s352) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/spill384.mir b/llvm/test/CodeGen/AMDGPU/spill384.mir
index 1ede641509a67..ad1bc7ae0d18f 100644
--- a/llvm/test/CodeGen/AMDGPU/spill384.mir
+++ b/llvm/test/CodeGen/AMDGPU/spill384.mir
@@ -17,7 +17,7 @@ body: |
; SPILLED-NEXT: successors: %bb.1(0x80000000)
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: S_NOP 0, implicit-def renamable $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15
- ; SPILLED-NEXT: SI_SPILL_S384_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s384) into %stack.0, align 4, addrspace 5)
+ ; SPILLED-NEXT: SI_SPILL_S384_SAVE killed $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s384) into %stack.0, align 4, addrspace 5)
; SPILLED-NEXT: S_CBRANCH_SCC1 %bb.1, implicit undef $scc
; SPILLED-NEXT: {{ $}}
; SPILLED-NEXT: bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir b/llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir
index 01509ffe08a1d..e8ea363183570 100644
--- a/llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir
+++ b/llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir
@@ -9,7 +9,7 @@
# CHECK: - { id: 0, name: '', type: spill-slot, offset: 0, size: 4, alignment: 4,
# CHECK-NEXT: stack-id: sgpr-spill,
-# CHECK: SI_SPILL_S32_SAVE killed renamable $sgpr5, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
+# CHECK: SI_SPILL_S32_SAVE killed renamable $sgpr5, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
# CHECK: renamable $sgpr5 = SI_SPILL_S32_RESTORE %stack.0, implicit $exec, implicit $sgpr32 :: (load (s32) from %stack.0, addrspace 5)
name: no_merge_sgpr_vgpr_spill_slot
diff --git a/llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir b/llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir
index c70327aaea993..588c75c0cf89f 100644
--- a/llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir
+++ b/llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir
@@ -49,7 +49,7 @@ body: |
$sgpr0_sgpr1 = V_CMP_EQ_U32_e64 1, killed $vgpr1, implicit $exec
$vgpr1 = V_CNDMASK_B32_e64 0, 0, 0, -1, killed $sgpr0_sgpr1, implicit $exec
$sgpr0_sgpr1 = COPY $exec, implicit-def $exec
- SI_SPILL_S64_SAVE $sgpr0_sgpr1, %stack.0, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE $sgpr0_sgpr1, %stack.0, -1, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
$sgpr2_sgpr3 = S_AND_B64 killed $sgpr0_sgpr1, killed $vcc, implicit-def dead $scc
$exec = S_MOV_B64_term killed $sgpr2_sgpr3
S_CBRANCH_EXECZ %bb.2, implicit $exec
@@ -84,7 +84,7 @@ body: |
# CHECK-LABEL: {{^}}name: undefined_physreg_sgpr_spill_reorder
# CHECK: $sgpr0_sgpr1 = COPY $exec, implicit-def $exec
# CHECK: $sgpr2_sgpr3 = S_AND_B64 $sgpr0_sgpr1, killed $vcc, implicit-def dead $scc
-# CHECK: SI_SPILL_S64_SAVE killed $sgpr0_sgpr1, %stack.0, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
+# CHECK: SI_SPILL_S64_SAVE killed $sgpr0_sgpr1, %stack.0, -1, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
# CHECK: $exec = COPY killed $sgpr2_sgpr3
name: undefined_physreg_sgpr_spill_reorder
alignment: 1
@@ -115,7 +115,7 @@ body: |
$vgpr1 = V_CNDMASK_B32_e64 0, 0, 0, -1, killed $sgpr0_sgpr1, implicit $exec
$sgpr0_sgpr1 = COPY $exec, implicit-def $exec
$sgpr2_sgpr3 = S_AND_B64 $sgpr0_sgpr1, killed $vcc, implicit-def dead $scc
- SI_SPILL_S64_SAVE killed $sgpr0_sgpr1, %stack.0, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
+ SI_SPILL_S64_SAVE killed $sgpr0_sgpr1, %stack.0, -1, implicit $exec, implicit $sgpr8_sgpr9_sgpr10_sgpr11, implicit $sgpr13, implicit-def dead $m0 :: (store (s64) into %stack.0, align 4, addrspace 5)
$exec = S_MOV_B64_term killed $sgpr2_sgpr3
S_CBRANCH_EXECZ %bb.2, implicit $exec
S_BRANCH %bb.1
diff --git a/llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir b/llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir
index 6c79c594771b9..7dcd0c7a32112 100644
--- a/llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir
+++ b/llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir
@@ -40,9 +40,9 @@ body: |
bb.0:
liveins: $sgpr34, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99
- SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr34 = SI_SPILL_S32_RESTORE %stack.2, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99 = SI_SPILL_S1024_RESTORE %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67 = SI_SPILL_S1024_RESTORE %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
diff --git a/llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir b/llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir
index 8647c6314c88e..1e12642fdc22e 100644
--- a/llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir
+++ b/llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir
@@ -32,9 +32,9 @@ body: |
liveins: $sgpr34, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99
%0:sreg_64 = ENTER_STRICT_WWM -1, implicit-def $exec, implicit-def $scc, implicit $exec
$exec = EXIT_STRICT_WWM killed %0
- SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_ENDPGM 0
...
@@ -63,9 +63,9 @@ machineFunctionInfo:
body: |
bb.0:
liveins: $sgpr34, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99
- SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_ENDPGM 0
...
@@ -94,8 +94,8 @@ machineFunctionInfo:
body: |
bb.0:
liveins: $sgpr34, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99
- SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
- SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51_sgpr52_sgpr53_sgpr54_sgpr55_sgpr56_sgpr57_sgpr58_sgpr59_sgpr60_sgpr61_sgpr62_sgpr63_sgpr64_sgpr65_sgpr66_sgpr67, %stack.0, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S1024_SAVE killed $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.1, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
+ SI_SPILL_S32_SAVE killed $sgpr34, %stack.2, -1, implicit $exec, implicit $sgpr0_sgpr1_sgpr2_sgpr3, implicit $sgpr32
S_ENDPGM 0
...
diff --git a/llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir b/llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir
index ce8aebe7458ed..442a653efc01c 100644
--- a/llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir
+++ b/llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir
@@ -42,7 +42,7 @@ body: |
bb.4:
liveins: $sgpr0_sgpr1, $sgpr2_sgpr3, $sgpr4_sgpr5
- SI_SPILL_S64_SAVE killed $sgpr2_sgpr3, %stack.0, implicit $exec, implicit $sgpr32
+ SI_SPILL_S64_SAVE killed $sgpr2_sgpr3, %stack.0, -1, implicit $exec, implicit $sgpr32
$sgpr56_sgpr57 = S_LOAD_DWORDX2_IMM $sgpr4_sgpr5, 48, 0
$sgpr2 = S_MOV_B32 124
%temp5:vreg_64 = COPY $sgpr56_sgpr57
diff --git a/llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir b/llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir
index 04a19868b0671..7345532dc0b89 100644
--- a/llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir
+++ b/llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir
@@ -157,8 +157,8 @@ body: |
; CHECK-NEXT: successors: %bb.3(0x40000000), %bb.2(0x40000000)
; CHECK-NEXT: liveins: $d2:0x0000000000000002, $d8:0x0000000000000001, $d13, $r20
; CHECK-NEXT: {{ $}}
- ; CHECK-NEXT: S2_storerd_io %stack.0, 0, renamable $d2 :: (store (s64) into %stack.0)
; CHECK-NEXT: ADJCALLSTACKDOWN 0, 0, implicit-def dead $r29, implicit-def dead $r30, implicit $r31, implicit $r30, implicit $r29
+ ; CHECK-NEXT: S2_storerd_io %stack.0, 0, renamable $d2 :: (store (s64) into %stack.0)
; CHECK-NEXT: J2_call @lrand48, implicit-def dead $d0, implicit-def dead $d1, implicit-def dead $d2, implicit-def dead $d3, implicit-def dead $d4, implicit-def dead $d5, implicit-def dead $d6, implicit-def dead $d7, implicit-def dead $r28, implicit-def dead $r31, implicit-def dead $p0, implicit-def dead $p1, implicit-def dead $p2, implicit-def dead $p3, implicit-def dead $m0, implicit-def dead $m1, implicit-def dead $lc0, implicit-def dead $lc1, implicit-def dead $sa0, implicit-def dead $sa1, implicit-def dead $usr, implicit-def $usr_ovf, implicit-def dead $cs0, implicit-def dead $cs1, implicit-def dead $w0, implicit-def dead $w1, implicit-def dead $w2, implicit-def dead $w3, implicit-def dead $w4, implicit-def dead $w5, implicit-def dead $w6, implicit-def dead $w7, implicit-def dead $w8, implicit-def dead $w9, implicit-def dead $w10, implicit-def dead $w11, implicit-def dead $w12, implicit-def dead $w13, implicit-def dead $w14, implicit-def dead $w15, implicit-def dead $q0, implicit-def dead $q1, implicit-def dead $q2, implicit-def dead $q3, implicit-def $r0
; CHECK-NEXT: ADJCALLSTACKUP 0, 0, implicit-def dead $r29, implicit-def dead $r30, implicit-def dead $r31, implicit $r29
; CHECK-NEXT: renamable $r18 = COPY $r0
More information about the llvm-commits
mailing list