[llvm] [InlineSpiller][AMDGPU] Skip undef lanes when spilling a partial-def super-register (PR #226689)
via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 26 05:57:37 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu
@llvm/pr-subscribers-llvm-regalloc
Author: xgxanq
<details>
<summary>Changes</summary>
When the greedy allocator live-range-splits a wide register (e.g. an sgpr_128 buffer descriptor) on a single sub-register, it produces sibling vregs each defined on one sub-lane and undef elsewhere. InlineSpiller shares one stack slot among all descendants of the Original value, so a full-width store of such a partially-defined sibling writes its undef lanes over a live sibling value in the same slot and corrupts it. On gfx950 this was observed (via asm data-flow analysis) as an in-loop buffer-descriptor lane being clobbered by an unrelated in-loop spill sharing the slot.
The hazard is target-independent (InlineSpiller's shared-slot logic is generic; SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs). Fix it in two places:
1. hoistSpillInsideBB: bail out unless every sub-range of the hoisted value is live at the hoist point, mirroring the cross-BB guard in isSpillCandBB (#<!-- -->177703) that the in-BB path was missing.
2. insertSpill: for a real (partial) spill, compute the defined-lane mask and hand it to the target via a new TargetInstrInfo::setSpillDefinedLaneMask hook (default no-op). AMDGPU records it in the SI_SPILL_S*_SAVE $lanemask operand; spillSGPR then writelanes only the defined dwords, leaving the slot's other lanes (a sibling value) intact. A -1 mask preserves the original full-width behavior.
Existing SGPR-spill tests are regenerated to carry the new -1 $lanemask operand. Adds inline-spiller-partial-subreg-shared-slot.ll as a gfx950 end-to-end regression check.
---
Patch is 95.32 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/226689.diff
40 Files Affected:
- (modified) llvm/include/llvm/CodeGen/TargetInstrInfo.h (+6)
- (modified) llvm/lib/CodeGen/InlineSpiller.cpp (+32-2)
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.cpp (+33-4)
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.h (+3)
- (modified) llvm/lib/Target/AMDGPU/SIInstructions.td (+8-2)
- (modified) llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp (+47-4)
- (modified) llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir (+1-1)
- (added) llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll (+257)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir (+4-14)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir (+6-6)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir (+5-5)
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill.mir (+11-11)
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll (+7-7)
- (modified) llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir (+6-6)
- (modified) llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir (+2-2)
- (modified) llvm/test/CodeGen/AMDGPU/spill192.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/spill224.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/spill288.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/spill320.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/spill352.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/spill384.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir (+1-1)
- (modified) llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir (+3-3)
- (modified) llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir (+3-3)
- (modified) llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir (+9-9)
- (modified) llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir (+1-1)
- (modified) llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir (+1-1)
``````````diff
diff --git a/llvm/include/llvm/CodeGen/TargetInstrInfo.h b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
index b013511d33313..022c39d0ca92d 100644
--- a/llvm/include/llvm/CodeGen/TargetInstrInfo.h
+++ b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
@@ -1226,6 +1226,12 @@ class LLVM_ABI TargetInstrInfo : public MCInstrInfo {
"TargetInstrInfo::storeRegToStackSlot!");
}
+ /// Tell the target which lanes of spill \p SpillMI hold a real value
+ /// (\p DefinedLanes); it can skip the undef ones when lowering, so a shared
+ /// slot's neighbor isn't stomped. Default: no-op.
+ virtual void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const {}
+
/// Load the specified register of the given register class from the specified
/// stack frame index. The load instruction is to be added to the given
/// machine basic block before the specified machine instruction. If \p
diff --git a/llvm/lib/CodeGen/InlineSpiller.cpp b/llvm/lib/CodeGen/InlineSpiller.cpp
index f3682a2e24808..a7a960a0cf441 100644
--- a/llvm/lib/CodeGen/InlineSpiller.cpp
+++ b/llvm/lib/CodeGen/InlineSpiller.cpp
@@ -454,6 +454,17 @@ bool InlineSpiller::hoistSpillInsideBB(LiveInterval &SpillLI,
if (DefMBB != CopyMI.getParent() || !SrcQ.isKill())
return false;
+ // Every sub-range of the hoisted value must be live at the hoist point. The
+ // store below is full-width, so a dead sub-range would write a stale sibling
+ // value into the shared spill slot and clobber it. This mirrors the sub-range
+ // liveness guard on the cross-BB path in isSpillCandBB (#177703), which the
+ // in-BB path here was missing.
+ if (SrcLI.hasSubRanges() &&
+ !all_of(SrcLI.subranges(), [&](const LiveInterval::SubRange &SR) {
+ return SR.getVNInfoAt(Idx) != nullptr;
+ }))
+ return false;
+
MachineBasicBlock *MBB = DefMBB;
MachineBasicBlock::iterator MII;
if (SrcVNI->isPHIDef())
@@ -1285,10 +1296,29 @@ void InlineSpiller::insertSpill(Register NewVReg, bool isKill,
MachineBasicBlock::iterator SpillBefore = std::next(MI);
bool IsRealSpill = isRealSpill(*MI);
- if (IsRealSpill)
+ if (IsRealSpill) {
TII.storeRegToStackSlot(MBB, SpillBefore, NewVReg, isKill, StackSlot,
MRI.getRegClass(NewVReg), Register());
- else
+
+ MachineInstr &SpillMI = *std::next(MI);
+
+ // Compute the lanes of NewVReg defined by MI. For a partial def (undef
+ // lanes remain in the full-width store above), pass the defined-lane mask
+ // so the target skips them and avoids clobbering a sibling in the shared
+ // stack slot. A fully-undef def never reaches here (filtered by
+ // isRealSpill).
+ LaneBitmask FullMask = MRI.getMaxLaneMaskForVReg(NewVReg);
+ LaneBitmask DefinedLanes;
+ for (const MachineOperand &MO : MI->all_defs()) {
+ if (MO.getReg() != NewVReg)
+ continue;
+ DefinedLanes |= MO.getSubReg()
+ ? TRI.getSubRegIndexLaneMask(MO.getSubReg())
+ : FullMask;
+ }
+ if (DefinedLanes != FullMask)
+ TII.setSpillDefinedLaneMask(SpillMI, DefinedLanes);
+ } else
// Don't spill undef value.
// Anything works for undef, in particular keeping the memory
// uninitialized is a viable option and it saves code size and
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
index 0346ebc254361..28a06d6921aa4 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
@@ -1800,10 +1800,11 @@ void SIInstrInfo::storeRegToStackSlotImpl(
}
BuildMI(MBB, MI, DL, OpDesc)
- .addReg(SrcReg, getKillRegState(isKill)) // data
- .addFrameIndex(FrameIndex) // addr
- .addMemOperand(MMO)
- .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
+ .addReg(SrcReg, getKillRegState(isKill)) // data
+ .addFrameIndex(FrameIndex) // addr
+ .addImm(-1) // lanemask (all lanes; refined by spiller)
+ .addMemOperand(MMO)
+ .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
return;
}
@@ -1837,6 +1838,34 @@ void SIInstrInfo::storeRegToStackSlotCFI(MachineBasicBlock &MBB,
MachineInstr::NoFlags, true);
}
+void SIInstrInfo::setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const {
+ // Only SGPR spill saves carry a $lanemask operand (see SI_SPILL_SGPR).
+ int Idx =
+ AMDGPU::getNamedOperandIdx(SpillMI.getOpcode(), AMDGPU::OpName::lanemask);
+ if (Idx == -1)
+ return;
+
+ // Build the per-dword mask consumed by spillSGPR: bit i is set when dword i
+ // of the spilled super-register intersects a defined lane. spillSGPR lowers
+ // the spill dword by dword and consults this mask to skip the undef ones.
+ const MachineRegisterInfo &MRI = SpillMI.getMF()->getRegInfo();
+ Register Data = SpillMI.getOperand(0).getReg();
+ const TargetRegisterClass *RC = MRI.getRegClass(Data);
+ unsigned NumDwords = RI.getRegSizeInBits(*RC) / 32;
+ // DwordMask is a u32; the widest SGPR tuple (SReg_1024) has 32 dwords, so
+ // this holds and keeps the 1u << i shift below well-defined.
+ assert(NumDwords <= 32 && "SGPR spill wider than DwordMask can represent");
+ unsigned DwordMask = 0;
+ for (unsigned i = 0; i != NumDwords; ++i) {
+ LaneBitmask DwordLanes =
+ RI.getSubRegIndexLaneMask(RI.getSubRegFromChannel(i));
+ if ((DefinedLanes & DwordLanes).any())
+ DwordMask |= 1u << i;
+ }
+ SpillMI.getOperand(Idx).setImm(DwordMask);
+}
+
static unsigned getSGPRSpillRestoreOpcode(unsigned Size) {
switch (Size) {
case 4:
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.h b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
index 48eba6b2a567d..42afa3cd8c60e 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.h
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
@@ -356,6 +356,9 @@ class SIInstrInfo final : public AMDGPUGenInstrInfo {
bool isKill, int FrameIndex, const TargetRegisterClass *RC, Register VReg,
MachineInstr::MIFlag Flags = MachineInstr::NoFlags) const override;
+ void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+ LaneBitmask DefinedLanes) const override;
+
void loadRegFromStackSlot(
MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, Register DestReg,
int FrameIndex, const TargetRegisterClass *RC, Register VReg,
diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 3cf20ffb1fcf4..3faae31302689 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -1178,14 +1178,20 @@ def V_INDIRECT_REG_READ_GPR_IDX_B32_V32 : V_INDIRECT_REG_READ_GPR_IDX_pseudo<VRe
multiclass SI_SPILL_SGPR <RegisterClass sgpr_class> {
let UseNamedOperandTable = 1, Spill = 1, SALU = 1, Uses = [EXEC] in {
+ // $lanemask: per-dword bitmask of which dwords of the spilled super-register
+ // are defined. Undef dwords must not be written into the (possibly shared)
+ // stack slot. -1 (all dwords defined) is the ordinary full-width spill; a
+ // partial mask is set by the spiller for partially-defined values. i32imm's
+ // 32 bits cover the widest SGPR tuple SReg_1024 (32 dwords).
def _SAVE : PseudoInstSI <
(outs),
- (ins sgpr_class:$data, i32imm:$addr)> {
+ (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
let mayStore = 1;
let mayLoad = 0;
}
- def _CFI_SAVE : PseudoInstSI<(outs), (ins sgpr_class:$data, i32imm:$addr)> {
+ def _CFI_SAVE : PseudoInstSI<(outs),
+ (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
let mayStore = 1;
let mayLoad = 0;
}
diff --git a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
index 1aa423e1da647..5a8dd05c239ab 100644
--- a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
@@ -2180,6 +2180,18 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
if (OnlyToVGPR && !SpillToVGPR)
return false;
+ // A partial-def spill (see InlineSpiller / setSpillDefinedLaneMask) records
+ // which dwords are real values in its $lanemask operand; a cleared bit means
+ // that dword is undef and must NOT be written to the (possibly shared) stack
+ // slot, so it does not clobber a sibling value living in the same slot. A
+ // mask of -1 (default) means "all lanes defined" = original behavior.
+ const MachineOperand *LaneMask =
+ SB.TII.getNamedOperand(*MI, AMDGPU::OpName::lanemask);
+ int64_t DefinedDwordMask = LaneMask ? LaneMask->getImm() : -1;
+ auto DwordDefined = [&](unsigned i) {
+ return (DefinedDwordMask & (int64_t(1) << i)) != 0;
+ };
+
const SIFrameLowering *TFL = ST.getFrameLowering();
assert(SpillToVGPR || (SB.SuperReg != SB.MFI.getStackPtrOffsetReg() &&
@@ -2195,15 +2207,29 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
"Num of SGPRs spilled should be less than or equal to num of "
"the VGPR lanes.");
+ // With a partial-def mask, find the first/last dword that is actually
+ // defined, so the ImplicitDefine of the super-register and the kill flag
+ // can be re-anchored onto emitted writelanes (undef dwords are skipped).
+ unsigned FirstDefined = 0, LastDefined = SB.NumSubRegs - 1;
+ while (FirstDefined < SB.NumSubRegs && !DwordDefined(FirstDefined))
+ ++FirstDefined;
+ while (LastDefined > 0 && !DwordDefined(LastDefined))
+ --LastDefined;
+
for (unsigned i = 0, e = SB.NumSubRegs; i < e; ++i) {
+ // Skip undef dwords: not writing them to a (possibly shared) stack slot
+ // preserves a sibling value living in the same slot.
+ if (!DwordDefined(i))
+ continue;
+
Register SubReg =
SB.NumSubRegs == 1
? SB.SuperReg
: Register(getSubReg(SB.SuperReg, SB.SplitParts[i]));
SpilledReg Spill = VGPRSpills[i];
- bool IsFirstSubreg = i == 0;
- bool IsLastSubreg = i == SB.NumSubRegs - 1;
+ bool IsFirstSubreg = i == FirstDefined;
+ bool IsLastSubreg = i == LastDefined;
bool UseKill = SB.IsKill && IsLastSubreg;
@@ -2259,13 +2285,30 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
// Per VGPR helper data
auto PVD = SB.getPerVGPRData();
+ // For a partial-def spill (undef dwords, see $lanemask), the whole TmpVGPR
+ // is stored at once, so first load the slot's current contents; the skipped
+ // dwords then keep the sibling value already there instead of writing
+ // undef.
+ bool IsPartialDef = SB.NumSubRegs > 1 && DefinedDwordMask != -1;
+
for (unsigned Offset = 0; Offset < PVD.NumVGPRs; ++Offset) {
RegState TmpVGPRFlags = RegState::Undef;
+ if (IsPartialDef) {
+ // Seed TmpVGPR with the slot's current contents, so the dwords we skip
+ // (undef in this partial def) keep the sibling value already living in
+ // the slot when the whole TmpVGPR is stored back below.
+ SB.readWriteTmpVGPR(Offset, /*IsLoad*/ true);
+ TmpVGPRFlags = {};
+ }
+
// Write sub registers into the VGPR
for (unsigned i = Offset * PVD.PerVGPR,
e = std::min((Offset + 1) * PVD.PerVGPR, SB.NumSubRegs);
i < e; ++i) {
+ if (IsPartialDef && !DwordDefined(i))
+ continue; // Undef dword: keep the sibling value seeded from the slot.
+
Register SubReg =
SB.NumSubRegs == 1
? SB.SuperReg
@@ -2286,8 +2329,8 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
Indexes->insertMachineInstrInMaps(*WriteLane);
}
- // There could be undef components of a spilled super register.
- // TODO: Can we detect this and skip the spill?
+ // Undef components of the spilled super-register are detected via the
+ // $lanemask operand and skipped above (see IsPartialDef).
if (SB.NumSubRegs > 1) {
// The last implicit use of the SB.SuperReg carries the "Kill" flag.
RegState SuperKillState = {};
diff --git a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
index 5c564decd1e7d..a026d13e0cce5 100644
--- a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
+++ b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
@@ -75,7 +75,7 @@ body: |
liveins: $sgpr12, $sgpr13, $sgpr14, $sgpr15
%45:vgpr_32 = IMPLICIT_DEF
- SI_SPILL_S32_SAVE $sgpr15, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
+ SI_SPILL_S32_SAVE $sgpr15, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
%16:vgpr_32 = V_AND_B32_e32 1, %45, implicit $exec
bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
index 94c22b1aa8664..5c186b9a0470b 100644
--- a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
+++ b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
@@ -57,7 +57,7 @@ body: |
; CHECK-NEXT: renamable $sgpr86 = S_MOV_B32 0
; CHECK-NEXT: renamable $sgpr87 = S_MOV_B32 0
; CHECK-NEXT: renamable $sgpr88 = S_MOV_B32 0
- ; CHECK-NEXT: SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
+ ; CHECK-NEXT: SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, 65535, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
; CHECK-NEXT: renamable $sgpr52_sgpr53 = IMPLICIT_DEF
; CHECK-NEXT: ADJCALLSTACKUP 0, 0, implicit-def dead $scc, implicit-def $sgpr32, implicit $sgpr32
; CHECK-NEXT: dead $sgpr30_sgpr31 = SI_CALL renamable $sgpr52_sgpr53, 0, csr_amdgpu, implicit $sgpr0_sgpr1_sgpr2_sgpr3
diff --git a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
index 63ee4473a56e5..43669338dd308 100644
--- a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
+++ b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
@@ -55,7 +55,7 @@ body: |
; CHECK-NEXT: liveins: $sgpr24_sgpr25_sgpr26_sgpr27:0x000000000000000F, $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19:0x000000000000FFFF, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51:0x000000000000FFFF
; CHECK-NEXT: {{ $}}
; CHECK-NEXT: renamable $sgpr12 = IMPLICIT_DEF
- ; CHECK-NEXT: SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
+ ; CHECK-NEXT: SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, 255, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
; CHECK-NEXT: renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = IMPLICIT_DEF
; CHECK-NEXT: dead undef [[IMAGE_SAMPLE_LZ_V1_V2_2:%[0-9]+]].sub0:vreg_96 = IMAGE_SAMPLE_LZ_V1_V2 undef [[DEF2]], killed renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43, renamable $sgpr12_sgpr13_sgpr14_sgpr15, 1, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s32), addrspace 8)
; CHECK-NEXT: renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = SI_SPILL_S512_RESTORE %stack.0, implicit $exec, implicit $sgpr32 :: (load (s512) from %stack.0, align 4, addrspace 5)
diff --git a/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
new file mode 100644
index 0000000000000..bbff66888c652
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
@@ -0,0 +1,257 @@
+; RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -stop-after=greedy %s -o - | FileCheck %s
+;
+; Regression test for the InlineSpiller shared-stack-slot partial-subregister
+; hazard.
+;
+; The hazard is target-independent: InlineSpiller's shared-slot logic is generic,
+; and SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs. When the
+; greedy allocator live-range-splits an sgpr_128 descriptor on one sub-register,
+; it produces sibling vregs each defined on a single sub-lane
+; (`undef %V.sub0:sgpr_128 = ...`), undef elsewhere. Since InlineSpiller shares
+; one stack slot among descendants of the Original value, a full-width store of
+; such a sibling writes its undef lanes over a live sibling in the SAME slot. This
+; input is a gfx950 reproducer (there the clobber corrupts an in-loop buffer
+; descriptor -- the kind of corruption that leads to an HSA memory fault), but the
+; fix and this check apply to all AMDGPU.
+;
+; The fix records in the save pseudo's $lanemask operand which dwords the store
+; defines; spillSGPR writelanes only those, leaving the slot's other lanes intact.
+; This pins that the two partial siblings sharing one slot each store with a
+; partial mask (matched as a non-negative immediate: the buggy full-width store
+; used -1, which the {{[0-9]+}} pattern cannot match).
+;
+; Driven from IR through -stop-after=greedy: a .mir + -run-pass=greedy form does
+; not reproduce, since the trigger needs pre-greedy PreRARemat pressure state that
+; MIR serialization drops. The CHECK block is vreg/slot-number independent.
+;
+; First partial sibling (only .sub0 defined), stored with a partial mask:
+; CHECK: undef [[DESC:%[0-9]+]].sub0:sgpr_128 = S_MOV_B32 0
+; CHECK-NEXT: SI_SPILL_S128_SAVE [[DESC]], [[SLOT:%stack\.[0-9]+]], {{[0-9]+}},
+; Second partial sibling into the SAME slot, also with a partial mask:
+; CHECK: SI_SPILL_S128_SAVE %{{[0-9]+}}, [[SLOT]], {{[0-9]+}},
+;
+target datalayout = "e-m:e-p:64:64-p1:64:64-p2:32:32-p3:32:32-p4:64:64-p5:32:32-p6:32:32-p7:160:256:256:32-p8:128:128:128:48-p9:192:256:256:32-i64:64-v16:16-v24:32-v32:32-v48:64-v96:128-v192:256-v256:256-v512:512-v1024:1024-v2048:2048-n32:64-S32-A5-G1-ni:7:8:9-p10:32:32-p11:32:32-p12:32:32-p13:32:32-p14:32:32-p15:32:32"
+target triple = "amdgcn-amd-amdhsa"
+
+define amdgpu_kernel void @_attn_fwd_IS_CAUSAL_1_NUM_Q_HEADS_8_NUM_K_HEADS_1_BLOCK_M_256_BLOCK_N_64_BLOCK_DMODEL_64_RETURN_SCORES_1_ENABLE_DROPOUT_1_IS_FP8_0_VARLEN_1_NUM_XCD_8_USE_INT64_STRIDES_1_ENABLE_SINK_0_SLIDING_WINDOW_0(ptr addrspace(1) inreg %0, ptr addrspace(1) inreg %1, i32 %2, i32 %3, i32 %4, i32 %5, i32 %6, <1 x i32> %7, i32 %8, i32 %9, i32 %10, i32 %11, i32 %12, i32 %13, i32 %14, i32 %15, i32 %16, i32 %17, i32 %18, i32 %19, i32 %20, i32 %21, i32 %22, i32 %23, i32 %24, i32 %25, i32 %26, i64 %27, i64 %28, i64 %29, i64 %30, i64 %31, i64 %32, i64 %33, i64 %34, i64 %35, i64 %36, i64 %37, i1 %38, i1 %39, i1 %40, i1 %41, i1 %42, i1 %43, i1 %44, i64 %sext452, i64 %.pn241.in795, i64 %sext496, i64 %.pn233.in799, i64 %sext500, i64 %.pn231.in800, i64 %sext501, i64 %.pn229.in801, i64 %sext503, i64 %.pn223.in804, i64 %sext505, i64 %.pn221.in805, i64 %sext506) #0 {
+..loopexit_crit_edge:
+ %45 = icmp slt i32 %8, 0
+ %46 = icmp slt i32 %6, 0
+ %47 = icmp slt i32 %9, 0
+ %48 = icmp slt i32 %10, 0
+ %49 = icmp slt i32 %12, 0
+ %50 = icmp slt i32 %2, 0
+ br label %51
+
+51: ; preds = ...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/226689
More information about the llvm-commits
mailing list