[llvm] [InlineSpiller][AMDGPU] Skip undef lanes when spilling a partial-def super-register (PR #226689)

via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 26 05:57:37 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu

@llvm/pr-subscribers-llvm-regalloc

Author: xgxanq

<details>
<summary>Changes</summary>

When the greedy allocator live-range-splits a wide register (e.g. an sgpr_128 buffer descriptor) on a single sub-register, it produces sibling vregs each defined on one sub-lane and undef elsewhere. InlineSpiller shares one stack slot among all descendants of the Original value, so a full-width store of such a partially-defined sibling writes its undef lanes over a live sibling value in the same slot and corrupts it. On gfx950 this was observed (via asm data-flow analysis) as an in-loop buffer-descriptor lane being clobbered by an unrelated in-loop spill sharing the slot.

The hazard is target-independent (InlineSpiller's shared-slot logic is generic; SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs). Fix it in two places:

1. hoistSpillInsideBB: bail out unless every sub-range of the hoisted value is live at the hoist point, mirroring the cross-BB guard in isSpillCandBB (#<!-- -->177703) that the in-BB path was missing.

2. insertSpill: for a real (partial) spill, compute the defined-lane mask and hand it to the target via a new TargetInstrInfo::setSpillDefinedLaneMask hook (default no-op). AMDGPU records it in the SI_SPILL_S*_SAVE $lanemask operand; spillSGPR then writelanes only the defined dwords, leaving the slot's other lanes (a sibling value) intact. A -1 mask preserves the original full-width behavior.

Existing SGPR-spill tests are regenerated to carry the new -1 $lanemask operand. Adds inline-spiller-partial-subreg-shared-slot.ll as a gfx950 end-to-end regression check.

---

Patch is 95.32 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/226689.diff


40 Files Affected:

- (modified) llvm/include/llvm/CodeGen/TargetInstrInfo.h (+6) 
- (modified) llvm/lib/CodeGen/InlineSpiller.cpp (+32-2) 
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.cpp (+33-4) 
- (modified) llvm/lib/Target/AMDGPU/SIInstrInfo.h (+3) 
- (modified) llvm/lib/Target/AMDGPU/SIInstructions.td (+8-2) 
- (modified) llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp (+47-4) 
- (modified) llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir (+1-1) 
- (added) llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll (+257) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-dead-frame-in-dbg-value.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value-list.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-fi-skip-processing-stack-arg-dbg-value.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-overlap-wwm-reserve.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-partially-undef.mir (+4-14) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber-unhandled.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-to-vmem-scc-clobber.mir (+6-6) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-vmem-large-frame.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill-wrong-stack-id.mir (+5-5) 
- (modified) llvm/test/CodeGen/AMDGPU/sgpr-spill.mir (+11-11) 
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-cycle-header.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-body.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-initial-insert-in-latch.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/si-lower-sgpr-spills-multi-entry-cycle.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/spill-reg-tuple-super-reg-use.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/spill-scavenge-offset.ll (+7-7) 
- (modified) llvm/test/CodeGen/AMDGPU/spill-sgpr-to-virtual-vgpr.mir (+6-6) 
- (modified) llvm/test/CodeGen/AMDGPU/spill-special-sgpr.mir (+2-2) 
- (modified) llvm/test/CodeGen/AMDGPU/spill192.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/spill224.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/spill288.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/spill320.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/spill352.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/spill384.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/stack-slot-color-sgpr-vgpr-spills.mir (+1-1) 
- (modified) llvm/test/CodeGen/AMDGPU/undefined-physreg-sgpr-spill.mir (+3-3) 
- (modified) llvm/test/CodeGen/AMDGPU/wwm-regalloc-partial-pool.mir (+3-3) 
- (modified) llvm/test/CodeGen/AMDGPU/wwm-regalloc-preallocation-guard.mir (+9-9) 
- (modified) llvm/test/CodeGen/AMDGPU/wwm-spill-superclass-pseudo.mir (+1-1) 
- (modified) llvm/test/CodeGen/Hexagon/regalloc-bad-undef.mir (+1-1) 


``````````diff
diff --git a/llvm/include/llvm/CodeGen/TargetInstrInfo.h b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
index b013511d33313..022c39d0ca92d 100644
--- a/llvm/include/llvm/CodeGen/TargetInstrInfo.h
+++ b/llvm/include/llvm/CodeGen/TargetInstrInfo.h
@@ -1226,6 +1226,12 @@ class LLVM_ABI TargetInstrInfo : public MCInstrInfo {
                      "TargetInstrInfo::storeRegToStackSlot!");
   }
 
+  /// Tell the target which lanes of spill \p SpillMI hold a real value
+  /// (\p DefinedLanes); it can skip the undef ones when lowering, so a shared
+  /// slot's neighbor isn't stomped. Default: no-op.
+  virtual void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+                                       LaneBitmask DefinedLanes) const {}
+
   /// Load the specified register of the given register class from the specified
   /// stack frame index. The load instruction is to be added to the given
   /// machine basic block before the specified machine instruction. If \p
diff --git a/llvm/lib/CodeGen/InlineSpiller.cpp b/llvm/lib/CodeGen/InlineSpiller.cpp
index f3682a2e24808..a7a960a0cf441 100644
--- a/llvm/lib/CodeGen/InlineSpiller.cpp
+++ b/llvm/lib/CodeGen/InlineSpiller.cpp
@@ -454,6 +454,17 @@ bool InlineSpiller::hoistSpillInsideBB(LiveInterval &SpillLI,
   if (DefMBB != CopyMI.getParent() || !SrcQ.isKill())
     return false;
 
+  // Every sub-range of the hoisted value must be live at the hoist point. The
+  // store below is full-width, so a dead sub-range would write a stale sibling
+  // value into the shared spill slot and clobber it. This mirrors the sub-range
+  // liveness guard on the cross-BB path in isSpillCandBB (#177703), which the
+  // in-BB path here was missing.
+  if (SrcLI.hasSubRanges() &&
+      !all_of(SrcLI.subranges(), [&](const LiveInterval::SubRange &SR) {
+        return SR.getVNInfoAt(Idx) != nullptr;
+      }))
+    return false;
+
   MachineBasicBlock *MBB = DefMBB;
   MachineBasicBlock::iterator MII;
   if (SrcVNI->isPHIDef())
@@ -1285,10 +1296,29 @@ void InlineSpiller::insertSpill(Register NewVReg, bool isKill,
   MachineBasicBlock::iterator SpillBefore = std::next(MI);
   bool IsRealSpill = isRealSpill(*MI);
 
-  if (IsRealSpill)
+  if (IsRealSpill) {
     TII.storeRegToStackSlot(MBB, SpillBefore, NewVReg, isKill, StackSlot,
                             MRI.getRegClass(NewVReg), Register());
-  else
+
+    MachineInstr &SpillMI = *std::next(MI);
+
+    // Compute the lanes of NewVReg defined by MI. For a partial def (undef
+    // lanes remain in the full-width store above), pass the defined-lane mask
+    // so the target skips them and avoids clobbering a sibling in the shared
+    // stack slot. A fully-undef def never reaches here (filtered by
+    // isRealSpill).
+    LaneBitmask FullMask = MRI.getMaxLaneMaskForVReg(NewVReg);
+    LaneBitmask DefinedLanes;
+    for (const MachineOperand &MO : MI->all_defs()) {
+      if (MO.getReg() != NewVReg)
+        continue;
+      DefinedLanes |= MO.getSubReg()
+                          ? TRI.getSubRegIndexLaneMask(MO.getSubReg())
+                          : FullMask;
+    }
+    if (DefinedLanes != FullMask)
+      TII.setSpillDefinedLaneMask(SpillMI, DefinedLanes);
+  } else
     // Don't spill undef value.
     // Anything works for undef, in particular keeping the memory
     // uninitialized is a viable option and it saves code size and
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
index 0346ebc254361..28a06d6921aa4 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.cpp
@@ -1800,10 +1800,11 @@ void SIInstrInfo::storeRegToStackSlotImpl(
     }
 
     BuildMI(MBB, MI, DL, OpDesc)
-      .addReg(SrcReg, getKillRegState(isKill)) // data
-      .addFrameIndex(FrameIndex)               // addr
-      .addMemOperand(MMO)
-      .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
+        .addReg(SrcReg, getKillRegState(isKill)) // data
+        .addFrameIndex(FrameIndex)               // addr
+        .addImm(-1) // lanemask (all lanes; refined by spiller)
+        .addMemOperand(MMO)
+        .addReg(MFI->getStackPtrOffsetReg(), RegState::Implicit);
 
     return;
   }
@@ -1837,6 +1838,34 @@ void SIInstrInfo::storeRegToStackSlotCFI(MachineBasicBlock &MBB,
                           MachineInstr::NoFlags, true);
 }
 
+void SIInstrInfo::setSpillDefinedLaneMask(MachineInstr &SpillMI,
+                                          LaneBitmask DefinedLanes) const {
+  // Only SGPR spill saves carry a $lanemask operand (see SI_SPILL_SGPR).
+  int Idx =
+      AMDGPU::getNamedOperandIdx(SpillMI.getOpcode(), AMDGPU::OpName::lanemask);
+  if (Idx == -1)
+    return;
+
+  // Build the per-dword mask consumed by spillSGPR: bit i is set when dword i
+  // of the spilled super-register intersects a defined lane. spillSGPR lowers
+  // the spill dword by dword and consults this mask to skip the undef ones.
+  const MachineRegisterInfo &MRI = SpillMI.getMF()->getRegInfo();
+  Register Data = SpillMI.getOperand(0).getReg();
+  const TargetRegisterClass *RC = MRI.getRegClass(Data);
+  unsigned NumDwords = RI.getRegSizeInBits(*RC) / 32;
+  // DwordMask is a u32; the widest SGPR tuple (SReg_1024) has 32 dwords, so
+  // this holds and keeps the 1u << i shift below well-defined.
+  assert(NumDwords <= 32 && "SGPR spill wider than DwordMask can represent");
+  unsigned DwordMask = 0;
+  for (unsigned i = 0; i != NumDwords; ++i) {
+    LaneBitmask DwordLanes =
+        RI.getSubRegIndexLaneMask(RI.getSubRegFromChannel(i));
+    if ((DefinedLanes & DwordLanes).any())
+      DwordMask |= 1u << i;
+  }
+  SpillMI.getOperand(Idx).setImm(DwordMask);
+}
+
 static unsigned getSGPRSpillRestoreOpcode(unsigned Size) {
   switch (Size) {
   case 4:
diff --git a/llvm/lib/Target/AMDGPU/SIInstrInfo.h b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
index 48eba6b2a567d..42afa3cd8c60e 100644
--- a/llvm/lib/Target/AMDGPU/SIInstrInfo.h
+++ b/llvm/lib/Target/AMDGPU/SIInstrInfo.h
@@ -356,6 +356,9 @@ class SIInstrInfo final : public AMDGPUGenInstrInfo {
       bool isKill, int FrameIndex, const TargetRegisterClass *RC, Register VReg,
       MachineInstr::MIFlag Flags = MachineInstr::NoFlags) const override;
 
+  void setSpillDefinedLaneMask(MachineInstr &SpillMI,
+                               LaneBitmask DefinedLanes) const override;
+
   void loadRegFromStackSlot(
       MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, Register DestReg,
       int FrameIndex, const TargetRegisterClass *RC, Register VReg,
diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 3cf20ffb1fcf4..3faae31302689 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -1178,14 +1178,20 @@ def V_INDIRECT_REG_READ_GPR_IDX_B32_V32 : V_INDIRECT_REG_READ_GPR_IDX_pseudo<VRe
 
 multiclass SI_SPILL_SGPR <RegisterClass sgpr_class> {
   let UseNamedOperandTable = 1, Spill = 1, SALU = 1, Uses = [EXEC] in {
+    // $lanemask: per-dword bitmask of which dwords of the spilled super-register
+    // are defined. Undef dwords must not be written into the (possibly shared)
+    // stack slot. -1 (all dwords defined) is the ordinary full-width spill; a
+    // partial mask is set by the spiller for partially-defined values. i32imm's
+    // 32 bits cover the widest SGPR tuple SReg_1024 (32 dwords).
     def _SAVE : PseudoInstSI <
       (outs),
-      (ins sgpr_class:$data, i32imm:$addr)> {
+      (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
       let mayStore = 1;
       let mayLoad = 0;
     }
 
-    def _CFI_SAVE : PseudoInstSI<(outs), (ins sgpr_class:$data, i32imm:$addr)> {
+    def _CFI_SAVE : PseudoInstSI<(outs),
+      (ins sgpr_class:$data, i32imm:$addr, i32imm:$lanemask)> {
       let mayStore = 1;
       let mayLoad = 0;
     }
diff --git a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
index 1aa423e1da647..5a8dd05c239ab 100644
--- a/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/SIRegisterInfo.cpp
@@ -2180,6 +2180,18 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
   if (OnlyToVGPR && !SpillToVGPR)
     return false;
 
+  // A partial-def spill (see InlineSpiller / setSpillDefinedLaneMask) records
+  // which dwords are real values in its $lanemask operand; a cleared bit means
+  // that dword is undef and must NOT be written to the (possibly shared) stack
+  // slot, so it does not clobber a sibling value living in the same slot. A
+  // mask of -1 (default) means "all lanes defined" = original behavior.
+  const MachineOperand *LaneMask =
+      SB.TII.getNamedOperand(*MI, AMDGPU::OpName::lanemask);
+  int64_t DefinedDwordMask = LaneMask ? LaneMask->getImm() : -1;
+  auto DwordDefined = [&](unsigned i) {
+    return (DefinedDwordMask & (int64_t(1) << i)) != 0;
+  };
+
   const SIFrameLowering *TFL = ST.getFrameLowering();
 
   assert(SpillToVGPR || (SB.SuperReg != SB.MFI.getStackPtrOffsetReg() &&
@@ -2195,15 +2207,29 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
            "Num of SGPRs spilled should be less than or equal to num of "
            "the VGPR lanes.");
 
+    // With a partial-def mask, find the first/last dword that is actually
+    // defined, so the ImplicitDefine of the super-register and the kill flag
+    // can be re-anchored onto emitted writelanes (undef dwords are skipped).
+    unsigned FirstDefined = 0, LastDefined = SB.NumSubRegs - 1;
+    while (FirstDefined < SB.NumSubRegs && !DwordDefined(FirstDefined))
+      ++FirstDefined;
+    while (LastDefined > 0 && !DwordDefined(LastDefined))
+      --LastDefined;
+
     for (unsigned i = 0, e = SB.NumSubRegs; i < e; ++i) {
+      // Skip undef dwords: not writing them to a (possibly shared) stack slot
+      // preserves a sibling value living in the same slot.
+      if (!DwordDefined(i))
+        continue;
+
       Register SubReg =
           SB.NumSubRegs == 1
               ? SB.SuperReg
               : Register(getSubReg(SB.SuperReg, SB.SplitParts[i]));
       SpilledReg Spill = VGPRSpills[i];
 
-      bool IsFirstSubreg = i == 0;
-      bool IsLastSubreg = i == SB.NumSubRegs - 1;
+      bool IsFirstSubreg = i == FirstDefined;
+      bool IsLastSubreg = i == LastDefined;
       bool UseKill = SB.IsKill && IsLastSubreg;
 
 
@@ -2259,13 +2285,30 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
     // Per VGPR helper data
     auto PVD = SB.getPerVGPRData();
 
+    // For a partial-def spill (undef dwords, see $lanemask), the whole TmpVGPR
+    // is stored at once, so first load the slot's current contents; the skipped
+    // dwords then keep the sibling value already there instead of writing
+    // undef.
+    bool IsPartialDef = SB.NumSubRegs > 1 && DefinedDwordMask != -1;
+
     for (unsigned Offset = 0; Offset < PVD.NumVGPRs; ++Offset) {
       RegState TmpVGPRFlags = RegState::Undef;
 
+      if (IsPartialDef) {
+        // Seed TmpVGPR with the slot's current contents, so the dwords we skip
+        // (undef in this partial def) keep the sibling value already living in
+        // the slot when the whole TmpVGPR is stored back below.
+        SB.readWriteTmpVGPR(Offset, /*IsLoad*/ true);
+        TmpVGPRFlags = {};
+      }
+
       // Write sub registers into the VGPR
       for (unsigned i = Offset * PVD.PerVGPR,
                     e = std::min((Offset + 1) * PVD.PerVGPR, SB.NumSubRegs);
            i < e; ++i) {
+        if (IsPartialDef && !DwordDefined(i))
+          continue; // Undef dword: keep the sibling value seeded from the slot.
+
         Register SubReg =
             SB.NumSubRegs == 1
                 ? SB.SuperReg
@@ -2286,8 +2329,8 @@ bool SIRegisterInfo::spillSGPR(MachineBasicBlock::iterator MI, int Index,
             Indexes->insertMachineInstrInMaps(*WriteLane);
         }
 
-        // There could be undef components of a spilled super register.
-        // TODO: Can we detect this and skip the spill?
+        // Undef components of the spilled super-register are detected via the
+        // $lanemask operand and skipped above (see IsPartialDef).
         if (SB.NumSubRegs > 1) {
           // The last implicit use of the SB.SuperReg carries the "Kill" flag.
           RegState SuperKillState = {};
diff --git a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
index 5c564decd1e7d..a026d13e0cce5 100644
--- a/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
+++ b/llvm/test/CodeGen/AMDGPU/bug-undef-spilled-agpr.mir
@@ -75,7 +75,7 @@ body:             |
     liveins: $sgpr12, $sgpr13, $sgpr14, $sgpr15
 
     %45:vgpr_32 = IMPLICIT_DEF
-    SI_SPILL_S32_SAVE $sgpr15, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
+    SI_SPILL_S32_SAVE $sgpr15, %stack.0, -1, implicit $exec, implicit $sgpr32 :: (store (s32) into %stack.0, addrspace 5)
     %16:vgpr_32 = V_AND_B32_e32 1, %45, implicit $exec
 
   bb.1:
diff --git a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
index 94c22b1aa8664..5c186b9a0470b 100644
--- a/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
+++ b/llvm/test/CodeGen/AMDGPU/greedy-alloc-fail-sgpr1024-spill.mir
@@ -57,7 +57,7 @@ body:             |
   ; CHECK-NEXT:   renamable $sgpr86 = S_MOV_B32 0
   ; CHECK-NEXT:   renamable $sgpr87 = S_MOV_B32 0
   ; CHECK-NEXT:   renamable $sgpr88 = S_MOV_B32 0
-  ; CHECK-NEXT:   SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
+  ; CHECK-NEXT:   SI_SPILL_S1024_SAVE renamable $sgpr68_sgpr69_sgpr70_sgpr71_sgpr72_sgpr73_sgpr74_sgpr75_sgpr76_sgpr77_sgpr78_sgpr79_sgpr80_sgpr81_sgpr82_sgpr83_sgpr84_sgpr85_sgpr86_sgpr87_sgpr88_sgpr89_sgpr90_sgpr91_sgpr92_sgpr93_sgpr94_sgpr95_sgpr96_sgpr97_sgpr98_sgpr99, %stack.0, 65535, implicit $exec, implicit $sgpr32 :: (store (s1024) into %stack.0, align 4, addrspace 5)
   ; CHECK-NEXT:   renamable $sgpr52_sgpr53 = IMPLICIT_DEF
   ; CHECK-NEXT:   ADJCALLSTACKUP 0, 0, implicit-def dead $scc, implicit-def $sgpr32, implicit $sgpr32
   ; CHECK-NEXT:   dead $sgpr30_sgpr31 = SI_CALL renamable $sgpr52_sgpr53, 0, csr_amdgpu, implicit $sgpr0_sgpr1_sgpr2_sgpr3
diff --git a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
index 63ee4473a56e5..43669338dd308 100644
--- a/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
+++ b/llvm/test/CodeGen/AMDGPU/infloop-subrange-spill-inspect-subrange.mir
@@ -55,7 +55,7 @@ body:             |
   ; CHECK-NEXT:   liveins: $sgpr24_sgpr25_sgpr26_sgpr27:0x000000000000000F, $sgpr4_sgpr5_sgpr6_sgpr7_sgpr8_sgpr9_sgpr10_sgpr11_sgpr12_sgpr13_sgpr14_sgpr15_sgpr16_sgpr17_sgpr18_sgpr19:0x000000000000FFFF, $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51:0x000000000000FFFF
   ; CHECK-NEXT: {{  $}}
   ; CHECK-NEXT:   renamable $sgpr12 = IMPLICIT_DEF
-  ; CHECK-NEXT:   SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
+  ; CHECK-NEXT:   SI_SPILL_S512_SAVE renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51, %stack.0, 255, implicit $exec, implicit $sgpr32 :: (store (s512) into %stack.0, align 4, addrspace 5)
   ; CHECK-NEXT:   renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = IMPLICIT_DEF
   ; CHECK-NEXT:   dead undef [[IMAGE_SAMPLE_LZ_V1_V2_2:%[0-9]+]].sub0:vreg_96 = IMAGE_SAMPLE_LZ_V1_V2 undef [[DEF2]], killed renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43, renamable $sgpr12_sgpr13_sgpr14_sgpr15, 1, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s32), addrspace 8)
   ; CHECK-NEXT:   renamable $sgpr36_sgpr37_sgpr38_sgpr39_sgpr40_sgpr41_sgpr42_sgpr43_sgpr44_sgpr45_sgpr46_sgpr47_sgpr48_sgpr49_sgpr50_sgpr51 = SI_SPILL_S512_RESTORE %stack.0, implicit $exec, implicit $sgpr32 :: (load (s512) from %stack.0, align 4, addrspace 5)
diff --git a/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
new file mode 100644
index 0000000000000..bbff66888c652
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/inline-spiller-partial-subreg-shared-slot.ll
@@ -0,0 +1,257 @@
+; RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -stop-after=greedy %s -o - | FileCheck %s
+;
+; Regression test for the InlineSpiller shared-stack-slot partial-subregister
+; hazard.
+;
+; The hazard is target-independent: InlineSpiller's shared-slot logic is generic,
+; and SGPR-spill-to-VGPR-lane exists on every AMDGPU that spills SGPRs. When the
+; greedy allocator live-range-splits an sgpr_128 descriptor on one sub-register,
+; it produces sibling vregs each defined on a single sub-lane
+; (`undef %V.sub0:sgpr_128 = ...`), undef elsewhere. Since InlineSpiller shares
+; one stack slot among descendants of the Original value, a full-width store of
+; such a sibling writes its undef lanes over a live sibling in the SAME slot. This
+; input is a gfx950 reproducer (there the clobber corrupts an in-loop buffer
+; descriptor -- the kind of corruption that leads to an HSA memory fault), but the
+; fix and this check apply to all AMDGPU.
+;
+; The fix records in the save pseudo's $lanemask operand which dwords the store
+; defines; spillSGPR writelanes only those, leaving the slot's other lanes intact.
+; This pins that the two partial siblings sharing one slot each store with a
+; partial mask (matched as a non-negative immediate: the buggy full-width store
+; used -1, which the {{[0-9]+}} pattern cannot match).
+;
+; Driven from IR through -stop-after=greedy: a .mir + -run-pass=greedy form does
+; not reproduce, since the trigger needs pre-greedy PreRARemat pressure state that
+; MIR serialization drops. The CHECK block is vreg/slot-number independent.
+;
+; First partial sibling (only .sub0 defined), stored with a partial mask:
+; CHECK:      undef [[DESC:%[0-9]+]].sub0:sgpr_128 = S_MOV_B32 0
+; CHECK-NEXT: SI_SPILL_S128_SAVE [[DESC]], [[SLOT:%stack\.[0-9]+]], {{[0-9]+}},
+; Second partial sibling into the SAME slot, also with a partial mask:
+; CHECK:      SI_SPILL_S128_SAVE %{{[0-9]+}}, [[SLOT]], {{[0-9]+}},
+;
+target datalayout = "e-m:e-p:64:64-p1:64:64-p2:32:32-p3:32:32-p4:64:64-p5:32:32-p6:32:32-p7:160:256:256:32-p8:128:128:128:48-p9:192:256:256:32-i64:64-v16:16-v24:32-v32:32-v48:64-v96:128-v192:256-v256:256-v512:512-v1024:1024-v2048:2048-n32:64-S32-A5-G1-ni:7:8:9-p10:32:32-p11:32:32-p12:32:32-p13:32:32-p14:32:32-p15:32:32"
+target triple = "amdgcn-amd-amdhsa"
+
+define amdgpu_kernel void @_attn_fwd_IS_CAUSAL_1_NUM_Q_HEADS_8_NUM_K_HEADS_1_BLOCK_M_256_BLOCK_N_64_BLOCK_DMODEL_64_RETURN_SCORES_1_ENABLE_DROPOUT_1_IS_FP8_0_VARLEN_1_NUM_XCD_8_USE_INT64_STRIDES_1_ENABLE_SINK_0_SLIDING_WINDOW_0(ptr addrspace(1) inreg %0, ptr addrspace(1) inreg %1, i32 %2, i32 %3, i32 %4, i32 %5, i32 %6, <1 x i32> %7, i32 %8, i32 %9, i32 %10, i32 %11, i32 %12, i32 %13, i32 %14, i32 %15, i32 %16, i32 %17, i32 %18, i32 %19, i32 %20, i32 %21, i32 %22, i32 %23, i32 %24, i32 %25, i32 %26, i64 %27, i64 %28, i64 %29, i64 %30, i64 %31, i64 %32, i64 %33, i64 %34, i64 %35, i64 %36, i64 %37, i1 %38, i1 %39, i1 %40, i1 %41, i1 %42, i1 %43, i1 %44, i64 %sext452, i64 %.pn241.in795, i64 %sext496, i64 %.pn233.in799, i64 %sext500, i64 %.pn231.in800, i64 %sext501, i64 %.pn229.in801, i64 %sext503, i64 %.pn223.in804, i64 %sext505, i64 %.pn221.in805, i64 %sext506) #0 {
+..loopexit_crit_edge:
+  %45 = icmp slt i32 %8, 0
+  %46 = icmp slt i32 %6, 0
+  %47 = icmp slt i32 %9, 0
+  %48 = icmp slt i32 %10, 0
+  %49 = icmp slt i32 %12, 0
+  %50 = icmp slt i32 %2, 0
+  br label %51
+
+51:                                               ; preds = ...
[truncated]

``````````

</details>


https://github.com/llvm/llvm-project/pull/226689


More information about the llvm-commits mailing list