[llvm] [PHIElim][AMDGPU]: shrink source subranges to lane-specific uses (PR #228073)

Alan Li via llvm-commits llvm-commits at lists.llvm.org
Thu Oct 1 06:20:23 PDT 2026


https://github.com/lialan created https://github.com/llvm/llvm-project/pull/228073

This patch corrects a lane-specific trimming flaw in the subrange-preservation implementation introduced by D158144 and subsequently relanded as PR #69429.

## Symptoms
While analyzing issue #227456 we discovered the bug where phi elimination trims every source subrange using the whole register's last-use position.  This can trigger a `LiveRange::removeSegment` assertion or leave invalid subrange endpoints which will be reported by machine verifier.

The issue indicated that when compiling an HSTU kernel with teh default gfx950 -O3 pipeline, where LiveIntervals are computed before PHI elimination.  We found that recent PR #227252 made the move to expose the issue.

To be specific, here is the MIR example. Two PHIs read the 32-bit halves of one 64-bit source:
```
%lo = PHI %src.sub0, %bb.1
%hi = PHI %src.sub1, %bb.1
```
PHI elimination inserts separate source copies:
```
%lo_in = COPY %src.sub0
%hi_in = COPY %src.sub1
```

The old code ends both source subranges at the second copy. The low-half subrange should end at the first copy; the second copy only reads the high half. Consequently, `-verify-machineinstrs` reports: `Instruction ending live segment doesn't read the register`.

## Fix
* Record each PHI source interval that has subranges, and
* after all PHIs in the function are lowered recompute its subranges with `LiveIntervals::shrinkToUses`
* remove any empty subranges.

## Validation
Without the fix, patch-included regression tests `phi_subrange_absent` and `phi_subrange_multiple_successors` would hit the assertion and `phi_subrange_separate_uses` fails the verifier.

## Previous discussions
* [D156872 — Verify LiveIntervals for PHIs](https://reviews.llvm.org/D156872) : this was probably the first time it was brought up.
* [D158144](https://reviews.llvm.org/D158144): added subrange updates during PHI elimination, but assumes every lane can share the whole register’s endpoint.

>From 9155852b1a6eaf21f49f4530a644c62aaa4f0cca Mon Sep 17 00:00:00 2001
From: Alan Li <me at alanli.org>
Date: Wed, 30 Sep 2026 18:28:21 -0500
Subject: [PATCH 1/3] PHIElimination: shrink source subranges to lane-specific
 uses

The whole-register last use is not a valid trimming endpoint for every
lane. Defer shrinking affected source subranges until all PHIs have been
deleted, then recompute them with LiveIntervals::shrinkToUses. Preserve
existing full-register trimming.

Add five AMDGPU MIR cases for absent lanes, earlier-dead lanes, fully live
lanes, separate lane uses, and PHIs in multiple successors.
---
 llvm/lib/CodeGen/PHIElimination.cpp           |  24 ++-
 ...phi-elimination-subrange-liveintervals.mir | 171 ++++++++++++++++++
 2 files changed, 191 insertions(+), 4 deletions(-)
 create mode 100644 llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir

diff --git a/llvm/lib/CodeGen/PHIElimination.cpp b/llvm/lib/CodeGen/PHIElimination.cpp
index de69ca8d918986a..b61e7d9ab280515 100644
--- a/llvm/lib/CodeGen/PHIElimination.cpp
+++ b/llvm/lib/CodeGen/PHIElimination.cpp
@@ -15,6 +15,7 @@
 #include "llvm/CodeGen/PHIElimination.h"
 #include "PHIEliminationUtils.h"
 #include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/SetVector.h"
 #include "llvm/ADT/SmallPtrSet.h"
 #include "llvm/ADT/Statistic.h"
 #include "llvm/Analysis/LoopInfo.h"
@@ -114,6 +115,9 @@ class PHIEliminationImpl {
   // Count the number of non-undef PHI uses of each register in each BB.
   VRegPHIUse VRegPHIUseCount;
 
+  // Source subranges must be shrunk to their own uses after all PHIs are gone.
+  SmallSetVector<LiveInterval *, 8> PHISrcIntervalsToShrink;
+
   // Defs of PHI sources which are implicit_def.
   SmallPtrSet<MachineInstr *, 4> ImpDefs;
 
@@ -306,6 +310,20 @@ bool PHIEliminationImpl::run(MachineFunction &MF) {
   }
 
   LoweredPHIs.clear();
+
+  // Different lanes may be used by different PHI source copies, or may already
+  // be dead in a predecessor. The main range's last use is therefore not a
+  // valid endpoint for every subrange. Wait until all PHIs have been removed
+  // before shrinking subranges to their remaining lane-specific uses.
+  if (LIS) {
+    for (LiveInterval *LI : PHISrcIntervalsToShrink) {
+      for (auto &SR : LI->subranges())
+        LIS->shrinkToUses(SR, LI->reg());
+      LI->removeEmptySubRanges();
+    }
+  }
+  PHISrcIntervalsToShrink.clear();
+
   ImpDefs.clear();
   VRegPHIUseCount.clear();
 
@@ -723,6 +741,8 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
       if (!SrcUndef &&
           !VRegPHIUseCount[BBVRegPair(opBlock.getNumber(), SrcReg)]) {
         LiveInterval &SrcLI = LIS->getInterval(SrcReg);
+        if (SrcLI.hasSubRanges())
+          PHISrcIntervalsToShrink.insert(&SrcLI);
 
         bool isLiveOut = false;
         for (MachineBasicBlock *Succ : opBlock.successors()) {
@@ -768,10 +788,6 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
           SlotIndex LastUseIndex = LIS->getInstructionIndex(*KillInst);
           SrcLI.removeSegment(LastUseIndex.getRegSlot(),
                               LIS->getMBBEndIdx(&opBlock));
-          for (auto &SR : SrcLI.subranges()) {
-            SR.removeSegment(LastUseIndex.getRegSlot(),
-                             LIS->getMBBEndIdx(&opBlock));
-          }
         }
       }
     }
diff --git a/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
new file mode 100644
index 000000000000000..dfdd4f9412320d5
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
@@ -0,0 +1,171 @@
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -verify-machineinstrs \
+# RUN:   -run-pass=liveintervals,phi-node-elimination -o - %s | FileCheck %s
+#
+# The high lane is dead before the PHI predecessor, but its segment belongs to
+# a block appearing later in layout order. It must not be trimmed using the
+# low lane's last-use slot.
+# CHECK-LABEL: name: phi_subrange_absent
+# CHECK: body:
+# CHECK-NOT: = PHI
+# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
+# CHECK: S_ENDPGM
+# CHECK-NOT: = PHI
+---
+name: phi_subrange_absent
+tracksRegLiveness: true
+body: |
+  bb.0:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %2:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %2
+    S_ENDPGM 0
+
+  bb.3:
+    successors: %bb.1
+    %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+    %1:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %1
+    S_BRANCH %bb.1
+...
+
+# The high lane's segment ends in an earlier layout block. It needs no trim.
+# CHECK-LABEL: name: phi_subrange_dead_earlier
+# CHECK: body:
+# CHECK-NOT: = PHI
+# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
+# CHECK: S_ENDPGM
+---
+name: phi_subrange_dead_earlier
+tracksRegLiveness: true
+body: |
+  bb.0:
+    successors: %bb.1
+    %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+    %1:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %2:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %2
+    S_ENDPGM 0
+...
+
+# Both lanes and the main live range still require normal PHI source trimming.
+# CHECK-LABEL: name: phi_subrange_full_live
+# CHECK: body:
+# CHECK-NOT: = PHI
+# CHECK: %{{[0-9]+}}:vreg_64 = COPY %{{[0-9]+}}
+# CHECK: S_ENDPGM
+---
+name: phi_subrange_full_live
+tracksRegLiveness: true
+body: |
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vreg_64 = PHI %0, %bb.1
+    $vgpr0_vgpr1 = COPY %3
+    S_ENDPGM 0
+...
+
+# Different PHI source copies read different lanes. Each lane must end at its
+# own copy, not at the main range's last use in the predecessor.
+# CHECK-LABEL: name: phi_subrange_separate_uses
+# CHECK: body:
+# CHECK-NOT: = PHI
+# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
+# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub1
+# CHECK: S_ENDPGM
+---
+name: phi_subrange_separate_uses
+tracksRegLiveness: true
+body: |
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    %4:vgpr_32 = PHI %0.sub1, %bb.1
+    $vgpr0 = COPY %3
+    $vgpr1 = COPY %4
+    S_ENDPGM 0
+...
+
+# PHIs in two different successor blocks share the source register. Shrinking
+# must wait for the whole function's PHIs to be removed.
+# CHECK-LABEL: name: phi_subrange_multiple_successors
+# CHECK: body:
+# CHECK-NOT: = PHI
+# CHECK: COPY %{{[0-9]+}}.sub0
+# CHECK: COPY %{{[0-9]+}}.sub1
+# CHECK: S_ENDPGM
+---
+name: phi_subrange_multiple_successors
+tracksRegLiveness: true
+body: |
+  bb.0:
+    successors: %bb.1, %bb.2
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    %5:sreg_32 = S_MOV_B32 0
+    S_CMP_EQ_U32 %5, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.2, implicit $scc
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.4
+    S_BRANCH %bb.4
+
+  bb.3:
+    successors: %bb.5
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %3
+    S_BRANCH %bb.5
+
+  bb.4:
+    successors: %bb.5
+    %4:vgpr_32 = PHI %0.sub1, %bb.2
+    $vgpr1 = COPY %4
+    S_BRANCH %bb.5
+
+  bb.5:
+    S_ENDPGM 0
+...

>From 92c021cdffafa66da56e271f9e9c970a4813c8f6 Mon Sep 17 00:00:00 2001
From: Alan Li <me at alanli.org>
Date: Thu, 1 Oct 2026 13:16:36 +0000
Subject: [PATCH 2/3] PHIElimination: strengthen subrange LiveIntervals test

Run the test under the new pass manager as well, matching the sibling
PHI elimination LiveIntervals tests, and autogenerate the checks with
update_mir_test_checks.py so the inserted copies and block structure
are pinned down instead of only being spot-checked.

Add two more cases that fail before the subrange shrinking fix:
- A lane that is live into another successor keeps the main range live
  out, so the edge to the PHI block is split and the split block
  inherits every lane. The lane not read by the PHI copy must be shrunk
  out of the split block rather than trimmed at that copy.
- A full-register PHI and a lane PHI reading the same source from the
  same predecessor, where the main range's kill is the lane copy but the
  other lane ends at the earlier full copy.
---
 ...phi-elimination-subrange-liveintervals.mir | 271 ++++++++++++++++--
 1 file changed, 243 insertions(+), 28 deletions(-)

diff --git a/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
index dfdd4f9412320d5..ce3d989bca7d94d 100644
--- a/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
+++ b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
@@ -1,19 +1,42 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
 # RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -verify-machineinstrs \
 # RUN:   -run-pass=liveintervals,phi-node-elimination -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 \
+# RUN:   --passes='require<live-intervals>,phi-node-elimination' -verify-each -o - %s | FileCheck %s
 #
 # The high lane is dead before the PHI predecessor, but its segment belongs to
 # a block appearing later in layout order. It must not be trimmed using the
 # low lane's last-use slot.
-# CHECK-LABEL: name: phi_subrange_absent
-# CHECK: body:
-# CHECK-NOT: = PHI
-# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
-# CHECK: S_ENDPGM
-# CHECK-NOT: = PHI
 ---
 name: phi_subrange_absent
 tracksRegLiveness: true
 body: |
+  ; CHECK-LABEL: name: phi_subrange_absent
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY %1.sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_BRANCH %bb.1
   bb.0:
     successors: %bb.3
     S_BRANCH %bb.3
@@ -38,15 +61,31 @@ body: |
 ...
 
 # The high lane's segment ends in an earlier layout block. It needs no trim.
-# CHECK-LABEL: name: phi_subrange_dead_earlier
-# CHECK: body:
-# CHECK-NOT: = PHI
-# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
-# CHECK: S_ENDPGM
 ---
 name: phi_subrange_dead_earlier
 tracksRegLiveness: true
 body: |
+  ; CHECK-LABEL: name: phi_subrange_dead_earlier
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY]]
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_ENDPGM 0
   bb.0:
     successors: %bb.1
     %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
@@ -67,15 +106,29 @@ body: |
 ...
 
 # Both lanes and the main live range still require normal PHI source trimming.
-# CHECK-LABEL: name: phi_subrange_full_live
-# CHECK: body:
-# CHECK-NOT: = PHI
-# CHECK: %{{[0-9]+}}:vreg_64 = COPY %{{[0-9]+}}
-# CHECK: S_ENDPGM
 ---
 name: phi_subrange_full_live
 tracksRegLiveness: true
 body: |
+  ; CHECK-LABEL: name: phi_subrange_full_live
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0_vgpr1 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
   bb.0:
     successors: %bb.1
     %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
@@ -95,16 +148,32 @@ body: |
 
 # Different PHI source copies read different lanes. Each lane must end at its
 # own copy, not at the main range's last use in the predecessor.
-# CHECK-LABEL: name: phi_subrange_separate_uses
-# CHECK: body:
-# CHECK-NOT: = PHI
-# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub0
-# CHECK: %{{[0-9]+}}:vgpr_32 = COPY %{{[0-9]+}}.sub1
-# CHECK: S_ENDPGM
 ---
 name: phi_subrange_separate_uses
 tracksRegLiveness: true
 body: |
+  ; CHECK-LABEL: name: phi_subrange_separate_uses
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_ENDPGM 0
   bb.0:
     successors: %bb.1
     %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
@@ -126,16 +195,50 @@ body: |
 
 # PHIs in two different successor blocks share the source register. Shrinking
 # must wait for the whole function's PHIs to be removed.
-# CHECK-LABEL: name: phi_subrange_multiple_successors
-# CHECK: body:
-# CHECK-NOT: = PHI
-# CHECK: COPY %{{[0-9]+}}.sub0
-# CHECK: COPY %{{[0-9]+}}.sub1
-# CHECK: S_ENDPGM
 ---
 name: phi_subrange_multiple_successors
 tracksRegLiveness: true
 body: |
+  ; CHECK-LABEL: name: phi_subrange_multiple_successors
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x40000000), %bb.2(0x40000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+  ; CHECK-NEXT:   S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.2, implicit $scc
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   successors: %bb.4(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.4
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   successors: %bb.5(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_BRANCH %bb.5
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.4:
+  ; CHECK-NEXT:   successors: %bb.5(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_BRANCH %bb.5
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.5:
+  ; CHECK-NEXT:   S_ENDPGM 0
   bb.0:
     successors: %bb.1, %bb.2
     %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
@@ -169,3 +272,115 @@ body: |
   bb.5:
     S_ENDPGM 0
 ...
+
+# Only sub1 is live into the other successor, so the main range is live-out and
+# the edge to the PHI block is split. The split block inherits every lane, but
+# the PHI copy placed there reads only sub0. sub1 must be shrunk out of the
+# split block rather than trimmed at the sub0 copy.
+---
+name: phi_subrange_lane_live_out_split_edge
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_lane_live_out_split_edge
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x40000000), %bb.4(0x40000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+  ; CHECK-NEXT:   S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.2, implicit $scc
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.4:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2, %bb.3
+    %5:sreg_32 = S_MOV_B32 0
+    S_CMP_EQ_U32 %5, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.2, implicit $scc
+    S_BRANCH %bb.3
+
+  bb.2:
+    %8:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %8
+    S_ENDPGM 0
+
+  bb.3:
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %3
+    S_ENDPGM 0
+...
+
+# One PHI reads the whole register and another reads only sub1 from the same
+# predecessor. The sub1 copy is the last use of the main range, but sub0 must
+# end at the earlier full-register copy.
+---
+name: phi_subrange_full_and_lane
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_full_and_lane
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0_vgpr1 = COPY [[COPY2]]
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr2 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vreg_64 = PHI %0, %bb.1
+    %4:vgpr_32 = PHI %0.sub1, %bb.1
+    $vgpr0_vgpr1 = COPY %3
+    $vgpr2 = COPY %4
+    S_ENDPGM 0
+...

>From 884df8ea2ec25b253f5c86ff241fb38b79be9b93 Mon Sep 17 00:00:00 2001
From: Alan Li <me at alanli.org>
Date: Thu, 1 Oct 2026 13:16:37 +0000
Subject: [PATCH 3/3] PHIElimination: track PHI source registers to shrink by
 Register

Store the registers whose subranges need shrinking instead of raw
LiveInterval pointers, and look the interval up when shrinking. Drop the
redundant LIS null check around the final loop: the set is only
populated on the LiveIntervals path, so the loop is empty otherwise.

No functional change.
---
 llvm/lib/CodeGen/PHIElimination.cpp | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/llvm/lib/CodeGen/PHIElimination.cpp b/llvm/lib/CodeGen/PHIElimination.cpp
index b61e7d9ab280515..0aa8a94021b665e 100644
--- a/llvm/lib/CodeGen/PHIElimination.cpp
+++ b/llvm/lib/CodeGen/PHIElimination.cpp
@@ -115,8 +115,9 @@ class PHIEliminationImpl {
   // Count the number of non-undef PHI uses of each register in each BB.
   VRegPHIUse VRegPHIUseCount;
 
-  // Source subranges must be shrunk to their own uses after all PHIs are gone.
-  SmallSetVector<LiveInterval *, 8> PHISrcIntervalsToShrink;
+  // PHI source registers whose subranges must be shrunk to their own uses once
+  // all PHIs are gone.
+  SmallSetVector<Register, 8> PHISrcRegsToShrink;
 
   // Defs of PHI sources which are implicit_def.
   SmallPtrSet<MachineInstr *, 4> ImpDefs;
@@ -315,14 +316,13 @@ bool PHIEliminationImpl::run(MachineFunction &MF) {
   // be dead in a predecessor. The main range's last use is therefore not a
   // valid endpoint for every subrange. Wait until all PHIs have been removed
   // before shrinking subranges to their remaining lane-specific uses.
-  if (LIS) {
-    for (LiveInterval *LI : PHISrcIntervalsToShrink) {
-      for (auto &SR : LI->subranges())
-        LIS->shrinkToUses(SR, LI->reg());
-      LI->removeEmptySubRanges();
-    }
+  for (Register Reg : PHISrcRegsToShrink) {
+    LiveInterval &LI = LIS->getInterval(Reg);
+    for (LiveInterval::SubRange &SR : LI.subranges())
+      LIS->shrinkToUses(SR, Reg);
+    LI.removeEmptySubRanges();
   }
-  PHISrcIntervalsToShrink.clear();
+  PHISrcRegsToShrink.clear();
 
   ImpDefs.clear();
   VRegPHIUseCount.clear();
@@ -742,7 +742,7 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
           !VRegPHIUseCount[BBVRegPair(opBlock.getNumber(), SrcReg)]) {
         LiveInterval &SrcLI = LIS->getInterval(SrcReg);
         if (SrcLI.hasSubRanges())
-          PHISrcIntervalsToShrink.insert(&SrcLI);
+          PHISrcRegsToShrink.insert(SrcReg);
 
         bool isLiveOut = false;
         for (MachineBasicBlock *Succ : opBlock.successors()) {



More information about the llvm-commits mailing list