[llvm] [PHIElim][AMDGPU]: shrink source subranges to lane-specific uses (PR #228073)

via llvm-commits llvm-commits at lists.llvm.org
Thu Oct 1 06:21:01 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: Alan Li (lialan)

<details>
<summary>Changes</summary>

This patch corrects a lane-specific trimming flaw in the subrange-preservation implementation introduced by D158144 and subsequently relanded as PR #<!-- -->69429.

## Symptoms
While analyzing issue #<!-- -->227456 we discovered the bug where phi elimination trims every source subrange using the whole register's last-use position.  This can trigger a `LiveRange::removeSegment` assertion or leave invalid subrange endpoints which will be reported by machine verifier.

The issue indicated that when compiling an HSTU kernel with teh default gfx950 -O3 pipeline, where LiveIntervals are computed before PHI elimination.  We found that recent PR #<!-- -->227252 made the move to expose the issue.

To be specific, here is the MIR example. Two PHIs read the 32-bit halves of one 64-bit source:
```
%lo = PHI %src.sub0, %bb.1
%hi = PHI %src.sub1, %bb.1
```
PHI elimination inserts separate source copies:
```
%lo_in = COPY %src.sub0
%hi_in = COPY %src.sub1
```

The old code ends both source subranges at the second copy. The low-half subrange should end at the first copy; the second copy only reads the high half. Consequently, `-verify-machineinstrs` reports: `Instruction ending live segment doesn't read the register`.

## Fix
* Record each PHI source interval that has subranges, and
* after all PHIs in the function are lowered recompute its subranges with `LiveIntervals::shrinkToUses`
* remove any empty subranges.

## Validation
Without the fix, patch-included regression tests `phi_subrange_absent` and `phi_subrange_multiple_successors` would hit the assertion and `phi_subrange_separate_uses` fails the verifier.

## Previous discussions
* [D156872 — Verify LiveIntervals for PHIs](https://reviews.llvm.org/D156872) : this was probably the first time it was brought up.
* [D158144](https://reviews.llvm.org/D158144): added subrange updates during PHI elimination, but assumes every lane can share the whole register’s endpoint.

---
Full diff: https://github.com/llvm/llvm-project/pull/228073.diff


2 Files Affected:

- (modified) llvm/lib/CodeGen/PHIElimination.cpp (+20-4) 
- (added) llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir (+386) 


``````````diff
diff --git a/llvm/lib/CodeGen/PHIElimination.cpp b/llvm/lib/CodeGen/PHIElimination.cpp
index de69ca8d91898..0aa8a94021b66 100644
--- a/llvm/lib/CodeGen/PHIElimination.cpp
+++ b/llvm/lib/CodeGen/PHIElimination.cpp
@@ -15,6 +15,7 @@
 #include "llvm/CodeGen/PHIElimination.h"
 #include "PHIEliminationUtils.h"
 #include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/SetVector.h"
 #include "llvm/ADT/SmallPtrSet.h"
 #include "llvm/ADT/Statistic.h"
 #include "llvm/Analysis/LoopInfo.h"
@@ -114,6 +115,10 @@ class PHIEliminationImpl {
   // Count the number of non-undef PHI uses of each register in each BB.
   VRegPHIUse VRegPHIUseCount;
 
+  // PHI source registers whose subranges must be shrunk to their own uses once
+  // all PHIs are gone.
+  SmallSetVector<Register, 8> PHISrcRegsToShrink;
+
   // Defs of PHI sources which are implicit_def.
   SmallPtrSet<MachineInstr *, 4> ImpDefs;
 
@@ -306,6 +311,19 @@ bool PHIEliminationImpl::run(MachineFunction &MF) {
   }
 
   LoweredPHIs.clear();
+
+  // Different lanes may be used by different PHI source copies, or may already
+  // be dead in a predecessor. The main range's last use is therefore not a
+  // valid endpoint for every subrange. Wait until all PHIs have been removed
+  // before shrinking subranges to their remaining lane-specific uses.
+  for (Register Reg : PHISrcRegsToShrink) {
+    LiveInterval &LI = LIS->getInterval(Reg);
+    for (LiveInterval::SubRange &SR : LI.subranges())
+      LIS->shrinkToUses(SR, Reg);
+    LI.removeEmptySubRanges();
+  }
+  PHISrcRegsToShrink.clear();
+
   ImpDefs.clear();
   VRegPHIUseCount.clear();
 
@@ -723,6 +741,8 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
       if (!SrcUndef &&
           !VRegPHIUseCount[BBVRegPair(opBlock.getNumber(), SrcReg)]) {
         LiveInterval &SrcLI = LIS->getInterval(SrcReg);
+        if (SrcLI.hasSubRanges())
+          PHISrcRegsToShrink.insert(SrcReg);
 
         bool isLiveOut = false;
         for (MachineBasicBlock *Succ : opBlock.successors()) {
@@ -768,10 +788,6 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
           SlotIndex LastUseIndex = LIS->getInstructionIndex(*KillInst);
           SrcLI.removeSegment(LastUseIndex.getRegSlot(),
                               LIS->getMBBEndIdx(&opBlock));
-          for (auto &SR : SrcLI.subranges()) {
-            SR.removeSegment(LastUseIndex.getRegSlot(),
-                             LIS->getMBBEndIdx(&opBlock));
-          }
         }
       }
     }
diff --git a/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
new file mode 100644
index 0000000000000..ce3d989bca7d9
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
@@ -0,0 +1,386 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -verify-machineinstrs \
+# RUN:   -run-pass=liveintervals,phi-node-elimination -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 \
+# RUN:   --passes='require<live-intervals>,phi-node-elimination' -verify-each -o - %s | FileCheck %s
+#
+# The high lane is dead before the PHI predecessor, but its segment belongs to
+# a block appearing later in layout order. It must not be trimmed using the
+# low lane's last-use slot.
+---
+name: phi_subrange_absent
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_absent
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY %1.sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  bb.0:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %2:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %2
+    S_ENDPGM 0
+
+  bb.3:
+    successors: %bb.1
+    %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+    %1:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %1
+    S_BRANCH %bb.1
+...
+
+# The high lane's segment ends in an earlier layout block. It needs no trim.
+---
+name: phi_subrange_dead_earlier
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_dead_earlier
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY]]
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+    %1:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %2:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %2
+    S_ENDPGM 0
+...
+
+# Both lanes and the main live range still require normal PHI source trimming.
+---
+name: phi_subrange_full_live
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_full_live
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0_vgpr1 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vreg_64 = PHI %0, %bb.1
+    $vgpr0_vgpr1 = COPY %3
+    S_ENDPGM 0
+...
+
+# Different PHI source copies read different lanes. Each lane must end at its
+# own copy, not at the main range's last use in the predecessor.
+---
+name: phi_subrange_separate_uses
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_separate_uses
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    %4:vgpr_32 = PHI %0.sub1, %bb.1
+    $vgpr0 = COPY %3
+    $vgpr1 = COPY %4
+    S_ENDPGM 0
+...
+
+# PHIs in two different successor blocks share the source register. Shrinking
+# must wait for the whole function's PHIs to be removed.
+---
+name: phi_subrange_multiple_successors
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_multiple_successors
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x40000000), %bb.2(0x40000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+  ; CHECK-NEXT:   S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.2, implicit $scc
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   successors: %bb.4(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.4
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   successors: %bb.5(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_BRANCH %bb.5
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.4:
+  ; CHECK-NEXT:   successors: %bb.5(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_BRANCH %bb.5
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.5:
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1, %bb.2
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    %5:sreg_32 = S_MOV_B32 0
+    S_CMP_EQ_U32 %5, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.2, implicit $scc
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.3
+    S_BRANCH %bb.3
+
+  bb.2:
+    successors: %bb.4
+    S_BRANCH %bb.4
+
+  bb.3:
+    successors: %bb.5
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %3
+    S_BRANCH %bb.5
+
+  bb.4:
+    successors: %bb.5
+    %4:vgpr_32 = PHI %0.sub1, %bb.2
+    $vgpr1 = COPY %4
+    S_BRANCH %bb.5
+
+  bb.5:
+    S_ENDPGM 0
+...
+
+# Only sub1 is live into the other successor, so the main range is live-out and
+# the edge to the PHI block is split. The split block inherits every lane, but
+# the PHI copy placed there reads only sub0. sub1 must be shrunk out of the
+# split block rather than trimmed at the sub0 copy.
+---
+name: phi_subrange_lane_live_out_split_edge
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_lane_live_out_split_edge
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x40000000), %bb.4(0x40000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+  ; CHECK-NEXT:   S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+  ; CHECK-NEXT:   S_CBRANCH_SCC1 %bb.2, implicit $scc
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.4:
+  ; CHECK-NEXT:   successors: %bb.3(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+  ; CHECK-NEXT:   S_BRANCH %bb.3
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   $vgpr1 = COPY [[COPY1]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.3:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0 = COPY [[COPY2]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2, %bb.3
+    %5:sreg_32 = S_MOV_B32 0
+    S_CMP_EQ_U32 %5, 0, implicit-def $scc
+    S_CBRANCH_SCC1 %bb.2, implicit $scc
+    S_BRANCH %bb.3
+
+  bb.2:
+    %8:vgpr_32 = COPY %0.sub1
+    $vgpr1 = COPY %8
+    S_ENDPGM 0
+
+  bb.3:
+    %3:vgpr_32 = PHI %0.sub0, %bb.1
+    $vgpr0 = COPY %3
+    S_ENDPGM 0
+...
+
+# One PHI reads the whole register and another reads only sub1 from the same
+# predecessor. The sub1 copy is the last use of the main range, but sub0 must
+# end at the earlier full-register copy.
+---
+name: phi_subrange_full_and_lane
+tracksRegLiveness: true
+body: |
+  ; CHECK-LABEL: name: phi_subrange_full_and_lane
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+  ; CHECK-NEXT:   [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+  ; CHECK-NEXT:   [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.1
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   successors: %bb.2(0x80000000)
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+  ; CHECK-NEXT:   S_BRANCH %bb.2
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.2:
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+  ; CHECK-NEXT:   $vgpr0_vgpr1 = COPY [[COPY2]]
+  ; CHECK-NEXT:   [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+  ; CHECK-NEXT:   $vgpr2 = COPY [[COPY3]]
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+    %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+    %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+    S_BRANCH %bb.1
+
+  bb.1:
+    successors: %bb.2
+    S_BRANCH %bb.2
+
+  bb.2:
+    %3:vreg_64 = PHI %0, %bb.1
+    %4:vgpr_32 = PHI %0.sub1, %bb.1
+    $vgpr0_vgpr1 = COPY %3
+    $vgpr2 = COPY %4
+    S_ENDPGM 0
+...

``````````

</details>


https://github.com/llvm/llvm-project/pull/228073


More information about the llvm-commits mailing list