[llvm] [PHIElim][AMDGPU]: shrink source subranges to lane-specific uses (PR #228073)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Oct 1 06:21:01 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu
Author: Alan Li (lialan)
<details>
<summary>Changes</summary>
This patch corrects a lane-specific trimming flaw in the subrange-preservation implementation introduced by D158144 and subsequently relanded as PR #<!-- -->69429.
## Symptoms
While analyzing issue #<!-- -->227456 we discovered the bug where phi elimination trims every source subrange using the whole register's last-use position. This can trigger a `LiveRange::removeSegment` assertion or leave invalid subrange endpoints which will be reported by machine verifier.
The issue indicated that when compiling an HSTU kernel with teh default gfx950 -O3 pipeline, where LiveIntervals are computed before PHI elimination. We found that recent PR #<!-- -->227252 made the move to expose the issue.
To be specific, here is the MIR example. Two PHIs read the 32-bit halves of one 64-bit source:
```
%lo = PHI %src.sub0, %bb.1
%hi = PHI %src.sub1, %bb.1
```
PHI elimination inserts separate source copies:
```
%lo_in = COPY %src.sub0
%hi_in = COPY %src.sub1
```
The old code ends both source subranges at the second copy. The low-half subrange should end at the first copy; the second copy only reads the high half. Consequently, `-verify-machineinstrs` reports: `Instruction ending live segment doesn't read the register`.
## Fix
* Record each PHI source interval that has subranges, and
* after all PHIs in the function are lowered recompute its subranges with `LiveIntervals::shrinkToUses`
* remove any empty subranges.
## Validation
Without the fix, patch-included regression tests `phi_subrange_absent` and `phi_subrange_multiple_successors` would hit the assertion and `phi_subrange_separate_uses` fails the verifier.
## Previous discussions
* [D156872 — Verify LiveIntervals for PHIs](https://reviews.llvm.org/D156872) : this was probably the first time it was brought up.
* [D158144](https://reviews.llvm.org/D158144): added subrange updates during PHI elimination, but assumes every lane can share the whole register’s endpoint.
---
Full diff: https://github.com/llvm/llvm-project/pull/228073.diff
2 Files Affected:
- (modified) llvm/lib/CodeGen/PHIElimination.cpp (+20-4)
- (added) llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir (+386)
``````````diff
diff --git a/llvm/lib/CodeGen/PHIElimination.cpp b/llvm/lib/CodeGen/PHIElimination.cpp
index de69ca8d91898..0aa8a94021b66 100644
--- a/llvm/lib/CodeGen/PHIElimination.cpp
+++ b/llvm/lib/CodeGen/PHIElimination.cpp
@@ -15,6 +15,7 @@
#include "llvm/CodeGen/PHIElimination.h"
#include "PHIEliminationUtils.h"
#include "llvm/ADT/DenseMap.h"
+#include "llvm/ADT/SetVector.h"
#include "llvm/ADT/SmallPtrSet.h"
#include "llvm/ADT/Statistic.h"
#include "llvm/Analysis/LoopInfo.h"
@@ -114,6 +115,10 @@ class PHIEliminationImpl {
// Count the number of non-undef PHI uses of each register in each BB.
VRegPHIUse VRegPHIUseCount;
+ // PHI source registers whose subranges must be shrunk to their own uses once
+ // all PHIs are gone.
+ SmallSetVector<Register, 8> PHISrcRegsToShrink;
+
// Defs of PHI sources which are implicit_def.
SmallPtrSet<MachineInstr *, 4> ImpDefs;
@@ -306,6 +311,19 @@ bool PHIEliminationImpl::run(MachineFunction &MF) {
}
LoweredPHIs.clear();
+
+ // Different lanes may be used by different PHI source copies, or may already
+ // be dead in a predecessor. The main range's last use is therefore not a
+ // valid endpoint for every subrange. Wait until all PHIs have been removed
+ // before shrinking subranges to their remaining lane-specific uses.
+ for (Register Reg : PHISrcRegsToShrink) {
+ LiveInterval &LI = LIS->getInterval(Reg);
+ for (LiveInterval::SubRange &SR : LI.subranges())
+ LIS->shrinkToUses(SR, Reg);
+ LI.removeEmptySubRanges();
+ }
+ PHISrcRegsToShrink.clear();
+
ImpDefs.clear();
VRegPHIUseCount.clear();
@@ -723,6 +741,8 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
if (!SrcUndef &&
!VRegPHIUseCount[BBVRegPair(opBlock.getNumber(), SrcReg)]) {
LiveInterval &SrcLI = LIS->getInterval(SrcReg);
+ if (SrcLI.hasSubRanges())
+ PHISrcRegsToShrink.insert(SrcReg);
bool isLiveOut = false;
for (MachineBasicBlock *Succ : opBlock.successors()) {
@@ -768,10 +788,6 @@ void PHIEliminationImpl::LowerPHINode(MachineBasicBlock &MBB,
SlotIndex LastUseIndex = LIS->getInstructionIndex(*KillInst);
SrcLI.removeSegment(LastUseIndex.getRegSlot(),
LIS->getMBBEndIdx(&opBlock));
- for (auto &SR : SrcLI.subranges()) {
- SR.removeSegment(LastUseIndex.getRegSlot(),
- LIS->getMBBEndIdx(&opBlock));
- }
}
}
}
diff --git a/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
new file mode 100644
index 0000000000000..ce3d989bca7d9
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/phi-elimination-subrange-liveintervals.mir
@@ -0,0 +1,386 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 -verify-machineinstrs \
+# RUN: -run-pass=liveintervals,phi-node-elimination -o - %s | FileCheck %s
+# RUN: llc -mtriple=amdgcn-amd-amdhsa -mcpu=gfx950 \
+# RUN: --passes='require<live-intervals>,phi-node-elimination' -verify-each -o - %s | FileCheck %s
+#
+# The high lane is dead before the PHI predecessor, but its segment belongs to
+# a block appearing later in layout order. It must not be trimmed using the
+# low lane's last-use slot.
+---
+name: phi_subrange_absent
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_absent
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vgpr_32 = COPY %1.sub0
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0 = COPY [[COPY1]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.3:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: $vgpr1 = COPY [[COPY2]]
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ bb.0:
+ successors: %bb.3
+ S_BRANCH %bb.3
+
+ bb.1:
+ successors: %bb.2
+ S_BRANCH %bb.2
+
+ bb.2:
+ %2:vgpr_32 = PHI %0.sub0, %bb.1
+ $vgpr0 = COPY %2
+ S_ENDPGM 0
+
+ bb.3:
+ successors: %bb.1
+ %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+ %1:vgpr_32 = COPY %0.sub1
+ $vgpr1 = COPY %1
+ S_BRANCH %bb.1
+...
+
+# The high lane's segment ends in an earlier layout block. It needs no trim.
+---
+name: phi_subrange_dead_earlier
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_dead_earlier
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: $vgpr1 = COPY [[COPY]]
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+ ; CHECK-NEXT: $vgpr0 = COPY [[COPY2]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ %3:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %4:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %3, %subreg.sub0, %4, %subreg.sub1
+ %1:vgpr_32 = COPY %0.sub1
+ $vgpr1 = COPY %1
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.2
+ S_BRANCH %bb.2
+
+ bb.2:
+ %2:vgpr_32 = PHI %0.sub0, %bb.1
+ $vgpr0 = COPY %2
+ S_ENDPGM 0
+...
+
+# Both lanes and the main live range still require normal PHI source trimming.
+---
+name: phi_subrange_full_live
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_full_live
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0_vgpr1 = COPY [[COPY1]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.2
+ S_BRANCH %bb.2
+
+ bb.2:
+ %3:vreg_64 = PHI %0, %bb.1
+ $vgpr0_vgpr1 = COPY %3
+ S_ENDPGM 0
+...
+
+# Different PHI source copies read different lanes. Each lane must end at its
+# own copy, not at the main range's last use in the predecessor.
+---
+name: phi_subrange_separate_uses
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_separate_uses
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0 = COPY [[COPY2]]
+ ; CHECK-NEXT: [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+ ; CHECK-NEXT: $vgpr1 = COPY [[COPY3]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.2
+ S_BRANCH %bb.2
+
+ bb.2:
+ %3:vgpr_32 = PHI %0.sub0, %bb.1
+ %4:vgpr_32 = PHI %0.sub1, %bb.1
+ $vgpr0 = COPY %3
+ $vgpr1 = COPY %4
+ S_ENDPGM 0
+...
+
+# PHIs in two different successor blocks share the source register. Shrinking
+# must wait for the whole function's PHIs to be removed.
+---
+name: phi_subrange_multiple_successors
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_multiple_successors
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x40000000), %bb.2(0x40000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+ ; CHECK-NEXT: S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.2, implicit $scc
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: successors: %bb.4(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: S_BRANCH %bb.4
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.3:
+ ; CHECK-NEXT: successors: %bb.5(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0 = COPY [[COPY2]]
+ ; CHECK-NEXT: S_BRANCH %bb.5
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.4:
+ ; CHECK-NEXT: successors: %bb.5(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+ ; CHECK-NEXT: $vgpr1 = COPY [[COPY3]]
+ ; CHECK-NEXT: S_BRANCH %bb.5
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.5:
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1, %bb.2
+ %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+ %5:sreg_32 = S_MOV_B32 0
+ S_CMP_EQ_U32 %5, 0, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.2, implicit $scc
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.3
+ S_BRANCH %bb.3
+
+ bb.2:
+ successors: %bb.4
+ S_BRANCH %bb.4
+
+ bb.3:
+ successors: %bb.5
+ %3:vgpr_32 = PHI %0.sub0, %bb.1
+ $vgpr0 = COPY %3
+ S_BRANCH %bb.5
+
+ bb.4:
+ successors: %bb.5
+ %4:vgpr_32 = PHI %0.sub1, %bb.2
+ $vgpr1 = COPY %4
+ S_BRANCH %bb.5
+
+ bb.5:
+ S_ENDPGM 0
+...
+
+# Only sub1 is live into the other successor, so the main range is live-out and
+# the edge to the PHI block is split. The split block inherits every lane, but
+# the PHI copy placed there reads only sub0. sub1 must be shrunk out of the
+# split block rather than trimmed at the sub0 copy.
+---
+name: phi_subrange_lane_live_out_split_edge
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_lane_live_out_split_edge
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x40000000), %bb.4(0x40000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[S_MOV_B32_:%[0-9]+]]:sreg_32 = S_MOV_B32 0
+ ; CHECK-NEXT: S_CMP_EQ_U32 [[S_MOV_B32_]], 0, implicit-def $scc
+ ; CHECK-NEXT: S_CBRANCH_SCC1 %bb.2, implicit $scc
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.4:
+ ; CHECK-NEXT: successors: %bb.3(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub0
+ ; CHECK-NEXT: S_BRANCH %bb.3
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: $vgpr1 = COPY [[COPY1]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.3:
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0 = COPY [[COPY2]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.2, %bb.3
+ %5:sreg_32 = S_MOV_B32 0
+ S_CMP_EQ_U32 %5, 0, implicit-def $scc
+ S_CBRANCH_SCC1 %bb.2, implicit $scc
+ S_BRANCH %bb.3
+
+ bb.2:
+ %8:vgpr_32 = COPY %0.sub1
+ $vgpr1 = COPY %8
+ S_ENDPGM 0
+
+ bb.3:
+ %3:vgpr_32 = PHI %0.sub0, %bb.1
+ $vgpr0 = COPY %3
+ S_ENDPGM 0
+...
+
+# One PHI reads the whole register and another reads only sub1 from the same
+# predecessor. The sub1 copy is the last use of the main range, but sub0 must
+# end at the earlier full-register copy.
+---
+name: phi_subrange_full_and_lane
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: phi_subrange_full_and_lane
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[V_MOV_B32_e32_:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ ; CHECK-NEXT: [[V_MOV_B32_e32_1:%[0-9]+]]:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:vreg_64 = REG_SEQUENCE [[V_MOV_B32_e32_]], %subreg.sub0, [[V_MOV_B32_e32_1]], %subreg.sub1
+ ; CHECK-NEXT: S_BRANCH %bb.1
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: successors: %bb.2(0x80000000)
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:vreg_64 = COPY [[REG_SEQUENCE]]
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY [[REG_SEQUENCE]].sub1
+ ; CHECK-NEXT: S_BRANCH %bb.2
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.2:
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vreg_64 = COPY [[COPY]]
+ ; CHECK-NEXT: $vgpr0_vgpr1 = COPY [[COPY2]]
+ ; CHECK-NEXT: [[COPY3:%[0-9]+]]:vgpr_32 = COPY [[COPY1]]
+ ; CHECK-NEXT: $vgpr2 = COPY [[COPY3]]
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ %1:vgpr_32 = V_MOV_B32_e32 1, implicit $exec
+ %2:vgpr_32 = V_MOV_B32_e32 2, implicit $exec
+ %0:vreg_64 = REG_SEQUENCE %1, %subreg.sub0, %2, %subreg.sub1
+ S_BRANCH %bb.1
+
+ bb.1:
+ successors: %bb.2
+ S_BRANCH %bb.2
+
+ bb.2:
+ %3:vreg_64 = PHI %0, %bb.1
+ %4:vgpr_32 = PHI %0.sub1, %bb.1
+ $vgpr0_vgpr1 = COPY %3
+ $vgpr2 = COPY %4
+ S_ENDPGM 0
+...
``````````
</details>
https://github.com/llvm/llvm-project/pull/228073
More information about the llvm-commits
mailing list