[llvm] AMDGPU: Invalidate VCC live ranges when lowering kill instructions (PR #226040)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 24 00:04:38 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu
Author: Matt Arsenault (arsenm)
<details>
<summary>Changes</summary>
lowerKillInstr replaces the kill pseudo with a v_cmp defining VCC plus
two instructions reading it, but left any previously computed VCC
regunit ranges alone. If something had already materialized them, the
new uses have no live segment:
*** Bad machine code: No live segment at use ***
- instruction: $exec = S_ANDN2_B64_term $exec, $vcc, implicit-def $scc
Drop the ranges so they are recomputed on demand, as is already done for
EXEC and SCC.
Co-Authored-By: Claude Opus 5 <noreply@<!-- -->anthropic.com>
---
Full diff: https://github.com/llvm/llvm-project/pull/226040.diff
2 Files Affected:
- (modified) llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp (+2)
- (modified) llvm/test/CodeGen/AMDGPU/wqm.mir (+38)
``````````diff
diff --git a/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp b/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
index 082fdf645ca26..a15a0b2924189 100644
--- a/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
+++ b/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
@@ -946,6 +946,8 @@ MachineInstr *SIWholeQuadMode::lowerKillF32(MachineInstr &MI) {
LIS->InsertMachineInstrInMaps(*EarlyTermMI);
LIS->InsertMachineInstrInMaps(*ExecMaskMI);
+ LIS->removeAllRegUnitsForPhysReg(AMDGPU::VCC);
+
return ExecMaskMI;
}
diff --git a/llvm/test/CodeGen/AMDGPU/wqm.mir b/llvm/test/CodeGen/AMDGPU/wqm.mir
index 4e5b5a6b43e7e..61acddb6b170b 100644
--- a/llvm/test/CodeGen/AMDGPU/wqm.mir
+++ b/llvm/test/CodeGen/AMDGPU/wqm.mir
@@ -42,6 +42,9 @@
define amdgpu_vs void @no_wqm_in_vs() {
ret void
}
+ define amdgpu_ps void @kill_f32_cond_imm_vcc_regunit() {
+ ret void
+ }
...
---
@@ -614,3 +617,38 @@ body: |
%4:vreg_128 = IMAGE_SAMPLE_V4_V2 %0:vreg_64, %100:sgpr_256, %101:sgpr_128, 15, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), align 4, addrspace 4)
...
+
+---
+# The live-in $vcc forces the VCC regunit ranges to be computed, so the
+# clobber introduced when lowering the kill must invalidate them.
+name: kill_f32_cond_imm_vcc_regunit
+tracksRegLiveness: true
+body: |
+ ; CHECK-LABEL: name: kill_f32_cond_imm_vcc_regunit
+ ; CHECK: bb.0:
+ ; CHECK-NEXT: successors: %bb.1(0x80000000)
+ ; CHECK-NEXT: liveins: $vgpr0, $vcc
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: [[COPY:%[0-9]+]]:sreg_64 = COPY $exec
+ ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY $vgpr0
+ ; CHECK-NEXT: [[COPY2:%[0-9]+]]:sreg_64 = COPY $vcc
+ ; CHECK-NEXT: V_CMP_NGT_F32_e32 0, [[COPY1]], implicit-def $vcc, implicit $mode, implicit $exec
+ ; CHECK-NEXT: dead [[COPY:%[0-9]+]]:sreg_64 = S_ANDN2_B64 [[COPY]], $vcc, implicit-def $scc
+ ; CHECK-NEXT: SI_EARLY_TERMINATE_SCC0 implicit $exec, implicit $scc
+ ; CHECK-NEXT: $exec = S_ANDN2_B64_term $exec, $vcc, implicit-def $scc
+ ; CHECK-NEXT: {{ $}}
+ ; CHECK-NEXT: bb.1:
+ ; CHECK-NEXT: $exec = S_AND_B64 $exec, [[COPY2]], implicit-def $scc
+ ; CHECK-NEXT: S_ENDPGM 0
+ bb.0:
+ successors: %bb.1
+ liveins: $vgpr0, $vcc
+
+ %0:vgpr_32 = COPY $vgpr0
+ %1:sreg_64 = COPY $vcc
+ SI_KILL_F32_COND_IMM_TERMINATOR %0, 0, 4, implicit-def $vcc, implicit $exec
+
+ bb.1:
+ $exec = S_AND_B64 $exec, %1, implicit-def $scc
+ S_ENDPGM 0
+...
``````````
</details>
https://github.com/llvm/llvm-project/pull/226040
More information about the llvm-commits
mailing list