[llvm] AMDGPU: Invalidate VCC live ranges when lowering kill instructions (PR #226040)

via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 24 00:04:38 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-backend-amdgpu

Author: Matt Arsenault (arsenm)

<details>
<summary>Changes</summary>

lowerKillInstr replaces the kill pseudo with a v_cmp defining VCC plus
two instructions reading it, but left any previously computed VCC
regunit ranges alone. If something had already materialized them, the
new uses have no live segment:

  *** Bad machine code: No live segment at use ***
  - instruction: $exec = S_ANDN2_B64_term $exec, $vcc, implicit-def $scc

Drop the ranges so they are recomputed on demand, as is already done for
EXEC and SCC.

Co-Authored-By: Claude Opus 5 <noreply@<!-- -->anthropic.com>

---
Full diff: https://github.com/llvm/llvm-project/pull/226040.diff


2 Files Affected:

- (modified) llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp (+2) 
- (modified) llvm/test/CodeGen/AMDGPU/wqm.mir (+38) 


``````````diff
diff --git a/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp b/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
index 082fdf645ca26..a15a0b2924189 100644
--- a/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
+++ b/llvm/lib/Target/AMDGPU/SIWholeQuadMode.cpp
@@ -946,6 +946,8 @@ MachineInstr *SIWholeQuadMode::lowerKillF32(MachineInstr &MI) {
   LIS->InsertMachineInstrInMaps(*EarlyTermMI);
   LIS->InsertMachineInstrInMaps(*ExecMaskMI);
 
+  LIS->removeAllRegUnitsForPhysReg(AMDGPU::VCC);
+
   return ExecMaskMI;
 }
 
diff --git a/llvm/test/CodeGen/AMDGPU/wqm.mir b/llvm/test/CodeGen/AMDGPU/wqm.mir
index 4e5b5a6b43e7e..61acddb6b170b 100644
--- a/llvm/test/CodeGen/AMDGPU/wqm.mir
+++ b/llvm/test/CodeGen/AMDGPU/wqm.mir
@@ -42,6 +42,9 @@
   define amdgpu_vs void @no_wqm_in_vs() {
     ret void
   }
+  define amdgpu_ps void @kill_f32_cond_imm_vcc_regunit() {
+    ret void
+  }
 ...
 ---
 
@@ -614,3 +617,38 @@ body:             |
 
     %4:vreg_128 = IMAGE_SAMPLE_V4_V2 %0:vreg_64, %100:sgpr_256, %101:sgpr_128, 15, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), align 4, addrspace 4)
 ...
+
+---
+# The live-in $vcc forces the VCC regunit ranges to be computed, so the
+# clobber introduced when lowering the kill must invalidate them.
+name:            kill_f32_cond_imm_vcc_regunit
+tracksRegLiveness: true
+body:             |
+  ; CHECK-LABEL: name: kill_f32_cond_imm_vcc_regunit
+  ; CHECK: bb.0:
+  ; CHECK-NEXT:   successors: %bb.1(0x80000000)
+  ; CHECK-NEXT:   liveins: $vgpr0, $vcc
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT:   [[COPY:%[0-9]+]]:sreg_64 = COPY $exec
+  ; CHECK-NEXT:   [[COPY1:%[0-9]+]]:vgpr_32 = COPY $vgpr0
+  ; CHECK-NEXT:   [[COPY2:%[0-9]+]]:sreg_64 = COPY $vcc
+  ; CHECK-NEXT:   V_CMP_NGT_F32_e32 0, [[COPY1]], implicit-def $vcc, implicit $mode, implicit $exec
+  ; CHECK-NEXT:   dead [[COPY:%[0-9]+]]:sreg_64 = S_ANDN2_B64 [[COPY]], $vcc, implicit-def $scc
+  ; CHECK-NEXT:   SI_EARLY_TERMINATE_SCC0 implicit $exec, implicit $scc
+  ; CHECK-NEXT:   $exec = S_ANDN2_B64_term $exec, $vcc, implicit-def $scc
+  ; CHECK-NEXT: {{  $}}
+  ; CHECK-NEXT: bb.1:
+  ; CHECK-NEXT:   $exec = S_AND_B64 $exec, [[COPY2]], implicit-def $scc
+  ; CHECK-NEXT:   S_ENDPGM 0
+  bb.0:
+    successors: %bb.1
+    liveins: $vgpr0, $vcc
+
+    %0:vgpr_32 = COPY $vgpr0
+    %1:sreg_64 = COPY $vcc
+    SI_KILL_F32_COND_IMM_TERMINATOR %0, 0, 4, implicit-def $vcc, implicit $exec
+
+  bb.1:
+    $exec = S_AND_B64 $exec, %1, implicit-def $scc
+    S_ENDPGM 0
+...

``````````

</details>


https://github.com/llvm/llvm-project/pull/226040


More information about the llvm-commits mailing list