[llvm] [AMDGPU][CodeGen] Allow remat with multiple users in same region (PR #214725)

Igor Wodiany via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 07:02:27 PDT 2026


IgWod wrote:

Under specific circumstances this causes a compiler crash on early architectures (tested `gfx600`). It surfaced in lit tests when I was working on unrelated `MachineSink` changes. The failing test was `llvm/test/CodeGen/AMDGPU/si-sgpr-spill.ll` hitting an assert in `VritRegMap` (`assert(PhysReg.isValid() && "Invalid SubReg for physical register");`). Repro based on the failure (crafted with a help of AI) below:

```
# RUN: llc -mtriple=amdgcn -mcpu=gfx600 -run-pass=machine-scheduler -verify-machineinstrs -o - %s

---
name:            remat_wide_sload_two_subreg_users
tracksRegLiveness: true
registers:
  - { id: 0, class: sgpr_64 }
  - { id: 1, class: vreg_64 }
  - { id: 2, class: sreg_64 }
  - { id: 3, class: sgpr_512 }
  - { id: 4, class: sgpr_512 }
  - { id: 5, class: sgpr_512 }
  - { id: 6, class: vreg_128 }
  - { id: 7, class: vgpr_32 }
  - { id: 8, class: vreg_128 }
  - { id: 9, class: sreg_64_xexec }
  - { id: 10, class: vgpr_32 }
machineFunctionInfo:
  isEntryFunction: true
  occupancy:       10
  psInputAddr:     4866
  psInputEnable:   4866
body:             |
  bb.0:
    successors: %bb.1(0x80000000)

    undef %0.sub1:sgpr_64 = IMPLICIT_DEF
    undef %1.sub1:vreg_64 = IMPLICIT_DEF
    %2:sreg_64 = IMPLICIT_DEF
    %3:sgpr_512 = S_LOAD_DWORDX16_IMM undef %0, 16, 0 :: (invariant load (s512), align 32, addrspace 4)
    %4:sgpr_512 = IMPLICIT_DEF
    %5:sgpr_512 = IMPLICIT_DEF

  bb.1:
    successors: %bb.2(0x04000000), %bb.1(0x7c000000)

    $exec = S_ANDN2_B64_term $exec, %2, implicit-def $scc
    S_CBRANCH_EXECNZ %bb.1, implicit $exec
    S_BRANCH %bb.2

  bb.2:
    successors: %bb.3(0x80000000)

    %6:vreg_128 = IMAGE_SAMPLE_V4_V2 undef %1, %3.sub0_sub1_sub2_sub3_sub4_sub5_sub6_sub7, %4.sub8_sub9_sub10_sub11, 15, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), addrspace 8)
    %7:vgpr_32 = IMAGE_SAMPLE_V1_V2 undef %1, %4.sub0_sub1_sub2_sub3_sub4_sub5_sub6_sub7, undef %5.sub8_sub9_sub10_sub11, 4, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), addrspace 8)
    %8:vreg_128 = IMAGE_SAMPLE_V4_V2 undef %1, %3.sub8_sub9_sub10_sub11_sub12_sub13_sub14_sub15, %4.sub12_sub13_sub14_sub15, 15, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), addrspace 8)
    %9:sreg_64_xexec = V_CMP_LT_F32_e64 0, 0, 0, %8.sub2, 0, implicit $mode, implicit $exec
    %10:vgpr_32 = V_CNDMASK_B32_e64 0, undef %8.sub0, 0, undef %8.sub1, undef %9, implicit $exec
    $exec = S_MOV_B64_term undef %2
    S_BRANCH %bb.3

  bb.3:
    S_ENDPGM 0
...
```

I tested this change (8a9c0755eb6c4f00a12bc128b7d4c9d8d0d08667) and verification passes before scheduling but fails after. I also confirmed that MIR verifies correctly after scheduling with (879a2e61e1efeeaa4fa6a60b63e8a036d478dd9a).

The problem seems to be that with this change `S_LOAD_DWORDX16_IMM` with 2 uses gets incorrectly rematerialized.

https://github.com/llvm/llvm-project/pull/214725


More information about the llvm-commits mailing list