[llvm] [AMDGPU] Legalize SGPR hi16 copies into S_PACK_HH_B32_B16 (PR #218371)

Arseniy Obolenskiy via llvm-commits llvm-commits at lists.llvm.org
Thu Aug 27 00:54:32 PDT 2026


aobolensk wrote:

> In that pattern could you please try changing `(EXTRACT_SUBREG VGPR_32:$src1, hi16)` to `(EXTRACT_SUBREG (i32 (COPY_TO_REGCLASS $src1, VGPR_32)), hi16)`? That's what we do in other places.

On that example:
```llvmir
define amdgpu_kernel void @f(float %sign) {
  %r = call float @llvm.copysign.f32(float 0.0, float %sign)
  %h = fptrunc float %r to half
  %c = fptosi half %h to i32
  store i32 %c, ptr addrspace(1) null
  ret void
}
```
on `llc -mtriple=amdgpu11.00 -mattr=+real-true16` it does not fix the issue if we go this way

`EXTRACT_SUBREG` dest reg class comes from the consuming operand `VS_16` constraint (SGPR-or-VGPR), not the forced `VGPR_32` source, so it still picks an SGPR hi16 and crashes the same way

Forcing a second `COPY_TO_REGCLASS` to `VGPR_16` avoids the crash but drops this pattern priority so it bassically stops matching for VGPR inputs too, regressing other copysign tests to the general BFI lowering

Do you have any suggestions on other alternatives?

https://github.com/llvm/llvm-project/pull/218371


More information about the llvm-commits mailing list