[llvm] [AMDGPU] Legalize SGPR hi16 copies into S_PACK_HH_B32_B16 (PR #218371)

Arseniy Obolenskiy via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 15 04:22:06 PDT 2026


aobolensk wrote:

> > > In that pattern could you please try changing `(EXTRACT_SUBREG VGPR_32:$src1, hi16)` to `(EXTRACT_SUBREG (i32 (COPY_TO_REGCLASS $src1, VGPR_32)), hi16)`? That's what we do in other places.
> > 
> > 
> > On that example:
> > ```
> > define amdgpu_kernel void @f(float %sign) {
> >   %r = call float @llvm.copysign.f32(float 0.0, float %sign)
> >   %h = fptrunc float %r to half
> >   %c = fptosi half %h to i32
> >   store i32 %c, ptr addrspace(1) null
> >   ret void
> > }
> > ```
> > 
> > 
> >     
> >       
> >     
> > 
> >       
> >     
> > 
> >     
> >   
> > on `llc -mtriple=amdgpu11.00 -mattr=+real-true16` it does not fix the issue if we go this way
> > `EXTRACT_SUBREG` dest reg class comes from the consuming operand `VS_16` constraint (SGPR-or-VGPR), not the forced `VGPR_32` source, so it still picks an SGPR hi16 and crashes the same way
> > Forcing a second `COPY_TO_REGCLASS` to `VGPR_16` avoids the crash but drops this pattern priority so it bassically stops matching for VGPR inputs too, regressing other copysign tests to the general BFI lowering
> > Do you have any suggestions on other alternatives?
> 
> I can't see why the isel stop selecting this pattern if the copy_to_regclass is added, since it's added on the output node. Can you elaborate it a bit? Also trying to make sure the destination class of `copy_to_regclass` is `vgpr_32`, but not `vgpr_16`

Not applicable anymore, this part is gone 

https://github.com/llvm/llvm-project/pull/218371


More information about the llvm-commits mailing list