[llvm] [AMDGPU] Legalize SGPR hi16 copies into S_PACK_HH_B32_B16 (PR #218371)
Arseniy Obolenskiy via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 15 04:22:06 PDT 2026
aobolensk wrote:
> > > In that pattern could you please try changing `(EXTRACT_SUBREG VGPR_32:$src1, hi16)` to `(EXTRACT_SUBREG (i32 (COPY_TO_REGCLASS $src1, VGPR_32)), hi16)`? That's what we do in other places.
> >
> >
> > On that example:
> > ```
> > define amdgpu_kernel void @f(float %sign) {
> > %r = call float @llvm.copysign.f32(float 0.0, float %sign)
> > %h = fptrunc float %r to half
> > %c = fptosi half %h to i32
> > store i32 %c, ptr addrspace(1) null
> > ret void
> > }
> > ```
> >
> >
> >
> >
> >
> >
> >
> >
> >
> >
> >
> > on `llc -mtriple=amdgpu11.00 -mattr=+real-true16` it does not fix the issue if we go this way
> > `EXTRACT_SUBREG` dest reg class comes from the consuming operand `VS_16` constraint (SGPR-or-VGPR), not the forced `VGPR_32` source, so it still picks an SGPR hi16 and crashes the same way
> > Forcing a second `COPY_TO_REGCLASS` to `VGPR_16` avoids the crash but drops this pattern priority so it bassically stops matching for VGPR inputs too, regressing other copysign tests to the general BFI lowering
> > Do you have any suggestions on other alternatives?
>
> I can't see why the isel stop selecting this pattern if the copy_to_regclass is added, since it's added on the output node. Can you elaborate it a bit? Also trying to make sure the destination class of `copy_to_regclass` is `vgpr_32`, but not `vgpr_16`
Not applicable anymore, this part is gone
https://github.com/llvm/llvm-project/pull/218371
More information about the llvm-commits
mailing list