[llvm] [AMDGPU] Fix uniform fcopysign pattern (PR #218644)
via llvm-commits
llvm-commits at lists.llvm.org
Fri Aug 28 01:31:14 PDT 2026
================
@@ -2479,8 +2479,12 @@ let True16Predicate = UseRealTrue16Insts in {
def : GCNPat <
(fcopysign fp16vt:$src0, fp16vt:$src1),
(EXTRACT_SUBREG (V_BFI_B32_e64 (S_MOV_B32 (i32 0x00007fff)),
- (REG_SEQUENCE VGPR_32, $src0, lo16, (i16 (IMPLICIT_DEF)), hi16),
- (REG_SEQUENCE VGPR_32, $src1, lo16, (i16 (IMPLICIT_DEF)), hi16)), lo16)
+ (REG_SEQUENCE VGPR_32,
+ (fp16vt (COPY_TO_REGCLASS $src0, VGPR_16)), lo16,
----------------
Shoreshen wrote:
Hi @broxigarchen , the reason of creating reg sequence of `%53:vgpr_32, %subreg.lo16` is due to isel create patterns like `%0:sreg_32, %subreg.lo16` and replace by VGPR in fix copy pass.
And the reason of `%0:sreg_32, %subreg.lo16` created is due to we put `f16` type in sreg_32:
```tablegen
def Reg16Types : RegisterTypes<[i16, f16, bf16]>;
...
def VS_16 : SIRegisterClass<"AMDGPU", Reg16Types.types, 16,
(add VGPR_16, SReg_32, LDS_DIRECT)> {
let isAllocatable = 0;
let HasVGPR = 1;
let Size = 16;
}
```
So in this patch, the choice is copy sreg_32 to vgpr_16 to fit the type and create:
```
%1:VGPR_16 = COPY_TO_REGCLASS %0:SReg_32
%34:vgpr_32 = REG_SEQUENCE %1:VGPR_16, %subreg.lo16, %35:sreg_32, %subreg.hi16
```
https://github.com/llvm/llvm-project/pull/218644
More information about the llvm-commits
mailing list