[llvm] [AMDGPU] Fix uniform fcopysign pattern (PR #218644)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 14 18:06:10 PDT 2026
================
@@ -2469,8 +2469,12 @@ let True16Predicate = UseRealTrue16Insts in {
def : GCNPat <
(fcopysign fp16vt:$src0, fp16vt:$src1),
(EXTRACT_SUBREG (V_BFI_B32_e64 (S_MOV_B32 (i32 0x00007fff)),
- (REG_SEQUENCE VGPR_32, $src0, lo16, (i16 (IMPLICIT_DEF)), hi16),
- (REG_SEQUENCE VGPR_32, $src1, lo16, (i16 (IMPLICIT_DEF)), hi16)), lo16)
+ (REG_SEQUENCE VGPR_32,
+ (fp16vt (COPY_TO_REGCLASS $src0, VGPR_16)), lo16,
----------------
Shoreshen wrote:
Hi @arsenm , I think maybe one of the method is to add a pesudo inst to allowing copy from SGPR32, VGPR32, VGPR16 to VGPR16, and register class can take effect in instruction definition.
https://github.com/llvm/llvm-project/pull/218644
More information about the llvm-commits
mailing list