[llvm] [AMDGPU][NFC] De-factorize scalar_to_vector bf16 patterns (PR #226847)
Changpeng Fang via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 27 15:34:26 PDT 2026
https://github.com/changpeng created https://github.com/llvm/llvm-project/pull/226847
For Fake16, an SGPR and a VGPR can share the same scalar_to_vector pattern because there is no 16-bit SGPR at all. The separate uniform SGPR pattern is therefore only needed for Real True16, where the divergent case uses a VGPR_16 REG_SEQUENCE.
This de-factorization is to prepare for follow-up PRs to separate SGPR and VGPR patterns of scalar_to_vector for other types.
>From 610cd647d153acb36c2536ab055a81e1aca53039 Mon Sep 17 00:00:00 2001
From: Changpeng Fang <changpeng.fang at amd.com>
Date: Sun, 27 Sep 2026 15:16:07 -0700
Subject: [PATCH] [AMDGPU][NFC] De-factorize scalar_to_vector bf16 patterns
For Fake16, an SGPR and a VGPR can share the same scalar_to_vector
pattern because there is no 16-bit SGPR at all. The separate uniform
SGPR pattern is therefore only needed for Real True16, where the
divergent case uses a VGPR_16 REG_SEQUENCE.
This de-factorization is to prepare for follow-up PRs to separate SGPR
and VGPR patterns of scalar_to_vector for other types.
---
llvm/lib/Target/AMDGPU/SIInstructions.td | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 3cf20ffb1fcf4..823a8b3a0c9f9 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -4256,8 +4256,8 @@ def : GCNPat <
>;
def : GCNPat <
- (v2bf16 (DivergentUnaryFrag<scalar_to_vector> (bf16 VGPR_32:$src0))),
- (COPY VGPR_32:$src0)
+ (v2bf16 (scalar_to_vector bf16:$src0)),
+ (COPY $src0)
>;
}
@@ -4286,14 +4286,14 @@ def : GCNPat <
(v2bf16 (DivergentUnaryFrag<scalar_to_vector> (bf16 VGPR_16:$src0))),
(REG_SEQUENCE VGPR_32, VGPR_16:$src0, lo16, (i16 (IMPLICIT_DEF)), hi16)
>;
-}
-// A bf16 value can only be in the low 16 bits of an SGPR, regardless of
-// Real-True16 or Fake-True16.
+// Unlike the divergent case above, a uniform bf16 value lives in the low 16
+// bits of a 32-bit SGPR even for Real-True16, so a plain COPY is enough.
def : GCNPat <
(v2bf16 (UniformUnaryFrag<scalar_to_vector> (bf16 SReg_32:$src0))),
(COPY SReg_32:$src0)
>;
+}
def : GCNPat <
(i64 (int_amdgcn_mov_dpp i64:$src, timm:$dpp_ctrl, timm:$row_mask,
More information about the llvm-commits
mailing list