[llvm] [AMDGPU][NFC] De-factorize scalar_to_vector bf16 patterns (PR #226847)

Changpeng Fang via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 27 15:34:26 PDT 2026


https://github.com/changpeng created https://github.com/llvm/llvm-project/pull/226847

For Fake16, an SGPR and a VGPR can share the same scalar_to_vector pattern because there is no 16-bit SGPR at all. The separate uniform SGPR pattern is therefore only needed for Real True16, where the divergent case uses a VGPR_16 REG_SEQUENCE.

This de-factorization is to prepare for follow-up PRs to separate SGPR and VGPR patterns of scalar_to_vector for other types.

>From 610cd647d153acb36c2536ab055a81e1aca53039 Mon Sep 17 00:00:00 2001
From: Changpeng Fang <changpeng.fang at amd.com>
Date: Sun, 27 Sep 2026 15:16:07 -0700
Subject: [PATCH] [AMDGPU][NFC] De-factorize scalar_to_vector bf16 patterns

For Fake16, an SGPR and a VGPR can share the same scalar_to_vector
pattern because there is no 16-bit SGPR at all. The separate uniform
SGPR pattern is therefore only needed for Real True16, where the
divergent case uses a VGPR_16 REG_SEQUENCE.

This de-factorization is to prepare for follow-up PRs to separate SGPR
and VGPR patterns of scalar_to_vector for other types.
---
 llvm/lib/Target/AMDGPU/SIInstructions.td | 10 +++++-----
 1 file changed, 5 insertions(+), 5 deletions(-)

diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 3cf20ffb1fcf4..823a8b3a0c9f9 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -4256,8 +4256,8 @@ def : GCNPat <
 >;
 
 def : GCNPat <
-  (v2bf16 (DivergentUnaryFrag<scalar_to_vector> (bf16 VGPR_32:$src0))),
-  (COPY VGPR_32:$src0)
+  (v2bf16 (scalar_to_vector bf16:$src0)),
+  (COPY $src0)
 >;
 }
 
@@ -4286,14 +4286,14 @@ def : GCNPat <
   (v2bf16 (DivergentUnaryFrag<scalar_to_vector> (bf16 VGPR_16:$src0))),
   (REG_SEQUENCE VGPR_32, VGPR_16:$src0, lo16, (i16 (IMPLICIT_DEF)), hi16)
 >;
-}
 
-// A bf16 value can only be in the low 16 bits of an SGPR, regardless of
-// Real-True16 or Fake-True16.
+// Unlike the divergent case above, a uniform bf16 value lives in the low 16
+// bits of a 32-bit SGPR even for Real-True16, so a plain COPY is enough.
 def : GCNPat <
   (v2bf16 (UniformUnaryFrag<scalar_to_vector> (bf16 SReg_32:$src0))),
   (COPY SReg_32:$src0)
 >;
+}
 
 def : GCNPat <
   (i64 (int_amdgcn_mov_dpp i64:$src, timm:$dpp_ctrl, timm:$row_mask,



More information about the llvm-commits mailing list