[llvm] [AMDGPU][GlobalISel] Add BF16 widening to V2BF16 for packed instructions (PR #214370)

Petar Avramovic via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 10 02:47:09 PDT 2026


================
@@ -0,0 +1,106 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py UTC_ARGS: --version 6
+# RUN: llc -global-isel=0 -mtriple=amdgpu12.5 -mattr=+real-true16 -run-pass=legalizer -verify-machineinstrs %s -o - | FileCheck %s
+# RUN: llc -global-isel=1 -mtriple=amdgpu12.5 -mattr=+real-true16 -run-pass=legalizer -verify-machineinstrs %s -o - | FileCheck %s
+
+# Test that fast-math flags are preserved when widening scalar BF16 operations to V2BF16
+# in both SelectionDAG and GlobalISel
+
+---
+name:            test_fadd_bf16_fast
+tracksRegLiveness: true
+body:             |
+  bb.0:
+    liveins: $vgpr0, $vgpr1
+
+    ; CHECK-LABEL: name: test_fadd_bf16_fast
+    ; CHECK: liveins: $vgpr0, $vgpr1
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: [[COPY:%[0-9]+]]:_(i32) = COPY $vgpr0
+    ; CHECK-NEXT: [[COPY1:%[0-9]+]]:_(i32) = COPY $vgpr1
+    ; CHECK-NEXT: [[TRUNC:%[0-9]+]]:_(bf16) = G_TRUNC [[COPY]](i32)
+    ; CHECK-NEXT: [[TRUNC1:%[0-9]+]]:_(bf16) = G_TRUNC [[COPY1]](i32)
+    ; CHECK-NEXT: [[DEF:%[0-9]+]]:_(bf16) = G_IMPLICIT_DEF
+    ; CHECK-NEXT: [[BUILD_VECTOR:%[0-9]+]]:_(<2 x bf16>) = G_BUILD_VECTOR [[TRUNC]](bf16), [[DEF]](bf16)
+    ; CHECK-NEXT: [[BUILD_VECTOR1:%[0-9]+]]:_(<2 x bf16>) = G_BUILD_VECTOR [[TRUNC1]](bf16), [[DEF]](bf16)
+    ; CHECK-NEXT: [[FADD:%[0-9]+]]:_(<2 x bf16>) = nnan ninf nsz arcp contract afn reassoc G_FADD [[BUILD_VECTOR]], [[BUILD_VECTOR1]]
+    ; CHECK-NEXT: [[UV:%[0-9]+]]:_(bf16), [[UV1:%[0-9]+]]:_(bf16) = G_UNMERGE_VALUES [[FADD]](<2 x bf16>)
+    ; CHECK-NEXT: [[COPY2:%[0-9]+]]:_(bf16) = COPY [[UV]](bf16)
+    ; CHECK-NEXT: [[ANYEXT:%[0-9]+]]:_(i32) = G_ANYEXT [[COPY2]](bf16)
+    ; CHECK-NEXT: $vgpr0 = COPY [[ANYEXT]](i32)
+    ; CHECK-NEXT: SI_RETURN implicit $vgpr0
+    %0:_(i32) = COPY $vgpr0
+    %1:_(i32) = COPY $vgpr1
+    %2:_(bf16) = G_TRUNC %0
----------------
petar-avramovic wrote:

I thought the same, but it works fine. There is no alternative to it it seems at the moment (new opcode).  We got it by not changing existing code while switching to extended LLTs. And it works fine.
Not sure how we should approach it, dedicated G_bitcasting_trunc that behaves like regular trunc?

https://github.com/llvm/llvm-project/pull/214370


More information about the llvm-commits mailing list