[llvm] [X86] Lower bf16->f32/f64 fpext without a GPR round-trip (PR #218359)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 24 07:49:37 PDT 2026
================
@@ -22973,6 +22975,39 @@ static SDValue LowerFP_TO_FP16(SDValue Op, SelectionDAG &DAG) {
return Res;
}
+SDValue X86TargetLowering::LowerBF16_TO_FP(SDValue Op,
+ SelectionDAG &DAG) const {
+ SDLoc DL(Op);
+ SDValue Src = Op.getOperand(0);
+ // Operand is usually already softened to i16 by type legalization.
+ if (Src.getValueType() == MVT::bf16)
+ Src = DAG.getBitcast(MVT::i16, Src);
+ else
+ Src = DAG.getAnyExtOrTrunc(Src, DL, MVT::i16);
+
+ // Peek through to bf16's underlying f16 register
+ MVT VecVT = MVT::v8i16;
+ SDValue VecSrc = Src;
+ if (Src.getOpcode() == ISD::BITCAST &&
+ Src.getOperand(0).getValueType() == MVT::f16) {
+ VecSrc = Src.getOperand(0);
+ VecVT = MVT::v8f16;
+ }
----------------
tfzee wrote:
Without the bitcast check there's some edgecases with when having AVX512FP16 which will fail and generate a GPR roundtrip. Maybe this isn't the cleanest way to handle it however.
For example fptosi_bf16_to_i32 will go from
```
; CHECK-LABEL: fptosi_bf16_to_i32:
-; CHECK: # %bb.0:
-; NEXT: vpslld $16, %xmm0, %xmm0
-; NEXT: vcvttss2si %xmm0, %eax
-; NEXT: retq
```
to
```
+; AVX10_2-LABEL: fptosi_bf16_to_i32:
+; NEXT: vmovw %xmm0, %eax
+; NEXT: vmovw %eax, %xmm0
+; NEXT: vpslld $16, %xmm0, %xmm0
+; NEXT: vcvttss2si %xmm0, %eax
+; NEXT: retq
```
Or add2 example will go from
```
-; NEXT: vpslld $16, %xmm1, %xmm1
-; NEXT: vpslld $16, %xmm0, %xmm0
-; NEXT: vaddss %xmm1, %xmm0, %xmm0
-; NEXT: vcvtneps2bf16 %xmm0, %xmm0
-; NEXT: retq
```
to
```
+; NEXT: vmovw %xmm0, %eax
+; NEXT: vmovw %xmm1, %ecx
+; NEXT: vmovw %ecx, %xmm0
+; NEXT: vpslld $16, %xmm0, %xmm0
+; NEXT: vmovw %eax, %xmm1
+; NEXT: vpslld $16, %xmm1, %xmm1
+; NEXT: vaddss %xmm0, %xmm1, %xmm0
+; NEXT: vcvtneps2bf16 %xmm0, %xmm0
+; NEXT: retq
```
https://github.com/llvm/llvm-project/pull/218359
More information about the llvm-commits
mailing list