[llvm] [X86] Lower bf16->f32/f64 fpext without a GPR round-trip (PR #218359)

Simon Pilgrim via llvm-commits llvm-commits at lists.llvm.org
Tue Aug 25 03:12:16 PDT 2026


================
@@ -22973,6 +22975,38 @@ static SDValue LowerFP_TO_FP16(SDValue Op, SelectionDAG &DAG) {
   return Res;
 }
 
+SDValue X86TargetLowering::LowerBF16_TO_FP(SDValue Op,
+                                           SelectionDAG &DAG) const {
+  SDLoc DL(Op);
+  SDValue Src = Op.getOperand(0);
+  // Operand is usually already softened to i16 by type legalization.
+  if (Src.getValueType() == MVT::bf16)
+    Src = DAG.getBitcast(MVT::i16, Src);
+
+  SDValue Res;
+  if (!Subtarget.hasFP16() && ISD::isNormalLoad(Src.getNode())) {
+    // Without AVX512FP16 we need a GPR to load Src anyway (there's no direct
+    // memory-to-XMM move for a 16-bit value), so do the shift in GPR
+    // instead of in the vector domain, since SHL has better throughput
+    // than a vector shift.
+    SDValue Wide = DAG.getZExtOrTrunc(Src, DL, MVT::i32);
+    Wide = DAG.getNode(ISD::SHL, DL, MVT::i32, Wide, DAG.getConstant(16, DL, MVT::i32));
----------------
RKSimon wrote:

getShiftAmountConstant

https://github.com/llvm/llvm-project/pull/218359


More information about the llvm-commits mailing list