[llvm] [X86] Lower bf16->f32/f64 fpext without a GPR round-trip (PR #218359)
Simon Pilgrim via llvm-commits
llvm-commits at lists.llvm.org
Tue Aug 25 03:12:16 PDT 2026
================
@@ -22973,6 +22975,38 @@ static SDValue LowerFP_TO_FP16(SDValue Op, SelectionDAG &DAG) {
return Res;
}
+SDValue X86TargetLowering::LowerBF16_TO_FP(SDValue Op,
+ SelectionDAG &DAG) const {
+ SDLoc DL(Op);
+ SDValue Src = Op.getOperand(0);
+ // Operand is usually already softened to i16 by type legalization.
+ if (Src.getValueType() == MVT::bf16)
+ Src = DAG.getBitcast(MVT::i16, Src);
+
+ SDValue Res;
+ if (!Subtarget.hasFP16() && ISD::isNormalLoad(Src.getNode())) {
+ // Without AVX512FP16 we need a GPR to load Src anyway (there's no direct
+ // memory-to-XMM move for a 16-bit value), so do the shift in GPR
+ // instead of in the vector domain, since SHL has better throughput
+ // than a vector shift.
+ SDValue Wide = DAG.getZExtOrTrunc(Src, DL, MVT::i32);
+ Wide = DAG.getNode(ISD::SHL, DL, MVT::i32, Wide, DAG.getConstant(16, DL, MVT::i32));
----------------
RKSimon wrote:
getShiftAmountConstant
https://github.com/llvm/llvm-project/pull/218359
More information about the llvm-commits
mailing list