[llvm] [X86] Lower bf16->f32/f64 fpext without a GPR round-trip (PR #218359)

Phoebe Wang via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 24 07:15:53 PDT 2026


================
@@ -46,8 +46,8 @@ define void @phi_vec1bf16_to_f32(ptr %src, ptr %dst) #0 {
 ; CHECK-NEXT:    .cfi_offset %rbx, -16
 ; CHECK-NEXT:    movq %rsi, %rbx
 ; CHECK-NEXT:    movzwl (%rdi), %eax
-; CHECK-NEXT:    shll $16, %eax
 ; CHECK-NEXT:    movd %eax, %xmm0
+; CHECK-NEXT:    pslld $16, %xmm0
----------------
phoebewang wrote:

SHL's throughput is better than pslld. It's better to keep it when we have to load it through GPR register.

https://github.com/llvm/llvm-project/pull/218359


More information about the llvm-commits mailing list