[llvm] [X86] Lower bf16->f32/f64 fpext without a GPR round-trip (PR #218359)
Phoebe Wang via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 24 07:15:53 PDT 2026
================
@@ -46,8 +46,8 @@ define void @phi_vec1bf16_to_f32(ptr %src, ptr %dst) #0 {
; CHECK-NEXT: .cfi_offset %rbx, -16
; CHECK-NEXT: movq %rsi, %rbx
; CHECK-NEXT: movzwl (%rdi), %eax
-; CHECK-NEXT: shll $16, %eax
; CHECK-NEXT: movd %eax, %xmm0
+; CHECK-NEXT: pslld $16, %xmm0
----------------
phoebewang wrote:
SHL's throughput is better than pslld. It's better to keep it when we have to load it through GPR register.
https://github.com/llvm/llvm-project/pull/218359
More information about the llvm-commits
mailing list