[llvm] [X86] Keep scalar bf16/f16 selects in vector registers (PR #224218)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 28 07:56:24 PDT 2026


================
@@ -2076,3 +2076,446 @@ define bfloat @PR115710(fp128 %0) nounwind {
   %2 = fptrunc fp128 %0 to bfloat
   ret bfloat %2
 }
+
+define bfloat @select_bf16(i1 %cond, bfloat %a, bfloat %b) nounwind {
+; X86-LABEL: select_bf16:
+; X86:       # %bb.0:
+; X86-NEXT:    testb $1, {{[0-9]+}}(%esp)
+; X86-NEXT:    leal {{[0-9]+}}(%esp), %eax
+; X86-NEXT:    leal {{[0-9]+}}(%esp), %ecx
+; X86-NEXT:    cmovnel %eax, %ecx
+; X86-NEXT:    vmovsh {{.*#+}} xmm0 = mem[0],zero,zero,zero,zero,zero,zero,zero
+; X86-NEXT:    retl
+;
+; SSE2-LABEL: select_bf16:
+; SSE2:       # %bb.0:
+; SSE2-NEXT:    andl $1, %edi
+; SSE2-NEXT:    negl %edi
+; SSE2-NEXT:    movd %edi, %xmm2
+; SSE2-NEXT:    pand %xmm2, %xmm0
+; SSE2-NEXT:    pandn %xmm1, %xmm2
+; SSE2-NEXT:    por %xmm2, %xmm0
+; SSE2-NEXT:    retq
+;
+; AVX512BF16-LABEL: select_bf16:
+; AVX512BF16:       # %bb.0:
+; AVX512BF16-NEXT:    kmovd %edi, %k1
+; AVX512BF16-NEXT:    vmovss %xmm0, %xmm1, %xmm1 {%k1}
+; AVX512BF16-NEXT:    vmovaps %xmm1, %xmm0
----------------
tfzee wrote:

I think that would be worse in this case since we would need to insert a negation instruction if we don't have access to the comparison itself while the vmovaps is effectively free with register renaming I would assume in most cases.
f32 does the same as well https://godbolt.org/z/r57bYnhne
I'm gonna double check for when we have access to the comparison operation and could in theory invert the comparisons check.

https://github.com/llvm/llvm-project/pull/224218


More information about the llvm-commits mailing list