[llvm] [X86] Keep scalar bf16/f16 selects in vector registers (PR #224218)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 28 07:56:24 PDT 2026
================
@@ -2076,3 +2076,446 @@ define bfloat @PR115710(fp128 %0) nounwind {
%2 = fptrunc fp128 %0 to bfloat
ret bfloat %2
}
+
+define bfloat @select_bf16(i1 %cond, bfloat %a, bfloat %b) nounwind {
+; X86-LABEL: select_bf16:
+; X86: # %bb.0:
+; X86-NEXT: testb $1, {{[0-9]+}}(%esp)
+; X86-NEXT: leal {{[0-9]+}}(%esp), %eax
+; X86-NEXT: leal {{[0-9]+}}(%esp), %ecx
+; X86-NEXT: cmovnel %eax, %ecx
+; X86-NEXT: vmovsh {{.*#+}} xmm0 = mem[0],zero,zero,zero,zero,zero,zero,zero
+; X86-NEXT: retl
+;
+; SSE2-LABEL: select_bf16:
+; SSE2: # %bb.0:
+; SSE2-NEXT: andl $1, %edi
+; SSE2-NEXT: negl %edi
+; SSE2-NEXT: movd %edi, %xmm2
+; SSE2-NEXT: pand %xmm2, %xmm0
+; SSE2-NEXT: pandn %xmm1, %xmm2
+; SSE2-NEXT: por %xmm2, %xmm0
+; SSE2-NEXT: retq
+;
+; AVX512BF16-LABEL: select_bf16:
+; AVX512BF16: # %bb.0:
+; AVX512BF16-NEXT: kmovd %edi, %k1
+; AVX512BF16-NEXT: vmovss %xmm0, %xmm1, %xmm1 {%k1}
+; AVX512BF16-NEXT: vmovaps %xmm1, %xmm0
----------------
tfzee wrote:
I think that would be worse in this case since we would need to insert a negation instruction if we don't have access to the comparison itself while the vmovaps is effectively free with register renaming I would assume in most cases.
f32 does the same as well https://godbolt.org/z/r57bYnhne
I'm gonna double check for when we have access to the comparison operation and could in theory invert the comparisons check.
https://github.com/llvm/llvm-project/pull/224218
More information about the llvm-commits
mailing list