[llvm] [X86] Keep scalar bf16/f16 selects in vector registers (PR #224218)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 28 23:54:29 PDT 2026
================
@@ -2076,3 +2076,446 @@ define bfloat @PR115710(fp128 %0) nounwind {
%2 = fptrunc fp128 %0 to bfloat
ret bfloat %2
}
+
+define bfloat @select_bf16(i1 %cond, bfloat %a, bfloat %b) nounwind {
+; X86-LABEL: select_bf16:
+; X86: # %bb.0:
+; X86-NEXT: testb $1, {{[0-9]+}}(%esp)
+; X86-NEXT: leal {{[0-9]+}}(%esp), %eax
+; X86-NEXT: leal {{[0-9]+}}(%esp), %ecx
+; X86-NEXT: cmovnel %eax, %ecx
+; X86-NEXT: vmovsh {{.*#+}} xmm0 = mem[0],zero,zero,zero,zero,zero,zero,zero
----------------
tfzee wrote:
On x86 the bf16 are handled as stack arguments. So instead of loading both and then conditionally moving, we instead do a conditional move on the pointers to the 2 stack locations and then loading the result. I haven't directly tested it but I would assume that to be better then loading both first.
This combiner only fires if there is only this one use of the load so in other cases where the loads result is reused it wont do this transformation.
In general this test is probably relatively useless since it requires AVX512_FP16 instructions like vmovsh while also being in 32bit mode.(maybe there exists use cases not sure)
https://github.com/llvm/llvm-project/pull/224218
More information about the llvm-commits
mailing list