[llvm] [AArch64] Remove superfluous bitconvert from bfdot lane pattern (PR #221745)

David Green via llvm-commits llvm-commits at lists.llvm.org
Tue Sep 8 00:33:12 PDT 2026


================
@@ -9206,8 +9206,8 @@ multiclass BaseSIMDThreeSameVectorBF16DotI<bit Q, bit U, string asm,
     def : Pat<(AccumType (int_aarch64_neon_bfdot
                 (AccumType RegType:$Rd), (InputType RegType:$Rn),
                 (InputType (bitconvert
-                  (DupTypes.VT0 (AArch64duplane32 (DupTypes.VT1
-                    (bitconvert (v8bf16 V128:$Rm))), VectorIndexS:$Idx)))))),
+                  (DupTypes.VT0 (AArch64duplane32 (DupTypes.VT1 V128:$Rm),
----------------
davemgreen wrote:

I see, is that where it comes from. I had honestly assumed that they were both (before and after) incorrect for BE, I hadn't considered that it might be this way on purpose. We would usually consider that BE should be correct, but it doesn't necessarily need to be performant (it doesn't have enough users in practice). The better approach would probably have been to disable the pattern for BE, than break it for LE. Considering that we do not generate a lane bfdot instruction from vbfdot_lane_f32 at the moment, that is probably fine and isn't making anything on BE worse in practice.

https://github.com/llvm/llvm-project/pull/221745


More information about the llvm-commits mailing list