[llvm] [LV][AArch64] Support partial reductions of extended compares (PR #212190)
Adam Scott via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 22:49:48 PDT 2026
================
@@ -973,3 +973,49 @@ exit:
!36 = distinct !{!36, !37, !38}
!37 = !{!"llvm.loop.interleave.count", i32 1}
!38 = !{!"llvm.loop.vectorize.width", i32 4}
+
+; zext(i1 mask)->i64 summed which should be costed at the compare's operand width and not at i1.
+define i64 @count_matches_i8_i64(ptr %src, i32 %n) {
+; NEON-LABEL: 'count_matches_i8_i64'
+; SVE-LABEL: 'count_matches_i8_i64'
+; SVE: Cost of 1 for VF 16: EXPRESSION vp<[[VP7:%[0-9]+]]> = ir<%acc> + partial.reduce.add (ir<%cmp> zext to i64)
+; SVE: Cost of 1 for VF vscale x 16: EXPRESSION vp<[[VP7]]> = ir<%acc> + partial.reduce.add (ir<%cmp> zext to i64)
----------------
as4230 wrote:
The 1 is InputLT.first * TCC_Basic https://github.com/llvm/llvm-project/blob/09ceea09385d09dd60fd7b42d4351fce0cdd90d7/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp#L6736, so it's one register of input for <16 x i8>. The i8 -> i64 case https://github.com/llvm/llvm-project/blob/09ceea09385d09dd60fd7b42d4351fce0cdd90d7/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp#L6770-L6777
has a FIXME saying it should be higher that we should adjust the cost for.
I've also been working on the neon lowering for this so hopefully that will have a entry soon https://github.com/llvm/llvm-project/pull/214636
https://github.com/llvm/llvm-project/pull/212190
More information about the llvm-commits
mailing list