[llvm] [LV][AArch64] Support partial reductions of extended compares (PR #212190)

Adam Scott via llvm-commits llvm-commits at lists.llvm.org
Wed Aug 26 22:49:48 PDT 2026


================
@@ -973,3 +973,49 @@ exit:
 !36 = distinct !{!36, !37, !38}
 !37 = !{!"llvm.loop.interleave.count", i32 1}
 !38 = !{!"llvm.loop.vectorize.width", i32 4}
+
+; zext(i1 mask)->i64 summed which should be costed at the compare's operand width and not at i1.
+define i64 @count_matches_i8_i64(ptr %src, i32 %n) {
+; NEON-LABEL: 'count_matches_i8_i64'
+; SVE-LABEL: 'count_matches_i8_i64'
+; SVE:  Cost of 1 for VF 16: EXPRESSION vp<[[VP7:%[0-9]+]]> = ir<%acc> + partial.reduce.add (ir<%cmp> zext to i64)
+; SVE:  Cost of 1 for VF vscale x 16: EXPRESSION vp<[[VP7]]> = ir<%acc> + partial.reduce.add (ir<%cmp> zext to i64)
----------------
as4230 wrote:

The 1 is InputLT.first * TCC_Basic https://github.com/llvm/llvm-project/blob/09ceea09385d09dd60fd7b42d4351fce0cdd90d7/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp#L6736, so it's one register of input for <16 x i8>. The i8 -> i64 case https://github.com/llvm/llvm-project/blob/09ceea09385d09dd60fd7b42d4351fce0cdd90d7/llvm/lib/Target/AArch64/AArch64TargetTransformInfo.cpp#L6770-L6777
has a FIXME saying it should be higher that we should adjust the cost for.


I've also been working on the neon lowering for this so hopefully that will have a entry soon https://github.com/llvm/llvm-project/pull/214636


https://github.com/llvm/llvm-project/pull/212190


More information about the llvm-commits mailing list