[libc-commits] [libc] [libc][MVE] implement MVE accelerated memcmp (PR #224857)
Schrodinger ZHU Yifan via libc-commits
libc-commits at lists.llvm.org
Sat Sep 19 17:45:06 PDT 2026
SchrodingerZhu wrote:
The codegen is not quite good if I use `cpp::simd` solely, especially when there is no `where` expression at all. However, I can try something like:
```
#include "src/__support/CPP/bit.h"
#include "src/__support/CPP/simd.h"
#include "src/__support/math_extras.h"
#include <arm_mve.h>
using namespace LIBC_NAMESPACE;
extern "C" int test_memcmp(const unsigned char *p1, const unsigned char *p2,
size_t count) {
using Vec = cpp::simd<uint8_t, 16>;
uintptr_t addr1 = cpp::bit_cast<uintptr_t>(p1);
uintptr_t addr2 = cpp::bit_cast<uintptr_t>(p2);
while (count != 0) {
mve_pred16_t pred = vctp8q(count);
auto active = cpp::bit_cast<cpp::simd<bool, 16>>(pred);
auto a = cpp::load_masked<Vec>(
active, cpp::bit_cast<const uint8_t *>(addr1), Vec{});
auto b = cpp::load_masked<Vec>(
active, cpp::bit_cast<const uint8_t *>(addr2), Vec{});
unsigned mismatches = vcmpneq_m_u8(cpp::bit_cast<uint8x16_t>(a),
cpp::bit_cast<uint8x16_t>(b), pred);
if (mismatches != 0) {
auto offset = cpp::countr_zero(mismatches);
return int(cpp::bit_cast<const uint8_t *>(addr1)[offset]) -
int(cpp::bit_cast<const uint8_t *>(addr2)[offset]);
}
addr1 += 16;
addr2 += 16;
if (sub_overflow(count, size_t{16}, count))
break;
}
return 0;
}
```
which has a stack adjustment overhead over the code the this patch. My concern is that we will need to use `vcmpneq_m_u8` and `vctp8q` additional bit cast, so it is not making the logic cleaner.
https://github.com/llvm/llvm-project/pull/224857
More information about the libc-commits
mailing list