[libc-commits] [libc] [libc][MVE] implement MVE accelerated memcmp (PR #224857)

Schrodinger ZHU Yifan via libc-commits libc-commits at lists.llvm.org
Sat Sep 19 17:45:06 PDT 2026


SchrodingerZhu wrote:

The codegen is not quite good if I use `cpp::simd` solely, especially when there is no `where` expression at all. However, I can try something like:
```
#include "src/__support/CPP/bit.h"
#include "src/__support/CPP/simd.h"
#include "src/__support/math_extras.h"
#include <arm_mve.h>
using namespace LIBC_NAMESPACE;
extern "C" int test_memcmp(const unsigned char *p1, const unsigned char *p2,
                           size_t count) {
  using Vec = cpp::simd<uint8_t, 16>;
  uintptr_t addr1 = cpp::bit_cast<uintptr_t>(p1);
  uintptr_t addr2 = cpp::bit_cast<uintptr_t>(p2);
  while (count != 0) {
    mve_pred16_t pred = vctp8q(count);
    auto active = cpp::bit_cast<cpp::simd<bool, 16>>(pred);
    auto a = cpp::load_masked<Vec>(
        active, cpp::bit_cast<const uint8_t *>(addr1), Vec{});
    auto b = cpp::load_masked<Vec>(
        active, cpp::bit_cast<const uint8_t *>(addr2), Vec{});
    unsigned mismatches = vcmpneq_m_u8(cpp::bit_cast<uint8x16_t>(a),
                                       cpp::bit_cast<uint8x16_t>(b), pred);
    if (mismatches != 0) {
      auto offset = cpp::countr_zero(mismatches);
      return int(cpp::bit_cast<const uint8_t *>(addr1)[offset]) -
             int(cpp::bit_cast<const uint8_t *>(addr2)[offset]);
    }
    addr1 += 16;
    addr2 += 16;
    if (sub_overflow(count, size_t{16}, count))
      break;
  }
  return 0;
}
```
which has a stack adjustment overhead over the code the this patch. My concern is that we will need to use `vcmpneq_m_u8` and `vctp8q` additional bit cast, so it is not making the logic cleaner.

https://github.com/llvm/llvm-project/pull/224857


More information about the libc-commits mailing list