[llvm] [Hashing] Replace CityHash mixers with xxh3 (PR #194567)
Fangrui Song via llvm-commits
llvm-commits at lists.llvm.org
Sat May 9 16:15:09 PDT 2026
================
@@ -362,63 +219,39 @@ template <typename T> auto get_hashable_data(const T &value) {
}
}
-/// Helper to store data from a value into a buffer and advance the
-/// pointer into that buffer.
-///
-/// This routine first checks whether there is enough space in the provided
-/// buffer, and if not immediately returns false. If there is space, it
-/// copies the underlying bytes of value into the buffer, advances the
-/// buffer_ptr past the copied bytes, and returns true.
-template <typename T>
-bool store_and_advance(char *&buffer_ptr, char *buffer_end, const T& value,
- size_t offset = 0) {
- size_t store_size = sizeof(value) - offset;
- if (buffer_ptr + store_size > buffer_end)
- return false;
- const char *value_data = reinterpret_cast<const char *>(&value);
- std::memcpy(buffer_ptr, value_data + offset, store_size);
- buffer_ptr += store_size;
- return true;
-}
-
/// Implement the combining of integral values into a hash_code.
///
/// This overload is selected when the value type of the iterator is
/// integral. Rather than computing a hash_code for each object and then
/// combining them, this (as an optimization) directly combines the integers.
+///
+/// xxh3 has no streaming entry point in libLLVMSupport, so the byte stream is
+/// flattened to a buffer and hashed in one shot. A 64-byte on-stack buffer
+/// covers the common cases; longer non-contiguous ranges (the prior chunked
+/// CityHash impl was streaming and never allocated) fall back to the heap.
template <typename InputIteratorT>
hash_code hash_combine_range_impl(InputIteratorT first, InputIteratorT last) {
- const uint64_t seed = get_execution_seed();
- char buffer[64], *buffer_ptr = buffer;
- char *const buffer_end = std::end(buffer);
- while (first != last && store_and_advance(buffer_ptr, buffer_end,
- get_hashable_data(*first)))
- ++first;
- if (first == last)
- return hash_short(buffer, buffer_ptr - buffer, seed);
- assert(buffer_ptr == buffer_end);
-
- hash_state state = state.create(buffer, seed);
- size_t length = 64;
- while (first != last) {
- // Fill up the buffer. We don't clear it, which re-mixes the last round
- // when only a partial 64-byte chunk is left.
- buffer_ptr = buffer;
- while (first != last && store_and_advance(buffer_ptr, buffer_end,
- get_hashable_data(*first)))
- ++first;
-
- // Rotate the buffer if we did a partial fill in order to simulate doing
- // a mix of the last 64-bytes. That is how the algorithm works when we
- // have a contiguous byte sequence, and we want to emulate that here.
- std::rotate(buffer, buffer_ptr, buffer_end);
-
- // Mix this chunk into the current state.
- state.mix(buffer);
- length += buffer_ptr - buffer;
- };
-
- return state.finalize(length);
+ alignas(uint64_t) char stack_buf[64];
+ std::unique_ptr<char[]> heap_buf;
+ char *buf = stack_buf;
+ size_t cap = sizeof(stack_buf);
+ size_t len = 0;
+ for (; first != last; ++first) {
+ auto data = get_hashable_data(*first);
+ if (len + sizeof(data) > cap) {
+ size_t new_cap = cap * 2;
+ while (new_cap < len + sizeof(data))
+ new_cap *= 2;
+ std::unique_ptr<char[]> new_buf(new char[new_cap]);
----------------
MaskRay wrote:
I'll bump the stack buffer size to 256 bytes, then there will be no heap allocation for `clang -c -O1 sqlite3.i`
```
● Stress-test results are conclusive:
┌────────────────────────────────────────────┬─────────────────────┬──────────────────┬──────────┐
│ Workload (sqlite3.i, 1.1MB preprocessed C) │ Iterator-path calls │ Heap allocations │ Max size │
├────────────────────────────────────────────┼─────────────────────┼──────────────────┼──────────┤
│ -O1 │ 23,573 │ 0 │ 64 B │
├────────────────────────────────────────────┼─────────────────────┼──────────────────┼──────────┤
│ -O3 │ 28,829 │ 0 │ 152 B │
├────────────────────────────────────────────┼─────────────────────┼──────────────────┼──────────┤
│ -O2 -flto │ 27,848 │ 0 │ 152 B │
└────────────────────────────────────────────┴─────────────────────┴──────────────────┴──────────┘
Distribution skews hard to [16,32) (≈91% — 2–3 pointer bundles, typical of operand_values()/successors). Tail at [128,256) only at -O3/LTO, peaking at 152 bytes — well under the 256 stack buffer.
For comparison, the original 64-byte stack would have hit the heap on every call ≥64 bytes:
- O1: 1 heap call (the single [64,128) sample)
- O3: ~720 heap calls (566 + 115 + 38)
- O2-LTO: ~709 heap calls
So 256 bytes is demonstrably sufficient for real workloads, and even small TUs would have driven the original 64-byte buffer to the heap at -O3. This data is strong evidence for the PR reply.
```
https://github.com/llvm/llvm-project/pull/194567
More information about the llvm-commits
mailing list