[llvm] [DenseMap] Share rehash and grow for relocatable bucket types. NFC (PR #225018)
Fangrui Song via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 01:35:12 PDT 2026
================
@@ -0,0 +1,101 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "llvm/ADT/DenseMap.h"
+#include "llvm/Support/MemAlloc.h"
+#include <cstring>
+
+using namespace llvm;
+using namespace llvm::densemap;
+using namespace llvm::densemap::detail;
+
+// A nonzero FixedSize turns the bucket copy into a couple of stores.
+template <size_t FixedSize, bool InlinePtrHash>
+static void rehashLoop(void *DstBuckets, UsedT *DstUsed, unsigned Mask,
+ const void *SrcBuckets, const UsedT *SrcUsed,
+ unsigned SrcNumBuckets, size_t RuntimeSize,
+ BucketHasher Hasher) {
+ const size_t BucketSize = FixedSize ? FixedSize : RuntimeSize;
+ char *Dst = static_cast<char *>(DstBuckets);
+ const char *Src = static_cast<const char *>(SrcBuckets);
+ forEachUsed(SrcUsed, SrcNumBuckets, [&](unsigned I) {
+ const char *SrcBucket = Src + static_cast<size_t>(I) * BucketSize;
+ unsigned Hash;
+ if constexpr (InlinePtrHash) {
+ void *Key;
+ std::memcpy(&Key, SrcBucket, sizeof(Key));
+ Hash = DenseMapInfo<void *>::getHashValue(Key);
+ } else {
+ Hash = Hasher(SrcBucket);
+ }
+ unsigned BucketNo = Hash & Mask;
+ while (used(DstUsed, BucketNo))
+ BucketNo = (BucketNo + 1) & Mask;
+ std::memcpy(Dst + static_cast<size_t>(BucketNo) * BucketSize, SrcBucket,
+ BucketSize);
+ setUsed(DstUsed, BucketNo);
+ });
+}
+
+template <bool InlinePtrHash>
+static void rehashBySize(void *Dst, UsedT *DstUsed, unsigned Mask,
+ const void *Src, const UsedT *SrcUsed,
+ unsigned SrcNumBuckets, size_t BucketSize,
+ BucketHasher Hasher) {
+ // The bucket sizes of 95% of the grow instantiations in an LLVM build.
+ switch (BucketSize) {
+#define REHASH_CASE(N) \
+ case N: \
+ return rehashLoop<N, InlinePtrHash>(Dst, DstUsed, Mask, Src, SrcUsed, \
+ SrcNumBuckets, BucketSize, Hasher);
+ REHASH_CASE(4)
+ REHASH_CASE(8)
+ REHASH_CASE(12)
+ REHASH_CASE(16)
+ REHASH_CASE(24)
+ REHASH_CASE(32)
+ REHASH_CASE(40)
+ REHASH_CASE(48)
+#undef REHASH_CASE
+ default:
+ return rehashLoop<0, InlinePtrHash>(Dst, DstUsed, Mask, Src, SrcUsed,
+ SrcNumBuckets, BucketSize, Hasher);
+ }
----------------
MaskRay wrote:
`x86-64: jump table; aarch64: compare tree, depth <= 4`.
```
Checked rather than assumed: I lowered the same optimized IR for four targets with llc -O3.
┌─────────┬────────────────────────────────────┬────────────────────────────────────────────────┐
│ target │ dispatch for the 8-case switch │ why │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ x86-64 │ 45-entry jump table, one jmp *%rax │ density 8/45 = 18% ≥ 10% (-jump-table-density) │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ RISC-V │ jump tables │ same │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ AArch64 │ compare tree, depth ≤ 4 │ -aarch64-min-jump-table-entries = 10 > 8 cases │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ PPC64 │ compare tree │ -ppc-min-jump-table-entries = 64 │
└─────────┴────────────────────────────────────┴────────────────────────────────────────────────┘
So his concern is real on AArch64 — but it's not "a long list"; it's a balanced tree:
cmp x24, #23
b.gt .LBB0_25 ; {24,32,40,48}
cmp x24, #11
b.gt .LBB0_47 ; {12,16}
cmp x24, #4
b.eq .LBB0_87
cmp x24, #8
b.ne .LBB0_119 ; default
... ; 8-byte body, inlined, follows
Eight instructions in the worst path, executed once per rehash call, and the bodies are still inlined behind it (3764 B for the whole function). The smallest rehash a grow can do is 48 entries (64 buckets at 3/4 load) at ~20 instructions each, so the tree is ≤ 0.8% of the smallest call and ~0.01% of a typical one. Halving the case count removes one tree level — two instructions per call — while the fixed-size body it drops is worth ~10 instructions per entry on that size (that's what the runtime arm measured: +17–39% on the whole fill benchmark). Even the rarest listed sizes (4 B 0.75%, 48 B 0.19% of instantiations) pay back a compare level after five entries.
So: no, don't decrease the cases. Two honest ways to answer him:
- Keep the switch and reply with the AArch64 codegen above plus the arithmetic. The list is a compare tree of depth 4 on targets that don't form the table, and the dispatch is per call, not per entry.
- Take the function-pointer table if he pushes: it is O(1) on every target, measured identical in instructions and within noise in cycles on x86, costs +1.4 KB of text in one TU, and is no more complex than the switch. The only thing it gives up is the inlined mega-function, which the measurements say is worth nothing.
```
https://github.com/llvm/llvm-project/pull/225018
More information about the llvm-commits
mailing list