[llvm] [DenseMap] Share rehash and grow for relocatable bucket types. NFC (PR #225018)

Fangrui Song via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 01:35:12 PDT 2026


================
@@ -0,0 +1,101 @@
+//===----------------------------------------------------------------------===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#include "llvm/ADT/DenseMap.h"
+#include "llvm/Support/MemAlloc.h"
+#include <cstring>
+
+using namespace llvm;
+using namespace llvm::densemap;
+using namespace llvm::densemap::detail;
+
+// A nonzero FixedSize turns the bucket copy into a couple of stores.
+template <size_t FixedSize, bool InlinePtrHash>
+static void rehashLoop(void *DstBuckets, UsedT *DstUsed, unsigned Mask,
+                       const void *SrcBuckets, const UsedT *SrcUsed,
+                       unsigned SrcNumBuckets, size_t RuntimeSize,
+                       BucketHasher Hasher) {
+  const size_t BucketSize = FixedSize ? FixedSize : RuntimeSize;
+  char *Dst = static_cast<char *>(DstBuckets);
+  const char *Src = static_cast<const char *>(SrcBuckets);
+  forEachUsed(SrcUsed, SrcNumBuckets, [&](unsigned I) {
+    const char *SrcBucket = Src + static_cast<size_t>(I) * BucketSize;
+    unsigned Hash;
+    if constexpr (InlinePtrHash) {
+      void *Key;
+      std::memcpy(&Key, SrcBucket, sizeof(Key));
+      Hash = DenseMapInfo<void *>::getHashValue(Key);
+    } else {
+      Hash = Hasher(SrcBucket);
+    }
+    unsigned BucketNo = Hash & Mask;
+    while (used(DstUsed, BucketNo))
+      BucketNo = (BucketNo + 1) & Mask;
+    std::memcpy(Dst + static_cast<size_t>(BucketNo) * BucketSize, SrcBucket,
+                BucketSize);
+    setUsed(DstUsed, BucketNo);
+  });
+}
+
+template <bool InlinePtrHash>
+static void rehashBySize(void *Dst, UsedT *DstUsed, unsigned Mask,
+                         const void *Src, const UsedT *SrcUsed,
+                         unsigned SrcNumBuckets, size_t BucketSize,
+                         BucketHasher Hasher) {
+  // The bucket sizes of 95% of the grow instantiations in an LLVM build.
+  switch (BucketSize) {
+#define REHASH_CASE(N)                                                         \
+  case N:                                                                      \
+    return rehashLoop<N, InlinePtrHash>(Dst, DstUsed, Mask, Src, SrcUsed,      \
+                                        SrcNumBuckets, BucketSize, Hasher);
+    REHASH_CASE(4)
+    REHASH_CASE(8)
+    REHASH_CASE(12)
+    REHASH_CASE(16)
+    REHASH_CASE(24)
+    REHASH_CASE(32)
+    REHASH_CASE(40)
+    REHASH_CASE(48)
+#undef REHASH_CASE
+  default:
+    return rehashLoop<0, InlinePtrHash>(Dst, DstUsed, Mask, Src, SrcUsed,
+                                        SrcNumBuckets, BucketSize, Hasher);
+  }
----------------
MaskRay wrote:

`x86-64: jump table; aarch64: compare tree, depth <= 4`. 

```
Checked rather than assumed: I lowered the same optimized IR for four targets with llc -O3.

┌─────────┬────────────────────────────────────┬────────────────────────────────────────────────┐
│ target  │   dispatch for the 8-case switch   │                      why                       │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ x86-64  │ 45-entry jump table, one jmp *%rax │ density 8/45 = 18% ≥ 10% (-jump-table-density) │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ RISC-V  │ jump tables                        │ same                                           │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ AArch64 │ compare tree, depth ≤ 4            │ -aarch64-min-jump-table-entries = 10 > 8 cases │
├─────────┼────────────────────────────────────┼────────────────────────────────────────────────┤
│ PPC64   │ compare tree                       │ -ppc-min-jump-table-entries = 64               │
└─────────┴────────────────────────────────────┴────────────────────────────────────────────────┘

So his concern is real on AArch64 — but it's not "a long list"; it's a balanced tree:

    cmp  x24, #23
    b.gt .LBB0_25         ; {24,32,40,48}
    cmp  x24, #11
    b.gt .LBB0_47         ; {12,16}
    cmp  x24, #4
    b.eq .LBB0_87
    cmp  x24, #8
    b.ne .LBB0_119        ; default
    ...                   ; 8-byte body, inlined, follows

Eight instructions in the worst path, executed once per rehash call, and the bodies are still inlined behind it (3764 B for the whole function). The smallest rehash a grow can do is 48 entries (64 buckets at 3/4 load) at ~20 instructions each, so the tree is ≤ 0.8% of the smallest call and ~0.01% of a typical one. Halving the case count removes one tree level — two instructions per call — while the fixed-size body it drops is worth ~10 instructions per entry on that size (that's what the runtime arm measured: +17–39% on the whole fill benchmark). Even the rarest listed sizes (4 B 0.75%, 48 B 0.19% of instantiations) pay back a compare level after five entries.

So: no, don't decrease the cases. Two honest ways to answer him:

- Keep the switch and reply with the AArch64 codegen above plus the arithmetic. The list is a compare tree of depth 4 on targets that don't form the table, and the dispatch is per call, not per entry.
- Take the function-pointer table if he pushes: it is O(1) on every target, measured identical in instructions and within noise in cycles on x86, costs +1.4 KB of text in one TU, and is no more complex than the switch. The only thing it gives up is the inlined mega-function, which the measurements say is worth nothing.

```

https://github.com/llvm/llvm-project/pull/225018


More information about the llvm-commits mailing list