[all-commits] [llvm/llvm-project] 40d671: [ADT][NFC] Use isEqual for ImmutableSet/Map tree c...

Balázs Benics via All-commits all-commits at lists.llvm.org
Mon Jul 6 02:02:07 PDT 2026


  Branch: refs/heads/main
  Home:   https://github.com/llvm/llvm-project
  Commit: 40d671f85d28a683fd5820b2e0a356f2d1cc58e3
      https://github.com/llvm/llvm-project/commit/40d671f85d28a683fd5820b2e0a356f2d1cc58e3
  Author: Balázs Benics <benicsbalazs at gmail.com>
  Date:   2026-07-06 (Mon, 06 Jul 2026)

  Changed paths:
    M llvm/benchmarks/ImmutableSetIteratorBM.cpp
    M llvm/include/llvm/ADT/ImmutableSet.h

  Log Message:
  -----------
  [ADT][NFC] Use isEqual for ImmutableSet/Map tree canonicalization (#207596)

`ImutAVLFactory::getCanonicalTree` deduplicates a newly built tree
against the trees already in its cache. On a digest collision it
confirmed structural equality with `compareTreeWithSection`, a plain
element-by-element in-order walk that is always linear in the tree size.

`ImutAVLTree` already provides `isEqual`, which performs the same
structural comparison but skips subtrees that are shared by pointer.
These persistent trees are heavily structurally shared -- a tree
produced by `add`/`remove` shares everything but the mutated spine with
its predecessor -- so switching `getCanonicalTree` to `isEqual` reduces
the confirmation from `O(tree size)` to `O(number of differing nodes)`
in the common case, and is never asymptotically worse. The comparison is
exact, so canonicalization behavior is unchanged.

This builds on the recent iterator rewrite in #205552 ("Rewrite
ImmutableSet/Map in-order iterator without per-node state"), which made
in-order traversal and `skipSubTree` ~2x faster. That change is a
constant-factor speedup of the iterator primitive and it deliberately
preserved `skipSubTree`/`isEqual` semantics; this patch is complementary
and algorithmic: it routes canonicalization through the pointer-skipping
`isEqual`, cutting the number of nodes visited rather than the per-node
cost. The two compose -- the remaining traversal here runs on the faster
iterator from #205552.

This confirmation is the dominant cost of `ExplodedGraph`
canonicalization in the Clang Static Analyzer whenever many equivalent
`ProgramState`s are kept live and re-derived on different paths (for
example with node reclamation disabled): each re-derivation of an
already-cached `Environment`/`Store` map paid a full in-order walk to
confirm the match. On several large reproducers the analyzer's
`ImutAVLFactory` node comparisons dropped by ~26-36x with identical
diagnostics and identical `ExplodedGraph` node/step counts. (These
measurements already include #205552, so this speedup is on top of that
iterator win.)

`compareTreeWithSection` had no other callers and is removed.

## Analyzer wall-clock

Measured on the three reproducers from

https://github.com/llvm/llvm-project/issues/105512#issuecomment-2389083480,
best-of-3 wall-clock, `RelWithDebInfo` on an Apple M-series host. "ON"
is the default; "OFF" disables node reclamation (`-analyzer-config
graph-trim-interval=0`), the configuration that exposes the cost:

| Reproducer   | Config | Before   | After   | Speedup |
| ------------ | ------ | -------- | ------- | ------- |
| ruby-unicode | ON     |  2.18 s  | 2.06 s  | ~1.0x   |
| ruby-unicode | OFF    | 11.97 s  | 2.28 s  | ~5.2x   |
| linux-topro  | ON     |  5.26 s  | 5.08 s  | ~1.0x   |
| linux-topro  | OFF    | 11.16 s  | 5.21 s  | ~2.1x   |
| MSP430       | ON     |  2.71 s  | 2.65 s  | ~1.0x   |
| MSP430       | OFF    |  7.58 s  | 2.71 s  | ~2.8x   |

The default (ON) path is unchanged within noise; the OFF path drops from
2.1-5.5x slower than ON to parity.

## Microbenchmark

Add a `CanonicalizeSharedEqual` case to `ImmutableSetIteratorBM`
covering `getCanonicalTree`/`isEqual`. It re-canonicalizes a freshly
built tree that is equal to, and shares subtrees with, a cached tree
(`RelWithDebInfo`, same host):

| Set size `N` | Before    | After   | Speedup |
| ------------ | --------- | ------- | ------- |
| 16           |    212 ns |  186 ns | ~1.1x   |
| 256          |   1354 ns |  347 ns | ~3.9x   |
| 4096         |  21603 ns |  547 ns | ~40x    |
| 65536        | 677240 ns | 1015 ns | ~667x   |

The existing `Iterate`/`Skip`/`CompareSamePos` cases are unchanged; this
does not touch the iterator.

Assisted by: Claude Opus 4.8



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list