[llvm] [AArch64] Cache repeated subtarget queries (NFC) (PR #222637)

Cullen Rhodes via llvm-commits llvm-commits at lists.llvm.org
Thu Sep 10 07:46:08 PDT 2026


================
@@ -436,6 +439,10 @@ AArch64TargetMachine::~AArch64TargetMachine() = default;
 
 const AArch64Subtarget *
 AArch64TargetMachine::getSubtargetImpl(const Function &F) const {
+  AttributeSet FnAttrs = F.getAttributes().getFnAttrs();
+  if (LastSubtarget && LastSubtargetAttrs == FnAttrs)
+    return LastSubtarget;
----------------
c-rhodes wrote:

it's a fair point, but I think the issue with that idea is not all function attributes are relevant to the subtarget, for example:
```
F1 { nounwind, target-cpu="generic" }
F2 { noinline, target-cpu="generic" }
```

nounwind/noinline are irrelevant here and the current approach takes that into account by not making them part of the `SubtargetMap` key, so the `LastSubtargetAttrs == FnAttrs` check will fail but it can at least reuse the subtarget after key construction.

I did try `mutable DenseMap<AttributeSet, std::unique_ptr<AArch64Subtarget>> SubtargetMap;` anyway but the results aren't great:
```
(venv) ➜  llvm-test-suite python compare.py build-compiletime-O0-g-baseline-87d9c3c24ff5 build-compiletime-O0-g-change-87d9c3c24ff5
                 instructions:u                 diff
                            old           new
7zip               131480854596  132043044022  0.43%
Bullet              57850566275   58151974077  0.52%
ClamAV              13119008079   13103945440 -0.11%
SPASS               13023144537   12972696119 -0.39%
consumer-typeset    11212639525   11205768950 -0.06%
kimwitu++           21194783161   21185996557 -0.04%
lencod              11990572480   11979059901 -0.10%
mafft                6416569903    6412749747 -0.06%
sqlite3              4238008011    4225284823 -0.30%
tramp3d-v4          18342632218   18224765870 -0.64%
geomean             16851905345   16839096463 -0.08%
```

there's regressions in 7zip/Bullet, likely explanation being these are the workloads with the most TUs (https://github.com/c-rhodes/llvm-tools/blob/8119159d5b5161bbdedfd2ad061ada9e3a028260/compile-time/gisel-irtranslator-stats/data/llvmorg-23.1.0-rc3/aarch64-O0-g/profile.md)

so the 2-level cache does seems to work quite well.

https://github.com/llvm/llvm-project/pull/222637


More information about the llvm-commits mailing list