[llvm] [AArch64] Use a frame record for non-leaf outlined functions on MachO (PR #213711)
Zhaoxuan Jiang via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 3 23:55:05 PDT 2026
nocchijiang wrote:
I have a downstream patch that does essentially the same thing as this PR. Some background on how we got there: when app teams adopted the Xcode 16 toolchain, the outliner started producing many more frame-creating outlined sequences, and our downstream ld64 fork couldn't fold them. The root cause was that ld64's ICF has no support for folding FDEs, so every FDE-bearing outlined function stayed distinct. My downstream patch works around that in codegen the same way this PR does: emit a frame record so the function gets a compact unwind encoding and no FDE.
>From that experience, a few things are missing from this PR:
1. FP availability is not checked. The outlined sequence can itself read or write FP - we caught real miscompiles from this downstream. My downstream fix gates the frame record on `Candidate::isAvailableInsideSeq(AArch64::FP, TRI)`. Something equivalent is needed here, plus a negative test with an FP read/write inside the repeated sequence.
2. The cost model doesn't reflect the actual size tradeoff. The +4 here is charged unconditionally on MachO, even when the frame record won't be emitted, and it should only apply when the frame record is actually chosen. Moreover, neither the current code nor this PR models the FDE that the `str x30` path forces - each outlined function today carries an unmodeled ~24–32 bytes of `__eh_frame`.
3. A size/performance concern: the extra `mov x29, sp` makes every non-leaf outlined function 4 bytes larger, including startup-path code. Since this trades `__text` size against overall binary size, I think it's worth gating behind a flag so users can choose which to prioritize - minimal `__text` for startup-sensitive binaries, or minimal total size.
One question about the description: it attributes the `__text` win to FDE-bearing functions being unfoldable under ICF, but as far as I can tell that's not the case with current linkers: I've verified LLD folds FDE-bearing functions, and ld-prime does as well. So I'm not sure where the `__text` delta actually comes from; it would be good to understand the mechanism before this lands.
https://github.com/llvm/llvm-project/pull/213711
More information about the llvm-commits
mailing list