[llvm] [ICP] Introduce a hot function cutoff threshold for ICP (PR #208060)

via llvm-commits llvm-commits at lists.llvm.org
Tue Jul 28 13:47:48 PDT 2026


modiking wrote:

> ICPing on [99%,99.99%] does not infer ~1% additional of executed counts. It means candidates that have counts within percentile [99%,99.99%] are promoted and that actually has no relationship with “~1%”, similar as if we did ICP on [0%,99%] (current default behavior) it would not mean ~99% additional of executed counts, but rather “candidates that have counts within percentile [0%,99%] are promoted” and the amount of those candidates is usually small.

Percentile mapping from counts to "executed cycles" is loose that is true. I'm thinking super high-level where if we correlate the sum of all executed counts as a proxy for performance adding [99%,99.99%] is akin to adding 1% more executed counts for optimization. That being said with ICP it's not at all a 1:1 relationship because ICP candidates are not evenly spread across executed counts.

Mainly I wanted to know more first level details of why performance is increasing to see if a potentially better heuristic can be used.
> My intuition is that there is a lot more variance in call targets counts for SamplePGO and tieing them to block counts may not be the right choice.

Something like the above where ICP candidates are ranked differently.

> The gain from ICP comes from the following two folds but not relevant to executed counts: 1) By promoting virtual dispatch to direct call, it gets rids of two loads that very often have cache misses which significantly stalls the pipeline since CPU does not know what instructions to fetch in terms of the target function. 2) It further enables more optimizations like I mentioned above.

IIRC we still load all the way to the function pointer when doing ICP. There is now an option to compare directly against the vtable ptr (https://github.com/llvm/llvm-project/pull/81442) but that's default off and needs `-fwhole-program-vtables`. I've found that the branch itself is also non-trivial (https://discourse.llvm.org/t/rfc-safer-whole-program-class-hierarchy-analysis/65144/11 showed ~0.5% from effectively removing these branches). So that leads me to believe most of the gains come from (2).

More information on what exactly is happening with (2) is also valuable because like before that can help us refine the ICP heuristic.

Going further here isn't necessary and adding the flag is reasonable. When I see a good test case that has measurable performance boosts I really want to poke at it to get more insights for broader improvements.

https://github.com/llvm/llvm-project/pull/208060


More information about the llvm-commits mailing list