[llvm] [Attributor] Skip the dead-internal-function walk for AAs that will not be updated (PR #227187)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 28 21:06:46 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-llvm-transforms

Author: Farid Zakaria (fzakaria)

<details>
<summary>Changes</summary>

AAIsDeadFunction::initialize calls isAssumedDeadInternalFunction, which runs checkForAllCallSites and so creates an AAIsDead for every internal caller, each of which repeats the walk. Initialization therefore recurses through the whole cone of internal callers.

When the AA will not be updated (outside of deduction, or for a function the Attributor is not run on), getOrCreateAAFor fixes it pessimistically right after initialize returns, discarding the walk's result. Skip the walk in that case and assume the entry block live.

This is the dominant cost of the CGSCC OpenMPOpt pass, whose cleanupIR queries the liveness of callers outside the current SCC on every SCC, giving (#SCCs) x (depth of the internal caller chain). #<!-- -->222226 removed the callee-seeding half of this cost; this removes the walk itself.

cleanup-no-seeding.ll drops a debug check that only the removed walk printed.

### Synthetic benchmark
A chain of n static noinline functions with one `#pragma omp parallel` at the top, compiled to IR with `-fopenmp`, then `opt -passes='default<O2>'` (Release+assertions, X86):

| n | before | after |
|---|---|---|
| 1000 | 0.57s | 0.27s |
| 4000 | 2.42s | 1.15s |
| 8000 | 5.04s | 2.38s |

At n=8000, OpenMPOptCGSCC alone goes from 2.49s to 0.18s. The optimized IR is unchanged on the repro and on the inputs of `llvm/test/Transforms/OpenMP`.

### Additional Context
I profiled the compilers of a large production C++ build that compiles all C/C++ with `-fopenmp` at Meta:
- **Overall:** OpenMPOpt is ~0.17% of all compiler cycles and shows up in ~1.2% of compiles.
- **Long compiles:** in compiles that had been running for 256s or longer, stacks under OpenMPOpt are ~14% of samples, versus ~0.19% across all compiles. They are spread across dozens of distinct translation units.
- **Where the time goes:** about 78% of the OpenMPOpt time is in `cleanupIR` / `identifyDeadInternalFunctions` / `checkForAllCallSites` / `getOrCreateAAFor`, and ~20% is tearing down the per-SCC Attributor state.
- For reference, before #<!-- -->222226 the same repro took 27.3s (clang-21); with both changes it's 2.9s.

---
Full diff: https://github.com/llvm/llvm-project/pull/227187.diff


2 Files Affected:

- (modified) llvm/lib/Transforms/IPO/AttributorAttributes.cpp (+6-1) 
- (modified) llvm/test/Transforms/Attributor/cleanup-no-seeding.ll (-1) 


``````````diff
The server is unavailable at this time. Please wait a few minutes before you try again.
``````````

</details>


https://github.com/llvm/llvm-project/pull/227187


More information about the llvm-commits mailing list