[llvm] [OpenMPOpt] Bound indirect-call specialization instead of paying for all of it (PR #219323)

Larry Meadows via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 5 18:06:17 PDT 2026


================
@@ -5830,6 +5911,20 @@ AAFoldRuntimeCall &AAFoldRuntimeCall::createForPosition(const IRPosition &IRP,
   return *AA;
 }
 
+/// Bound the if-cascade AAIndirectCallInfo builds for an indirect call. Device
+/// code routes many parallel regions through a single runtime dispatcher, so a
+/// call site there can see every outlined region in the module; specializing
+/// all of them costs more in code size and compile time than the direct calls
+/// are worth.
+static constexpr unsigned MaxIndirectCallSpecializations = 3;
+
+static bool shouldSpecializeIndirectCallee(Attributor &,
+                                           const AbstractAttribute &,
+                                           CallBase &, Function &,
+                                           unsigned NumAssumedCallees) {
+  return NumAssumedCallees <= MaxIndirectCallSpecializations;
----------------
lfmeadow wrote:

You're right, and the name is the problem. `NumAssumedCallees` is
`AssumedCallees.size()` for every callee in that loop, so this is a threshold on
the call site, not a per-callee cap: four callees gets you zero specializations,
not three. I'll rename it to `MaxCalleesForSpecialization` and say so in the
comment and the test.

I'd rather keep the threshold than specialize the first K. Skipped callees keep
the fallback indirect call, so a real cap pays for K cascade levels *and* the
indirect call, and the K it picks are whatever order `AssumedCallees` came in.
For the dispatcher case — one site seeing every region in the module — that
looks worse than leaving it indirect. If your downstream case has a few hot
callees among many, that argues for ordering by profile, which I'd do
separately.

On the double inlining, I haven't reproduced it. `__kmpc_parallel_60` is
always_inline and invokes the region pointer on the serialized, SPMD and
inactive paths, so several sites through that pointer in one kernel is expected,
but I don't see why *suppressing* specialization would cause more inlining
rather than less. I'll measure and report back. If you have the downstream
reproducer, that would save me guessing at the shape.

https://github.com/llvm/llvm-project/pull/219323


More information about the llvm-commits mailing list