[llvm] [Attributor] Drop norecurse when specializing an indirect call closes a cycle (PR #218637)
Larry Meadows via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 01:36:31 PDT 2026
lfmeadow wrote:
You were right, thanks - traced it, and the bad annotation is clang's.
Clang marks every outlined OpenMP region function `norecurse`, in
`emitOutlinedFunctionPrologue` and `emitOutlinedFunctionPrologueAggregate`, and
`CGOpenMPRuntimeGPU` does the same for the GPU parallel wrapper. With nested
parallel regions that is false: the region body is reached through the function
pointer the device runtime is handed, and the body opens the next parallel
region, so `__kmpc_parallel_60` is re-entered while the first call is still live
and both it and the region body occur inside a cycle in the dynamic call graph.
`__kmpc_parallel_60` itself carries no `norecurse` in the module the device LTO
pipeline starts from. `rpo-function-attrs` derives it, correctly, from those
false caller annotations once internalization makes the entry internal; running
`opt -passes=rpo-function-attrs` on the internalized module alone reproduces it.
Specialization then only turned the pre-existing wrong attribute into a visible
direct self-recursive call.
Dropping the three clang annotations fixes the offload reproducer with no
Attributor change, so closing this in favour of #218862.
https://github.com/llvm/llvm-project/pull/218637
More information about the llvm-commits
mailing list