[llvm] [lit] Migrate lit to ProcessPoolExecutor (PR #202681)

via llvm-commits llvm-commits at lists.llvm.org
Thu Jun 25 20:46:04 PDT 2026


prasoon054 wrote:

Thanks for the detailed report. That was really helpful to identify the root cause for this.

The `spawn_main` entries in `ps` confirm workers are already on the `spawn` start method, which is the macOS default since Python 3.8. This means https://github.com/python/cpython/issues/84559 isn't the cause here. The `--multiprocessing-fork` flag in the command line is just an internal python naming artifact.

The hang is actually due to a lock-order inversion inside CPython's `ProcessPoolExecutor.join_executor_internals()` (python <= 3.11). You can see the problematic sequence directly in [join_executor_internals()](https://github.com/python/cpython/blob/79f66145c028666006fa1f76513db2846a148807/Lib/concurrent/futures/process.py#L561).

On macOS, `Queue.join_thread()` requires worker processes to be joined first, but CPython's implementation calls `p.join()` after it. The feeder thread blocks waiting for workers to drain the pipe, and the workers wait on `call_queue.get()` waiting for sentinels that the feeder can never deliver. 

Our abort path (`_abort_executors`) already calls `ex._call_queue.cancel_join_thread()` to avoid exactly this scenario. The success path was just missing the same call.

I have already sent a fix PR for review: https://github.com/llvm/llvm-project/pull/205961 Can you confirm whether this resolves the hang on your bot?

https://github.com/llvm/llvm-project/pull/202681


More information about the llvm-commits mailing list