Why can multiprocessing's fork start method deadlock a child of a threaded service?
answer
- Only one thread survives the copy
- The state of the world is copied too
- Whoever would release it is gone
- First log line, then silence
- The fix is a different start method
basics
~20 sForking duplicates only the calling thread but copies every lock exactly as it was. A lock another thread held at fork time arrives in the child locked forever, with no owner left to release it, so the child hangs.
solid answer
~50 sA fork gives the child a copy of the parent's memory but only one thread — the one that called it. Locks are just memory, so a mutex another thread was holding arrives in the child **already locked**, and the thread that would have released it does not exist there. The child blocks forever the first time it touches that subsystem. The usual victims are ubiquitous low-level locks: a logging handler's lock, an allocator or connection-pool lock inside a C library, an internal lock in a database client. The symptom is the worst kind — a worker that never produces output and never raises, so a pool `join()` hangs with no traceback. The fix is to stop forking: use `forkserver` or `spawn`, the defaults on 3.14, or create the pool before any thread starts. `os.register_at_fork()` repairs only state you own.
code
python · 13 linesimport os
import threading
cache_lock = threading.Lock()
os.register_at_fork(
before=cache_lock.acquire,
after_in_parent=cache_lock.release,
after_in_child=cache_lock.release,
)
if __name__ == "__main__":
print("fork hooks registered", cache_lock.locked())go deeper
Recall the headline: forking a program that has threads is unsafe because the child gets only one thread but all the locks. Knowing to prefer the default start method is enough at this level.
Explain why a copied lock can never be released in the child, and name the ordinary subsystems — logging, allocators, client connection pools — where the held lock usually comes from.
Show the diagnosis path for a worker that hangs with no traceback and no exit code, and rank the fixes: change the start method, fork before threads exist, hook the fork for state you own, rebuild connections in an initializer.
Own the policy: forbid forking from request handlers in threaded services, make the start method an explicit, reviewed choice, and treat any library that forks under the hood as a supportability risk when the fleet upgrades.
This is the failure that drove Python's whole retreat from `fork`, and it is worth being able to explain from first principles. ### The mechanism A fork produces a child whose address space is a copy of the parent's, but whose thread table contains exactly one thread: the one that called fork. Every other thread is gone — not stopped, not joined, simply absent. A lock, however, is not a thread. It is a small piece of state in memory saying "held" or "free", plus a wait queue. It gets copied like everything else. So if a background thread happened to hold a lock at the instant of the fork, the child inherits a lock in the held state whose owner does not exist and never will. The first time code in the child tries to acquire it, it blocks forever. Nothing raises; nothing times out. The interval during which this can happen is not a rare window. Locks are everywhere in a real process: the lock inside a logging handler, the internal locks of a C allocator, the mutex protecting a connection pool in a database or HTTP client library, buffer locks inside an I/O layer. A single background heartbeat, metrics or log-flush thread is enough to make the hazard permanent, because the moment of forking is uncorrelated with what that thread is doing. ### What it looks like in production Take a route-optimisation service that runs a background thread emitting progress logs while the request handler forks a pool of workers to score candidate routes. It works for months, then a change to logging or a library upgrade widens the window a little and workers begin to hang — not always, not reproducibly, and typically under load, when the background thread is busiest. The pool's `join()` never returns, the queue never drains, `Process.exitcode` stays `None`, and there is no traceback because the child is not crashing, it is waiting. Chasing it across a three-week release train is miserable precisely because the trigger is timing, not code path. When you suspect it, the useful moves are: check whether the parent is multi-threaded at the moment of the fork; take a stack of the stuck child from outside the process (CPython 3.14's remote debugging support via `sys.remote_exec()`, or a native debugger) and look for a thread parked in a lock acquire during library initialisation; and confirm that the same workload with `spawn` does not hang. A child stuck before it prints anything at all, on its very first log line or first client call, is the signature. ### Why fork is retreating rather than being fixed POSIX only guarantees that a child of a multi-threaded fork may safely call async-signal-safe functions until it execs. Python code is nowhere near that restriction. There is no way for the interpreter to know which library locks exist or how to reinitialise them, so the language has moved the default instead: macOS switched to `spawn` in 3.8 because forking after touching higher-level system frameworks crashes outright; Python 3.12 added a `DeprecationWarning` when the implicit default `fork` was used from a multi-threaded parent; and 3.14 changed the default on Unix other than macOS to `forkserver`. `forkserver` is the structural fix — the process that does the forking is a small, single-threaded server created before your threads exist, so no application thread can ever be holding a lock at fork time. ### The mitigations, ranked First, change the start method: `forkserver` keeps fork's cheap worker creation, `spawn` is the portable choice, and on 3.14 you are already on one of them unless you asked for `fork`. Second, if you must fork, do it *before* starting any threads — create the pool at start-up, keep it, and never create one from a request handler in a threaded server. Third, `os.register_at_fork()` lets you register callbacks that run before the fork, in the parent after it, and in the child after it, which is how you take your own locks around the fork and reinitialise your own state; libraries can use it for the state they own, but you cannot register on behalf of code you did not write. Fourth, re-establish anything connection-shaped in a pool initializer rather than inheriting it — an inherited socket or client handle is a second, independent hazard on the same fork. What does **not** work: adding a timeout and retrying (the lock never becomes free), catching an exception (there is none), or reducing thread count until the race "goes away" (it narrows the window and hides the bug).
- Why does forkserver avoid this problem when fork does not?Workers are forked from a dedicated server process that the runtime starts on first use and keeps single-threaded, not from your application process. Your background threads never exist in the process that performs the fork, so no application lock can be held at fork time. Worker creation is still a fork, so it stays cheap.
- Can os.register_at_fork() make fork safe in general?No. It fixes state you own: acquire your locks in the `before` callback, release them in the parent and reinitialise in the child. It cannot register handlers for locks buried in third-party or C libraries, and it cannot revive threads. It is a targeted mitigation, not a guarantee, which is why the default moved instead.
- How do you get evidence out of a child that is hung rather than crashed?Inspect it from outside: CPython 3.14 supports attaching to a running process through `sys.remote_exec()`, and a native debugger can show a thread parked in a lock acquire. Faulthandler dumps armed in the child before the risky call also help. A stack parked during library initialisation, before any work started, is the tell.
It is like photographing a room where someone is holding the only key, then stepping into the photograph: the door is locked, the key is in the picture, and the person who had it never existed here.
saying these in an interview costs you the question
- Says all threads are duplicated into the child
- Suggests a retry or timeout will clear it
- Blames pickling for a silent worker hang
- Thinks the GIL prevents this class of deadlock
- Claims register_at_fork makes fork universally safe
- Proposes reducing thread count as the fix