skip to content

Workers forked with os.fork() from an ETL exporter hang on their first log call. Why?

level: seniorimportance: should knowfreq 35%

answer

  1. Parent healthy, children stuck at one point
  2. Load-dependent, so it hides in tests
  3. Something was held at the copy instant
  4. Background threads log more than you think
  5. Get a stack, do not add a timeout

basics

~20 s

A background thread was holding a logging handler's lock at the instant os.fork() ran. The child inherits that lock already locked, and the thread that would release it does not exist there, so the child's first log call blocks forever.

solid answer

~50 s

The child is deadlocked on a lock that was already held when the address space was copied. At a 1,200-jobs-per-minute peak the exporter's metrics and log-shipping threads are logging constantly, so the chance that one is inside a `logging.Handler`'s critical section at the moment of the fork is high — and in the child that thread does not exist, so nothing will ever release the lock. The tells are that the parent stays healthy while children hang at the same point, that the hang is load-dependent, and that a colleague's shared-mutable-default theory does not survive contact with the code. Confirm it by pulling a traceback out of a stuck child — `faulthandler.register` on a signal, or `sys.remote_exec` on Python 3.14 — and you will find it parked in a lock acquire. The fix is to stop forking a threaded process: use `forkserver`, the Python 3.14 Linux default, or `spawn`.

code

python · 20 lines
python
import os
import threading
import time

lock = threading.Lock()

def holder():
    with lock:
        time.sleep(5)

threading.Thread(target=holder, daemon=True).start()
time.sleep(0.2)

pid = os.fork()
if pid == 0:
    # The holder thread does not exist here, so nothing will release the lock.
    print("child acquired:", lock.acquire(timeout=1), flush=True)
    os._exit(0)

os.waitpid(pid, 0)

go deeper

for a junior

Learn the shape of this bug: a child process that hangs right after os.fork() is usually waiting on something the parent's other threads left locked, not doing slow work.

for a middle

Explain the mechanism and its timing - the lock is copied in whatever state it happened to be in, so the failure is probabilistic and its rate tracks how busy the parent's threads are.

for a senior

Show the diagnosis end to end: pull a traceback out of the stuck child, rule out I/O from CPU and connection state, resist the timeout reflex, and change the start method rather than patching around it.

for a principal

Turn it into a platform rule: services that run threads do not fork, the default start method is pinned deliberately, and dev-mode warnings run in CI so the pattern cannot re-enter the codebase unnoticed.

### The symptom The exporter runs on a schedule, forks a handful of worker processes to push batches to the warehouse, and at the 1,200-jobs-per-minute peak some children stop producing output entirely. They are not slow — they never emit their first log line. The parent stays healthy, the children hold their memory, consume no CPU, and sit there until something kills them. Rerun the same job off-peak and everything works, which is what sends teams looking for a data problem instead of a runtime one. A colleague's first theory is usually a shared mutable default — a list or dict created once at function definition and quietly accumulating across calls. It is a good instinct and it is wrong here: a shared default corrupts *results*, it does not stop a process before its first statement executes, and it would misbehave identically off-peak. ### The mechanism `os.fork()` copies the parent's address space but gives the child only one thread of control: the caller. Every other thread of the parent is absent from the child while all of its memory arrives intact and mid-flight. The exporter, like most services, runs background threads — a metrics reporter, a health check, a log shipper. Those threads log, and every emitted record takes the lock inside a `logging.Handler` so two threads cannot interleave a line. If the fork lands while such a thread is inside that critical section, the child receives the handler lock **already acquired**, and the thread that would have released it does not exist there. The child's first `logging` call blocks on `acquire()` and never returns. That explains every part of the symptom: - **Load-dependent:** the busier the parent's threads, the larger the fraction of wall-clock time they spend holding that lock, and the higher the chance a fork lands inside the window. At 1,200 jobs a minute the window is hit regularly; at ten a minute, almost never. - **Only some children:** each fork is an independent roll of the dice. - **Parent unaffected:** in the parent the holder thread is alive and releases normally. - **No traceback and no CPU:** it is a blocking lock acquisition, not an exception and not a loop. The same story can be told with the import machinery's per-module locks, a buffered writer's lock, a connection pool's internal lock, or the memory allocator's — logging is merely the most common because background threads log constantly. ### Confirming it rather than guessing Get a stack out of a stuck child; do not reason from the outside. - Install `faulthandler.register(signal.SIGUSR1)` in the worker entry point, then send that signal to a hung child and read the traceback it dumps. A post-fork deadlock parks in an `acquire` frame with no I/O beneath it. A slow warehouse write parks in a socket send or receive. - On **Python 3.14**, remote debugging (PEP 768) lets you attach to a running process without preparing it in advance, via `sys.remote_exec`. - Corroborate from the OS: the process is sleeping, its CPU time is flat, and it has no open connection to the warehouse — a slow write would have one. Also check whether the `DeprecationWarning` was there all along. Since **Python 3.12** `os.fork()` raises one when the process is multi-threaded, and services routinely discard it because `DeprecationWarning` is silent by default outside `__main__`. ### The fix Not a longer timeout, and not a retry — a deadlocked child never recovers, so a retry only converts a hang into a slower hang. 1. **Stop forking a threaded process.** On Python 3.14 the `multiprocessing` default is already `forkserver` on Linux and other Unix (macOS and Windows use `spawn`), which forks workers from a small single-threaded helper started early rather than from your threaded exporter. If the code calls `os.fork()` directly or asks for `fork` explicitly, that is the line to change. 2. **Or fork before the threads exist.** Build the worker pool during startup, ahead of the metrics and log-shipping threads. This shrinks the window rather than closing it, and it is fragile: whoever adds a thread earlier next quarter reopens the hole. 3. **Or exec immediately in the child**, replacing the address space so the inherited locks are discarded with it. Then make the class of bug visible so it cannot return quietly: run tests and staging with `-X dev` or `PYTHONDEVMODE=1`, and `logging.captureWarnings(True)` so the fork warning reaches your logs. ### What this question is really testing Whether you diagnose from evidence — a real stack out of a real stuck process — instead of from the most recently-read blog post, and whether you know that a load-dependent hang with no traceback in a freshly forked child has a specific, well-understood cause rather than being "flaky".

  • How would you prove the hang is a lock rather than a slow write to the warehouse?
    Get a stack instead of guessing. Call `faulthandler.register(signal.SIGUSR1)` in the worker entry point, signal a stuck child and read the traceback: a post-fork deadlock parks in an `acquire` frame with no I/O beneath it, while a slow write parks in a socket call. Corroborate from the OS — a deadlocked child burns no CPU and holds no connection to the warehouse, and on Python 3.14 you can attach to it with `sys.remote_exec`.
  • If you cannot change the start method right away, what makes the window smaller?
    Fork before the process grows threads — create the worker pool during startup, ahead of the metrics, health-check and log-shipping threads — and keep the child's post-fork path to an immediate `exec` or to code that touches no inherited subsystem. That shrinks the race rather than closing it, and it is fragile: the next thread someone starts earlier reopens it. The durable fix is not forking a multi-threaded process.
  • Why would a retry with a longer timeout make this worse?
    Because the child is not slow, it is stopped: no thread in it will ever release that lock, so no amount of waiting changes the outcome. A retry converts a fast, visible failure into a long stall that ties up scheduler slots and delays the batch, and it hides the load correlation that would otherwise point straight at the fork. Fail the worker fast and fix the fork.

The fork photographs the building at one instant and hands the child a copy in which one storeroom is locked and the only person with the key was never copied across.

saying these in an interview costs you the question

  • Blames slow warehouse writes without getting a stack
  • Expects the child to recover once the parent proceeds
  • Proposes a longer timeout or a retry as the fix
  • Thinks the child re-runs the parent's logging thread
  • Dismisses a load-dependent hang as flaky infrastructure
  • Recommends fork for speed regardless of threads

context