skip to content

A forked child reruns the parent's atexit cleanup in a 6-hour nightly reconciliation job, causing intermittent timeouts -- how do you find and fix it?

level: seniorimportance: should knowfreq 30%

answer

  1. The cleanup happened more often than the work
  2. A copied list nobody marked as the parent's
  3. Look at which pid logged the release
  4. The exit path of the child branch
  5. Pid-guard the callback, hard-exit the child

basics

~20 s

The child inherited the parent's atexit registry at fork time and ran it on its way out, releasing a lease the parent still holds. Tag cleanup callbacks with the pid that registered them, and end the child with os._exit().

solid answer

~50 s

`atexit.register` stores callbacks in ordinary interpreter state, so `os.fork()` copies the whole registry into the child. Any child that leaves through the normal exit path -- `sys.exit()`, falling off the end, or an uncaught exception -- runs interpreter shutdown and fires those callbacks a second time, from a process that does not own the resources they release. In a batch job that shows up as a lease or advisory lock being released early, so the next batch blocks on a resource it believes is held and eventually times out; it looks intermittent because it only bites when a child finishes near the moment the parent needs the resource again. Confirm it by logging `os.getpid()` inside each cleanup callback and looking for the cleanup line under a pid that is not the main process. Fix it by ending the child with `os._exit()` after flushing, and defensively by having every cleanup callback record the pid that registered it and return immediately if `os.getpid()` no longer matches.

code

python · 15 lines
python
import atexit, os, sys

OWNER_PID = os.getpid()

def release_lease():
    if os.getpid() != OWNER_PID:
        return                      # an inherited copy in a child: not ours
    print("lease released by", os.getpid())

atexit.register(release_lease)

if os.fork() == 0:
    sys.stdout.flush()
    os._exit(0)
os.wait()

go deeper

for a junior

Recall that os.fork() copies the interpreter's atexit registry into the child, so cleanup the parent registered can run a second time when the child exits normally.

for a middle

Explain the chain: normal exit in the child means interpreter shutdown, shutdown runs the inherited atexit callbacks, and those callbacks release resources the parent still holds. Name os._exit() as the way to skip it.

for a senior

Show the diagnosis, not just the fix -- log os.getpid() from every cleanup callback, notice the release logged under a child's pid and at the wrong moment, then read the child's exit path and any broad except between the fork and it.

for a principal

Decide the policy: cleanup callbacks that are ownership claims must be pid-guarded, fork sites must go through one reviewed helper, and you should be able to justify why raw os.fork() remains in a job that a higher-level process API could run instead.

### Why an inherited registry is a production hazard `atexit.register(func)` appends to a list held in the interpreter. `os.fork()` copies the interpreter's memory, so the child starts life with an identical list. Nothing marks those entries as belonging to the parent, because at registration time there was only one process. When the child later shuts down normally, it runs them in last-in-first-out order exactly as the parent eventually will. The callbacks that get registered in a long-running batch job are precisely the dangerous ones: release a database advisory lock, delete a run-lock or pid file, return a pooled connection, mark a lease free in a coordination store, write a "finished cleanly" marker. Every one of those is a claim of ownership over something exactly one process owns. Run it twice and the second run is a lie. ### Why it presents as an intermittent timeout A nightly reconciliation job that forks a child per batch and runs for six hours has a race with a wide window and a narrow trigger. The child's early release of the lease is only visible when something else tries to acquire it before the parent finishes -- the next batch, a retry, a monitoring probe. Most nights the ordering is harmless; on the nights it is not, some later step blocks on a resource that the coordination store now believes is free (or, worse, has handed to another waiter) and eventually gives up on its acquire timeout. The timeout is reported far from the fork, in a step that has no obvious relationship to it, which is what makes the failure look flaky rather than causal. The tell that separates this from an ordinary lock-contention problem is *count*: the cleanup was performed more times than the job started units of work. If the resource was acquired once and released twice, contention is not the explanation. ### Confirming it The cheapest confirmation is to make the callbacks self-identifying. Print or log `os.getpid()` from inside every cleanup callback, and record the main process's pid once at startup. If a cleanup line appears under a pid that is not the main one, an inherited callback fired in a child and the diagnosis is done. A second confirming signal is ordering: the duplicate cleanup is logged long before the job ends, at the moment a batch child finished, not at shutdown. From there, read the fork site and answer two questions. First, how does the child leave -- does the branch end in `os._exit()`, or does it `sys.exit()`, `return`, or simply fall off the end? Second, is there a broad `except` between the fork and the exit that could swallow a `SystemExit` and let the child continue as a second copy of the program? Those two answers usually explain everything. ### The fix, in layers **Layer one: the child must hard-exit.** Flush anything the child wrote, then `os._exit(0)`, and wrap the child's body so that an escaping exception ends in `os._exit(1)` instead of in the normal exception path. This is the actual fix; the child now cannot reach interpreter shutdown, so no inherited callback can fire. **Layer two: make the callbacks refuse to run in the wrong process.** Record the owning pid when you register, and have the callback compare it with `os.getpid()` and return if they differ. This is a few lines and it converts a class of future bugs -- someone adds a fork somewhere else, a library forks under you -- from silent corruption into a no-op. It is the layer worth arguing for in review, because it survives changes to code you do not own. **Layer three: give the child its own cleanup deliberately.** If the child genuinely has something to clean up, drop the inherited entries with `atexit.unregister` first and register the child's own afterwards. Then the normal exit path in the child would be safe -- though the hard exit is still simpler and does not depend on anyone remembering. ### What not to reach for Do not "fix" it by removing the cleanup callback from the parent. The parent's release is real work; the bug is that a second process performed it. Do not fix it by widening the acquire timeout either -- that hides a correctness bug behind a longer wait, and on a long nightly run it converts a fast failure into a slow one. And do not assume that because the job usually succeeds the race is rare enough to live with: a lease released by a process that does not own it can also mean two processes reconciling the same batch, which is a data problem rather than a scheduling one. ### Version note The mechanics here are stable across Python 3.10 through 3.14. What changed around them is the default that steers people away from raw forking: since 3.14 the `multiprocessing` default start method is `forkserver` on Unix other than macOS, with macOS and Windows on `spawn` and `fork` requested explicitly. A job that has been calling `os.fork()` directly since long before that default moved is exactly the kind of code where this bug is still living.

  • How would you tell this apart from ordinary lock contention between the batches?
    Count the operations. Contention means the resource was acquired and released the same number of times and someone waited; this bug means it was released more times than it was acquired. Logging `os.getpid()` inside the cleanup callback settles it in one run: a release recorded under a pid that is not the main process, at the moment a child finished rather than at shutdown, is an inherited callback firing, not a queueing problem.
  • The child also needs to release something of its own. Does hard-exiting it break that?
    It would if you relied on `atexit` for it, because `os._exit()` skips the registry entirely. Do the child's cleanup explicitly at the end of its body -- or, if you want the registry, call `atexit.unregister` on the inherited callbacks and register the child's own after the fork, then let it exit normally. The hard exit is still preferable because it does not depend on anyone remembering the unregister step.
  • Why is pid-guarding the callbacks worth adding if the child already hard-exits?
    Because the hard exit only protects the fork site you fixed. A future fork elsewhere in the job, or a library that forks under you, reintroduces the same failure silently. A callback that records the pid at registration and returns when `os.getpid()` no longer matches degrades that from resource corruption to a no-op, and it costs two lines in code you do control.

saying these in an interview costs you the question

  • Blames lock contention and raises the acquire timeout
  • Removes the parent's cleanup callback instead of the double run
  • Thinks a fork gives the child an empty atexit registry
  • Only checks the parent's logs, never the child's pid
  • Assumes sys.exit() in the child skips interpreter shutdown
  • Treats an intermittent timeout as unrelated to the fork site

context