What does a forked worker inherit from a preloaded parent, and which of it is unsafe to reuse?
answer
- The descriptor table is duplicated too
- Two writers on one connection
- Listening socket is the intended exception
- Flush before you fork
- Your own generator instance is copied
basics
~20 sThe child gets a duplicate of the parent's memory and of every open file descriptor: connected sockets, database sessions, buffered file objects and any generator state built at import. The listening socket is meant to be shared; connected descriptors and pre-seeded generators are not.
solid answer
~50 s`os.fork()` duplicates the address space and the descriptor table, so anything the preload opened now exists in every worker. A connected socket is the dangerous case: two processes writing the same TCP stream interleave bytes and desynchronise the protocol, and whichever child closes first affects the shared connection. So any client session opened at import must be discarded and reopened per worker. Buffered writes made before the fork are duplicated and flushed twice, which is why you flush before forking. A `random.Random` instance you created yourself is copied wholesale, so every worker produces the identical stream — the module-level `random` functions are the exception, reseeded in the child by a hook the stdlib has registered since 3.7. The listening socket is the one thing you do want inherited: the kernel hands each accepted connection to exactly one worker.
code
python · 15 linesimport os
import random
import sys
sampler = random.Random(20260903) # built once, before the fork
for _ in range(2):
if os.fork() == 0:
# sampler.random() matches in both children; random.random() does not
print(os.getpid(), sampler.random(), random.random())
sys.stdout.flush()
os._exit(0)
os.wait()
os.wait()go deeper
Remember the two-line version: a forked child gets a copy of the parent's memory and of its open file descriptors. Anything opened before the fork exists in every worker afterwards.
Be able to sort inherited state into safe and unsafe: the listening socket is meant to be shared, while connected sockets, client sessions, buffered writes and hand-built generator instances must be recreated per worker. Explain why two writers on one TCP stream break a protocol.
An interviewer expects the failure stories: interleaved query traffic from a pooled connection built at import, duplicated log lines from an unflushed buffer, workers whose retry jitter is identical, and a child deadlocked on a lock held by a thread that no longer exists.
Own the rule that makes this unnecessary to rediscover. Decide what the preload phase is allowed to construct at all — pure code and read-mostly data, never sockets, threads or seeded state — and encode it so no service has to relearn the inherited-descriptor lesson in production.
## What fork actually duplicates `os.fork()` gives the child a **copy-on-write duplicate of the parent's address space** and a **copy of its file-descriptor table**. Two consequences follow, and almost every preloading bug is one of them. 1. First, every Python object the parent had built at import time is visible in the child at the same address, with the same contents. 2. Second, every descriptor the parent had open — sockets, files, pipes, connections to other services — is open in the child too, and both descriptors refer to the *same* underlying kernel object, sharing one file offset and one connection state. Only the process is duplicated; the thing on the other end of the wire is not. ## Connected sockets are the sharp edge If the preload opened a session to a database, a cache or a remote API, every worker now holds a descriptor to that one connection. Both can write, and their bytes interleave in whatever order the kernel accepts them, which desynchronises any stateful wire protocol immediately: one worker reads another worker's response, a prepared statement belongs to nobody, a transaction is half owned. Closing is equally shared — the connection is torn down when the last descriptor closes, so one worker's cleanup can surprise the others, or the peer sees a connection that never seems to go away. The rule is simple and absolute: **connections are per-worker state**. Either: - open them lazily, on first use inside the worker, - or discard and reopen them at the top of the child. Connection pools built at import are the same defect wearing a nicer name. ## The listening socket is the exception, and it is the point A pre-fork server deliberately creates the **listening socket** in the parent, before forking, so that every worker inherits the same descriptor and calls `accept()` on it. Here sharing is correct because the kernel arbitrates: a pending connection is handed to exactly one accepting process. That is the whole trick that lets N workers serve one port with no coordination between them. ## Buffered I/O A buffered file object holds unflushed bytes in user space, and fork duplicates that buffer. If the parent had written to a log or to standard output without flushing, both parent and child hold the same pending bytes and both will eventually write them, so the line appears twice, or a file ends up with interleaved fragments from several workers. **Flush anything buffered immediately before forking.** The mirror-image bug is calling `os._exit()` in a child, which skips interpreter shutdown and therefore skips flushing — output written just before it silently disappears. ## Generator state Pseudo-random generators are inherited state that people forget is state. A `random.Random` instance created during the preload is copied bit for bit, so every worker walks the identical sequence: identical jitter, identical sampling decisions, identical synthetic identifiers. If those values decide when a worker retries or which subset it samples, the workers move in lockstep, which is the opposite of what the jitter was for. - The module-level functions in `random` are the one case handled for you: since 3.7 the module registers an after-fork hook with `os.register_at_fork` that reseeds the module-level generator in the child, so `random.random()` diverges per worker while your own instance does not. - Anything drawing from `os.urandom`, and therefore `secrets` and `uuid4`, is safe regardless, because it reads the operating system on each call rather than carrying state. ## Threads and locks Only the calling thread survives into the child, but the memory of the others does not vanish: a lock another thread held at fork time **stays held forever** in the child, because the thread that would release it does not exist there. A preload that started a background thread, a metrics flusher or an internal pool therefore risks a child that deadlocks on the first use of a library that touches that lock. Start threads after the fork, in the worker. ## Which start method you are on matters Only `fork` inherits any of this. - `spawn` starts a fresh interpreter that re-imports, - and `forkserver` forks from a small, deliberately under-imported server process, so neither child sees connections the parent opened later; `multiprocessing.set_forkserver_preload()` exists precisely to choose what that server does import. In 3.14 the `multiprocessing` default became `forkserver` on Unix other than macOS, with macOS and Windows on `spawn`, so `fork` — and with it the whole inherit-the-preloaded-app model — must now be asked for explicitly.
- Why is it safe for every worker to inherit the listening socket but not a database connection?Because the kernel arbitrates one and not the other. Several processes may block in `accept()` on the same listening socket, and each pending connection is delivered to exactly one of them. A connected socket has no such arbitration: both processes read from and write to a single byte stream, so responses land in the wrong worker and the protocol desynchronises. Listening descriptors are a rendezvous point; connected descriptors are private conversation.
- What happens to data sitting in a buffered file object at the moment of the fork?It is duplicated. Parent and child each hold the same pending bytes and each will eventually write them, so a log line appears twice or a file gets interleaved fragments from several workers. Flush every buffered stream immediately before forking. The opposite failure is calling `os._exit()` in the child, which skips interpreter shutdown and therefore skips flushing, so recent output is lost.
- Why can a background thread started during preload leave a worker deadlocked?Because fork carries only the calling thread into the child while carrying all of the memory. A lock that another thread held at fork time is still marked held in the child, and the thread that would have released it does not exist there, so the first call that needs that lock blocks forever. Start threads, pools and flushers inside the worker after the fork, never during the preload.
saying these in an interview costs you the question
- Says fork gives each worker its own copy of the connection
- Reuses a module-level database session opened at import time
- Believes only one worker may accept on the inherited listening socket
- Assumes every pseudo-random generator is reseeded after a fork
- Forks without flushing and blames duplicate log lines on the logger
- Starts background threads during preload and expects them in the child