How would you use `os.register_at_fork` to stop forked log-ingest workers from writing into a database connection inherited from the parent?
answer
- The child did not get its own connection
- One socket, two writers, interleaved bytes
- Which of the three phases repairs the copy?
- Politely closing it harms the other process
- Detach the descriptor, then clear the cache
basics
~10 sRegister an after_in_child callback that abandons the inherited handle — detach and close the duplicated descriptor, then clear whatever cached it so the child reconnects on first use. Never call the driver's polite close().
solid answer
~50 s`os.fork()` duplicates the descriptor table, so parent and child hold two descriptors onto **one** TCP connection and one server session. Both write into the same byte stream, and replies land one statement out of step — a worker parses the previous batch's reply and commits an off-by-one boundary, reproducible in only a few of a 340-case regression pack. The repair is an `after_in_child` hook that detaches the socket and closes the descriptor, then clears the module-level or pool-held reference so the next use dials a fresh connection. Crucially the child must not call the connection object's ordinary `close()`: that sends a goodbye or rollback down the shared socket and destroys the parent's session. Pair it with `before=pool_lock.acquire` and `release` in both `after` hooks so the fork never captures a half-updated pool. These hooks fire only for real forks — not for `spawn` or `forkserver` workers.
code
python · 19 linesimport os
import socket
conn, peer = socket.socketpair()
def drop_inherited_connection():
os.close(conn.detach())
os.register_at_fork(after_in_child=drop_inherited_connection)
if os.fork() == 0:
print("child fileno after reset:", conn.fileno(), flush=True)
os._exit(0)
os.wait()
conn.sendall(b"parent still owns this connection")
print("peer received:", peer.recv(64))go deeper
Remember the core fact: a forked child inherits copies of the parent's open descriptors, so both processes are talking over the same socket. Do not create connections before forking and expect each side to have its own.
Explain the mechanics — descriptor table duplication, one shared TCP stream, duplicated driver-side protocol state — and show the after_in_child reset that detaches the descriptor and clears the cached handle rather than closing the connection politely.
Show you have diagnosed this live: interleaved or off-by-one replies, a pid check on every use, one server session fed by several client pids, and a before hook taking the pool lock so the fork never captures a half-updated pool.
Own the standard: which layer is responsible for fork safety, whether the platform mandates a start method so the hazard cannot recur, and how you keep worker initialization correct under fork, forkserver and spawn without duplicating cleanup logic in three places.
## What the child actually inherits `os.fork()` duplicates the address space *and* the file-descriptor table. The child does not get its own connection to the database; it gets a second descriptor referring to the **same** open socket — the same TCP connection, the same kernel send and receive buffers, the same server-side session. The Python object graph is duplicated with it: the driver's connection object, its buffered protocol state, its notion of "the next reply belongs to the query I just sent", and any pool bookkeeping. Two processes now write into one byte stream. In a log-ingest service that forks a worker per shard, the symptom is rarely a clean error. Replies arrive shifted by one — a worker parses the response to the *previous* statement, reads the wrong row count for its batch, and commits a boundary one record off. A 340-case regression pack that exercises the workers reproduces it in a handful of cases and only under the fork start method, which is exactly the profile of a fork-inheritance bug: nondeterministic, load-dependent, and invisible in single-process tests. ## The fix, phase by phase ```python import os import socket conn, peer = socket.socketpair() def drop_inherited_connection(): os.close(conn.detach()) os.register_at_fork(after_in_child=drop_inherited_connection) if os.fork() == 0: print("child fileno after reset:", conn.fileno()) os._exit(0) os.wait() conn.sendall(b"parent still owns this connection") print("peer received:", peer.recv(64)) ``` **`after_in_child` — abandon, do not close politely.** The single most important detail is that the child must *not* call the driver's ordinary `close()` or `disconnect()`. Those methods speak the protocol: they send a goodbye frame, or a rollback, over a socket the parent is still using, and they destroy the parent's session while corrupting the parent's stream on the way out. What the child owns is only its copy of the descriptor, so the correct action is to detach the handle from the Python object and close the descriptor number — or, if even that is unsafe because a background reaper might reuse it, simply drop the reference and let the child exec or exit. Then clear whatever cached the handle (the pool's free list, a module-level `_connection`) so the next use in the child dials a fresh connection of its own. **`before` — quiesce, so the copy is not torn.** If a pool is guarded by a `threading.Lock`, take that lock in `before` and release it in both `after_in_parent` and `after_in_child`. Without it the fork can be taken while another parent thread is halfway through handing out a connection, and the child inherits a half-mutated pool plus a lock that is permanently held by a thread that does not exist in the child. **`after_in_parent` — put the parent back.** Release what `before` acquired. Nothing else: the parent's connection is still valid and must not be reset. ## The traps around it * **Hooks fire only for `os.fork()` and `os.forkpty()`.** `multiprocessing` triggers them only with the fork start method. Since Python 3.14 the default on Linux is `forkserver` (macOS and Windows use `spawn`), and those workers are fresh interpreter startups that never inherit your registration — the inheritance bug disappears, but so does your handler. Never let a hook be the *only* place a worker's connection is established. * **Reconnecting inside the hook is usually wrong.** Dialing a new connection in `after_in_child` runs before the child has done any of its own setup and costs a handshake in every child that may never issue a query. Reset to "no connection" and let the first real use reconnect lazily. * **The hook cannot fail loudly.** An exception in an at-fork callback is reported through `sys.unraisablehook` and swallowed; the fork continues with a child holding the descriptor you meant to drop. Keep the callback small and total. * **It is not only the database.** Any inherited descriptor has the same problem: a metrics or log-shipping socket, a lock file, a pipe to a supervisor. The same `after_in_child` should reset them all. ## How to prove it in production Fork inheritance is confirmable rather than guessable. Compare `os.getpid()` recorded when the connection was created against the current pid at use time — that single check turns an intermittent off-by-one into a clean error, and many drivers ship exactly that guard for this reason. Watching the server's session list is the other half: one session receiving statements from several client pids is the whole diagnosis. The hook then makes the fix automatic instead of relying on every worker entry point to remember to reset.
- Why not simply call the connection object's `close()` inside the child?Because `close()` speaks the protocol. It writes a goodbye or a rollback onto a socket the parent is still using, terminating the parent's session and injecting bytes into its stream. The child owns only its copy of the descriptor, so the correct action is to detach the handle and close the descriptor number — or drop the reference entirely and let the child exit.
- Do these hooks run for workers started with the spawn or forkserver start method?No. Those workers are fresh interpreter startups that never inherited the registration, so the callback never fires — and it does not need to, because nothing was duplicated. The trap is relying on the hook as the *only* place a worker resets or establishes its connection; that code must live in per-worker initialization to survive a start-method change.
- How would you prove in production that a connection is being shared across forked processes?Record `os.getpid()` when the connection is created and compare it against the current pid on every use — that check converts an intermittent off-by-one into a clean, attributable error, and several drivers ship exactly this guard. From the server side, one session receiving statements from multiple client process ids is the same diagnosis seen from the other end.
- Should the `after_in_child` hook open a replacement connection immediately?Usually not. It runs before the child has done any of its own setup, and it pays a handshake in every child, including ones that never issue a query. Reset the state to "no connection" and let the first real use reconnect lazily; that also keeps the hook small, which matters because an exception inside it is swallowed by the unraisable hook.
Forking a process with an open connection is like photocopying a page that lists a phone call already in progress: both copies now think they are the one holding that handset, and whichever speaks first confuses the person on the other end.
saying these in an interview costs you the question
- Calls the driver's normal close() in the forked child
- Assumes each child automatically gets its own connection
- Thinks a duplicated descriptor is an independent socket
- Relies on at-fork hooks under the spawn start method
- Reconnects in the before hook, in the wrong process
- Blames the server for interleaved or shifted replies