skip to content

After os.fork(), does a change the child makes to a Python list reach the parent?

level: juniorimportance: should knowfreq 40%

answer

  1. The child starts from a snapshot
  2. Copy-on-write, not shared memory
  3. First write gets a private page
  4. Nothing propagates in either direction

basics

~20 s

No. os.fork() gives the child a copy-on-write copy of the parent's address space, so the child sees the parent's objects exactly as they were at fork time, but every write it makes stays private to it.

solid answer

~50 s

`os.fork()` duplicates the calling process's address space. The kernel does not copy it eagerly: it maps the same physical pages into both processes read-only, and the first write from either side traps so the writer gets a private copy of that page. So the child starts with the parent's entire object graph at the same addresses — every module, global and preloaded table — with no re-import and no unpickling, which is what makes pre-fork worker pools attractive. But a `list.append` in the child writes into the list, that write copies the page, and the parent's list is untouched. The relationship is symmetric: anything the parent builds after the fork is equally invisible to the child. Real sharing needs an explicit channel — a pipe, a `multiprocessing.Queue`, a socket, a file, or a shared-memory block — never a module-level global.

code

python · 11 lines
python
import os, sys

rows = ["LH441", "BA912"]
pid = os.fork()
if pid == 0:
    rows.append("AF017")
    print("child:", len(rows))
    sys.stdout.flush()
    os._exit(0)
os.waitpid(pid, 0)
print("parent:", len(rows))

go deeper

for a junior

Be ready to state plainly that os.fork() gives the child a private copy of memory: both sides begin from identical objects, and every change made after that instant stays in the process that made it.

for a middle

Explain the mechanism, not just the rule: pages are shared read-only and copied on the first write, so a child's list append quietly produces a private page instead of a shared update.

for a senior

Show that you design for it. Results leave a forked worker through a pipe, queue, socket, shared-memory block or file, and per-worker caches are budgeted as per-worker memory rather than assumed to be shared.

for a principal

Own the choice of model. Argue when a pre-fork pool earns its keep against spawn, threads or separate services, and account for everything the snapshot silently inherits alongside the data.

## What `os.fork()` actually duplicates `os.fork()` is a thin wrapper over the Unix system call of the same name. The kernel creates a second process whose address space is logically a full copy of the caller's, but nothing is physically copied up front. Instead both processes' page tables are pointed at the same physical frames and those frames are marked read-only. The first *write* to such a page traps into the kernel, which allocates a fresh frame, copies the page (typically 4 KB on x86-64 Linux, 16 KB on Apple silicon) and points the writer at the private copy. That is copy-on-write, and it is why forking a process holding a gigabyte of parsed data returns in microseconds. For a Python program this is unusually convenient. The child wakes up inside the same interpreter state as the parent: every entry in `sys.modules`, every module-level global, every object your startup code built, all at the same addresses, with no import cost and no deserialization. A service that parses one large table at startup and then forks sixteen workers is exploiting exactly this. ## Why the child's mutation stays in the child Say the parent built `rows = ['LH441', 'BA912']` and the child calls `rows.append('AF017')`. The append writes into the list object's internal array of pointers — and, if the array has to grow, into a freshly allocated block plus the list header. Those are writes to copy-on-write pages, so the kernel hands the *child* private copies and the child's list now describes three elements. The parent's list object still lives on the original page and still describes two. Nothing in the mechanism propagates a change; there is no shared mapping and no synchronization to make one visible. The symmetry matters as much as the rule. Configuration the parent reloads *after* forking is invisible to the children, which is why signalling workers to reload usually means replacing them rather than mutating the parent. The Python-level summary is the one an interviewer wants stated plainly: **after `os.fork()` the two processes share nothing observable at the object level.** They start from identical values, and from that instant they diverge. ## What genuinely is shared One category really is shared, because it does not live in the copied address space at all: open file descriptors. The descriptor numbers are duplicated, but they refer to the same kernel open-file description, so the read/write *offset* is shared. Two forked children writing to the same inherited descriptor advance one common offset and interleave, rather than overwriting each other from position zero. That is a genuinely different behaviour from the object graph, and it is worth being able to name the difference. Sockets, database connections and locks inherited across a fork carry their own hazards, which are a topic of their own. ## How to get results back Because nothing propagates, every design that computes something in a worker needs an explicit channel: - a pipe or `multiprocessing.Queue` for structured results; - a socket, when the worker is answering a request anyway; - a file or a database, when the result must outlive the worker; - an `mmap.mmap` region or a shared-memory block, when the result is a large buffer and copying it would dominate; - the exit status, for a single small integer. The classic bug is a worker that computes an expensive value, stores it in a module-level `CACHE` dict and expects the pool to benefit. Sixteen workers produce sixteen private caches, each warming independently, and the memory cost is sixteen times what the author budgeted. ## Where the cheapness ends The snapshot is free at fork time, but it does not stay free. CPython writes to an object's header whenever a reference to it is taken or dropped, and page-granularity dirt tracking turns each of those tiny writes into a private copy of a whole page. A worker that only *reads* the parent's preloaded table will therefore still convert most of it into private memory over its lifetime. Understanding the snapshot semantics is the junior half of this topic; understanding why the sharing decays is the next layer, and it is the reason people reach for tricks like freezing the heap before forking. ## The check to remember If a design depends on two processes seeing each other's Python objects, `os.fork()` is the wrong tool. Fork gives you a fast start from a common state, not a shared heap. Threads give you a shared heap; forked processes give you isolation with a cheap warm start, and the cost of that isolation shows up later as memory rather than at the moment of the call.

  • Which resources does a forked child genuinely share with its parent rather than copy?
    Open file descriptors. The numbers are duplicated but they point at the same kernel open-file description, so the read/write offset is shared and two processes writing to it interleave instead of overwriting. The address space, by contrast, is copied lazily and diverges on first write, so no Python object is shared.
  • How do you get a result computed in a forked child back to the parent?
    Through an explicit channel: a pipe or `multiprocessing.Queue` for structured values, a socket if the worker is already answering a request, a file or database if it must outlive the worker, a shared-memory block or `mmap.mmap` region for a large buffer, or the exit status for one small integer. Writing to a module global reaches nobody.

Fork is a photocopy of a whiteboard, not a window onto it: both rooms start with the same diagram, and every mark either room adds afterwards is invisible to the other.

saying these in an interview costs you the question

  • Says parent and child share objects after os.fork()
  • Thinks copy-on-write means writes propagate both ways
  • Expects the child to see config the parent loads later
  • Claims fork eagerly copies all memory and is slow
  • Passes results back through a module-level global

context