skip to content

Why does a forked Python worker's memory climb even when it only reads inherited data?

level: seniorimportance: should knowfreq 30%

answer

  1. Copy-on-write only survives if nobody writes
  2. Reading an object still writes to it
  3. The count lives inside the object
  4. The collector touches every tracked container
  5. Freeze before forking, or store bytes

basics

~20 s

Copy-on-write shares pages until something writes to them, and CPython writes constantly: touching any object updates its reference count, and the cyclic collector writes to object headers as it traverses. Those writes dirty pages, so read-only access privately copies the heap anyway.

solid answer

~50 s

`fork` gives the child copy-on-write pages, so in theory a 2.4 GB dataset costs nothing extra per worker. In practice CPython stores each object's reference count inside the object itself, and merely reading an object increments and decrements that count — a write to the page holding it. The cyclic garbage collector makes it worse by walking every tracked container and writing to its GC header during a collection. Both dirty pages the kernel then copies, so a worker that never mutates anything still ends up with a largely private copy of the heap. The mitigations are `gc.freeze()` after loading and before forking, which moves existing objects into a permanent generation the collector skips, disabling the collector in short-lived children, and above all keeping bulk data out of Python objects — in one large buffer, a `memoryview`-backed block or a memory-mapped file, where the pages carry no reference counts at all.

code

python · 13 lines
python
import gc, os

corpus = [{"ticket": i, "tags": ["triage"]} for i in range(300_000)]

gc.collect()
gc.freeze()
print("objects moved to the permanent generation:", gc.get_freeze_count())

if os.fork() == 0:
    gc.disable()
    print("child reads", len(corpus), "tickets")
    os._exit(0)
os.wait()

go deeper

for a junior

Recall that forked children share memory only until something writes, and that in CPython even reading an object writes to its reference count. That single fact explains most surprise memory growth in worker processes.

for a middle

Explain the two write sources — reference counts stored inside each object, and the cyclic collector writing headers as it traverses — and describe what gc.freeze() before forking removes and what it leaves behind.

for a senior

Diagnose it in production: measure private versus shared resident memory rather than Python allocations, freeze and collect before forking, disable the collector in short-lived children, and move bulk data into one buffer or a memory-mapped file.

for a principal

Own the capacity model. Decide whether per-worker memory scales with worker count or stays flat, price that against throughput, and note that 3.14's forkserver default removes inheritance as a strategy — so the durable architecture shares an explicit block rather than betting on copy-on-write.

## The promise and the leak `fork` is cheap because the kernel does not copy the parent's memory; it marks the pages shared and read-only and copies a page only when someone writes to it. The appealing conclusion is that you can load a 2.4 GB working set once in a parent process, fork eight workers, and pay for it once. Watch the resident memory of a real triage service doing this and you see something else: each worker's RSS climbs steadily over the first minutes and settles at a large fraction of the parent's, even though no worker ever mutates the data. ## Why reading is a write in CPython CPython manages object lifetime with reference counting, and the count lives in the object's own header, right next to its data. Every time a name is bound to an object, an object is pushed on the evaluation stack, or a container is iterated, that count is incremented and later decremented. From the kernel's point of view those are ordinary writes to ordinary pages. Reading one dictionary out of a list of millions touches the list's page, the dictionary's page, the pages of its keys and values, and each of those touches converts a shared page into a private copy. Pages are typically 4 KB, and Python's small objects are scattered across them, so touching a modest fraction of the objects can dirty most of the pages. That is the mechanism behind the slow climb: it is not a leak, it is copy-on-write doing exactly what it promises, applied to a runtime that writes everywhere. ## The garbage collector makes it systematic Reference counting alone cannot free cycles, so CPython also runs a generational cyclic collector. When it runs, it walks every tracked container — every list, dict, instance, tuple-of-containers — and writes bookkeeping into each object's GC header as it computes reachability. A single full collection in the child can therefore touch the entire inherited heap and dirty essentially all of it at once, even if the application never read those objects at all. This is what `gc.freeze()` addresses. Called after your data is loaded and immediately before forking, it moves every currently tracked object into a permanent generation that the collector does not traverse. The inherited objects then stay untouched by collections in the children, and only objects allocated after the fork are collected normally. `gc.unfreeze()` reverses it and `gc.get_freeze_count()` reports how many objects are parked there. Pair it with `gc.disable()` in short-lived workers when the work allocates little and exits quickly. Since Python 3.12 some objects are immortal — `None`, `True`, `False`, small integers, interned strings — and their reference counts are never modified, so they no longer dirty pages. That helps at the margins; it does nothing for your own millions of dicts and lists. ## The real fix: fewer objects Freezing the collector removes one source of writes; reference counting remains. The durable fix is to stop representing bulk data as millions of Python objects at all. One large `bytes` object, an `array.array`, or a `memoryview` over a `multiprocessing.shared_memory.SharedMemory` block holds its payload in a single allocation with exactly one reference count, so reading element three million writes nothing. A file mapped with `mmap.mmap` is better still: the pages are backed by the kernel page cache, genuinely shared between processes, and evictable under memory pressure rather than counting against every worker. ## Two things that make this a 3.14 conversation First, you may not be forking at all any more. On Python 3.14 the default start method is `forkserver` on Unix other than macOS, and `spawn` on macOS and Windows, so a child no longer inherits the parent's heap unless you explicitly ask for `fork` via `multiprocessing.get_context("fork")`. A design that depended on inherited data silently stops working on upgrade — it does not get slower, the data is simply not there. Second, requesting `fork` explicitly is increasingly a liability: forking a process that already has threads is unsafe and has emitted a `DeprecationWarning` since 3.12, and plenty of libraries start threads behind your back. ## Measuring it Python-level tools mislead here. `tracemalloc` accounts for allocations the interpreter made, not for pages the kernel copied, so a worker that dirties inherited pages looks innocent. Measure at the OS level — resident and proportional set size per process — and compare shared against private mappings. A worker whose private memory grows while its Python-level allocations stay flat is the signature of copy-on-write being undone.

  • Exactly what does gc.freeze do, and when in the process lifetime must you call it?
    It moves every currently tracked object into a permanent generation that the cyclic collector never traverses. Call it after all the long-lived data is loaded and immediately before forking, ideally after a `gc.collect()` so you are not freezing garbage. Objects allocated afterwards are collected normally, so the children keep working collectors for their own short-lived allocations.
  • Why does tracemalloc show flat memory while the worker's RSS keeps growing?
    `tracemalloc` accounts for allocations the interpreter performs. Copy-on-write growth is the kernel privatizing pages the process already had mapped — no new Python allocation happens, so nothing is recorded. Diagnose it at the OS level instead, by comparing shared and private resident memory per worker.
  • Your service upgrades to 3.14 and workers no longer see the preloaded dataset at all. What changed?
    The default start method on Unix other than macOS is now `forkserver`, so workers are forked from a clean server process rather than from your main process and never inherit its heap. Either request `fork` explicitly with `multiprocessing.get_context("fork")`, accepting that forking alongside threads is deprecated, or stop relying on inheritance and load the data in each worker or from a shared block.

Sharing a library book that stamps a counter on every page you open: you only meant to read, but each page you touch is now marked, and the librarian has to make you your own copy of it.

saying these in an interview costs you the question

  • Believes read-only access keeps copy-on-write pages shared
  • Calls it a memory leak rather than page privatization
  • Thinks gc.freeze stops reference counts from being written
  • Relies on inherited globals without pinning the start method
  • Uses tracemalloc to diagnose copy-on-write growth
  • Keeps millions of small objects instead of one large buffer

context