skip to content

Why does preloading a Python app before forking workers save less memory than copy-on-write suggests?

level: middleimportance: must knowfreq 55%

answer

  1. Sharing lasts only until a write
  2. Reading a Python object still writes
  3. The count lives in the object header
  4. The cyclic collector dirties pages too
  5. Freeze the heap before the fork

basics

~20 s

Copy-on-write shares pages only until something writes to one. CPython stores a reference count inside every object header, so merely touching a preloaded object dirties its page and the sharing quietly decays over the worker's lifetime.

solid answer

~50 s

After `os.fork()` the child maps the parent's pages instead of copying them; the kernel copies a page only on the first write to it. The trouble is that in CPython almost nothing is a pure read: every object carries a reference count in its header, and binding a name, iterating a container or calling a method increments and decrements it. Those writes are scattered, so a few touched entries dirty whole pages. The cyclic collector makes it worse, because a collection walks tracked containers and updates their GC bookkeeping in place. Mitigate it by calling `gc.freeze()` after import and before the fork, and by holding bulk data as `bytes`, an `array.array` or a memory map instead of millions of small objects. Since 3.12 some objects are immortal and skip refcount updates entirely. Then measure: preloading is a real but decaying saving, not a guarantee.

code

python · 8 lines
python
import sys

app_config = {'targets': 6800}
before = sys.getrefcount(app_config)
alias = app_config
after = sys.getrefcount(app_config)

print(after - before, alias is app_config)  # 1 True: a 'read' wrote

go deeper

for a junior

Recall the one-line mechanism: after a fork the two processes share pages until one writes, and then the kernel copies that page. Knowing that preloading exists to share the imported application is enough at this level.

for a middle

Be ready to explain why the sharing decays: CPython writes a reference count into every object header, so language-level reads are memory-level writes, and the cyclic collector adds more writes on top. Name gc.freeze() as the mitigation.

for a senior

An interviewer expects you to have measured this on a real service: proportional set size rather than summed RSS, private-dirty growth over a worker's lifetime, and a decision about data shape (bytes or arrays over millions of small objects) driven by those numbers.

for a principal

Own the tradeoff rather than the trick. Decide when the decaying saving from preloading justifies the inherited-state discipline it imposes, what memory headroom the fleet must keep once sharing decays, and whether the money is better spent on data representation than on worker count.

## Copy-on-write is a promise about pages, not about objects When a pre-fork server imports the entire application in the parent and only then calls `os.fork()` once per worker, the child does not receive a copy of the parent's heap. It receives the same physical pages, mapped into both processes and marked read-only in both page tables. The first time either process writes to such a page the CPU traps, the kernel copies that one page (typically 4 KiB) into a private frame, and only then lets the write land. That is **copy-on-write**, and it is the whole reason "import once, fork N times" can host twenty workers for far less than twenty times the memory of one. The Python-specific catch is that CPython's heap is not read-mostly. Every object starts with a header holding a **reference count**, incremented whenever a new reference is taken and decremented when one goes away. Binding a name, appending the object to a list, passing it as an argument, iterating a container that holds it, calling a method on it — each of those is an ordinary read at the language level and a write at the memory level. The writes are also badly scattered: a table loaded at import time can have its keys and values spread across hundreds of pages, so touching a handful of entries dirties a handful of entire pages, not a handful of bytes. The **cyclic garbage collector** compounds this. Every container CPython tracks — lists, dicts, instances with a `__dict__`, tuples containing containers — carries collector bookkeeping alongside the refcount, and a collection walks the generation updating that bookkeeping in place. One full collection inside a worker can therefore touch nearly every preloaded container it still reaches, un-sharing a large slice of the supposedly shared heap in a single sweep. The classic symptom is a fleet whose workers look wonderfully cheap for the first minute and then converge, over an hour, on roughly the footprint they would have had with no preloading at all. ## What actually helps - `gc.freeze()`, available since 3.7, is the direct mitigation. Called after the application is imported and before the fork, it moves every currently tracked object into a **permanent generation** that collections skip, so the collector stops writing to those headers and the parent's constant data really does stay shared. The price is that frozen objects are never collected while frozen, which is why the call belongs immediately after imports rather than after the process has been running and churning; a child that genuinely needs those objects collected can call `gc.unfreeze()`. - **Data shape** matters more than any flag. A million small objects means a million refcounted headers spread over a lot of pages. The same payload held as `bytes`, an `array.array`, or a memory-mapped file has one header and a contiguous block, and nothing in the interpreter writes into that block just because you read it. If the point of preloading is to share a large read-mostly table, choosing a representation with few object headers beats every piece of GC tuning available. - Since 3.12, PEP 683 made a set of objects **immortal** — `None`, `True`, `False`, small integers, interned strings — meaning their counts are pinned and no longer updated on every reference. That removes a genuine source of dirtying, because those are precisely the objects every code path touches. On the free-threaded build, officially supported from 3.14, reference counting is implemented differently again, so the only honest answer to "how much did preloading save" is a measurement taken on the build you actually ship. ## Measuring it Adding up per-process **resident set size** is the standard mistake: a shared page is counted once for every process that maps it, so the total is inflated and preloading looks like it did nothing. On Linux, **proportional set size** divides each shared page among its sharers, and the per-process smaps rollup separates shared-clean from private-dirty pages. The number worth graphing is **private-dirty growing over a worker's lifetime** — that is the shared heap being un-shared, one page fault at a time. `resource.getrusage(resource.RUSAGE_SELF)` reports a peak in its `ru_maxrss` field, which answers "how big" and never "how shared". ## Where this sits in the preload tradeoff Preloading buys three things: - shared pages, - a fast worker start with no re-import, - and a single place where an import error surfaces. It costs the discipline of never carrying an inherited connection into a child, and it means the sharing benefit decays rather than holding. Treat it as a measured optimisation with a number attached, not as an architectural guarantee.

  • How would you actually measure whether preloading saved memory in a pre-fork service?
    Not by summing per-process RSS, which counts every shared page once per process. On Linux read proportional set size, or the per-process smaps rollup, and compare total PSS with preloading on and off. Then watch private-dirty per worker over an hour: if it climbs steadily toward the un-preloaded footprint, the sharing is decaying and the interesting work is reducing header churn, not adding workers.
  • Does calling gc.freeze() have a downside?
    Yes. Frozen objects are moved to a permanent generation and are never collected while frozen, so cycles among them are never reclaimed. Call it once, right after imports and before forking, when the heap is at its most static. Do not freeze a heap that has been serving traffic and is full of churn, and let a child call `gc.unfreeze()` if it genuinely needs those objects collectable.
  • Why does holding bulk data as bytes or an array help more than GC tuning?
    Because the cost is per object header, not per byte. A million small objects means a million refcounts spread over many pages, all of which get dirtied as workers touch them. The same payload as `bytes`, an `array.array` or a memory map is one header over a contiguous block: reading it never writes into it, so those pages stay genuinely shared for the worker's whole life.

Copy-on-write is a shared library book that is duplicated the moment anyone scribbles in it. CPython scribbles a tally mark on every page it opens, so most of the library ends up duplicated anyway.

saying these in an interview costs you the question

  • Says copy-on-write means forked workers never copy anything
  • Believes reading an object cannot dirty a memory page
  • Claims preloading always halves memory, with no measurement
  • Sums per-process RSS and calls the total real usage
  • Thinks the cyclic collector leaves preloaded objects untouched
  • Confuses virtual size with resident, shared or private memory

context