skip to content

Why do Python pre-fork workers lose the copy-on-write sharing they start with?

level: middleimportance: must knowfreq 45%

answer

  1. Reading is not free in CPython
  2. Every object header holds a counter
  3. Dirt is tracked per page, not per object
  4. The collector writes headers too
  5. Private_Dirty climbs, sharing decays

basics

~20 s

Because CPython writes an object's reference count whenever a reference to it is taken or dropped. Merely reading a shared object dirties the whole page holding it, so a worker's private memory grows until little is really shared.

solid answer

~50 s

Every CPython object carries a reference count in its header, and the count is written on every read: binding a name, iterating a list, passing an argument all take and drop references. The MMU tracks dirt per page, so one counter update copies a whole 4 KB page and every unrelated object packed onto it. The cyclic collector makes it worse — each pass walks the tracked container objects and writes their GC headers even when nothing is garbage. The result is that a worker which never mutates the parent's preloaded table still converts most of it to private memory within minutes. Sum of RSS across workers hides this; read per-process PSS and `Private_Dirty` instead. The durable fixes are fewer, larger objects — one `bytes` blob, an `array.array`, an `mmap.mmap` region — plus `gc.freeze()` before forking.

code

python · 7 lines
python
import sys

schedule = [["LH441", 27], ["BA912", 41]]
row = schedule[0]
print(sys.getrefcount(schedule[0]))
del row
print(sys.getrefcount(schedule[0]))

go deeper

for a junior

Recall that reading a Python object is not free: CPython adjusts a counter stored inside the object every time a reference to it is taken or dropped, so nothing about access is purely passive.

for a middle

Explain refcount writes plus page granularity together: one counter update dirties a whole 4 KB page and every unrelated object packed onto it, which is why read-only access still costs memory.

for a senior

Demonstrate measurement and remedy: read per-process PSS and Private_Dirty instead of summing RSS, then restructure bulk data into few large buffers so the sharing actually survives the workers' lifetime.

for a principal

Frame the capacity decision: how many workers a host really supports once sharing has decayed, and whether the answer is fewer larger objects, an out-of-process store, or a different concurrency model entirely.

## Reading an object writes to it Every object in CPython begins with a header holding at least a reference count and a type pointer; container types that the cyclic collector tracks carry an extra GC header in front of that. The reference count is maintained eagerly: taking a reference increments it and dropping one decrements it. Crucially, *taking a reference is what reading does*. Binding a name to an existing object, iterating a list, passing an argument to a function, putting a value on the interpreter's stack — each of these increments a counter inside the object and later decrements it again. In the default build these are plain memory writes. So the intuitive model "the child only reads the parent's data, therefore the pages stay shared" is wrong at the first line of the loop. There is no read path through the object model that does not write. ## Page granularity turns a counter into a kilobyte The hardware and the kernel track dirt at page granularity — typically 4 KB, and 16 KB on Apple silicon. A page holds a lot of Python: dozens of small integers, tuples or strings, or the headers of several dozen larger objects. Updating one reference count therefore copies the entire page, including every neighbouring object the worker never looked at. That is worse than it sounds, because CPython's allocator packs small objects created at the same time into the same pools and arenas. The rows you parsed at startup are physically adjacent, so walking any part of the structure drags its neighbours into private memory with it. Locality, which normally helps, works against sharing here. ## The collector adds a second write path Even a worker that touched nothing would eventually dirty the heap, because the cyclic garbage collector traverses every tracked container object and writes into its GC header — the doubly linked list pointers, and the field the collector temporarily stashes a tentative reference count in. It does this whether or not anything turns out to be garbage, and survivors that get promoted between generations have their list links rewritten again. One full collection over an old generation full of startup data is enough to dirty most of it. This is precisely why `gc.freeze()` before forking is a real optimization rather than a micro-optimization. ## Measuring it honestly RSS counts every resident page in full, in every process, so adding up worker RSS double-counts shared pages and makes a fleet look enormous while sharing is still good — and looks identical after the sharing is gone. The metric that answers the question is PSS, the proportional set size, which divides each shared page among the processes mapping it. On Linux the per-process `smaps_rollup` file reports `Pss`, `Private_Dirty` and `Shared_Clean`; watch `Private_Dirty` on one worker over its lifetime. That curve *is* the copy-on-write erosion, and it usually flattens out startlingly close to the size of the data you thought you were sharing. ## What genuinely stays shared Three things survive: 1. **Read-only mappings** — the interpreter binary and shared libraries, which are never written. 2. **Immortal objects** — since Python 3.12 a fixed set of interpreter objects has a saturated reference count that increment and decrement skip, so their headers are never written. 3. **Payload bytes that are not object headers.** This is the practical lever. Refcounting writes the *header*, not the contents. One 200 MB `bytes` object has one header: touching it dirties a single page and leaves the remaining bytes clean and shared forever. A list of five million small objects has five million headers spread across the heap and will be almost entirely private within minutes. ## The design rule Move bulk data out of the Python object graph. Store a table as an `array.array`, as one `bytes` object read through a `memoryview` with `struct` accessors, or as an `mmap.mmap` of a prebuilt file, and the sharing holds because there is almost no header to dirty. Where the object graph must stay, `gc.freeze()` after startup removes the collector's write path, and reducing object count reduces the refcount write path. If neither is enough, the honest conclusion is that pre-fork is the wrong shape for that workload — a shared-memory design, an out-of-process store, or threads in one address space will hold the data once instead of once per worker. ## The sentence that shows you understand it "Copy-on-write shares memory until something writes, and in CPython reading is a write." Everything else — page granularity, allocator locality, GC traversal, the blob trick — follows from that.

  • Which parts of a large bytes object stay shared after a fork, and which do not?
    Only the header page is dirtied. Reference counting writes the object's header, never its payload, so a 200 MB `bytes` object costs one copied page and keeps the rest shared for the life of the worker. That asymmetry is the whole argument for holding bulk data as one large buffer rather than millions of small objects.
  • Does calling gc.disable() in the workers prevent copy-on-write erosion?
    It removes one of the two write paths — the collector no longer traverses tracked objects and rewrites their GC headers — but reference-count writes continue on every read, so erosion continues. It also lets reference cycles accumulate unbounded. `gc.freeze()` before forking is the targeted version of the same idea and keeps collection working for objects created afterwards.
  • How would you confirm on a running worker that sharing has decayed rather than guess?
    Read the per-process memory summary the kernel exposes for that PID rather than RSS: `Pss` tells you the worker's fair share of physical memory, `Shared_Clean` what is still genuinely shared, and `Private_Dirty` what has been copied. Sample one worker over its lifetime; a rising `Private_Dirty` that converges on the size of the preloaded data is the erosion, measured.

Like a shared reference book whose margin is initialled every time anyone glances at a page: the initial is tiny, but the library's rule is to reprint the whole sheet the moment one appears.

saying these in an interview costs you the question

  • Claims read-only access never copies a shared page
  • Sums worker RSS and calls it real memory used
  • Thinks only mutation copies an object's page
  • Believes gc.disable() stops reference-count writes
  • Expects a list of millions of small objects to stay shared

context