skip to content

What do gc.freeze() and gc.unfreeze() do, and when do you call them around a fork?

level: seniorimportance: should knowfreq 28%

answer

  1. A generation the collector never scans
  2. Stop the collector touching startup objects
  3. Collect first, freeze, then fork
  4. Frozen cycles are never reclaimed

basics

~20 s

gc.freeze() moves every currently tracked object into a permanent generation the cyclic collector never scans, so later collections stop writing their GC headers and dirtying pages shared after a fork. Call it right before forking; gc.unfreeze() puts them back.

solid answer

~40 s

CPython's cyclic collector writes into each tracked object's GC header as it traverses, so a collection in a forked child dirties pages that were still shared — even when nothing is garbage. `gc.freeze()` moves everything tracked at that moment into a permanent generation that no collection ever examines, so those writes stop happening. The recipe is: finish imports and warm-up, call `gc.collect()` so you do not make garbage permanent, call `gc.freeze()`, then fork. `gc.get_freeze_count()` reports how many objects are in there and `gc.unfreeze()` returns them to the oldest generation. The price is real: cycles among frozen objects are never reclaimed for the life of the process, so freeze a startup graph you intend to keep, never a churning one. It does not touch reference counting, so it removes one erosion path, not both.

code

python · 10 lines
python
import gc, os

gc.collect()          # do not make import-time garbage permanent
gc.freeze()           # nothing tracked now will be scanned again
print(gc.get_freeze_count() > 0)

pid = os.fork()
if pid == 0:
    os._exit(0)
os.waitpid(pid, 0)

go deeper

for a junior

Know that gc.freeze() exists, that it concerns CPython's cyclic garbage collector rather than reference counting, and that it is called in the parent before forking worker processes.

for a middle

Explain the permanent generation: frozen objects are skipped by every later collection, so the collector stops writing their GC headers and the pages inherited by forked children stay clean.

for a senior

Show the whole recipe and its price: collect first, freeze at the end of startup, fork, and accept that cycles among frozen objects are never reclaimed for the life of the process.

for a principal

Judge whether the saving justifies the permanently retained memory and the startup discipline it imposes, and whether the workload should move to a shared-memory or threaded design instead.

## The permanent generation `gc.freeze()` walks the objects the cyclic collector is currently tracking and moves all of them into a fourth, permanent generation. No collection ever scans that generation. `gc.unfreeze()` moves them back into the oldest generation, where they become collectable again, and `gc.get_freeze_count()` reports how many objects are currently frozen. Nothing else changes: the objects are still ordinary, still mutable, still reference-counted, still reachable in exactly the same way. ## Why that matters at a fork The collector is a second, independent source of page dirt in a pre-forked worker. As it traverses, it writes into each tracked object's GC header: the doubly linked list pointers that thread objects onto a generation's list, and the field where it temporarily stashes a tentative reference count while it subtracts internal references. Survivors promoted from one generation to the next get their links rewritten again. None of this depends on anything being garbage — a heap of perfectly live startup data is fully traversed and fully dirtied. In a single process that cost is invisible. In sixteen forked workers each doing their own collections over an inherited heap, it is the difference between a shared preloaded table and sixteen private copies of it. Freezing before the fork removes the write path entirely for that data: the children inherit the frozen state, their collectors never look at those objects, and the pages stay clean unless something else touches them. ## The recipe The order is the part interviewers probe: 1. Finish all imports and every warm-up allocation, so the data you want shared exists. 2. Call `gc.collect()` — freezing without collecting first makes the garbage produced during import permanent, and it is never coming back. 3. Call `gc.freeze()`. 4. Fork the workers. Some deployments also disable collection inside short-lived workers with `gc.disable()`, which is a blunter version of the same idea and risks unbounded cycle growth in a long-lived process. Freezing is the targeted move, because objects created *after* the freeze are still tracked and still collected normally — the worker keeps a working garbage collector for its own request-scoped garbage while the startup graph is left alone. ## The cost Frozen objects are never examined again, so any reference cycle among them is never reclaimed. For a startup graph — modules, classes, functions, a parsed configuration, a preloaded table you intend to keep for the process's whole life — that is not a leak in any meaningful sense; the memory was never going to be released anyway. For anything that churns, it is exactly a leak, and freezing halfway through a request loop is a real bug. The judgement call is simply: is this set of objects the process's permanent furniture, or is it working data? The second limit is scope. Freezing addresses the collector's writes only. Reference-count writes continue on every read, so a worker that walks the frozen structure still dirties its pages one by one. Freezing buys you the pages you would have lost to collection *without* touching anything; it does not make a hot, heavily traversed object graph stay shared. Combining it with a data layout that has few headers — one large buffer rather than millions of small objects — is what actually keeps the sharing. ## Unfreezing `gc.unfreeze()` exists for the case where the frozen set must become collectable again: a long-lived process that froze a graph it later replaces, or a child whose job is different enough that it genuinely needs that memory back. Note what it costs when you call it in a worker: putting the objects back into the oldest generation means the next collection will traverse them and write their headers, which dirties precisely the pages the freeze was protecting. So the normal answer for a pre-fork pool is that workers *stay* frozen; you unfreeze only when reclaiming beats sharing. ## Verifying it did anything `gc.get_freeze_count()` confirms the freeze took, but the measurement that matters is the one on the workers: per-process `Private_Dirty` over a worker's lifetime, before and after introducing the freeze. If the curve barely moves, the erosion in that service was refcount-driven rather than collector-driven, and the fix belongs in the data layout instead. Freezing is cheap enough to be worth trying and specific enough that it should be justified by a number, not by folklore.

  • Where exactly in a pre-fork server does the gc.freeze() call belong?
    At the end of startup: after every import and warm-up allocation, immediately following a `gc.collect()`, and before the first fork. Earlier and you miss the objects you meant to protect; later and the children have already been created, so each of them would need its own freeze and would already have started dirtying pages.
  • When should a worker call gc.unfreeze()?
    Only when it needs that memory reclaimable again. Unfreezing returns the objects to the oldest generation, so the next collection traverses them and writes their GC headers — dirtying exactly the pages the freeze was protecting. In a pre-fork pool the normal answer is that workers stay frozen for their whole life.

Freezing is moving the startup furniture into a room the cleaners never enter: nothing in there is tidied away, and nothing in there is disturbed.

saying these in an interview costs you the question

  • Thinks gc.freeze() makes objects read-only or immutable
  • Calls gc.freeze() without collecting first
  • Believes freezing also stops reference-count writes
  • Assumes frozen cycles are still reclaimed eventually
  • Confuses gc.freeze() with gc.disable(), which stops collection entirely

context