A pickle.Pickler reused for millions of genome records makes memory grow without bound — what is holding them?
answer
- Growth tracks records, not concurrency
- Collector cannot help with strong references
- Something remembers what it already wrote
- Ids must stay valid, so objects stay alive
- The table belongs to the Pickler
basics
~20 sThe pickler's memo. A pickle.Pickler holds a strong reference to every object it has already written, and that table lives as long as the Pickler, not as long as one dump() call. Call clear_memo(), or use a fresh Pickler.
solid answer
~40 s`pickle.Pickler` keeps a *memo*: a table mapping the id of every object it has serialized to that object, so a second occurrence is written as a back-reference instead of a second copy. It must hold a **strong** reference, otherwise an object could be freed and its id reused by a different object, corrupting the stream. The memo belongs to the Pickler instance, not to one `dump()` call, so a long-lived Pickler streaming records into one open file pins every record it has ever written — memory climbs linearly with records processed and never falls. The fix is to bound the memo: `Pickler.clear_memo()` between records or every N records, or construct a fresh Pickler per record. The cost is that cross-record sharing is lost, so genuinely shared sub-objects get re-emitted and the file grows.
code
python · 24 linesimport gc, io, pickle, weakref
class Record:
__slots__ = ("gene", "__weakref__")
def __init__(self, gene):
self.gene = gene
buf = io.BytesIO()
pickler = pickle.Pickler(buf, protocol=5)
seen = []
for i in range(1000):
record = Record(f"gene-{i}")
seen.append(weakref.ref(record))
pickler.dump(record)
del record
gc.collect()
print("still alive:", sum(1 for ref in seen if ref() is not None)) # 1000
pickler.clear_memo()
gc.collect()
print("after clear_memo:", sum(1 for ref in seen if ref() is not None)) # 0go deeper
Know that pickling the same object twice writes it once and refers back the second time, and that this bookkeeping means the pickler remembers objects it has already written.
Explain the memo: keyed on identity, strongly referenced so ids stay valid, and scoped to the Pickler rather than to a single dump() call — then name clear_memo() and a per-record Pickler as the two ways to bound it.
Demonstrate the diagnosis end to end: linear growth immune to collection, weakrefs plus referrer walking to find the holder, clear_memo() as the decisive experiment, and a judgement call on per-record dumps versus a bounded shared memo.
Frame it as a format decision: streaming a million objects through an object-graph serializer buys sharing you do not need and costs memory and coupling, so set the standard for record-oriented artefacts and put the guardrail in the shared writer.
### The symptom A genome-annotation pipeline writes one long stream: a helper opens the output file once, constructs a `pickle.Pickler` over it, and calls `dump()` once per annotated record. Resident memory rises smoothly with the number of records processed, is never released, and a `gc.collect()` changes nothing. The records themselves are dropped by the loop as soon as they are written, so on a first reading nothing should be retaining them. That shape — linear growth, tracking work done rather than concurrency, immune to garbage collection — says *something holds strong references*, not *reference cycles* and not *fragmentation*. Reference cycles would be reclaimed by the collector; the growth would also plateau if the culprit were an interpreter arena that could be reused. ### What the memo is and why it retains `pickle` guarantees that a shared or recursive object graph round-trips with its sharing intact: if the same sub-object appears twice, the stream carries it once and refers back to it, and a self-referential structure terminates rather than recursing forever. The mechanism is the **memo**, a table keyed on `id()` of each object the pickler has already emitted, mapping to the memo slot and the object. The object has to be in that table, not just its id. Ids in CPython are addresses, and an address is only unique while the object is alive; if the pickler let a memoized object be freed, a later object could be allocated at the same address and be mistakenly written as a back-reference to the wrong value. Holding a strong reference is the only correct implementation. The scope of that table is the part teams get wrong. Reading the API — `Pickler.dump(obj)` — it is natural to assume the memo is per call. It is per **Pickler**. That is deliberate: reusing one Pickler across many `dump()` calls is how you get sharing *between* records in a single stream. It is also why a Pickler that lives for the whole run accumulates every object of the whole run. ```python pickler = pickle.Pickler(open("out.pkl", "wb"), protocol=5) for record in records: # each record stays alive in the memo pickler.dump(record) ``` The unpickling side has the mirror-image problem: a long-lived `pickle.Unpickler` reading the same stream keeps its own memo, so a streaming reader that loads a million records holds a million of them even if the caller discards each after use. ### Confirming it rather than guessing The diagnosis that separates this from an application-level cache is a retention test rather than an allocation test. Take `tracemalloc` snapshots at two points far apart and diff them: the growth localizes to the code that builds records, which is misleading on its own — it tells you where they were allocated, not who keeps them. Then prove retention directly: hold `weakref.ref` objects to a sample of records, drop your own references, run `gc.collect()`, and check how many are still alive. If they survive while your code holds nothing, walk referrers with `gc.get_referrers` and the chain ends at the pickler's memo. Calling `clear_memo()` and re-checking the weakrefs is the decisive experiment — the count drops to zero. ### Fixes and their trade-offs * **Fresh Pickler per record** (or `pickle.dumps` per record, writing each blob with a length prefix). Simplest, and the natural shape for a record-oriented file you want to read incrementally. Every record becomes independently loadable, which is usually a feature. * **`Pickler.clear_memo()` between records, or every N records.** Keeps one Pickler and one stream, and bounds the memo to a window. Note that clearing loses back-references, so shared sub-objects are re-emitted and the output file grows; and you must clear at a record boundary, never mid-object. * **Do not stream a million objects through pickle at all.** If the artefact is a table of annotations, a record format with an explicit schema costs less memory, is readable by other tools, and does not make your class names part of the file format. The organizational fix matters too. Where an eleven-person team shares one serialization helper, a comment on the shared writer — "the memo is per Pickler; clear it or make a new one per record" — prevents the same bug being reintroduced by whoever next adds an output path, which is exactly how it usually returns.
- Why does the pickler's memo hold a strong reference rather than a weak one?It is keyed on `id()`, which in CPython is the object's address and is only unique while the object lives. With weak references a memoized object could be freed and a new object allocated at the same address, so the pickler would emit a back-reference to the wrong value and silently corrupt the stream. Strong references make the key valid for the memo's whole lifetime; bounding the lifetime is the caller's job.
- What do you lose by calling clear_memo() between records?Back-references across the boundary. Any object shared by two records — an interned lookup table, a shared parent node — is written once before the clear and again after it, so the stream grows and the two copies come back as separate objects rather than one shared object on load. That identity change matters if downstream code compares with `is` or mutates a shared structure.
- How would you tell this apart from an ordinary application-level cache leak?Look at who holds the references, not where memory was allocated. Weakref a sample of the records, drop your own references, collect, and see if they survive; then walk referrers. A cache leak's chain terminates in your own dict or list, this one terminates in the pickler's memo, and calling clear_memo() releases them — an experiment a cache leak will not respond to.
The memo is a guest book that also keeps hold of each guest's coat: it can only recognize a returning visitor while it still has the coat, so nobody ever leaves.
saying these in an interview costs you the question
- Blames reference cycles the collector cannot reach
- Assumes the memo resets on every dump() call
- Says flushing or closing the file frees the objects
- Reaches for gc tuning before finding the referrer
- Thinks a lower protocol version reduces memory