How do you confirm reference cycles are behind a Python transcript archiver's memory growth across 6,800-row batches?
answer
- First decide which of two problems it is
- Force a pass at a quiet point
- The return value is the signal
- Keep what the collector found, then tally
- Referrers close the loop; sample, do not sweep
basics
~20 sForce a gc.collect() between batches and watch its return value and the object counts. If a forced pass reclaims a lot, the growth is cyclic garbage awaiting collection; if it reclaims nothing, the objects are still reachable and the collector is irrelevant.
solid answer
~50 sThe first measurement splits the problem in two. At a quiet point between batches, call `gc.collect()` and record its return value alongside the process footprint. A large non-zero return with memory dropping means the archiver manufactures cycles faster than the collector reclaims them — real, but a tuning and design problem. A near-zero return with memory unchanged means the objects are still reachable from something live, and no collector will ever free them. Only in the first case do you dig into the `gc` module: set `gc.set_debug(gc.DEBUG_SAVEALL)`, run one batch, collect, and inspect `gc.garbage`, tallying the types you find. `gc.get_referrers()` on a sample then shows what closes each loop. In archiving code the usual culprits are a per-row object storing a back-pointer to its batch, a callback held by the object it calls, and exception objects kept on an error list, since an exception holds its traceback which holds the frame that holds the exception.
code
python · 12 linesimport gc
def probe(label):
unreachable = gc.collect()
tracked = len(gc.get_objects())
print(f"{label}: collected={unreachable} tracked={tracked}")
for batch in range(3):
probe(f"before batch {batch}")
rows = [{"row": i} for i in range(6800)]
del rows
probe(f"after batch {batch}")go deeper
Know that gc.collect() can be called explicitly and that its return value counts unreachable objects. Being able to say that reachable objects are never collected, however much memory they use, is the key idea to carry.
Explain the diagnostic split — forced collection reclaims a lot versus reclaims nothing — and name the tools for each side. Be ready to describe what gc.get_referrers() returns and why it is expensive.
Show a repeatable procedure on a live process: measure between batches, take a bounded diagnostic window, tally types, sample referrers, then fix the object graph at the source rather than reaching for thresholds.
Frame it as a class of defect to design out: decide whether back-pointers, retained exceptions and callback registries are allowed in hot data paths, and require the per-batch memory probe as standing instrumentation rather than something rebuilt during each incident.
**Step one: split reachable growth from cyclic garbage.** These have opposite fixes and the same symptom, so never guess. Instrument the batch loop at a quiet point: ```python import gc unreachable = gc.collect() print("collected:", unreachable, "tracked:", len(gc.get_objects())) ``` Run that between batches and watch three numbers over a dozen batches: the value `gc.collect()` returns, the size of `gc.get_objects()`, and the process footprint. - Return value large and steady, footprint flat after each forced collection → the archiver *is* creating cycles. They are being collected, just later than you would like. This is a real cost — memory sits high between passes and the collector burns CPU — but nothing leaks. - Return value near zero and both the tracked-object count and the footprint still climbing → nothing is garbage. Something live holds those objects: a module-level list, a cache with no eviction, a logging handler holding formatted records, a registry of callbacks. The `gc` module cannot help and every knob in it is a distraction. - Return value near zero, tracked-object count flat, footprint climbing anyway → the objects are being freed and the allocator is simply holding the pages, or the growth is outside the Python heap entirely. Only the first branch is a cycle problem, and it is the branch worth ruling in explicitly before spending an afternoon in the collector. **Step two: see what the cycles are made of.** `gc.set_debug(gc.DEBUG_SAVEALL)` makes the collector append everything it finds unreachable to `gc.garbage` instead of freeing it. Run exactly one batch with it on, collect, and tally: ```python import gc from collections import Counter gc.set_debug(gc.DEBUG_SAVEALL) run_one_batch() gc.collect() print(Counter(type(o).__name__ for o in gc.garbage).most_common(10)) gc.set_debug(0) gc.garbage.clear() ``` The tally names the shape immediately: thousands of one row class, or a pile of `frame` and `traceback` objects, or `cell` objects from closures. Turn the flag off and clear `gc.garbage` afterwards, or you have converted a diagnostic into a genuine leak. **Step three: find who closes the loop.** `gc.get_referrers(obj)` returns the objects that refer to a sample object; `gc.get_referents(obj)` walks the other way. Two cautions come from experience. It is slow — it scans every tracked object — so sample two or three, never the whole batch. And your own diagnostic frames and lists appear in the results, so read past them. **What the usual culprits look like in batch archiving code.** Three recur. A per-row object that keeps a back-pointer to its batch, while the batch keeps a list of its rows, is a cycle per row — with a 6,800-row batch that is 6,800 objects the collector must find rather than the counter freeing them instantly. A handler or callback stored on the object it is meant to call is the same shape one level up. And the sharpest one: exception objects retained past their `except` block. An exception references its traceback, the traceback references the frame, and the frame's locals reference the exception — a cycle every time. Python unbinds the `except ... as` name at the end of the block precisely to break it, so code that appends the caught exception to a failures list, or stashes it on `self`, reintroduces it deliberately. Keep the message and a formatted string instead of the live exception object, and the frames go away with it. **Step four: fix at the source, not at the collector.** Break the back-edge when a batch finishes, hold the back-pointer as a non-owning reference, or restructure so rows do not know their batch. Raising the collection thresholds is a legitimate second-order tuning once the cycles are as few as you can make them, but it is a cost-shifting move, not a fix: it trades collector CPU for a bigger resident set. **One trap worth naming.** A circular import at startup is not this problem, however similar the words sound. Modules referencing each other stay in `sys.modules` for the process lifetime, so they are permanently reachable, never garbage, and contribute nothing to per-batch growth. If growth is proportional to batches processed, look at what each batch retains, not at import structure.
- A forced `gc.collect()` frees nothing but memory still climbs each batch — where do you look next?At live references, not the collector. Something reachable is accumulating: an unbounded cache, a module-level list, a logging handler retaining records, a registry that appends and never removes. Snapshot object counts by type across batches to see what grows, then find the owner. No collector setting will help, because nothing you are looking at is garbage.
- Why does keeping caught exception objects in a failures list grow memory more than keeping their messages?An exception instance references its traceback, the traceback references the frames, and each frame keeps its locals alive — potentially a whole batch of rows per failure. It is also a cycle, since the frame's locals refer back to the exception, so nothing is freed at zero. Storing the formatted message, or the type and string, drops all of it.
- What is the risk of leaving `gc.set_debug(gc.DEBUG_SAVEALL)` on in production?It converts every collection into a leak: unreachable objects are appended to `gc.garbage` and kept alive instead of being freed, so memory climbs for as long as it is set. It is a short diagnostic window on one process, followed by `gc.set_debug(0)` and clearing `gc.garbage` — never a default.
saying these in an interview costs you the question
- Tunes gc thresholds before proving the garbage is cyclic
- Assumes any memory growth means a reference cycle
- Calls gc.get_referrers on thousands of objects in production
- Leaves DEBUG_SAVEALL enabled and creates a real leak
- Blames a circular import for per-batch growth
- Expects gc.collect to free objects a cache still holds