Why does keeping caught exceptions in a list hold whole call frames alive, and what do you store instead?
answer
- The object is not just a message
- It reaches the frames it travelled
- Frames own their local variables
- Not a cycle — a live reference
- Capture a frameless view instead
basics
~10 sAn exception references its traceback, which references every frame it passed through, which references those frames' locals. Keeping the exception keeps all of it. Store traceback.TracebackException.from_exception(exc), or the rendered text, instead.
solid answer
~40 sThe retention chain is exception → `__traceback__` → each node's `tb_frame` → that frame's locals → whatever those locals point at. A worker that appends failures to a list for an end-of-run report therefore holds every argument, buffer and partial result of every failed call until the run ends, and RSS climbs in proportion to the failure count rather than to the message sizes. The garbage collector cannot help: these are live references, not cycles. The fix is to capture a rendered, frameless view at the moment of the failure — `traceback.TracebackException.from_exception(exc)` keeps the type, message and resolved frame summaries and can be formatted later, or simply store `"".join(traceback.format_exception(exc))`. `traceback.clear_frames(exc.__traceback__)` is the blunter option when you must keep the exception object itself.
code
python · 14 linesimport traceback
failures = []
def parse_legs(rows):
for row in rows:
try:
yield int(row)
except ValueError as exc:
failures.append(traceback.TracebackException.from_exception(exc))
print(sum(parse_legs(["3", "x", "7", "y"])))
print(len(failures))
print("".join(failures[0].format()), end="")go deeper
Take away one fact: an exception you catch drags its whole call path with it, so do not stash exception objects in module-level lists or caches. Log them or format them to text instead.
Be able to trace the chain out loud — exception, traceback, frame, locals — and name traceback.TracebackException.from_exception and traceback.format_exception as the frameless alternatives to storing the object.
Recognise the shape in production: memory tracking failure count rather than workload, snapshots showing surviving frames, a collector that frees nothing. Then argue the two-line fix at the accumulation site rather than raising the memory limit.
Own the policy for what a failure record contains. Frame locals can hold secrets and personal data, retention costs memory in every worker that aggregates, and error-reporting integrations differ in what they ship — set the default and make it the easy path.
### The chain that keeps memory alive An exception instance looks small — a type and a message. It is not. Its `__traceback__` links one node per frame it propagated through, each node holds `tb_frame`, and a frame object owns its local variables. So retaining one exception retains, transitively, every local of every frame between the `try` and the `raise`: the request payload it was parsing, the batch it was halfway through, the file buffer it had read. This is ordinary reachability, not a cycle. The cycle collector is irrelevant, and `gc.collect()` will not shrink the process. As long as your list holds the exception, everything behind it is live by definition. ### Where it bites The pattern that produces it is a good one otherwise: a long-running batch worker that keeps going after a failure and reports at the end. Say a nightly route-optimisation job processes tens of thousands of jobs and appends each failure to a list so the 11-person team gets one digest instead of a flood of alerts. Each retained exception pins the frame that was building a route, including the partially built result. A run with a few thousand failures grows steadily and can end in an out-of-memory kill *after* the useful work is done — and, because it dies at report time, the digest is never written either. A second, quieter case is any object that stores an exception for later inspection: a result wrapper, a retry record, a future's captured error, a cache of "last failure per key". If it outlives the handler, it carries frames. ### What to store instead **`traceback.TracebackException.from_exception(exc)`** is the purpose-built answer. It walks the traceback once, resolves each node into a summary with filename, line number, function name and source line, keeps the exception's type name and message, and holds **no** frame references. Its `format()` method yields the same lines the interpreter would print, so you can render the digest at the end without keeping anything live. Pass `capture_locals=True` if you deliberately want the locals — note that it stores their `repr()` strings, so the objects still are not retained, but anything sensitive in them is now in your report. **Rendered text.** `"".join(traceback.format_exception(exc))` is the crudest and often the right answer: a string is inert, bounded and trivially loggable. You lose the ability to re-render differently later, which usually does not matter. **`traceback.clear_frames(tb)`** clears the local variables of each frame in a traceback in place. Use it when something outside your control insists on keeping the exception object; the frames remain but their contents are dropped, so the rendered stack still works while the payloads are released. ### A related subtlety The language already protects the common case: the name bound by `except ... as exc:` is deleted when the block ends, precisely because the handler's own frame would otherwise reference the exception whose traceback references that frame. That protection is why casual handling does not leak — and why the leak reappears the moment you deliberately store the exception somewhere longer-lived. If you need it after the block, bind it to another name knowingly, and prefer storing a captured view rather than the object. ### How you would find it RSS that climbs with the *failure count* rather than the workload is the tell. Confirm it with allocation snapshots taken before and after a batch and diffed, and look for frame or traceback objects among the survivors; an object-graph walk from your accumulator list will show the traceback chain hanging off each stored exception. The fix is a two-line change at the append site, which is why this is worth recognising by shape rather than rediscovering by profiler each time.
- Would running the garbage collector release the memory held by a stored exception?No. The accumulator list holds a live reference, so the exception, its traceback, the frames and their locals are all reachable — the collector's job is unreachable cycles, not reachable objects. `gc.collect()` returning nothing while RSS keeps climbing is a useful signal that you are looking at retention, not a cycle.
- What does capture_locals=True change about a captured TracebackException?It records each frame's locals as `repr()` strings at capture time, so the rendered report shows values rather than just lines. The objects themselves are still not retained, so the memory benefit stands — but the reprs may contain credentials or personal data, so it is a deliberate choice for a debug path, not a default for shipped error reports.
- How do you tell this apart from an ordinary object leak while diagnosing a growing worker?Look at what growth tracks. Retained tracebacks make RSS scale with the *failure* count, not the item count, and it stops climbing when the input is clean. Allocation snapshots diffed across a batch will show frame and traceback objects surviving, and the referrer chain leads straight back to whatever collection is storing exceptions.
Filing an incident report is cheap; impounding the entire vehicle so you can re-photograph it later is not — storing the exception impounds every frame it drove through.
saying these in an interview costs you the question
- Thinks an exception object is only a type and a message
- Blames the cycle collector for a live reference
- Says clearing the message would free the frames
- Assumes the traceback stores text rather than frames
- Stores exception objects in a long-lived cache without thought
- Believes formatting later is always equivalent to capturing now