skip to content

Why does storing a caught exception keep its traceback frames and their locals alive?

level: middleimportance: nice to knowfreq 25%

answer

  1. An exception is bigger than it looks
  2. It carries a slice of the stack
  3. Frames keep their local variables
  4. Ask why the as-name disappears
  5. Store formatted text instead

basics

~10 s

An exception holds its traceback in traceback, each traceback entry holds the frame it was raised from, and every frame holds its locals. Keeping the exception keeps the whole call chain's locals alive.

solid answer

~40 s

The retention chain is `exception -> __traceback__ -> traceback object -> frame -> locals`, one link per stack level between the raise and the handler. So appending caught exceptions to an errors list, or stashing one on `self.last_error`, pins every local in every frame along the raise path -- including large payloads and any resource the frame had opened but not closed. Python 3 deletes the `except ... as name` binding at the end of the handler block precisely to stop this happening by accident, which is why that name is unbound afterwards. If you must keep the failure, keep something small: the formatted text from `traceback.format_exception()`, or the exception with its frames dropped via `traceback.clear_frames()` or `exc.with_traceback(None)`.

code

python · 18 lines
python
class Payload:
    def __del__(self):
        print("payload released")

errors = []

def load_record():
    payload = Payload()
    raise ValueError("bad record")

try:
    load_record()
except ValueError as exc:
    errors.append(exc)

print("exception stored; frame locals still alive")
errors.clear()
print("error list cleared")

go deeper

for a junior

Recall that a caught exception carries its traceback, and that a traceback points at the stack frames it came from. That is enough to know why holding on to exception objects is heavier than holding on to strings.

for a middle

Explain the chain link by link -- exception, traceback, traceback entry, frame, locals -- and why Python 3 unbinds the except-as name at the end of the block. Name at least one way to keep the failure without the stack.

for a senior

Show you have met this in production: failures accumulating in a retry or report list, RSS and open descriptors climbing together because a frame held a resource that was never closed, and the fix being to store rendered text rather than objects.

for a principal

Own the convention. Decide once, for the codebase, what an error-collection API is allowed to hold -- rendered text and a small structured record, not live exception objects -- so that retry queues, batch reports and error middleware cannot each rediscover this leak on their own.

### The chain, link by link When an exception propagates, the interpreter builds a traceback: a linked list with one entry per stack frame between the `raise` and the handler that caught it. That list is attached to the exception instance as `__traceback__`, which is what makes a printed traceback possible at the point of handling rather than at the point of failure. Each traceback entry keeps a reference to its frame object. Each frame object keeps the local variables that were live in it. So the full retention chain from a single caught exception is: ``` exception -> __traceback__ -> traceback entry -> frame -> every local in that frame -> next traceback entry -> frame -> every local in that frame -> ... ``` An exception is not a small object. It is a handle on an entire slice of the call stack as it stood at the moment of failure. ### Why this becomes a leak The pattern that produces it is deliberate and looks harmless. A batch loader collects failures so it can report them at the end: ```python errors = [] for sample_id in batch: try: load(sample_id) except LoadError as exc: errors.append(exc) ``` Every appended exception pins the locals of `load` and everything `load` called. If `load` had read a multi-megabyte record into a local, that record is retained. If `load` had opened a file or a socket and failed before closing it, that resource is retained too -- and because the object is still referenced it is never finalized, so the descriptor stays open. In a long-lived clinical-lab result loader this shows as two curves rising together: RSS, and open file descriptors, the second usually tripping a limit first and giving you a much cheaper alarm than the memory graph. The same shape appears wherever an exception outlives its handler: `self.last_error = exc` on a long-lived service object, a retry queue of failed work items carrying their exceptions, a cache of failures keyed by input, a logging call that stores the exception object instead of rendering it. ### The implicit del you did not write Python 3 makes the handler name disappear at the end of the block. This code raises `NameError` on the last line: ```python try: load(sample_id) except LoadError as exc: handle(exc) print(exc) # NameError: exc is not defined ``` The compiler wraps the handler body so that the name is deleted when the block exits. That is not tidiness for its own sake: the handler's own frame is part of the chain the exception's traceback references, so a surviving `exc` local would create a reference cycle -- frame refers to exception, exception refers to traceback, traceback refers to frame -- keeping the entire raise path alive until the cyclic collector happened to notice. Deleting the name breaks the cycle at the end of the block, immediately and deterministically. Knowing *why* that name is unbound is the part of this question interviewers actually look for. Such cycles are collectable, so this is not an unfixable leak; it is a retention with an unpredictable and possibly long delay, which in a service handling many failures is indistinguishable from a leak on any dashboard. ### Keeping the failure without keeping the stack There are three good options, in rough order of preference: 1. **Keep text, not objects.** `traceback.format_exception(exc)` renders the whole chain -- type, message, frames, source lines -- to a list of strings. Strings retain nothing. This is what almost every reporting path really wants. 2. **Drop the frames, keep the exception.** `traceback.clear_frames(exc.__traceback__)` clears the local variables out of every frame in the traceback while leaving the shape of the traceback intact, so you can still print file names and line numbers. 3. **Drop the traceback entirely.** `exc.with_traceback(None)` returns the exception with no traceback attached, which is the smallest thing that is still an exception object you can re-raise or inspect by type. Also worth naming: a `__cause__` or `__context__` chain means one stored exception can pin *several* stacks, since each chained exception carries its own traceback; and an `ExceptionGroup` holds every exception it groups, multiplying the effect. When you find frames holding memory in a leak hunt -- a walk back from a leaking object that terminates at a `frame` or a `traceback` -- a stored exception is one of the two usual explanations, the other being a generator that was never exhausted or closed.

  • Why is the except-as name unbound after the handler block ends?
    Because the compiler deletes it there. The handler's frame refers to the exception, the exception refers to its traceback, and the traceback refers back to that frame -- a cycle that would keep the whole raise path alive until the cyclic collector noticed. Deleting the name at block exit breaks the cycle immediately and deterministically. It is a memory-management decision, and it is why you must copy the exception to another name if you need it afterwards.
  • What would you store instead if you need to report failures at the end of a batch?
    Formatted text. `traceback.format_exception(exc)` renders type, message, frames and source lines into strings that retain nothing, which is what a report actually needs. If you need to keep the object -- to re-raise it, or to branch on its type -- use `traceback.clear_frames()` on its traceback to drop the locals, or `exc.with_traceback(None)` to drop the traceback entirely.
  • How does a stored exception show up when you are walking referrers during a leak hunt?
    The climb from a leaking object terminates at a `frame` or a `traceback` object rather than at a container you can name. That is the signature: something is holding a stack, not a collection. The two usual causes are an exception stored beyond its handler and a generator that was started but never exhausted or closed, and you distinguish them by looking at what the frame belongs to.

Filing a caught exception is like keeping the whole crime scene rather than the photograph: the report you wanted is a page of text, but the object you stored is every room, still cordoned off.

saying these in an interview costs you the question

  • Thinks an exception object is just a type and a message
  • Stores exception objects in a retry list without clearing frames
  • Calls the unbound except-as name a language quirk
  • Assumes the frame dies when the function returns, regardless
  • Believes only the failing frame is retained, not its callers
  • Says the cyclic collector makes stored exceptions harmless

context