Why capture a failure as a traceback.TracebackException instead of keeping the exception?
answer
- Holding it keeps more than you think
- Frames keep locals, locals keep everything
- Capture on one thread, render on another
- A flat report, not a live object
- Source lines resolved at capture time
basics
~10 straceback.TracebackException.from_exception() pre-renders everything a report needs — frames, source lines and the chained exceptions — and keeps no live frames. Holding the exception instead pins every frame's locals in memory until you release it.
solid answer
~40 sAn exception object references its traceback through `__traceback__`, and that traceback references live frames, whose locals reference whatever the failing call was working on. Park a few hundred of those in a list and you have pinned a few hundred request-sized object graphs. `traceback.TracebackException.from_exception(exc)` walks the traceback once at capture time, records each frame as a flat summary, resolves the source lines, and follows `__cause__` and `__context__` into nested `TracebackException` objects — then holds **no frame references at all**, so the failing call's memory becomes collectable immediately. Later, `.format()` yields the same report text the interpreter would have printed and `.format_exception_only()` yields the summary line. The trade is that you have a report, not an exception: you cannot re-raise it, and the source lines are frozen as they were at capture.
code
python · 20 linesimport queue
import threading
import traceback
failures = queue.Queue()
def worker(batch):
try:
raise TimeoutError(f"batch {batch} stalled")
except TimeoutError as exc:
failures.put(traceback.TracebackException.from_exception(exc))
for batch in range(3):
t = threading.Thread(target=worker, args=(batch,))
t.start()
t.join()
while not failures.empty():
captured = failures.get()
print("".join(captured.format_exception_only()), end="")go deeper
Know that an exception carries its traceback and its frames with it, and that there is a traceback-module object that turns a failure into a plain report you can keep and print later.
Explain the capture: frames flattened, source lines resolved, chain followed, no frame references kept. Show the format call that renders it and say why capture and render are separate steps.
Demonstrate the production reasoning — retained exceptions pinning request-sized object graphs, capture on the worker and render on the writer, and the honest costs of line lookups on a hot error path.
Own the policy: whether every failure is captured or sampled, what the retention ceiling on collected reports is, and what structured facts are extracted at capture time so later stages can decide without the live exception.
### The retention problem An exception is not a small object once it has been raised. `exc.__traceback__` points at a chain of traceback objects, each pointing at a live frame, each frame holding its locals — the request body being parsed, the batch being ingested, whatever was in flight. That is by design and it is exactly what makes a traceback useful. It also means **an exception you keep is a call stack you keep**. CPython already works around this in the small: `except SomeError as exc:` deletes the name `exc` when the block ends, precisely so a handler does not leave a reference cycle and a pinned stack behind. The moment you defeat that — by appending the exception to a list, putting it on a `queue.Queue`, or storing it on `self` for a report later — you have taken ownership of the retention. In a long-running collector this shows up as steadily climbing memory with no obvious owner. A batch job in a log-ingest pipeline that gathers every failure from a 27-minute run and formats them at the end is the archetype: three hundred failures, each holding the frames that were parsing a record, is three hundred records that never leave the heap. ### What TracebackException does instead `traceback.TracebackException.from_exception(exc)` performs the whole capture at the point of failure: - it summarises each frame into a flat record — file, line number, function name, source text — with no reference to the frame object; - it resolves the source lines then and there, so the report is correct even if the file is later replaced or the process moves on; - it captures the exception's type and message as text; - it follows `__cause__` and `__context__` recursively, building nested `TracebackException` objects, and honours `__suppress_context__`, so the chain that would have been printed is preserved; - for an `ExceptionGroup` it captures the contained exceptions the same way, so the group tree survives. What comes back is a self-contained, pure-data description of a failure. Holding a thousand of them costs the text of a thousand reports — large, but bounded and predictable, and unrelated to how big the objects in the failing frames were. ### Rendering later `.format()` returns an iterable of strings; joining them gives the report the interpreter would have printed, chain and all. `.format_exception_only()` gives just the type-and-message part, which is what you want for a summary line or a grouping key. Rendering is a separate step from capture, which is the whole point: capture on the failing thread where the exception exists, render on whatever thread, process or later phase actually writes the report. That separation is what makes it the right tool for a worker pool. Each worker catches, converts and puts the `TracebackException` on a queue; one writer thread drains the queue and renders. The workers never hold a stack, and the writer never needs access to the frames. ### The costs, honestly Capture is not free. Resolving source lines means file reads through the line cache, on the failing path, for every frame. On a hot error path that is real work, and if failures are the common case you should be sampling rather than capturing every one. You also lose the exception. A `TracebackException` cannot be re-raised, has no `args`, and cannot be caught by type by a caller downstream — it is a report. If a later stage needs to *decide* something based on the failure, it needs the class or a code you extracted at capture time, not the rendered text. And the source lines are frozen. Usually that is the point — a report generated an hour later still shows the code that actually ran. But it means the capture reflects the file as of capture time, and reading the report will not show you an edit made since. ### When you do not need it If you are formatting the failure immediately inside the handler — log it and let it go — then `format_exc()` is simpler and there is nothing to hold. `TracebackException` earns its place exactly when the report must outlive the handler: crossing a thread boundary, waiting in a batch, or being rendered by code that will never see the exception itself.
- What exactly does holding the exception object keep alive?The exception references its traceback through `__traceback__`, the traceback references each live frame, and every frame keeps its locals — and transitively everything those locals point at. So one retained exception can pin an entire request- or record-sized object graph, which is why an `except E as exc` binding is deleted automatically when the block ends.
- What do you give up by converting a failure into a TracebackException?You give up the exception itself: it cannot be re-raised, it has no args, and downstream code cannot catch or branch on its class. Anything a later stage must decide on has to be extracted at capture time. Capture also costs source-line lookups on the failing path, which matters if failures are frequent.
- Does the capture preserve chained exceptions?Yes. It follows `__cause__` and `__context__` recursively into nested TracebackException objects and honours `__suppress_context__`, so calling `.format()` later reproduces the same chained report, banners included. Exception groups are captured the same way, so the group tree survives to render time.
- When is capturing overkill?When the report never outlives the handler. If you log the failure right there and move on, `traceback.format_exc()` is simpler, holds nothing and costs less. The capture object earns its keep only when the report must cross a thread, a queue or a phase boundary away from the exception.
saying these in an interview costs you the question
- Sees no memory cost in keeping exception objects
- Thinks a TracebackException can be re-raised
- Claims capture is free on the failing path
- Says the chain is lost unless captured separately
- Assumes source lines are stored inside the traceback