skip to content

How do you capture post-mortem state from a headless nightly ad-auction bidder that crashes once a week?

level: seniorimportance: should knowfreq 30%

answer

  1. No terminal, so no prompt
  2. Turn the crash into an artefact
  3. Format the frames, values included
  4. One stdlib class does the capture
  5. Locals in logs is a disclosure risk

basics

~20 s

There is no terminal to hold a debugger prompt, so capture instead of attach: at the top-level handler, serialize the exception's frames with their locals - traceback.TracebackException with capture_locals=True - and write that artefact where you can read it later.

solid answer

~40 s

A prompt is useless in an unattended job, so the goal is to turn the one crash you get into a readable artefact. Wrap the entry point (or install a `sys.excepthook`, plus `threading.excepthook` if worker threads can die) and, from the caught exception, build `traceback.TracebackException.from_exception(exc, capture_locals=True)`; formatting it yields the usual traceback with **each frame's local variables inlined**, which is exactly what you would have printed at a post-mortem prompt. Emit it to your log or an artefact file, then reproduce offline and use `pdb.post_mortem()` there. Two cautions: locals get `repr()`-ed, so secrets, tokens and raw payloads land in logs and huge objects bloat them; and holding the exception keeps every frame alive, so capture, emit, and drop the reference.

code

python · 13 lines
python
import traceback


def apply_bid(bids, index):
    factor = 0.83
    return bids[index] * factor


try:
    apply_bid([12, 30], 5)
except IndexError as exc:
    report = traceback.TracebackException.from_exception(exc, capture_locals=True)
    print("".join(report.format()))

go deeper

for a junior

Take away the core distinction: a logged traceback shows lines, not values. Knowing that Python can print a traceback with each frame's variables included, and that this is what you want for a crash nobody watched, is enough here.

for a middle

Be able to write the capture: catch at the entry point or install a hook, build a TracebackException with capture_locals=True, format it, emit it. Explain why the repr() is taken eagerly at capture time.

for a senior

Show that you have operated this. Cover both excepthooks, the disclosure risk of locals in logs, the size and custom-__repr__ hazards, and the retention trap of holding exceptions - then how the capture drives an offline reproduction.

for a principal

Own the standard: what every service captures on failure, where those artefacts live and for how long, who may read them, and how you keep interactive debuggers structurally impossible in unattended paths.

### Why the interactive answer does not apply A nightly bidder runs under a scheduler with no controlling terminal. `pdb.post_mortem()` there reads from a standard input that is not a tty: at best the job hangs holding a slot until something reaps it, at worst it dies on end-of-input with a confusing secondary error and you lose the original failure. Any interactive debugger in an unattended path is an outage waiting to happen. The technique that survives is the same idea applied without a human: extract from the traceback, at the moment of the crash, everything the human would have typed at the prompt. ### The scenario decides what you need Take the concrete failure: the bidder assumes the bid stream arrives in priority order, and with an 83% cache-hit rate the cached path preserves that order almost always. The 17% that miss the cache occasionally interleave, and one interleaving trips an index error deep in the settlement code. Once a week, in the middle of the night. Re-running does not reproduce it, because the input that produced the ordering is gone. The only thing that will ever answer 'why did this fail' is the state that existed when it failed - the ordering as the raising frame saw it. That is a capture problem, not a debugging-session problem. ### The capture `traceback.TracebackException.from_exception(exc, capture_locals=True)` walks the exception's frames and, for each one, records the `repr()` of every local variable at that moment. Its `format()` produces the familiar traceback with those values inlined under each frame: ``` File "bidder.py", line 5, in apply_bid return bids[index] * factor bids = [12, 30] factor = 0.83 index = 5 ``` That is a post-mortem inspection performed by the program on your behalf. It is deliberately eager: the `repr()` is taken at capture time, while the frames are still referenced, so the artefact is safe to keep after the exception is dropped. ### Where to hook it Wrapping the entry point in `try/except/raise` is the simplest and keeps the failure semantics intact. `sys.excepthook` catches anything that escapes the main thread even from code you did not wrap; `threading.excepthook` is its per-thread counterpart, and forgetting it is why 'the job logged nothing' happens when the crash was in a worker. Whichever you use, emit the artefact and then let the exception do what it was going to do, so exit status and alerting still work. ### The costs, which an interviewer is listening for **Disclosure.** Frame locals are where credentials, tokens, personal data and raw request bodies live. Capturing them writes all of it to wherever the log goes. Realistic mitigations: capture to a restricted artefact store rather than the general log stream, redact by name, or capture only for specific exception types. **Size.** `repr()` of a large container can be enormous, and one crash can produce megabytes. Cap the depth, and be aware that a custom `__repr__` runs your code during failure handling - it can itself raise, or be slow. **Retention.** As long as you hold the exception, its traceback holds the frames, which hold everything they reference. Capturing into a long-lived structure ('keep the last 50 failures') quietly retains 50 whole stacks. Format, emit, drop. ### What else the artefact needs A capture that is not findable is not useful. For a job that fails once a week, the capture must trigger an alert of its own, carry the run identifier and timestamp, and record enough about the inputs - which batch, which shard, which offsets - that you can reconstruct the conditions rather than just admire the values. It is worth writing the artefact before doing anything else in the handler, because cleanup code that runs first can itself fail and lose the original evidence. ### Then reproduce The captured locals normally tell you the input shape that broke; you write a test that constructs it, run that locally, and at that point the interactive tools are appropriate again - `pdb.post_mortem()` on the reproduction, with a terminal, is fine. The production side stays capture-only. If you want an interactive escape hatch for a developer-run invocation, gate it on both an explicit environment switch and a tty check, so no scheduled run can ever open a prompt.

  • What is the main risk of shipping captured frame locals to your normal log stream?
    Disclosure. Locals are exactly where credentials, tokens, session identifiers and raw payloads sit, and capture takes the `repr()` of all of them indiscriminately, so a crash exports secrets into a log that has far wider read access than production data does. Send captures to a restricted artefact store, redact by variable name, or capture only for the exception types you actually need.
  • Why can holding onto captured exceptions cause a memory problem?
    An exception's traceback holds its frame objects, and each frame holds every object its locals reference. Keeping a ring buffer of recent exceptions therefore keeps whole stacks - request bodies, buffers, connections - alive indefinitely. Format the exception into text, emit it, and drop the reference; keep the rendered artefact, never the live exception.
  • If the crash happens in a worker thread, does a `sys.excepthook` see it?
    No. `sys.excepthook` handles exceptions that escape the main thread; an exception that escapes a worker's target function goes to `threading.excepthook` instead, which by default just prints. Install both if worker threads can die, otherwise the symptom is a job that logs nothing useful and looks like it simply stopped doing work.

You cannot interview a witness who has left, so you photograph the scene while it is intact: the capture is the photograph, and the offline reproduction is where you finally get to ask questions.

saying these in an interview costs you the question

  • Puts an unguarded `pdb.post_mortem()` in a scheduled job
  • Assumes a plain logged traceback shows variable values
  • Ignores that captured locals can contain secrets
  • Keeps exception objects around and leaks whole frames
  • Plans to just re-run until the rare failure reappears
  • Installs `sys.excepthook` only and misses worker threads

context