skip to content

What does a Python traceback stop telling you after a recursive walk becomes a `while` loop?

level: seniorimportance: should knowfreq 32%

answer

  1. a traceback lists the active frames
  2. one frame now, same lines every iteration
  3. the identity is no longer carried for free
  4. continue without recording is a bug
  5. attach a note, collect, raise a group

basics

~20 s

Recursion put one frame per level into the traceback, so it showed the route to the failing record. A loop runs every level in one frame, so you learn where the code failed but not which item.

solid answer

~50 s

A traceback is a list of the active frames, so a recursive walk records its whole descent — and CPython collapses repeated frames into a `[Previous line repeated N more times]` line, so even a deep one stays readable. The loop version has a single frame and processes every level on the same source lines, so the traceback identifies the code and nothing about the record. It gets worse when the loop wraps each item in `except Exception: continue` to survive bad input: a museum-catalogue importer running through a 1,200-record-per-minute peak then finishes clean, reports success and silently omits an object. Rebuild the diagnostics deliberately: push the node's identity and parent path onto the stack entry, attach it with `BaseException.add_note()` or by re-raising with `from`, log it with `logging.exception()`, and collect the failures so you can raise an `ExceptionGroup` at the end instead of dropping them.

code

python · 19 lines
python
import traceback

children = {"root": ["ceramics", "textiles"], "ceramics": ["urn:bad"],
            "textiles": [], "urn:bad": []}

stack = [("root", ("root",))]
failures = []
while stack:
    node, path = stack.pop()
    try:
        if node.startswith("urn:"):
            raise ValueError("unparsable accession number")
    except ValueError as exc:
        exc.add_note("record path: " + "/".join(path))
        failures.append(exc)          # collected, not swallowed
        continue
    stack.extend((child, path + (child,)) for child in children[node])

print("".join(traceback.format_exception_only(failures[0])))

go deeper

for a junior

Remember that a traceback lists the frames that were active. Recursion creates one frame per level, so it shows the path; a loop has a single frame, so its traceback repeats the same few lines.

for a middle

Explain what to push onto the stack entry so a failure can identify itself, and why a bare except Exception: continue inside the loop leaves you with nothing at all to debug.

for a senior

Demonstrate the recovery design end to end: identity on the entry, a note attached to the exception, a record identifier in the log, and failures collected and re-raised as a group rather than swallowed.

for a principal

Own the policy. Decide which failures abort a batch and which are collected, what the diagnostic contract for an unattended job is, and at what failure rate someone gets paged rather than reading it in a report.

### What a traceback is made of A Python traceback is a list of the frames that were active when the exception propagated, innermost last. That structure is the reason a recursive failure is so easy to debug: each level of the recursion is a separate frame, so the traceback literally prints the path the walk took to reach the record that blew up. Even when the recursion is a thousand levels deep the output stays readable, because CPython's traceback formatting collapses runs of identical frames into a single `[Previous line repeated N more times]` line — behaviour that has been there since 3.6 — and you still get the innermost frames, which is where the bad data usually is. ### What the loop version shows instead Convert the walk to a `while` loop over an explicit stack and all of that disappears. There is exactly one frame now, and every level of the structure is processed by the same handful of source lines. The traceback tells you *where in the code* the failure happened — which is often obvious anyway — and nothing at all about *which record* caused it. The path from the root, the depth, the parent node: those used to be frame-local variables that the traceback carried for free, and after the rewrite they are just values in a tuple that the traceback knows nothing about. The failure mode compounds when the loop is written to survive bad input. A batch importer normally must not die on one malformed record, so the loop grows a `try` around the per-record work — and the tempting shape is `except Exception: continue`. Now the exception is not merely under-described, it is gone. A museum-catalogue importer walking nested collections into objects, running through a 1,200-record-per-minute peak, will finish cleanly, report success, and quietly omit an object whose accession number failed to parse. Nobody notices until a curator asks why a piece is missing from the collection page, months later. ### Putting the context back deliberately The frames were doing diagnostic work you now have to do yourself, and it is a small amount of work: **Push the identity.** Whatever names the node — its key, its accession number, the tuple of parents that got you there — goes on the stack entry next to the node and is unpacked on pop. If it is not on the entry, it does not exist when the failure happens. **Attach it to the exception.** `BaseException.add_note()` (3.11, PEP 678) appends a line that every formatter prints beneath the exception message, which is ideal for a record path; the alternative on older releases is re-raising a wrapping exception with `raise WrapperError(...) from exc`, which keeps the original as the cause. Either way the identifying context travels *with* the exception rather than only in a log line somewhere else. **Never discard.** Catching an exception to keep the batch going is a legitimate policy; catching it and dropping it is a bug. Append the caught exception to a list. At the end, if the list is non-empty, either raise an `ExceptionGroup` (3.11, PEP 654) carrying every failure — a caller can then handle subsets with `except*` — or log each one and fail the job's exit status. Both make the swallowed failure visible; the bare `continue` makes it invisible. **Log with the identifier.** `logging.exception()` inside the handler writes the message plus the traceback at error level, and the record path belongs in that call's message or in its structured extra fields. A log line reading "ValueError: unparsable accession number" with no record identifier, repeated forty times, is worth almost nothing at 3 a.m. ### A shape that keeps all of it ```python failures = [] while stack: node, path = stack.pop() try: ingest(node) except Exception as exc: exc.add_note("record path: " + "/".join(path)) failures.append(exc) continue stack.extend((child, path + (child,)) for child in children_of(node)) if failures: raise ExceptionGroup("catalogue import failed", failures) ``` The import still finishes. The failures still surface. Each one names the record. That is the whole contract the frames used to provide. ### Judgement, not ceremony Not every loop needs this. A short, self-contained iteration whose failures are already obvious from the message does not. The rewrite that deserves the treatment is the one that replaced a recursion over data you did not author — a nested import, a crawl, a config expansion — because that is exactly the case where the failing item is one of very many and the code's own line number tells you nothing. The question to ask of any recursion-to-loop rewrite is simply: when this raises on the ten-thousandth record, will the output say which record? If the answer is no, the diagnostics have to be rebuilt by hand before the rewrite ships.

  • The loop catches per-record failures and continues so the batch finishes — how do you keep that without losing them?
    Collect instead of discard. Append each caught exception, with the record path attached via `add_note()`, to a list, and when the walk finishes either raise an `ExceptionGroup` carrying all of them or log each at exception level and fail the job's exit status. Continuing past a bad record is a policy decision; silently dropping the exception is a bug that turns a data error into missing data.
  • Why did the recursive version's traceback stay readable a thousand levels deep?
    CPython's traceback formatting collapses consecutive identical frames into a `[Previous line repeated N more times]` line, so a deep uniform recursion prints its head and tail rather than a thousand entries. That has been the behaviour since 3.6. You still get the innermost frames, which is normally where the failing data is.
  • What exactly should go on the stack entry so a failure is self-describing?
    Whatever names the item to a human: its key or accession number, the tuple of parents that led to it, the depth, the source file it came from. Push a tuple of the node plus that context, unpack it on pop, and put it into the note, the log record and any wrapping exception. The frames used to carry that for free; after the rewrite it is your job.

Recursion leaves a trail of breadcrumbs, one per turn taken; the loop walks the same maze without dropping any, so when it trips you know it fell over in the maze but not where.

saying these in an interview costs you the question

  • Assumes the loop's traceback still shows the path to the node
  • Writes except Exception: continue and records nothing
  • Thinks a deep recursion prints one traceback line per level
  • Logs the exception message with no record identifier
  • Aborts the whole batch on the first bad record by reflex
  • Puts the context only in a log line, never on the exception

context