skip to content

Why can catching RecursionError inside a recursive walk silently truncate its output?

level: middleimportance: nice to knowfreq 22%

answer

  1. The interpreter survives, which is the problem
  2. Partial results look exactly like complete ones
  3. The handler itself runs near the ceiling
  4. A RuntimeError subclass, so broad catches grab it
  5. Handle at the unit boundary, never mid-recursion

basics

~20 s

The handler swallows the failure at whatever depth it struck and returns whatever was collected so far. CPython unwinds cleanly, so nothing crashes and no error surfaces — the caller simply receives a short, wrong result.

solid answer

~40 s

`RecursionError` is an ordinary exception: it propagates outward, frames unwind, the depth counter falls back, and the process carries on. That recoverability is exactly what makes a handler placed *inside* the recursion dangerous. A collector that wraps its recursive calls in `except RecursionError: pass` returns the partial list it had built when the limit hit, and the caller cannot tell a truncated result from a complete one. It gets worse: `RecursionError` subclasses `RuntimeError`, so a broad `except Exception` written for something else swallows it too, and the handler itself runs while still deep on the stack, so logging or cleanup inside it may immediately re-raise. Either let it propagate, or catch it at the boundary of a unit of work and fail that unit loudly.

code

python · 22 lines
python
import sys

sys.setrecursionlimit(120)

def collect(node, out):
    out.append(node["id"])
    try:
        for kid in node["replies"]:
            collect(kid, out)
    except RecursionError:
        pass          # "handled" -- and silently truncated
    return out

chain = {"id": 0, "replies": []}
cur = chain
for i in range(1, 500):
    nxt = {"id": i, "replies": []}
    cur["replies"].append(nxt)
    cur = nxt

out = collect(chain, [])
print("archived", len(out), "of 500 messages; no error surfaced")

go deeper

for a junior

Take away one rule: do not wrap a recursive call in a handler that swallows the error. If the walk cannot finish, the caller must find out rather than receive a short list that looks complete.

for a middle

Be ready to explain the mechanics — the exception propagates, frames unwind, the depth counter recovers — and why that clean recovery is what turns a caught RecursionError into silently wrong data rather than a crash.

for a senior

An interviewer expects the operational view: broad except Exception handlers in retry loops swallowing it, handlers that fail while logging because they run near the ceiling, and the practice of translating it at the unit-of-work boundary into a domain error naming the depth.

for a principal

Own the standard: which failures a service is permitted to absorb silently, how truncation must be represented in output when it is allowed at all, and how you keep broad catch-all handlers from converting resource exhaustion into quiet data loss.

## Why the recovery is what hurts you When the frame counter crosses the ceiling, CPython raises `RecursionError` from the call that would have crossed it. From there it is an exception like any other: it propagates up until some handler catches it, `finally` blocks and context managers run as the frames unwind, and the per-thread depth counter comes back down with them. Nothing about the interpreter is left damaged, and the process keeps running. That clean recovery is precisely the trap. A crash is loud; a caught exception is silent, and if the handler is sitting *inside* the recursion, the value that comes back out is a partial result wearing the shape of a complete one. Concretely, a transcript archiver that walks a reply chain and collects message ids, wrapping its recursive descent in `except RecursionError: pass`, returns however many ids it managed before the limit struck. There is no exception at the call site, no log line unless someone wrote one, and the archive it writes looks structurally valid. The data loss is discovered weeks later, if at all — a silent truncation is the worst outcome of the three available (crash, error, wrong data), because it is the only one that leaves no trace. ## Three mechanics worth being able to explain **The handler runs while still deep.** The `except` block executes in the frame that caught the exception, which is still near the ceiling. Anything it calls — a logger, a formatter, a metrics client, a cleanup routine — needs frames it may not have, so it can immediately raise a second `RecursionError` from inside the handler. This is why recovery code placed deep in a recursion behaves erratically: sometimes it logs, sometimes it fails while logging, and the second failure carries a confusing traceback with the first one chained onto it. **Broad handlers catch it by accident.** `RecursionError` subclasses `RuntimeError`, which subclasses `Exception`. A worker loop written as a bare `except Exception` that logs a warning and continues, or a retry decorator that catches `Exception`, swallows a genuine stack overflow and retries an operation that will fail identically every time. The result is a job that appears to be working — it logs, it retries, it never succeeds — rather than one that fails visibly. **The counter does not stay pinned.** People sometimes assume the interpreter is permanently degraded after a `RecursionError`. It is not: depth is restored as frames pop, and subsequent work at normal depth runs normally. Older CPython versions tracked an overflow flag and could turn a mishandled overflow into a fatal interpreter error; modern CPython simply recovers, which removes the crash and leaves the quiet failure. ## What to do instead 1. **Do not catch it inside the recursion.** Let it propagate to the boundary of a unit of work — one archive, one request, one file — and fail that unit with an error that says the input was too deeply nested. 2. **Translate it at that boundary.** Convert it into a domain-level exception naming the artefact and the observed depth, so operators see *"transcript thread exceeded the supported nesting depth"* rather than a bare stack traceback. 3. **Count depth explicitly if partial results are genuinely acceptable.** Pass a depth argument, stop at a documented maximum, and return the truncation as *data*: a flag, a count of skipped items, a marker in the output. A truncation you can see is a design decision; a truncation you cannot see is a bug. 4. **Narrow the broad handlers.** In retry loops and worker pools, either re-raise `RecursionError` explicitly before the generic `except Exception`, or catch the specific exceptions you actually intended to handle. 5. **Test with adversarial depth.** A suite whose fixtures nest a few dozen levels will never reach the limit no matter how long it runs; add at least one input deeper than the ceiling and assert the loud failure you designed. The underlying principle is the one every silent-data-loss bug teaches: an exception you catch without changing the caller's understanding of the result is not error handling, it is error hiding.

  • Why can a retry loop written as `except Exception` hide a recursion overflow?
    Because `RecursionError` subclasses `RuntimeError` and therefore `Exception`. The loop catches it, logs a generic warning, and retries an operation whose input is unchanged, so it fails at exactly the same depth every attempt. The job looks alive and makes no progress. Either re-raise `RecursionError` explicitly ahead of the broad handler, or catch only the exceptions the loop was written for.
  • If catching it inside the recursion is wrong, where should the handler live?
    At the boundary of one unit of work — a single archive, request or file — where failing is meaningful. There you can translate it into a domain error that names the artefact and the depth observed, record the failure, and move to the next unit. The distinction is that the caller learns the unit failed, instead of receiving a short result it cannot distinguish from a complete one.
  • Is the interpreter left in a bad state after a RecursionError is caught?
    No. The frames unwind, `finally` blocks and context managers run, and the per-thread depth counter falls back with them, so work at ordinary depth continues normally. The one caveat is that the handler itself executes while still near the ceiling, so calls it makes — logging, formatting, cleanup — can raise a second RecursionError before the stack has unwound.

It is a scanner that jams halfway through a document and quietly staples the pages it managed: the output is neat, plausible and missing the ending.

saying these in an interview costs you the question

  • Catches RecursionError inside the recursion and returns partial data
  • Assumes a broad except Exception cannot catch RecursionError
  • Claims the interpreter is corrupted after recovering from one
  • Thinks the depth counter stays pinned once the handler runs
  • Treats a short result as valid because no error was raised
  • Logs inside the deep handler and is surprised when logging fails

context