Why is catching RecursionError an unreliable way to recover from runaway recursion?
answer
- You catch it while still deep
- The margin for handlers is small
- Cleanup code can overflow again
- A second overflow is not catchable
- It inherits from a broader error
basics
~20 sWhen RecursionError is raised you are still at maximum depth, with only a small margin of extra frames. Handlers, finally blocks and finalizers run inside that margin, so cleanup can overflow again and the interpreter aborts outright.
solid answer
~50 s`RecursionError` is raised at the moment the ceiling is crossed, so the handler machinery, `finally` blocks and `__exit__` methods all run while the stack is still essentially full. CPython grants a small headroom of extra frames so the exception can propagate; exceed it and the interpreter gives up with a fatal error and aborts the process, with no traceback and no cleanup. That means the code most likely to fail is exactly your resource release: a `finally` that serializes a message or a finalizer that logs through a deep call chain. It is also easy to swallow by accident, since `RecursionError` subclasses `RuntimeError` and a broad `except Exception` in a request loop will catch it. Treat it as a defect signal, catch it only at a shallow frame if at all, and keep cleanup trivial.
code
python · 18 linesclass Connection:
def __init__(self):
self.closed = False
def close(self):
self.closed = True
def walk(node, seen=0):
return walk(node, seen + 1)
conn = Connection()
try:
try:
walk(None)
finally:
conn.close()
except RecursionError:
print("closed after RecursionError:", conn.closed)go deeper
Know that RecursionError is catchable but signals a bug, and that catching an exception is not the same as fixing the code that raised it.
Explain that handlers run while the stack is still nearly full, that only a small headroom of frames is granted, and that RecursionError subclasses RuntimeError so broad handlers catch it.
Demonstrate the operational reasoning: keep release paths trivial, catch only where the stack has unwound, and connect a swallowed RecursionError in a worker loop to a slow resource leak nobody alerts on.
Own the policy: whether services may catch broadly at all, where depth-bounded input validation belongs, and how to make an aborted process observable when no traceback is produced.
## Where you are standing when it fires `RecursionError` is raised at the call that would have crossed the ceiling — which means the thousand frames beneath you are all still there. Nothing has unwound yet. Every `except` clause, every `finally` block, every `__exit__` and every finalizer that runs as the exception propagates executes at whatever depth its own frame sits at, and near the bottom of the propagation that depth is still at or near the limit. CPython allows a small headroom of extra frames past the ceiling precisely so this propagation can happen at all — otherwise raising the exception would itself be impossible. That headroom is tens of frames, not hundreds. If the code running inside it goes deep again, the interpreter has nowhere to go: it stops with a fatal error along the lines of "Cannot recover from stack overflow." and aborts the process. There is no exception on that path, no traceback, no `finally`, and nothing registered for interpreter shutdown runs. ## Why that lands on cleanup specifically The handlers are where cleanup lives, so the failure mode is precisely inverted from what you want: * a `finally` block that closes a connection is fine, because closing is shallow; * a `finally` block that first formats a diagnostic by walking a nested structure is not, because formatting recurses; * an `__exit__` that flushes a buffer through a serializer is a coin toss; * a `__del__` that logs through a delegating wrapper can re-enter deep code at the worst moment, and exceptions in finalizers are reported and discarded rather than propagated. Consider an ad-auction bidder whose per-request handler wraps a broad `except Exception`, returns "no bid" and moves on. A deep input trips `RecursionError`; the handler catches it; but the `finally` that should have returned the upstream connection to the pool did its own deep formatting first and raised again. One descriptor leaks per bad request. Nothing alerts, because the service keeps answering, and a 340-case regression pack never reproduces it — none of those cases is deep enough. Hours later the process cannot open sockets. ## The accidental catch `RecursionError` subclasses `RuntimeError`. Any `except RuntimeError` or `except Exception` in a worker loop swallows it silently. That is worse than crashing, because a runaway recursion is a defect: the loop keeps accepting work while each iteration leaves something behind. If you catch broadly in a worker loop, at minimum log the exception type, and treat `RecursionError` in the log as a bug ticket rather than as noise. ## What to do instead * **Treat it as a defect signal, never as control flow.** Do not use `RecursionError` to discover how deep a structure is, and never retry the same work after catching it: the input has not changed, so the second attempt fails identically. * **Catch shallow.** If you catch it at all, catch it at the top of the worker or request handler, where the stack has already unwound and you have room to log, release and fail the unit of work cleanly. Once the exception has reached a shallow frame the interpreter is entirely healthy again; the danger window is only while you are still deep. * **Keep release paths trivial.** Resource release belongs in context managers whose `__exit__` does nothing but close. Push formatting, serialization and reporting outside the cleanup path so a cleanup can never be the thing that overflows. * **Bound the input, not the ceiling.** If depth is a function of data you did not write, validate the depth on the way in. A rejected oversized document is an operable failure; an aborted process is not. ## Making the abort visible Because the fatal path produces no exception, none of your usual instrumentation sees it: no handler runs, no error is reported to your logging pipeline, and no shutdown hook fires. What you get is a message on standard error and a process that is simply gone. Operationally that means the supervisor's exit status is the only signal, so capture the child's stderr rather than discarding it, alert on abnormal termination as a distinct condition from a clean non-zero exit, and treat a worker that vanished without a traceback as a strong hint to look for recursion depth rather than for a memory limit. ## The distinction to state out loud There are two failures wearing the same name. `RecursionError` raised and propagated cleanly is a well-behaved, catchable exception, and CPython works hard to make that the outcome. A second overflow *while handling the first* is not an exception at all: it is the interpreter conceding. Everything in the advice above is aimed at staying in the first case, and the recognition that catching does not by itself guarantee it is what separates a considered answer from a confident one.
- Is it ever legitimate to catch RecursionError deliberately?Yes, at a shallow frame: a request handler or worker loop that logs the failure, releases the unit of work and moves on, or a driver deliberately probing how deep a structure goes. What is not legitimate is retrying the same input, or catching it deep inside the recursion where there is no room to act.
- Why can a __del__ method make this considerably worse?Finalizers run at an unpredictable point, possibly while the stack is still deep, and exceptions raised inside them are reported and discarded rather than propagated. A `__del__` that logs through a deep call chain can therefore overflow again during unwinding, and the failure that hides the original one leaves no usable trace.
- What is the difference between the exception and the fatal error?Crossing the ceiling raises `RecursionError`, an ordinary exception that unwinds and runs cleanup. Overflowing again inside the small headroom granted for that unwinding is not an exception: the interpreter aborts the process immediately, with no traceback, no `finally` and no shutdown hooks.
It is like being handed a fire extinguisher while still standing in the doorway of the burning room: you can act, but only briefly and only if the action is small.
saying these in an interview costs you the question
- Uses RecursionError as ordinary control flow
- Retries the same input after catching it
- Assumes cleanup always runs after RecursionError
- Thinks a second overflow raises another exception
- Swallows it in a broad except Exception without logging
- Believes catching it makes the recursion safe