Is sys.setrecursionlimit a safe fix for a RecursionError raised while walking a deep chat transcript?
answer
- Ask whether the depth is bounded first
- The setting is process-wide, not local
- Another guard sits under the Python one
- C-level recursion ignores the Python ceiling
- Bound the input, not just the ceiling
basics
~20 sOnly once you have proved the depth is bounded by the data rather than by a cycle. A higher ceiling lifts the Python frame count, but recursion that runs through C code still hits a separate stack guard.
solid answer
~50 sFirst establish whether the depth is a property of the input or a runaway: a transcript archiver that recurses per reply is bounded by the longest chain, whereas a back-reference makes it infinite and a bigger number only buys a slower failure. If it is genuinely deep data, measure the worst real chain, then set the limit once at process start — never inside a library or a request path, because `sys.setrecursionlimit` is interpreter-wide and everything in the process shares it. Understand what it does not cover: since 3.12 CPython guards the C stack separately, so recursion through C code — the `repr` of a deeply nested structure, deep copying, serialising — raises `RecursionError` with a stack-overflow message no matter how high you set the Python limit. The durable fix is to bound nesting depth at ingest, or to walk the structure with an explicit stack.
code
python · 17 linesimport sys
sys.setrecursionlimit(60_000)
def depth(n=1):
return n if n == 50_000 else depth(n + 1)
print("pure Python frames:", depth()) # fine: no C frame per Python call
nested = []
for _ in range(50_000):
nested = [nested]
try:
repr(nested) # this recursion runs in C
except RecursionError as exc:
print("C-level guard:", exc)go deeper
Know that sys.setrecursionlimit exists and that reaching for it first is the wrong instinct: a RecursionError usually means a missing base case or a cycle, and only sometimes means genuinely deep data.
Explain what the call changes — an interpreter-wide ceiling on the per-thread Python frame count — and why it belongs in the entry point rather than a library. Mention that memory is allocated only as frames are actually created.
Show the diagnosis you would run: reproduce with the offending input, read the collapsed traceback, instrument actual depth, then decide between an input bound, an explicit stack and a raised ceiling. Name the C-level guard that a higher Python limit does not lift.
Own the boundary decision: which component is responsible for rejecting pathologically nested input, what the documented maximum is, and how you avoid a process-wide interpreter setting becoming a load-bearing dependency for one subsystem.
## Diagnose before you tune The first question is not *how high* but *is the depth bounded at all*. A chat-transcript archiver that recurses once per reply in a thread has depth equal to the longest reply chain, which is a real, measurable, data-dependent number. The same archiver walking a transcript where one message quotes an ancestor has unbounded depth, and raising the ceiling converts a fast, loud failure into a slow one that burns memory first. So: reproduce with the offending archive, read the collapsed traceback (a single repeated line means straight runaway; alternating lines mean mutual recursion), and instrument the walk with an explicit depth counter so you know the real distribution rather than the one your fixtures happen to have. This class of bug is characteristically invisible in testing — a 27-minute suite whose deepest fixture nests a few dozen replies proves nothing about a production thread nesting tens of thousands. ## What sys.setrecursionlimit actually changes It raises the interpreter-wide ceiling that the per-thread Python frame counter is compared against. Three properties matter operationally: 1. **It is global to the interpreter, not scoped.** Every thread and every library in the process shares the new value. Setting it from a library or from a request handler silently changes behaviour for code that never asked, and a component that relies on `RecursionError` firing early as a guard against hostile input loses that guard. Set it once, in the application entry point, with a comment saying why. 2. **It does not allocate anything.** Frames are heap-allocated as needed, so raising the limit does not reserve memory up front; the cost shows up as the memory a genuinely deep recursion consumes while running. 3. **It does not lift the C-stack guard.** This is the part people miss, and it is covered below. The historical fear — that a high limit crashes the interpreter — was accurate on older versions. Through 3.10 each Python-to-Python call also nested a C frame in the evaluation function, so a limit far above the default really could walk off the C stack and take the process down with a segfault. 3.11 inlined Python-to-Python calls, 3.12 introduced a separate C-level recursion accounting, and 3.13 and 3.14 check remaining C stack space directly and raise `RecursionError` instead of crashing. On 3.14 a pure-Python recursion 50,000 frames deep runs fine on the main thread once the limit allows it. ## The guard that a raised limit does not lift Recursion that passes through C code still consumes real C stack per level, and CPython guards that separately. Building the `repr` of a deeply nested container, comparing or hashing deeply nested structures, deep-copying them, serialising them, and compiling deeply nested source all recurse in C. On 3.14 they raise `RecursionError` carrying a message like `Stack overflow (used 3906 kB) while getting the repr of an object` — the same exception class, a different mechanism, and completely indifferent to your Python-level limit. For an archiver this is exactly where the bug reappears: the walk itself is now fine at depth 50,000, and then serialising the structure it produced fails. If your pipeline builds a deeply nested object and then hands it to any C-implemented serialiser, comparison or copy, the nesting depth is a hard constraint on the *data shape*, not merely on your traversal code. ## threading.stack_size, and why it buys less than it used to `threading.stack_size(n)` requests a stack size for threads started *afterwards*; it has no effect on threads that already exist, including the main thread. Combined with a raised limit, it was the standard workaround in the era when every Python call nested a C frame: run the deep work in a worker thread with a large stack. On 3.14 that reasoning is much weaker, because pure-Python recursion no longer grows the C stack per call, and the interpreter's C-stack guard does not simply scale with the size you requested — a measurement on a 3.14 build showed the C-recursion ceiling essentially unchanged between the main thread and a thread created with a far larger stack. Treat it as something to measure on your platform and build, not as a lever you can assume works. ## The fixes that hold Ranked by how well they survive contact with hostile input: - **Bound the depth where the data enters.** Validate nesting at ingest and reject or flatten beyond a documented maximum, with a domain error that names the limit. This is the only option that also protects the C-level operations downstream. - **Walk with an explicit stack** so traversal depth stops being a stack-space question at all. - **Raise the limit**, with a measured margin over the worst observed depth, set once at startup, and paired with an input bound so the ceiling is a safety net rather than the mechanism. Whatever you choose, make the failure loud. A truncated archive that no one noticed is worse than an archive that failed with a clear error.
- Where in an application should sys.setrecursionlimit be called?Once, in the entry point, with a comment recording the measured depth it is sized for. The setting is interpreter-wide, so calling it from a library or a request handler changes behaviour for every other component in the process — including ones relying on the default as a guard against hostile input — and makes the effective limit depend on import order or request timing.
- The walk now runs 50,000 deep, but serialising the result still raises RecursionError. Why?Because that recursion is happening in C, not in Python frames. Building a repr, deep-copying, comparing or serialising a deeply nested structure recurses in the interpreter's own C code, which CPython guards by checking remaining C stack space. It raises the same exception class with a stack-overflow message regardless of the Python-level limit, so the nesting depth of the data itself is the constraint.
- Does running the walk in a thread created after threading.stack_size(64 << 20) solve it?It must be called before the thread is created and affects only threads started afterwards, so the main thread is never covered. It was the standard companion to a raised limit when every Python call nested a C frame. On 3.14 pure-Python recursion no longer grows the C stack per call and the C-stack guard does not scale straightforwardly with the requested size, so measure before depending on it.
saying these in an interview costs you the question
- Raises the limit to a huge number without measuring real depth
- Calls sys.setrecursionlimit from a library or a request handler
- Believes a higher Python limit also lifts the C-level stack guard
- Thinks threading.stack_size affects threads that already exist
- Ships a raised ceiling with no bound on input nesting depth
- Assumes a suite of shallow fixtures proves the depth is safe