skip to content

When would you use threading.stack_size to run a deeply recursive routine?

level: seniorimportance: nice to knowfreq 14%

answer

  1. The main thread's stack is not yours
  2. Applies to threads created afterwards
  3. Two limits, both must be lifted
  4. Platform minimum and page-size granularity
  5. No help at all against a cycle

basics

~20 s

When depth is genuine and the recursive code cannot be rewritten. threading.stack_size sets the stack for threads created afterwards, so you give a worker thread a large stack, raise the recursion limit, and run the deep work there rather than on the main thread.

solid answer

~50 s

`threading.stack_size(size)` sets the stack size used for threads created *after* the call; with no argument it returns the current setting, and 0 means the platform default. It is the only way from Python to enlarge the real stack a call chain runs on — the main thread's stack is fixed by the OS and the shell's limits, not by Python. So the idiom is: call `threading.stack_size` with a generous value, raise `sys.setrecursionlimit` because the frame counter is a separate constraint, then run the deep work in a `threading.Thread` and join it. It is worth doing when the depth is real and bounded and the recursive code is not yours to restructure — a recursive-descent parser over machine-generated input, a one-shot batch. It is not a fix for a cycle, it applies to every thread created afterwards, and platforms impose a minimum and may require the size to be a multiple of the page size.

code

python · 17 lines
python
import sys
import threading

def deep(n=1):
    if n >= 30_000:
        return n
    return deep(n + 1)

def run():
    sys.setrecursionlimit(40_000)
    print("reached depth", deep())

threading.stack_size(64 * 1024 * 1024)
worker = threading.Thread(target=run)
worker.start()
worker.join()
threading.stack_size(0)

go deeper

for a junior

Know that the recursion limit and the thread's actual stack are two different constraints, and that threading.stack_size touches the second one for threads created after the call.

for a middle

Explain the full idiom — set the stack size, create and join a worker thread, raise the interpreter-wide frame limit — and why the main thread cannot be reconfigured from inside Python.

for a senior

Show when it is justified: bounded, measured depth in code you cannot restructure, with the reservation scoped tightly and reset afterwards, and never as a remedy for a cycle.

for a principal

Weigh it against the durable options — restructuring the traversal, a documented depth bound, or isolating the work in a separate process — and decide which failure mode your platform is willing to carry.

## The lever it actually pulls Two separate things bound recursion in CPython: the frame counter you set with `sys.setrecursionlimit`, and the real stack the operating system gave the thread. `threading.stack_size` is the only lever in the standard library that moves the second one, and it moves it only for threads created *after* the call: ```python import threading threading.stack_size() # current setting; 0 means platform default threading.stack_size(64 * 1024 * 1024) ``` Called with no argument it reports the setting. Called with a size it records it for subsequently created threads; existing threads, and the main thread, are unaffected. The main thread's stack was allocated by the OS before Python started, so from inside the process there is nothing to change — which is exactly why the pattern is to move the deep work onto a worker thread rather than to attempt it on the main one. Platforms constrain the value: there is a minimum (32 KiB), some systems require a multiple of the page size, and a platform that does not support setting it raises an error rather than silently ignoring the request. Handle that rather than assuming it took. ## The full idiom Both constraints have to be lifted, because raising either one alone leaves the other binding: ```python import sys import threading def deep(n=1): if n >= 30_000: return n return deep(n + 1) def run(): sys.setrecursionlimit(40_000) print(deep()) threading.stack_size(64 * 1024 * 1024) worker = threading.Thread(target=run) worker.start() worker.join() ``` Note that `sys.setrecursionlimit` is interpreter-wide, so raising it inside the worker raises it for every thread — the stack size is the only part of this that is genuinely scoped to the new thread. ## When it is the right call It earns its place in a narrow band: * the depth is **real**, not a cycle — a cycle exhausts any stack, and a big one merely makes the failure slower and the crash larger; * the depth is **bounded and known**, so you can size the stack against a measured maximum rather than guessing; * the recursive code is **not yours to restructure** — a third-party recursive-descent parser, a legacy traversal with subtle post-recursion work — or the job is one-shot and not worth reshaping; * the workload is **batch-shaped**, so an extra worker thread and a large stack reservation are acceptable. If you own the traversal and it will keep growing with the data, restructuring so depth lives in a heap structure rather than the call stack is the durable answer, and the thread trick is the stopgap that buys time for it. ## What it costs and what it does not fix The setting is process-global for future threads, so a pool created later inherits the large stack for every worker. On most systems the stack is reserved as virtual address space and only becomes resident as it is touched, so the cost of an unused reservation is modest — but a pool of many threads each reserving tens of megabytes is a real address-space and accounting concern, and it is easy to set once at import and forget. Set it, create the thread you need, and set it back to 0 if other threads follow. It also does not change the accounting split: since **3.12** recursion that passes through C code answers to an internal guard rather than to your `sys.setrecursionlimit` value, so a larger stack helps that recursion survive physically while the guard may still stop it. ## The alternatives to weigh against it * **Restructure the traversal** so pending work is held in a heap data structure — the only fix that scales with the data indefinitely. * **Bound the depth explicitly** and reject input beyond a documented maximum, which is often the correct product answer for machine-generated or untrusted input. * **Run the work in a subprocess** whose stack limit is raised by the launching environment, which isolates both the memory and the crash risk from the parent service. The thread trick is a legitimate, well-understood tool, but it is a capacity adjustment, not a correctness fix — reach for it once you have established that the depth is genuine.

  • Why not simply call threading.stack_size once at import and be done with it?
    Because it applies to every thread created afterwards, not just the one you had in mind. A pool spun up later inherits the large reservation per worker, which inflates address-space use and makes the process harder to reason about. Set it immediately before creating the thread that needs it and reset it to 0 afterwards, so the rest of the process keeps platform defaults.
  • Can you enlarge the main thread's stack from inside Python?
    No. The main thread's stack is allocated by the operating system before the interpreter starts and is governed by the launching environment's limits, so nothing in the standard library can change it mid-process. The options are to run the deep work on a worker thread created with a larger stack, or to launch the process from an environment configured with a bigger stack limit.
  • If you give the thread a large stack, do you still need sys.setrecursionlimit?
    Yes. They are independent constraints: the frame counter still trips at its ceiling no matter how much stack is available, and the stack still overflows no matter how high the counter is set. The idiom lifts both — the stack for the thread that runs the work, and the interpreter-wide frame limit — and lifting only one leaves the other binding.

You cannot widen the corridor you are already standing in, but you can build a wider one next door and do the walking there.

saying these in an interview costs you the question

  • Thinks threading.stack_size resizes the main thread's stack
  • Expects it to affect threads that already exist
  • Uses a bigger stack to survive a cyclic reference chain
  • Forgets to raise the frame limit as well
  • Sets a huge stack globally at import for every future thread
  • Assumes any size is accepted on every platform

context