skip to content

In Ruby, a recursive walk over a deep undo history raises SystemStackError; why does rescue => e miss it, and how do you fix it?

level: seniorimportance: should knowfreq 40%

answer

  1. stack level too deep
  2. subclass of Exception, not StandardError
  3. RUBY_THREAD_VM_STACK_SIZE
  4. fibers get a smaller VM stack
  5. explicit Array stack instead of recursion

basics

~20 s

SystemStackError (stack level too deep) inherits directly from Exception, so a bare rescue, which catches StandardError, misses it. The fix is to make depth independent of data: replace the recursion with a loop over an explicit Array stack.

solid answer

~40 s

Each Ruby method call takes space on the VM stack, and when recursion exhausts it CRuby raises `SystemStackError` with the message `stack level too deep`. `SystemStackError` is a direct subclass of `Exception`, so `rescue => e` and a bare `rescue`, which catch only `StandardError`, let it through; `rescue SystemStackError` would catch it, but that only hides a depth that grows with input. `RUBY_THREAD_VM_STACK_SIZE` raises the VM stack size for threads at start-up, and fibers get a smaller stack from `RUBY_FIBER_VM_STACK_SIZE`, so the same recursion can fail sooner inside a fiber. The durable fix is iteration: keep pending nodes in an Array, `push` children and `pop` the next one, so depth becomes heap memory instead of stack frames.

code

ruby · 22 lines
ruby
Entry = Data.define(:action, :previous)

# Recursive: depth grows with the history length.
def replay_recursive(entry, canvas)
  return unless entry
  replay_recursive(entry.previous, canvas)
  canvas << entry.action
end

# Iterative: collect with a loop, then replay oldest first.
def replay(entry, canvas)
  stack = []
  while entry
    stack.push(entry)
    entry = entry.previous
  end
  canvas << stack.pop.action until stack.empty?
end

history = (1..200_000).reduce(nil) { |prev, i| Entry.new(:"stroke#{i}", prev) }
replay(history, canvas = [])
canvas.first(2)   # => [:stroke1, :stroke2]

go deeper

for a junior

Recall the message stack level too deep, that it means runaway or very deep recursion, and that a plain rescue does not catch it.

for a middle

Explain where SystemStackError sits under Exception, the stack-size environment variables, and the smaller fiber stack.

for a senior

Find recursion whose depth follows input size, convert it to an explicit Array stack, and guard nesting depth on untrusted data.

for a principal

Set limits on accepted input depth across services, since stack exhaustion from nested data is a denial-of-service vector.

## Where SystemStackError comes from Every Ruby method or block call pushes a **frame** onto the thread's **VM stack**, and C functions use the **machine stack**. Both are fixed-size regions. When recursion is deep enough to exhaust them, CRuby raises **`SystemStackError`** with the message **`stack level too deep`**. ```ruby def me_myself_and_i me_myself_and_i end me_myself_and_i # SystemStackError: stack level too deep ``` Infinite recursion is the obvious cause. The subtler one is **correct recursion over deep data**: a method that walks a linked chain of undo entries, a deeply nested JSON document or a long parent chain works in tests and fails on production-sized input. ## Why rescue => e misses it Ruby's exception tree has `Exception` at the top and `StandardError` below it for ordinary application errors. `SystemStackError` sits **directly under `Exception`**, beside `NoMemoryError`, `SystemExit` and `SignalException`, not under `StandardError`. - A bare `rescue` and `rescue => e` catch **`StandardError`** and its subclasses only. - So `SystemStackError` passes straight through them and usually ends the request or the program. - `rescue SystemStackError` does catch it, and so does `rescue Exception`, which also swallows `SystemExit` and `Interrupt` and is almost always wrong. Catching it is rarely a fix. The depth that caused it depends on the input, so the next larger input fails the same way. ## Stack size knobs CRuby reads stack sizes from environment variables at start-up, in bytes: | Variable | Applies to | 64-bit default | |---|---|---| | `RUBY_THREAD_VM_STACK_SIZE` | VM stack of each thread | 1048576 | | `RUBY_THREAD_MACHINE_STACK_SIZE` | machine stack of threads Ruby creates | 1048576 | | `RUBY_FIBER_VM_STACK_SIZE` | VM stack of each fiber | 131072 | | `RUBY_FIBER_MACHINE_STACK_SIZE` | machine stack of each fiber | 524288 | Two consequences: 1. Raising the thread VM stack size buys a constant factor of depth for every thread, at a memory cost per thread. 2. **Fibers have a much smaller VM stack**, so recursion that works on a thread can overflow inside a fiber, for example inside an external `Enumerator#next` or a fiber-based server. The man page warns that these variables are implementation-dependent and may change between versions. ## The fix: an explicit stack Turning recursion into iteration moves the pending work from call frames into an **Array on the heap**, where the limit is memory, not a fixed stack region: 1. Start with `stack = [root]`. 2. Loop `until stack.empty?`. 3. `pop` a node, process it, and `push` its children. This is depth-first order. Using `push` with `shift` instead gives breadth-first order. Either way, no single call gets deeper as the data grows. ## Diagnosing it A `SystemStackError` backtrace is long and repetitive, which is itself the clue: - Look for the **frame that repeats**; that method is the recursion. - Check whether its depth follows **input size**, such as history length or nesting level, rather than a bug in the base case. - Reproduce with production-sized input in a test, so the fix is proven before it ships. ## Alternatives and their limits - **Tail-call optimisation** exists in CRuby only as an opt-in compile option that is off by default, so it is not a general fix. - **Memoisation** reduces repeated work but not depth. - **Limiting input depth** is a sensible guard for untrusted nested input, such as a maximum nesting level when parsing. ## A drawing app example A drawing app stores history as linked entries, each pointing to its previous state. A recursive `replay(entry)` that calls `replay(entry.previous)` before applying the entry works for a short session and raises `SystemStackError` after a long one. Collecting entries with a loop into an Array, then replaying them in order, keeps the depth constant however long the session runs.

  • Why can recursion that works in a thread fail inside a Fiber?
    CRuby gives fibers a smaller VM stack than threads: `RUBY_FIBER_VM_STACK_SIZE` defaults to 131072 bytes on 64-bit, against 1048576 for `RUBY_THREAD_VM_STACK_SIZE`. Code run through an external `Enumerator#next` or a fiber-based server therefore overflows at a shallower depth.
  • Is rescue SystemStackError ever reasonable?
    At a boundary, yes: a job runner can catch it to mark one job failed instead of crashing the worker, and log the input size. Inside the recursive code it is a mistake, because the stack is already exhausted there and the next larger input fails again.

saying these in an interview costs you the question

  • A bare rescue catches SystemStackError like any other error
  • SystemStackError inherits from StandardError
  • Raising RUBY_THREAD_VM_STACK_SIZE removes the depth limit
  • CRuby optimises tail calls by default, so tail recursion is safe
  • Recursion depth is the same inside a Fiber as in a Thread