Your Go query parser dies with 'fatal error: stack overflow' on hostile input — what now?
answer
- read the first two lines of the crash, not the last
- this one is not a panic
- a bigger ceiling only buys more memory burned
- the parser should count, not the runtime
basics
~20 sDeeply nested input drove the recursive-descent parser past the 1 GB per-goroutine stack limit. That is a fatal runtime error, not a panic, so recover cannot save the process. The fix is a depth budget in the parser plus an input size limit.
solid answer
~40 sThe runtime prints `runtime: goroutine stack exceeds 1000000000-byte limit` followed by `fatal error: stack overflow`, and the traceback shows the same handful of parse functions repeating for thousands of frames. Read that as unbounded recursion driven by attacker-controlled nesting, not as a memory leak. The critical operational fact is that a stack overflow is a fatal error rather than a panic: no deferred `recover` catches it, running the parse in its own goroutine does not contain it, and the whole process exits. Raising the ceiling with `runtime/debug.SetMaxStack` only changes how much memory you burn before dying. The real fix is a depth counter threaded through the recursive functions that returns a normal error past a fixed limit, plus a cap on request body size, plus a regression test that feeds deeply nested input.
code
text · 8 linesruntime: goroutine stack exceeds 1000000000-byte limit
fatal error: stack overflow
goroutine 42 [running]:
main.(*parser).parseExpr(...)
main.(*parser).parseTerm(...)
main.(*parser).parseExpr(...)
...additional frames elided...go deeper
Recall the distinction that matters: a Go stack overflow is a fatal error, not a panic, so recover cannot catch it and the process exits. Deep recursion on untrusted input is the usual cause.
Explain the mechanism end to end: goroutine stacks double on demand up to a per-goroutine limit of 1 GB on 64-bit, and the crash output names both the limit and the repeating frames of the recursion cycle.
Demonstrate the production judgment: reject the non-fixes, add a depth budget that returns an error, cap input size ahead of parsing, add a regression test with pathological nesting, and audit every other recursion over untrusted structure.
Own it as a policy rather than a patch: untrusted structure gets an explicit depth and size budget everywhere it is parsed, and the standard is enforced in review and in tests rather than rediscovered after each outage.
## What the crash is telling you A recursive-descent parser turns nesting in the input into nesting in the call stack: each `(` in `((((…))))`, each level of a nested boolean expression, is another frame. Because Go grows goroutine stacks on demand, the parser does not fail at some small fixed depth the way a fixed-stack runtime would — it keeps going, doubling the stack, until it hits the ceiling. The output looks like this: ```text runtime: goroutine stack exceeds 1000000000-byte limit fatal error: stack overflow runtime stack: ... goroutine 42 [running]: main.(*parser).parseExpr(...) main.(*parser).parseTerm(...) main.(*parser).parseExpr(...) ... ``` Two things in that output are the whole diagnosis. The **limit line** names the per-goroutine maximum stack size — 1 GB on 64-bit, 250 MB on 32-bit. The **repeating block of frames** identifies the recursion cycle: a small set of mutually recursive functions alternating for thousands of entries, with the traceback eliding the middle. You do not need a profiler for this one; the crash output is the diagnostic. ## The fact that changes how you respond **A stack overflow in Go is a fatal error, not a panic.** The runtime throws rather than panicking, which means: - A deferred `recover()` anywhere up the chain does **not** catch it. - Wrapping the parse in its own goroutine does not contain the blast radius; the process still dies. - There is no handler-level mitigation, no middleware you can add, no per-request isolation inside the process. So unlike a nil dereference or an out-of-range index — both ordinary panics that a handler-level recover can turn into a 500 — this one takes the entire service down for every in-flight request. If the input is attacker-supplied, one small request is a remote denial of service, and it is cheap to send: a few kilobytes of `(((((` can force the process to allocate up to a gigabyte of stack before it dies. ## What not to do `runtime/debug.SetMaxStack` raises or lowers the limit. Raising it is the wrong instinct: it means a hostile input consumes more memory before killing the process, and on a memory-constrained host you may now get killed by the operating system's out-of-memory killer instead, which is strictly harder to debug. *Lowering* it is occasionally defensible as a blast-radius control during triage — you fail faster and with less memory churn — but it still fails. Also resist "just make the parser iterative" as the first move. Rewriting a recursive-descent parser to carry an explicit stack is a real technique and sometimes correct, but it is a big change under incident pressure, and by itself it does not bound anything: an explicit stack grows on the heap until you run out of memory instead. **Bounding depth is the fix; the recursion style is orthogonal.** ## The fix Thread a depth through the recursive functions and refuse to exceed it: ```go const maxDepth = 200 func (p *parser) parseExpr(depth int) (any, error) { if depth > maxDepth { return nil, fmt.Errorf("expression nested deeper than %d levels", maxDepth) } // ... recursive calls pass depth+1 return p.parseTerm(depth + 1) } ``` The pieces that make it hold up in production: - **Pick the limit from the domain, not from the stack.** What is the deepest legitimate query anyone has written? Take that, multiply by a comfortable factor, and make the limit a named constant. 200 levels of nesting is far beyond human-authored queries and nowhere near dangerous. - **Return an error, do not panic.** The caller turns it into a 400-class rejection with a clear message. A rejected query is a good outcome; a dead process is not. - **Cap the input independently.** Limit the request body size before parsing, so a multi-megabyte payload cannot even reach the tokenizer. - **Test it.** Add a regression test that builds an input with, say, 100,000 nested parentheses and asserts a returned error rather than a crash. Fuzzing this input shape finds neighbouring cases — deeply nested lists, unbalanced brackets — that the same budget must cover. - **Audit the neighbours.** Any other recursion over caller-supplied structure has the same hole: a JSON-to-internal-model converter, a schema validator, a rule-tree evaluator. Depth budgets belong on all of them. ## How to talk about it Name the failure mode (unbounded recursion on untrusted nesting), name the Go-specific consequence (fatal error, unrecoverable, whole process), reject the tempting non-fixes (bigger stack limit, isolating goroutine, recover in the handler), and land on the bounded-depth fix with an input-size cap and a regression test. That sequence is what distinguishes someone who has actually shipped this fix from someone reasoning about it for the first time.
- Why does a deferred recover in the HTTP handler not save the process here?Because exceeding the goroutine stack limit is a fatal runtime error, not a panic. The runtime throws, prints the tracebacks and exits; the deferred functions that a panic would run are never given the chance. Recover only handles panics, so nil dereferences and index-out-of-range are catchable while stack overflow is not.
- Would running the parse in a separate goroutine limit the damage?No. The stack limit is per goroutine, but the failure is process-wide: the runtime treats it as fatal and the whole program exits with status 2. Isolation would require a separate process, which is a much bigger design decision than simply bounding the parser's recursion depth.
- How do you choose the depth limit?From the domain, not from the runtime. Find the deepest nesting any legitimate query has ever used, multiply by a healthy factor, and freeze it as a named constant with a clear error message. A limit in the low hundreds is far above human-written input and orders of magnitude below anything that threatens the stack.
- What else in a service usually has the same hole?Anything that recurses over caller-supplied structure: decoding deeply nested documents into an internal model, validating a nested schema, evaluating a rule tree, or walking a directory tree with symlinked cycles. Each needs its own depth budget and its own regression test with pathological input.
saying these in an interview costs you the question
- Suggests recovering from the stack overflow in the handler
- Raises debug.SetMaxStack to make the crash go away
- Says running the parse in a goroutine isolates the failure
- Blames a memory leak instead of unbounded recursion
- Treats a rewrite to an explicit stack as the bound itself
- Assumes only enormous inputs can reach the limit