How do you defend json.loads against a deeply nested untrusted document?
answer
- Each level of nesting costs stack
- The error is not a decode error
- It inherits from RuntimeError
- Raising the limit only moves it
- Bound depth before you parse
basics
~20 sBound the input before parsing: cap the payload size and reject documents nested deeper than the contract allows. json.loads recurses once per level and raises RecursionError on deep input, and raising sys.setrecursionlimit() only moves the failure, it does not remove it.
solid answer
~40 s`json.loads` is a recursive-descent parser, so every `[` or `{` costs a level of stack. A few hundred kilobytes of nothing but opening brackets exhausts it and raises `RecursionError` — on CPython 3.14 with a message naming the stack it consumed. Two traps follow. First, `RecursionError` inherits from `RuntimeError`, so a broad `except Exception: continue` swallows it and the record disappears with no crash and no alert. Second, raising `sys.setrecursionlimit()` does not help: since 3.12 the C-level recursion guard is tracked separately and is not what that function sets, so the C scanner still stops. The real defence is at the edge — a hard byte cap on the body, an explicit depth check or a streaming parser before the document reaches `json.loads`, and precise `except (json.JSONDecodeError, RecursionError)` handling that records the rejection.
code
python · 12 linesimport json
import sys
payload = "[" * 100_000 + "]" * 100_000
sys.setrecursionlimit(200_000) # bounds Python frames, not the C scanner
try:
json.loads(payload)
except RecursionError as exc:
print("rejected:", exc)
print(issubclass(RecursionError, RuntimeError)) # True
print(issubclass(RecursionError, Exception)) # True -> bare except swallows itgo deeper
Recall that json.loads recurses once per level of nesting and that very deep input raises RecursionError, not a decode error. Know that catching Exception catches it too.
Explain the class hierarchy that makes RecursionError a RuntimeError, why sys.setrecursionlimit bounds Python frames rather than the C scanner since 3.12, and how you would catch json.JSONDecodeError and RecursionError separately.
Show the production judgment: a byte cap plus an explicit depth limit taken from the contract, precise exception handling that counts and reports rejected records, and the reasoning against ever raising the recursion limit in a service.
Own the ingest policy across services — where document bounds are enforced, whether arbitrary-document parsing gets its own isolated process, and how you make the same limits apply to every parser the platform accepts input into rather than one at a time.
## Why nesting is a cost at all `json.loads` is a recursive-descent parser: to read an array it reads values, and a value may itself be an array, so the parser calls itself. Depth of nesting in the document becomes depth of recursion in the parser, and depth of recursion becomes stack. The attacker's leverage is that nesting is *cheap to write and expensive to read* — one byte per level. A 200 KB body consisting of nothing but `[` buys 200,000 levels. ```python import json, sys print(sys.getrecursionlimit()) # 1000 json.loads("[" * 100_000 + "]" * 100_000) # RecursionError: Stack overflow (used 3906 kB) while decoding a JSON # array from a unicode string ``` Two things in that transcript matter. It raises rather than crashing — CPython now checks how much C stack remains and turns exhaustion into an exception on supported platforms, which is a hardening improvement over older releases where deep C recursion could take the process down. And the number it reports is *kilobytes of stack*, not a count of Python frames, which is the clue that `sys.getrecursionlimit()` is not the bound in play. ## Why raising the recursion limit is not the fix `sys.setrecursionlimit()` sets the limit on **Python** frames. Since 3.12 CPython maintains a separate guard for **C-level** recursion, and that function does not raise it. The C accelerated scanner in `json` is bounded by the C guard, so pushing the Python limit up changes nothing: ```python sys.setrecursionlimit(200_000) json.loads("[" * 100_000 + "]" * 100_000) # still RecursionError ``` That separation is deliberate and is a good thing. Before it existed, raising the limit far enough genuinely could let a C-level parser run off the end of the real thread stack and segfault — trading a catchable exception for a process death, and on a worker that handles other requests, that is strictly worse. Treat any `sys.setrecursionlimit(100000)` in a service as a defect: it is either useless or dangerous, and never a bound on attacker-controlled input. Threads make it sharper still. A thread gets whatever stack `threading.stack_size()` was set to before it started, typically smaller than the main thread's, so a document that parses on the main thread can fail on a worker thread. Depth limits you rely on must be *your* limits, checked in your code, not whatever the platform's stack happens to allow. ## The swallowed-exception trap `RecursionError` is a subclass of `RuntimeError`, and therefore of `Exception`. Picture a CSV import for a payroll system where each row carries a JSON blob of allowance details, and the row loop is written defensively: ```python for row in reader: try: record = json.loads(row["details"]) except Exception: continue # "one bad row shouldn't kill the import" apply(record) ``` A crafted or merely corrupt cell now vanishes silently. Nothing is logged, the import reports success, and the effect surfaces a fortnight later as a missing allowance on somebody's payslip. The bug is not the parser — the parser did exactly the right thing and raised — it is that a bare `except Exception` cannot tell a malformed record from an attack from a bug in your own code. Catch precisely and record what you dropped: ```python except (json.JSONDecodeError, RecursionError) as exc: log.warning("rejected row %s: %s", row["id"], exc) rejected += 1 ``` and fail the import if `rejected` crosses a threshold. A silent skip is a decision to lose data; make it a visible one. ## The defences that actually bound the attack **Cap the bytes first.** A hard limit on request-body size, enforced before you have the string in memory, is the cheapest control and the one that bounds every other parser cost too. Note it does not bound depth on its own: a 1 MB cap still permits a million levels, which is far past anything legitimate. **Cap the depth explicitly.** Decide what your contract allows — most real schemas are under ten levels deep — and enforce it. Either pre-scan the bytes counting brackets that lie outside string literals, or use a streaming/iterative parser that reports depth as it goes and lets you abort early. Enforcing your own small number is what makes the behaviour identical on every platform, thread and interpreter build. **Isolate what you cannot bound.** If a feature must accept arbitrary documents, parse them in a child process with a wall-clock and memory bound, so the worst case is a killed child and a rejected request rather than a stalled worker. **Remember the output side.** `json.dumps` recurses too. Serializing a deeply nested structure you built from untrusted input hits the same wall, and a structure containing a reference back to itself raises `ValueError` for a circular reference rather than looping forever. The general lesson generalizes past JSON: any recursive-descent parser turns one byte of input into one frame of stack, so untrusted input into such a parser always needs an explicit, self-imposed depth bound.
- A colleague fixes the RecursionError by calling sys.setrecursionlimit(100000). What do you tell them?That it is either useless or dangerous. Since 3.12 that function bounds Python frames only, and `json`'s C scanner is guarded separately, so the parse still fails. On older interpreters it worked just well enough to let deep C recursion run off the real thread stack and kill the process — swapping a catchable exception for a segfault on a worker serving other requests. The bound has to be your own depth check on the input.
- Does a request-body size limit make nesting depth a non-issue?No. Nesting costs one byte per level, so even a 1 MB cap permits roughly a million levels — orders of magnitude past any legitimate schema. A size cap is necessary and cheap, and it bounds the total work, but depth needs its own explicit limit, chosen from the contract rather than from whatever the platform's stack happens to survive.
- Why can a document that parses fine in one process fail in a worker thread?Because the guard is ultimately about available C stack, and a thread gets whatever `threading.stack_size()` was set to before it started — usually less than the main thread. That platform-dependent variability is the argument for enforcing your own small depth limit: it makes the behaviour identical across threads, platforms and interpreter builds instead of depending on the stack you happened to get.
Deep nesting is a letter with a thousand envelopes inside envelopes: writing it takes seconds, and the clerk who has to open them all runs out of desk.
saying these in an interview costs you the question
- Raises sys.setrecursionlimit and calls the issue fixed
- Thinks json.loads can only raise JSONDecodeError
- Catches Exception and continues the loop silently
- Says a body size cap alone bounds nesting depth
- Believes deep nesting must crash the interpreter
- Treats the depth limit as the parser's job, not the contract's