Your envelope reader is a stack machine, but the model's unbounded store meets finite memory; how do you set and enforce a nesting-depth limit?
answer
- depth is untrusted input
- one entry per open level
- depth bounded only by input length
- check at the push, before allocating
- publish the limit as part of the contract
basics
~20 sTreat depth as untrusted input: measure the deepest legitimate document, set a limit with headroom above it, check it at the push rather than waiting for memory to run out, reject over-deep input with a distinct reason, and publish the number as part of the format contract.
solid answer
~50 sThe store is unbounded only in the model. In a reader it costs one entry per open level, and depth is bounded by nothing except input length, so a payload consisting entirely of openers forces a store **linear in the payload** - a cheap request buying expensive memory. A limit converts that unbounded allocation into a bounded one, so it is an input-validation control and not a tuning knob. Set it by measuring real documents and adding headroom, then cost it as `limit x entry size x concurrent readers` against the memory budget. Enforce it at the push, before the entry exists, so the failure is deterministic and attributable rather than an allocation failure raised somewhere unrecoverable. Then publish it: two readers with different limits disagree about the same document, and a producer cannot tell a limit from a bug without a distinct rejection reason.
go deeper
Know that a reader's store is real memory, and that one entry per open level means deeply nested input costs memory in proportion to its depth.
Explain that depth is bounded only by input length, so the store is linear in the payload unless a limit caps it, and that the check belongs before the entry is allocated.
Demonstrate the enforcement: an explicit store rather than a frame per level, a check at the push, a distinct rejection reason, and a metric so a limit set too low is visible instead of mysterious.
The call you own is the number and its contract - memory budget per concurrent reader weighed against the deepest document you must accept, then published so producers and other readers agree instead of each guessing.
The model says the store is unbounded. A reader runs in a process with a fixed memory budget and, often, a fixed call-stack budget as well. The gap between those two statements is where a real incident lives, and closing it is a decision someone has to own explicitly. ## Where the idealisation breaks - Each open level costs one entry, so memory is proportional to **depth**. - Depth is bounded only by input length: nothing stops a payload from being openers all the way down, so worst-case depth grows with every byte. - A reader that recurses once per level spends a call frame per level instead, and the ceiling there is usually lower and the failure less recoverable. - Nothing in the format's structure caps depth unless the format says so; well-formedness and affordability are different questions. The consequence is that depth is an **input dimension under someone else's control**, which is the definition of untrusted. A payload measured in kilobytes can demand memory measured in megabytes, and no individual step of the reader looks wrong while it happens. ## Choosing the number 1. **Measure.** Take the deepest documents that legitimately exist today across every producer you know of, and note the distribution, not just the maximum. 2. **Add headroom** over the observed maximum, enough that an unusual but honest document is not rejected, and not so much that the limit stops meaning anything. 3. **Cost it.** Multiply limit by entry size by the number of readers that can be live at once. That product is what the limit actually promises the memory budget. 4. **Compare** that product against the budget you are willing to spend on hostile input, not against the budget for normal traffic. 5. **Make it configurable and observable.** Emit a metric on rejections at the limit, so a limit set too low announces itself instead of manifesting as one producer's mysterious failures. The number is rarely interesting in itself; what matters is that it was derived from a measurement and a budget rather than chosen because it looked large. ## Enforcement that actually holds - **Check before the push**, not after. A check that runs before the entry is allocated caps memory deterministically; a check that runs afterwards has already paid for what it is refusing. - **Prefer an explicit store to one frame per level.** An explicit store makes the limit a number you control and the failure an ordinary rejection; per-level recursion makes the real limit an implementation detail and the failure a process-level event. - **Do no work proportional to depth before the check.** Anything allocated or copied per level must sit behind the limit, or the limit protects only part of the cost. - **Count depth, not bytes.** A size limit is a different control with a different purpose; deep nesting can be small, and large payloads can be flat. - **Reject with a distinct reason** that names depth, so the failure is diagnosable from the outside. | | with no limit | with an enforced limit | |---|---|---| | worst-case store | linear in payload size | bounded by the limit | | failure mode | allocation failure, possibly mid-operation | ordinary rejection at a known point | | attribution | unclear, may strike unrelated work | the offending input, named | | recovery | uncertain | continue serving the next request | | producer experience | a hang or an opaque error | a specific, actionable rejection | One honest caveat: the limit bounds the **store**, not the reader's whole memory profile. Payload buffers, decoded values and whatever the reader hands downstream are separate budgets with their own controls. ## The trade-off no number settles A depth limit is part of the format's contract whether or not anyone writes it down. Too low and legitimate documents are rejected, which producers experience as an incompatibility. Too high and the exhaustion vector is still open, just more expensive to exercise. Worse, different readers of the same format that pick different limits create documents that one consumer accepts and another rejects, and that divergence is discovered in production by whoever wrote the deepest document. So the lead-level position is: pick the number from evidence, write it into the specification, give it a named rejection reason, keep it configurable for the operator, and measure how often it fires. A limit that is documented and observed is a contract. A limit that is an accident of the reader's implementation is a trap - and the fact that the theoretical model promises an unbounded store is exactly why people forget to set one.
- Why is a depth limit a security control rather than a tuning knob?Because depth grows with input length, a small payload made only of openers forces a store proportional to it, and a reader that recurses per level to spend a frame per level. Without a limit, a cheap request buys expensive memory from the server, which is exactly the shape of a resource-exhaustion attack rather than a performance concern.
- Why check the limit at the push instead of catching the memory failure?A check at the push is deterministic, attributable to the specific input, and leaves the reader in a state it can report from. A memory failure raised deep inside per-level recursion may not be safely recoverable, can strike unrelated work sharing the process, and gives no useful diagnosis to the producer.
- How should the limit be surfaced to producers?As a documented part of the format contract, with a rejection reason that names depth rather than a generic malformed-input error. Otherwise a producer whose document works against one reader and fails against another has no way to tell a limit from a bug, and differing limits across readers become an interoperability trap.
- Does a payload size limit make a depth limit unnecessary?No - they bound different things. A size limit caps total bytes, but a small payload can still be nested to its full length, so it constrains depth only at the crudest level. Deep and small is the dangerous combination, and only a depth limit addresses it directly.
saying these in an interview costs you the question
- Says the model's unbounded store means no limit is needed
- Relies on an allocation failure as the enforcement mechanism
- Sets the limit by guesswork without measuring real documents
- Assumes a small payload cannot produce deep nesting
- Keeps the limit internal and never tells producers about it
- Counts bytes instead of depth when applying the limit