skip to content

On a public upload endpoint, why must a decoder's size and nesting ceilings be enforced during the parse rather than after it?

level: middleimportance: must knowfreq 62%

answer

  1. hostile bytes reach the parser first
  2. a later check is a post-mortem
  3. the cost is spent before it is measured
  4. counters tested at the point of growth
  5. abort mid-parse and discard the partial value

basics

~20 s

By the time a parse finishes, the memory, CPU and stack the sender asked for have already been spent, so a later check only reports the damage. A ceiling has to be a counter the parser itself tests as it reads, aborting mid-stream.

solid answer

~50 s

The decoder is the first code that touches sender-controlled bytes on an upload endpoint. It runs before field validation, before business rules, and on a large body long before anything has decided whether the request is even allowed. If the ceiling is a check applied to the finished value, the recursion has already run and the objects already exist: the check is a post-mortem, not a defence. So each ceiling is a counter the parser tests at the point of growth — bytes as they are consumed, depth incremented on entry to each nested value, elements counted as each one is appended — and the parse aborts the instant one is exceeded, discarding what it built. Two things follow: the cost of a rejected request is bounded by your limit rather than by the sender's choice, and rejection stays cheap enough that a flood of hostile bodies does not become the outage.

go deeper

for a junior

Recall that decoding untrusted input is work the sender controls, and that limits on body size and structure exist to bound that work before your code sees a value.

for a middle

Explain the counters — bytes consumed, nesting depth, elements per collection — and why each is tested at the moment the value grows rather than once the parse has finished.

for a senior

Show you have set these on a real endpoint: which numbers, chosen against measured traffic, what the service returns when one trips, and the metric that reveals a limit set in the wrong place.

for a principal

Weigh a ceiling that is safe against one that silently breaks a legitimate large caller, and decide who owns the number, how it is reviewed, and what bounds the fleet's aggregate decode load.

## Where the decoder sits relative to everything else On a public upload endpoint the bytes arrive from someone you have not vetted, and they pass through several layers before any of your own logic sees a value: a framing layer decides where the message ends, an optional decompression stage expands it, and a structural decoder turns the bytes into values in memory. Field validation, authorization rules and business logic all operate on the **result** of that decode. That ordering is the whole problem. An attacker does not need the request to succeed. Making the service **build** something expensive is already the win, and everything that would have said no runs afterwards. ## Why a check on the finished value is not a ceiling A check applied to a fully decoded value answers the question *was it too big?* — after the bill has been paid. Three costs are already sunk by then: - **Memory**: every node, string and collection the document described now exists on the heap. - **CPU**: tokenising and building them took real time, on a request that was always going to be rejected. - **Stack**: a decoder that recurses once per nesting level may not reach the check at all; the thread dies first. And a server handles many requests at once, so each of those costs is multiplied by the number of concurrent decodes. A limit that is only correct for one request in isolation is not a limit on the service. ## The ceilings worth setting, and where each is tested | Ceiling | What it stops | Where it is enforced | |---|---|---| | Total bytes consumed for one message | Unbounded buffering of a single body | Byte counter in the read loop, before the buffer grows | | Maximum nesting depth | Stack or heap exhaustion from deep structure | Depth counter checked on entry to each nested value | | Elements per collection | A huge array or map materialising as objects | Counter per container, incremented as each element is appended | | Maximum length of one scalar | A single enormous string or byte field | Length check while the scalar is accumulated | | Total node or token count | Many small nodes — breadth, not depth | Global counter across the whole parse | | Decompressed output and expansion ratio | A tiny body expanding to gigabytes | Output counter inside the decompression loop | The right-hand column is the point of this question. Every one of these is a comparison against a counter the parser already maintains, costing a few instructions per step — which is why *inside the parser* is not merely the safest place but also an affordable one. ## What enforcement actually looks like 1. Maintain the counter at the point where the value **grows** — not at the start, not at the end. 2. Compare **before** the expensive act: before appending, before allocating, before recursing. 3. On breach, abort the parse, release the partial value, and stop reading the body. Draining or closing the connection is a policy choice, but continuing the parse to produce a nicer error message is not. 4. Record which ceiling fired and what value tripped it. ## What the endpoint returns, and what it records The response should be terse and deterministic: a payload-too-large status when a byte, element or output ceiling trips, a bad-request status when the document is structurally out of bounds. Do not echo the submitted bytes back, and do not include a fragment of the document in the error. Publishing the limit values themselves in your API documentation is fine and helpful — they are not secrets, and a caller who knows them can stay inside them. Telemetry matters more than the status code. Count breaches per ceiling, per caller and per endpoint. A spike is either an attack or a limit set below real traffic, and nothing else in the system can tell you which. ## The traps - **Unknown defaults.** Most decoding stacks ship some ceilings and omit others, and the ones they ship differ between ecosystems. Knowing which are set, and to what, is part of owning the endpoint. - **A proxy that buffers first.** A size limit enforced by a component that has already read the whole body into memory protects the layer behind it, not itself. - **Treating a crash as a limit.** Catching the error raised by stack exhaustion and carrying on leaves a thread in an unclear state; it is not a substitute for a depth cap. - **Nested payloads.** A field that itself carries an encoded document — a base64 blob, an embedded envelope — is a second decode of untrusted bytes and needs its own ceilings, budgeted inside the outer ones. - **Per-request only.** Ceilings bound one decode. The number of decodes running at once needs a bound too, or the per-request limit simply sets the price per connection.

  • What should the endpoint return and log when a decode ceiling trips?
    A terse, deterministic client error: a payload-too-large status for a byte, element or output ceiling, a bad-request status for a structurally out-of-bounds document. Echo none of the submitted bytes. Log which ceiling fired, the observed value against the limit, the caller and the request id, and increment a counter per ceiling — a spike is either an attack or a limit set below real traffic, and the metric is the only thing that distinguishes them.
  • Does a streaming decoder remove the need for these ceilings?
    No. Streaming bounds how much input is buffered at once, not what the input asks you to build: a streamed document can still nest without end, declare a collection of a billion elements, or feed a handler that accumulates every element. Streaming changes where a ceiling is checked — incrementally, and cheaply enough to abort early — but it does not supply the ceiling.
  • One service decodes a payload and forwards it to another. Where do the ceilings belong?
    In both. A ceiling upstream bounds what the upstream built, not what the downstream will build from the same bytes, and the two decoders may differ in cost per element. Treat each decode as its own trust boundary with its own numbers, and make the downstream limits no looser than the upstream ones so a rejection surfaces at the edge rather than deep in the chain.

saying these in an interview costs you the question

  • Thinks validating the decoded object is early enough
  • Assumes a request byte cap also bounds nesting depth
  • Believes an authenticated caller cannot send a hostile body
  • Leaves ceilings to the framework without knowing its defaults
  • Catches the stack-exhaustion error and keeps serving
  • Bounds one request but never the number of concurrent decodes