You own a public upload API used by many client teams: how do you choose its decode ceilings and handle a caller that legitimately exceeds one?
answer
- measure the distribution before choosing
- per endpoint, not per fleet
- report only, then enforce
- change the shape, not the ceiling
- concurrency times the cap is the real number
basics
~20 sDerive each ceiling from measured traffic rather than a round number, set it per endpoint, roll it out in report-only mode before enforcing, and answer a legitimate outlier with a different request shape instead of a raised cap or a per-caller exemption.
solid answer
~50 sThe mechanism is the easy part; the numbers are where this gets decided badly. Start from the observed distribution of body size, nesting depth, element counts and expansion ratio per endpoint, take a high percentile of real traffic and add headroom — a ceiling should sit above everything legitimate and far below what hurts. Set them per endpoint, because a bulk-ingest path and a settings path have nothing in common. Ship them in report-only mode first: log what *would* have been rejected, per caller, then contact the outliers and only then enforce. When a caller genuinely needs more, change the **shape** rather than the limit — chunked or multi-request upload, or an out-of-band transfer the API only references — so each individual decode stays bounded. A per-caller exemption makes a memory ceiling a property of identity, and identity can be stolen or misused.
go deeper
Understand that these limits are chosen numbers rather than defaults from nowhere, and that a limit set badly either breaks real callers or protects nothing.
Be able to say which ceilings an endpoint needs and roughly why its numbers differ from another endpoint's on the same service.
Show the rollout: instrument, observe the distribution, run report-only through a real peak, contact the outliers, then enforce and alert on breaches.
Own the trade-off end to end — per-endpoint numbers, the aggregate capacity arithmetic, the refusal to grant identity-based exemptions, and the review a change in either direction goes through.
## Start from measurement, not from a round number Every decode ceiling is a number, and the common failure is that nobody measured before picking it. Too low and you break legitimate callers, at which point someone disables the limit in an incident and it never returns. Too high and it is decoration. The usable procedure is the same for every ceiling: 1. Instrument the decode path to record the observed value — body bytes, maximum nesting depth reached, largest collection, decompressed output, expansion ratio — per request and per endpoint. 2. Look at the distribution, not the mean. A high percentile of real traffic is the floor for the limit. 3. Add headroom over that percentile, and check the result is still far below what actually hurts. A good ceiling has a wide gap on both sides; if the legitimate high percentile and the dangerous value are close, the endpoint's design is the problem, not the number. ## One size does not fit the fleet | Endpoint shape | What dominates | Where its ceilings sit | |---|---|---| | Small configuration or command payloads | Fixed, shallow documents | Tight everywhere; a low element cap is realistic | | Bulk record ingest | Element count and total bytes | High element cap, modest depth, strict per-element size | | Document or file upload | Total bytes and expansion ratio | High byte cap, low structural limits, compression bounded | | Machine-to-machine internal traffic | Predictable, generated payloads | Tightest of all, because the producer is known and stable | A single global number satisfies the loosest endpoint and therefore protects none of the others. Ceilings belong to the endpoint, with a platform default that applies when an endpoint has not chosen. ## Rolling a ceiling out without causing the outage you were preventing Enforcing a new limit on live traffic is itself a change that can break callers. The sequence that works: 1. **Report only.** Evaluate the ceiling, record the breach, let the request through. Run it long enough to cover weekly and monthly peaks. 2. **Read it per caller.** The output is a list of who would break and by how much — the only input that makes the next decision honest. 3. **Talk to the outliers**, or move the ceiling if the data says your number was wrong. 4. **Enforce**, keeping the counter. Post-enforcement breaches are now either attacks or regressions, and both are worth an alert. Skipping step 1 is how a limit gets rolled back permanently after one bad morning. ## The legitimate outlier Someone will genuinely need more than the ceiling allows. Three responses, in descending order of quality: - **Change the shape.** Chunked upload, pagination of the batch, or an out-of-band transfer the API merely references. Each individual decode stays bounded, which means the ceiling still means something and the caller is unblocked. - **Move the ceiling for that endpoint**, if the measurement says the original number was simply wrong for the traffic it serves. This is a legitimate outcome of the report-only data. - **Exempt that caller's credentials.** Avoid this. It makes a memory ceiling a property of identity, so a stolen or misused credential carries the exemption with it, and the exemption list becomes permanent because nobody can prove it is safe to remove. A related trap is exempting traffic that arrives over an internal network path. Internal services relay bytes that originated outside, and they are compromised like anything else; the network path is not a trust boundary for this purpose. ## Per-request limits are not a capacity plan Ceilings bound one decode. What a lead owns is the aggregate: peak concurrent decodes multiplied by the per-decode ceiling is the honest worst case for the instance, and it should fit in the memory you actually provisioned with room for everything else the process does. If it does not, the missing control is a bound on concurrent decodes or a shared decode budget, not a smaller per-request number that breaks callers. The same arithmetic decides whether an endpoint that spools decompressed output to disk can share a volume with anything that matters. ## Who owns the number, and how it changes Treat a ceiling as part of the published contract: documented, returned as a specific error, versioned with the endpoint, and changed through the same review as a schema change. Loosening one is a security-relevant change and should read like one in the diff. Tightening one is a compatibility-relevant change and needs the report-only cycle again. Owning both directions explicitly is what stops the limits drifting upward one incident at a time.
- How do you roll out a tighter ceiling without breaking existing callers?Run it in report-only mode first: evaluate the limit, record the breach, let the request through, and cover the weekly and monthly peaks. The output is a per-caller list of who would break and by how much. Contact them, or move the number if the data says yours was wrong, then enforce and keep the counter so later breaches raise an alert.
- Why is raising the ceiling for one demanding caller a poor default?Because the ceiling stops being a property of the endpoint. Raised globally, every decode on that path now budgets for the largest body anyone ever sent. Raised per credential, a memory limit becomes identity-dependent and travels with a stolen or misused credential, and the exemption is never removed because nobody can prove it is safe to. Changing the request shape unblocks the caller without either cost.
- What number tells you the fleet is actually safe, rather than one request?Peak concurrent decodes multiplied by the per-decode ceiling, compared against the memory you provisioned and the disk any spooling shares. If that product does not fit, the control you are missing is a bound on simultaneous decodes or a shared decode budget, not a smaller per-request cap that breaks legitimate callers.
saying these in an interview costs you the question
- Picks a round number with no measurement behind it
- Applies one global ceiling to every endpoint in the fleet
- Raises the ceiling permanently the first time a caller complains
- Enforces a new ceiling immediately with no report-only period
- Exempts trusted callers or internal network paths from the ceiling
- Bounds each request but never the number of concurrent decodes