Your platform caps every request body at 1 MB and a team's legitimate 5 MB payload now returns 413. How do you decide?
answer
- the default protects endpoints nobody is discussing
- price it before you argue about the number
- three factors, and one is unbounded
- an exception with an owner beats a global raise
- sometimes the payload should not be a body at all
basics
~20 sDo not raise the global default. Price the request — cap times decoded expansion times requests in flight — then grant a per-route exception with its own cap and concurrency bound, or change the payload shape so the large case never travels in a request body.
solid answer
~50 sTreat it as a capacity and blast-radius decision, not a config edit. The global default protects every unauthenticated endpoint, so raising it spends memory on routes that never needed it and widens what one anonymous poster can cost. Instead, price the specific route: the cap times the expansion factor of its decoded payload times the requests you allow in flight, against the memory headroom the service actually has. If that number is affordable, grant a per-route cap — one handler wrapped with a larger limit while the default stands — with a bound on concurrency for that route and a documented owner. If it is not, change the shape rather than the number: upload to object storage and post a reference, split the batch, or stream and process incrementally instead of buffering a whole value. And keep the server-side cap even if the edge proxy has one, because internal callers and any path that bypasses the edge still reach the handler.
code
go · 3 linesmux := http.NewServeMux()
mux.Handle("/webhooks/", http.MaxBytesHandler(webhookHandler, 1<<20)) // default
mux.Handle("/imports/", http.MaxBytesHandler(importHandler, 8<<20)) // documented exceptiongo deeper
Understand that a size limit is a chosen number with consequences on both sides — too low rejects real users, too high lets one caller consume the service's memory.
Be able to implement the exception cleanly: a per-route cap rather than a changed global default, and a 413 whose message names the limit.
Bring the measurement. Payload distribution, expansion factor, concurrency bound and memory headroom turn the debate into arithmetic instead of preference.
Own the default, the exception procedure and the escalation path, and be willing to answer a large-payload request by changing its shape — reference upload, batch split, incremental processing — rather than the limit.
## What is actually being decided The request looks like a number change. It is really three decisions: what an unauthenticated caller may cost you, who bears the memory cost of the new limit, and whether a large payload should travel in a request body at all. A platform owner who answers only the first question by editing a default has made the other two silently. ## Price the route before you argue about the number The worst-case memory a route can hold is approximately: cap × expansion × in-flight requests - **cap** is what you set. - **expansion** is how much larger the decoded Go value is than the encoded bytes. It is a property of the payload's shape, and it is measured, not assumed: a 5 MB payload of long strings into a typed struct expands modestly, while an open-ended nested structure decoded into an `any` expands by a large multiple. - **in-flight requests** is the term teams forget. Unbounded by default, it means no per-request cap produces a peak figure. Run the numbers with the team's real payloads and the service's real memory headroom. Very often the answer is not "5 MB is fine" or "5 MB is too much" but "5 MB is fine at eight concurrent, and catastrophic at four hundred" — which converts the argument from a preference into a bound. ## Prefer a per-route exception to a global raise A global default exists because every endpoint inherits it, including the ones nobody is thinking about today. Raising it from 1 MB to 8 MB grants that budget to every public handler in the fleet, and the routes that most need protection — anonymous webhook receivers — are exactly the ones nobody will revisit. Wrap the one handler that needs more with its own larger cap, leave the default alone, and record who owns the exception. Keep the exception visible in code review rather than in an environment variable that drifts. ## Consider changing the shape, not the limit Many 5 MB request bodies are a design accident. The alternatives are usually cheaper and more durable: - **Reference instead of payload**: the caller uploads to object storage and posts a URL and a checksum. The request stays small, the transfer is resumable, and the size ceiling stops being your problem. - **Batch splitting**: if the payload is a list, the caller sends several requests. Now the limit is a documented batch size instead of a byte count, and retries are finer-grained. - **Incremental processing**: read the body as a stream and handle elements as they arrive rather than holding the whole decoded value. This changes the expansion factor rather than the cap. If none of these fit — a genuinely atomic document, a signed payload that cannot be split — that is a real argument for the exception, and it should be written down as such. ## Defence in depth still applies "The edge proxy already limits bodies" is not a reason to remove the server-side cap. The handler is reachable from inside the network, from health checks and internal callers, from a misrouted deployment, and from any future path that does not traverse the edge. The Go server holds its own limit because it is the last thing that is definitely in the request path. The edge limit is a bandwidth saving; the handler limit is the correctness guarantee. ## Make the rejection usable Whatever number wins, the 413 must be actionable: name the limit in the response body, log the route and the observed size so you can see whether the cap is mis-sized, and publish the per-route limits alongside the API documentation. A caller who gets "body too large" with no number will guess, retry, and fill your logs with the same rejection. ## Decide the escalation path once The healthy version of this conversation is not per-incident. Publish the default, the procedure for requesting an exception, the evidence an exception needs — payload distribution, expansion measurement, concurrency bound — and who signs it off. Then this week's 413 is a form to fill in rather than an argument between a product team and a platform team about a number neither has priced. ## What a strong answer sounds like "Keep the default. Measure their payload's expansion, cap that route separately at a size with headroom over their p99.9, bound its concurrency so the peak is a number I chose, and ask whether the payload should be a reference to object storage instead. Keep the server-side cap regardless of the edge. Put the limit in the 413 body, and write the exception down with an owner."
- The team argues the edge proxy already caps bodies, so the Go server's cap is redundant. What do you say?The proxy saves bandwidth; the handler holds the guarantee. Internal callers, health checks, a misrouted deploy and any future path that skips the edge all reach the handler directly. A limit that is only true when traffic takes the expected route is not a limit.
- What evidence would make you grant the exception rather than push back?A measured payload size distribution showing the 5 MB case is real and not a batching accident, a measured expansion factor for their decode path, a concurrency bound for that route, and headroom in the service's memory budget for the resulting product. Plus a reason the payload cannot be split or referenced.
- How do you keep this from becoming a per-incident negotiation every quarter?Publish the default, the exception procedure, the evidence required and the owner who signs it off, and put the per-route limits in the API documentation. Then a 413 is a known process rather than an argument between teams about a number neither has priced.
- What belongs in the 413 response itself?The limit, in bytes, and enough wording for the caller to act — split the batch, or use the upload endpoint. Log the route and the observed size so you can tell a mis-sized cap from an abusive caller. A bare "request entity too large" guarantees repeated retries and no fix.
saying these in an interview costs you the question
- Raises the global default because one team asked
- Removes the server-side cap since the edge proxy has one
- Quotes capacity as the per-request cap with no concurrency term
- Grants the exception with no owner and no documentation
- Never asks whether the payload should be a body at all
- Returns 413 without naming the limit the caller must respect