skip to content

How would you stop an oversized request body from exhausting a service's memory, and what should the client see?

level: seniorimportance: should knowfreq 54%

answer

  1. cap times concurrency is your worst case
  2. count while reading, never after
  3. a declared length is a claim
  4. 413, and let the client read it
  5. compressed bodies expand after decoding

basics

~20 s

Count bytes while reading and abort the moment the cap is passed, rather than buffering the body and checking afterwards. Answer 413, and handle a client that is still uploading so it receives the status instead of a connection reset.

solid answer

~50 s

An unbounded body turns one client into an outage: worst-case memory is roughly the body cap multiplied by concurrent in-flight requests, so with no cap there is no bound. Enforce it while reading — a counter over the input, aborting as soon as it crosses the limit — instead of reading the body and checking its size afterwards, which is the very allocation you were trying to prevent. A declared `Content-Length` larger than the cap is a useful early rejection, but it is a claim, not a guarantee, so it never replaces the counter. Answer **413**, and either drain a bounded amount of the remaining body or close the connection deliberately, because a client still uploading may see a reset before it reads the response. Apply the cap at the edge and again in the framework, with per-route overrides where a large body is legitimate.

go deeper

for a junior

Know that request bodies must have a maximum size, that it is configured rather than written per handler, and that exceeding it produces a 413 response.

for a middle

Explain the enforcement mechanism: a counter over the body read that aborts at the cap, with a declared Content-Length used only as an early rejection rather than as the check itself.

for a senior

Show the operational judgment: size the cap from memory budget divided by concurrency, layer edge and framework limits, and make sure the client actually reads the 413 instead of a reset.

for a principal

Treat it as a platform control. Decide where limits are owned, how per-route overrides are granted and reviewed, and how decompressed size, concurrency and read timeouts are bounded alongside body size.

## Why this is an availability problem, not a validation problem If a framework buffers a request body before anyone checks its size, the worst case is simple arithmetic: **maximum body size multiplied by the number of concurrent in-flight requests**. With no maximum, the product is unbounded, and the cost to the attacker is a slow upload from a single connection. This is why "add a body size limit" appears on every hardening checklist and why the question is asked at senior level — it is an operational control, not input validation. The same arithmetic gives you the number to configure. Decide how much memory the service may devote to in-flight bodies, divide by the concurrency you allow, and that is your cap. A cap chosen without reference to concurrency is a guess. ## Enforce during the read The mechanism that works is a counting read: 1. wrap the body input so every chunk read increments a counter; 2. compare the counter against the cap after each chunk; 3. on the first chunk that crosses it, **stop reading** and raise a limit error; 4. map that error to a response centrally, so every route behaves alike. The property that matters is that the bytes are never all resident at once. A check written as "read the body, then compare its length" performs exactly the allocation the limit exists to prevent, and it is the most common wrong answer to this question. A declared `Content-Length` is a helpful optimisation in front of the counter: when it is present and already exceeds the cap, you can reject before reading a single byte of body. It is not a substitute. The declared length may be absent, and a value on the wire is a claim by the sender rather than something the server has verified. ## Where the limit lives | Layer | What it protects | Typical shape | |---|---|---| | Edge proxy or gateway | the whole fleet, and the network behind it | one generous global cap | | Framework, global | this service's memory and worker pool | the real working limit | | Route override | endpoints that legitimately accept more | narrow, explicit, documented | Configure the layers so the innermost is the one that normally fires, and keep the edge cap at or above it — an edge cap *below* the framework cap produces rejections your service never sees and cannot explain to a caller. Route-level overrides should raise the limit only for the endpoints that need it; a global cap sized for the largest endpoint gives every other endpoint that same exposure. ## Responding without losing the response The status is **413**, the code meaning the body is larger than the server is willing to process. Getting the status right is the easy half. The awkward half is timing: you decide to reject while the client is still sending. If the server writes the response and immediately closes, the client may still be writing, receive a connection reset, and never read the status you carefully chose — the symptom being a client-side "connection reset" where your logs clearly show 413. The mitigations are: - **drain a bounded remainder** of the request before closing, so the client finishes its write and reads the response; - **close deliberately** after the response is flushed, signalling that the connection will not carry more of this request; - **include what the client needs** — the limit that was exceeded — so the caller can fix its request rather than retry it identically. ## Things that make a limit less effective than it looks - **Compressed bodies expand.** A cap applied to the bytes on the wire does not bound the decoded size, so a small compressed body can decode into a very large one. Apply a limit to the decoded stream as well, counted as it is produced. - **Many small bodies still add up.** A per-request cap bounds one request; total exposure is still cap times concurrency, so connection and concurrency limits are part of the same control. - **Slow uploads hold resources.** A body under the cap that trickles in occupies a worker or connection for as long as the client likes, which is why read timeouts belong next to size caps. - **A limit that only exists at the edge is one misroute from irrelevant.** Anything reaching the service directly — internal callers, a health path, a sidecar — bypasses it entirely. - **Silent truncation is worse than rejection.** Reading up to the cap and processing what fits accepts a corrupted request as a valid one. Reject; never truncate. ## What to check when you inherit a service Ask for the configured cap, the maximum concurrency, and the product of the two. If nobody can answer, the service has no bound. Then check that a body over the cap actually produces a 413 in a real client rather than a reset, and that an endpoint accepting large payloads has an explicit override rather than a raised global default.

  • Why can a client see a connection reset instead of your 413?
    Because the rejection happens while the client is still uploading. If the server writes the response and closes immediately, the client's outstanding write fails and it may never read what was sent. Draining a bounded remainder before closing, or closing only after the response is flushed, lets the client finish writing and read the status.
  • How do you choose the actual number for a body size limit?
    Work from memory, not intuition. Decide the memory budget for in-flight bodies, divide by the concurrency the service permits, and use that as the global cap. Then raise it only on the specific routes that genuinely need more, so one large-payload endpoint does not set the exposure for every other route.
  • A limit is enforced on the wire bytes and the service still runs out of memory. What did you miss?
    Most likely decompression. A compressed body that fits under the wire cap can expand many times over once decoded, so the limit must also be applied to the decoded stream as it is produced. Slow trickled uploads and unbounded concurrency are the other two usual suspects.

saying these in an interview costs you the question

  • Reads the whole body into memory and then checks its length
  • Trusts the declared Content-Length instead of counting actual bytes
  • Lets an oversized body surface as a 500 from a failed parse
  • Sets one global limit sized for the largest endpoint in the service
  • Applies a limit to the compressed bytes but not to the decoded stream
  • Truncates the body at the cap and processes what fits