skip to content

In multipart upload handling, why do servers buffer small parts in memory but spill larger ones to a temporary file?

level: middleimportance: should knowfreq 62%

answer

  1. the caller chooses the work
  2. memory, temp file, or straight through
  3. threshold changes destination, not acceptance
  4. concurrency multiplies every buffer
  5. spilling buys a cleanup obligation

basics

~20 s

Memory is the scarce shared resource: small parts are cheap to hold, so past a configured byte threshold the parser writes the part to a temporary file instead, trading disk and copying for a bounded memory footprint.

solid answer

~40 s

There are three places a part's bytes can go: memory, a temporary file, or straight through to wherever the handler is sending them. Full in-memory buffering is simplest and re-readable, but the worst case is concurrent uploads multiplied by the maximum part size, which is a memory exhaustion risk. Streaming through uses near-constant memory but is single-pass: the handler consumes bytes before it knows the request is complete or within limits. The spill threshold is the compromise — hold a part in memory while it is small, and once it crosses a configured number of bytes, write the remainder to a temporary file and hand the handler a file-backed handle. The cost is disk space, an extra copy, and a temporary file that something must delete.

go deeper

for a junior

Know that an uploaded file does not necessarily sit in memory: past a configured size the server writes it to a temporary file and gives the handler a file-backed handle instead.

for a middle

Explain the three destinations and their costs, and be clear that the spill threshold changes where bytes go while a separate maximum decides whether they are accepted at all.

for a senior

Show the arithmetic: threshold x parts x concurrent uploads, which volume the temporary directory lives on, and how the handler undoes partial work when a streamed upload aborts.

for a principal

Own the policy question — which services stream, which buffer, and what the platform guarantees about the temporary volume — so each team is not rediscovering the same exhaustion mode.

## The resource being protected An upload endpoint is unusual among request handlers: the amount of work it is asked to do is chosen by the caller. A server that holds each uploaded part wholly in memory has a worst case of *concurrent uploads x maximum part size*, and that product is what falls over first. Buffering strategy is the lever that bounds it. ## The three destinations for a part's bytes | Strategy | Memory used | Disk used | Re-readable | Main failure it invites | |---|---|---|---|---| | Fully in memory | the whole part, per concurrent request | none | yes | memory exhaustion under concurrency or one large upload | | Spill to a temporary file above a threshold | the threshold, per part | the whole part | yes | the temporary file is never deleted; disk fills | | Streamed straight through | a small fixed buffer | none | no | handler has committed to bytes before the request is known good | Most server-side frameworks default to the middle row, because it fails in the least damaging way: disk is cheaper than memory, easier to monitor, and the failure is gradual rather than instant. ## What the threshold actually does The threshold is a byte count, applied per part: 1. The parser accumulates the part's bytes in an in-memory buffer. 2. If the part ends before the buffer reaches the threshold, it stays in memory — no file is ever created. 3. If the buffer reaches the threshold, the parser creates a temporary file, writes out what it already holds, and streams the remainder to that file. 4. The handler receives a handle that hides which of the two happened, and typically exposes both a byte view and a stream view. Two details follow that candidates often miss. First, the threshold is **not** a limit: crossing it changes where bytes go, it does not reject anything. The rejection lever is a separate maximum size. Second, the threshold applies per part, so the real memory bound is the threshold multiplied by the number of parts being buffered at once, not the threshold alone. ## Why not stream everything Streaming is the cheapest option and it is the right one for large media, but it changes the shape of the handler in ways that matter: - The bytes are **single-pass**. Once consumed they cannot be read again unless the handler itself kept a copy. - The handler starts writing to its destination before the request has finished arriving, so it must be able to undo that work when the upload turns out to be truncated, aborted, or over a limit. - Anything the handler needs to decide *before* writing — a field carried in another part — must arrive earlier in the body, because the parse is sequential. - Work that depends on the whole payload at once, such as computing a digest over the entire file, is still possible, but only as bytes flow past, not by re-reading. ## Why not buffer everything in memory It is attractive for small forms and it is genuinely simplest, but the exposure is poor: - The worst case is set by the client, not the server. - A single request holding many parts multiplies the cost. - Memory pressure degrades the entire process, including requests that have nothing to do with uploads, whereas disk pressure degrades the upload path first. - Under an execution model that dedicates a thread to each request, a slow upload also pins that thread for as long as it holds the buffer; under an event-loop model it holds the buffer without pinning a thread, but the memory cost is identical. ## Choosing and tuning the numbers - Set the threshold at, or slightly above, the size of the largest **field** you expect — the point is to keep ordinary form values out of the filesystem. - Size the temporary directory deliberately and know which volume it is on; a default that points at a small system partition is a classic production surprise. - Compute the bound explicitly: threshold x parts buffered concurrently x concurrent uploads, and compare it against the memory you are prepared to lose. - Prefer streaming when the destination is another sink entirely and the handler does not need to re-read; prefer spilling when the handler must inspect the payload, re-read it, or hand it to code that expects a file. - Remember that a spilled part is a new lifecycle: something must delete it, on the error path as well as the success path. ## The summary answer Memory is shared and unbounded demand for it is the failure everyone wants to avoid; disk is cheap, measurable and degrades locally. The threshold buys the ergonomics of an in-memory value for the common small part while capping the exposure of the rare large one — and the price of that trade is a temporary file with a lifecycle of its own.

  • Does crossing the in-memory threshold ever reject the upload?
    No. The threshold decides where bytes are kept; it is not an acceptance decision. Rejection comes from separate maximum sizes for a part and for the whole request, which produce a 413 response. Conflating the two leaves a server with a small memory footprint and no actual cap.
  • What does a handler give up when it consumes a part as a stream instead of a buffered handle?
    Re-readability and hindsight. The bytes pass once, so anything needing a second pass must be done inline, and the handler may have already written data to its destination when the upload turns out to be truncated, aborted or over a limit — so it needs a compensating cleanup path.
  • How do you compute the memory this setting can cost you?
    Multiply the in-memory threshold by the number of parts buffered at the same time within one request, then by the number of uploads you allow to run concurrently. That product, not the threshold on its own, is the figure to compare against the memory budget.

It is the same trade an external sort makes: keep the small runs in memory where they are fast, and spill the big ones to disk rather than fail outright. The spill is not a limit, it is a change of address.

saying these in an interview costs you the question

  • Thinks the buffer threshold is the maximum upload size
  • Sizes the threshold per upload and forgets concurrency multiplies it
  • Believes streaming removes the need for any size limit
  • Assumes spilling to disk is free because disk is cheap
  • Thinks a streamed part can simply be read a second time