skip to content

Multipart Upload Handling

Parsing multipart bodies: streaming parts vs buffering to memory or disk, size caps, temp-file cleanup, and the untrusted filename. Interviewers ask it because uploads leak disk and memory.

on this pageshow

questions

5

How does a server-side web framework turn a multipart/form-data request body into the parts a handler works with?

level: juniorimportance: must knowfreq 70%

answer

  1. one body, many envelopes
  2. boundary token from the content type
  3. each part has its own headers
  4. name means field, filename means file
  5. sequential scan, no look-ahead

basics

~20 s

A multipart/form-data body is a run of boundary-delimited parts, each with its own headers naming the form field and, for files, a filename. The framework walks them in arrival order as field values or file handles.

solid answer

~40 s

The request's `Content-Type` header carries a `boundary` token, and the body is a run of parts separated by that boundary, ending with a closing boundary. Each part opens with its own headers — normally a `Content-Disposition` header with a `name` parameter (the form field) and, for file parts, a `filename` parameter — then a blank line, then raw bytes. A parser scans the stream, splits on the boundary, and surfaces each part either as a text field or as a file handle. Nothing prefixes a part with its length, so the parser only learns a part has ended when it meets the next boundary. Parts therefore arrive strictly in the order the client wrote them, and the same `name` may legitimately appear more than once.

go deeper

for a junior

Recall the three structural facts: a boundary from the content type separates parts, each part has its own headers, and a filename parameter marks a file part.

for a middle

Explain why parsing is sequential and single-pass — no length prefixes, no index — and what that implies for ordering, repeated field names and re-reading a part.

for a senior

Show you have designed around it: required field ordering, deciding between collecting and iterating, and what the handler is allowed to assume about parts it has not reached yet.

for a principal

Frame it as a contract with clients. Requiring small fields before large ones, or capping part counts, is an API decision you must document and version, not a private parser detail.

## Why a separate body format exists A form whose fields are all short text can be sent as one encoded string of name/value pairs. That stops working the moment a field holds arbitrary bytes — an image, an archive, a document — because those bytes can contain the separator characters themselves, and escaping every occurrence would inflate the payload badly. `multipart/form-data` solves it differently: each field gets its own envelope, and the envelopes are separated by a delimiter the client picks and asserts does not occur in any of the data. ## The shape on the wire The request's `Content-Type` header names the media type and carries a `boundary` parameter — a random-looking token chosen by the client. The body is then, in order: 1. A line consisting of two hyphens and the boundary token. 2. One part: a small header block, a blank line, then the raw bytes of that field's value. 3. The boundary line again, then the next part — repeated for as many parts as the form has. 4. A final boundary line with two trailing hyphens, which marks the end of the body. A part's header block is tiny. It almost always holds a `Content-Disposition: form-data` header with a `name` parameter giving the form field name, plus a `filename` parameter when the part is a file. A part may also declare its own `Content-Type` describing its bytes. ```http POST /documents HTTP/1.1 Content-Type: multipart/form-data; boundary=ab17 --ab17 Content-Disposition: form-data; name="title" Quarterly report --ab17 Content-Disposition: form-data; name="document"; filename="q3.pdf" Content-Type: application/pdf <raw bytes of the file> --ab17-- ``` ## Field parts and file parts | | field part | file part | |---|---|---| | `Content-Disposition` | has `name`, no `filename` | has `name` **and** `filename` | | expected size | small, a form value | unbounded, whatever was selected | | how a framework surfaces it | a string in the request's form values | a handle: a stream, a temporary file, or a byte array | | its own `Content-Type` | usually absent | often present, chosen by the client | The distinguishing signal is the presence of the `filename` parameter, not the size of the part and not the declared media type. A file part can be empty, and a field part can be large. ## How the parser actually walks it A multipart parser is a sequential scanner, not a random-access reader. For each part it reads the header block, then streams bytes onward while watching for the next boundary. Several consequences follow, and they are what interviewers are usually probing: - **No part carries its length.** The format has no length prefix, so the end of a part is discovered, not announced. A per-part size cap can therefore only be enforced as bytes accumulate. - **Order is the client's order.** Parts arrive in the sequence the client wrote them, which normally mirrors the field order in the form. - **Names repeat legally.** The same `name` may appear in several parts (a multi-select, or several files under one field), so the natural in-memory shape is a multi-map, not a map. - **Each part has its own type.** The type declared on a part describes only that part; the request's own `Content-Type` describes the multipart envelope. - **The header block is untrusted input too.** Its size, the number of parameters and the encoding of the values all come from the client. ## Two ways a handler receives the parts Frameworks differ here, and many offer both: - **Collected.** The framework parses the whole body before invoking the handler, which then sees a finished collection of fields and files. Simple to use, at the cost of having buffered every part somewhere first. - **Iterated.** The handler pulls parts one at a time as the parser produces them, and must consume or explicitly discard a part's bytes before requesting the next. Cheaper in memory, but order-dependent and single-pass. ## The ordering trap in an iterated parse An iterated parse cannot look ahead. If the handler wants a small `token` field that the client wrote *after* a very large file part, the only way to reach that field is to get past the file part's bytes first — into memory, onto disk, or discarded. This is why services that want to check a small field before accepting a large payload either require clients to send the small fields first, or fall back to collecting the whole body. It is also why a consumed part cannot simply be read again: unless the framework kept a copy, those bytes are gone. ## What to say in an interview Name the three facts and you have answered it: the body is boundary-delimited, every part carries its own headers with `name` and optionally `filename`, and parsing is sequential so ordering and single-pass consumption are real constraints on the handler.

  • What tells a server that a part is a file rather than an ordinary form field?
    The presence of a `filename` parameter on the part's `Content-Disposition` header. Size does not decide it and neither does the part's declared media type: a file part may be empty, and a text field may be large. A missing `filename` means the part is treated as a form value.
  • Why can a handler iterating parts not read a later field before an earlier file part?
    Because the format has no index and no length prefixes: the position of a later part is only known once the bytes before it have been scanned for the next boundary. Reaching it means consuming, buffering or discarding everything in front of it.
  • What should a server do when the same field name appears on several parts?
    Treat it as legal and decide explicitly. The natural representation is a collection per name; taking only the first or only the last is a choice that should be deliberate, because clients legitimately send repeated fields for multi-value inputs and multiple file selections.

saying these in an interview costs you the question

  • Thinks a multipart body is a JSON document with files attached to it
  • Assumes parts form a random-access collection readable in any order
  • Believes a part's field name and its filename are the same thing
  • Expects every part to be fully in memory before the handler starts
  • Thinks a part without a filename is malformed rather than a form field
open as a page

Which limits should a server enforce while parsing a multipart upload, and at what moment must it answer 413?

level: seniorimportance: must knowfreq 58%

basics

~10 s

At least four: a per-part cap, a whole-request cap, a maximum part count, and caps on header blocks and in-memory field values. Counters run as bytes arrive, so 413 goes out mid-body, not after.

open as a page

In multipart upload handling, why do servers buffer small parts in memory but spill larger ones to a temporary file?

level: middleimportance: should knowfreq 62%

basics

~20 s

Memory is the scarce shared resource: small parts are cheap to hold, so past a configured byte threshold the parser writes the part to a temporary file instead, trading disk and copying for a bounded memory footprint.

open as a page

What do the filename and content type declared on a multipart part actually tell a server about the uploaded bytes?

level: middleimportance: should knowfreq 52%

basics

~20 s

Almost nothing reliable. Both are strings the client wrote into the part's headers and the server copies through unverified: claims, not facts. Treat them as untrusted display metadata and derive storage identity and real type server-side.

open as a page

A service that accepts uploads slowly fills its disk with temporary files. How does multipart temp-file cleanup work, and where does it fail?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A spilled part's temporary file is deleted by a hook tied to the end of the request. Disk fills on the endings that skip it: client aborts, parses stopped by a limit, handlers that move the file, and crashes.

open as a page