What do the filename and content type declared on a multipart part actually tell a server about the uploaded bytes?
answer
- written by the client, copied through
- a claim, never a verification
- presence is structural, value is input
- generate identity, keep name for display
- derive real type from the content
basics
~20 sAlmost nothing reliable. Both are strings the client wrote into the part's headers and the server copies through unverified: claims, not facts. Treat them as untrusted display metadata and derive storage identity and real type server-side.
solid answer
~40 sA part's `filename` and its declared `Content-Type` arrive inside the part's own header block, chosen entirely by the client. Nothing in the format checks the bytes against either one, so a part labelled as an image can hold anything. The practical rules follow directly: generate your own identifier for stored objects rather than reusing the supplied name, determine the media type from the content whenever a decision depends on it, keep the original name only as display metadata attached to a record, and bound its length and character set because it is arbitrary text. The one structural thing `filename` does tell you is that the part is a file part rather than a form field — that signal is reliable because it is about the envelope, not the content.
go deeper
Remember that the filename and the type on a part are typed by the client and copied through unchecked. They describe intent, not content.
Distinguish the structural signal from the value: the presence of a filename reliably marks a file part, while the string itself and the declared type are untrusted input.
Demonstrate the split you enforce in production — generated identity for storage, original name as bounded display metadata, real type derived from content where a decision depends on it.
Make it a platform default rather than a per-service habit: if every upload path generates identity the same way, the next team that ships an endpoint is far less likely to reintroduce the failure.
## Where these two values come from Both values live in the part's own header block, which the client wrote: ```http --ab17 Content-Disposition: form-data; name="document"; filename="invoice.pdf" Content-Type: application/pdf <bytes> ``` A parser reads `filename` out of the `Content-Disposition` header and the media type out of the part's `Content-Type` header, and passes them to the handler as-is. No stage of that pipeline compares either string against the bytes that follow. They are assertions made by whoever composed the request, which for a public endpoint means anyone. ## What each one is actually good for | Value | What it reliably tells you | What it does not tell you | |---|---|---| | presence of `filename` | that this part is a file part, not a form field | anything about the content | | the `filename` string itself | what the client wants the file called | that it is safe to use as an identifier, a path segment, or a type hint | | the part's `Content-Type` | what the client claims the bytes are | what the bytes actually are | | absence of a part `Content-Type` | nothing; it is frequently omitted | that the content is text | The asymmetry is the point. The **presence** of the `filename` parameter is a structural fact about the envelope and is trustworthy. Its **value** is user input in the fullest sense. ## Why the extension is the weakest signal of all An extension is just the tail of a string a client chose. Three separate problems compound: - Extensions and media types do not map one-to-one; several types share one extension and several extensions map to one type. - The client is often not even lying deliberately — operating systems and browsers guess the type themselves, so a genuine upload can arrive mislabelled. - Consumers downstream may make their own determination from the bytes and disagree with the label you stored, which is how a file ends up treated as one thing by your service and another thing by the next one. Where a decision genuinely depends on the type — whether to transcode, whether to render, which pipeline to route it to — determine it from the content and treat the declared type as a hint at best. ## Treating the name as data, not identity The durable rule is to separate the two jobs the filename appears to do: 1. **Identity** — what your system calls the stored object. Generate it yourself: an opaque identifier, with an extension you derived, in a location your code chose. Nothing carried in the request should be able to influence it. 2. **Display** — what a user sees when the file is listed or downloaded again. Keep the original string for this, stored as an ordinary attribute of the record, escaped and bounded like any other user-supplied text. Once identity is generated server-side, a whole family of problems about what the supplied name might contain simply stops being reachable from the storage path. That does not make the string safe everywhere else: it is still untrusted text wherever it is rendered, logged, or echoed back, and building a filesystem path out of it is a well-known vulnerability class in its own right rather than something to get clever about. ## The awkward edges of the string itself A robust handler assumes the supplied name may be any of these, because all of them occur in the wild: - **Empty, or absent entirely** on a part that is clearly a file. - **Very long** — long enough to be worth capping before it is stored or logged. - **Non-ASCII**, including scripts and encodings your storage layer may normalise differently than your database does. - **Duplicated** across several parts in the same request, so a naive name-keyed map silently drops files. - **Containing separators, control characters or padding whitespace**, none of which are rejected by the format. The corresponding defensive posture is dull and effective: cap the length, normalise the encoding once at the boundary, never key anything on it, and always store it beside a generated identifier rather than as one. ## A note on size The declared type and name are also the cheapest thing to reject on: they arrive in the part's headers, before any of its bytes. A part whose declared type is far outside what the endpoint accepts can be refused after a few hundred bytes have been read. That is an optimisation, not a control — a client can declare anything, so the real decision still has to be made once the content is known. ## The interview answer in one breath The presence of `filename` is structural and trustworthy; the value of `filename` and the declared type are untrusted claims. Generate identity, derive the type from the content when it matters, keep the original name as display metadata, and bound it like any other free-text field.
- If the declared type is untrustworthy, what is it still useful for?Cheap early rejection and routing hints. It arrives in the part's headers before any bytes, so an obviously unacceptable declaration can end the request early. Because a client can declare anything, the real determination still has to come from the content once it is available.
- Why store the original filename at all if you generate your own identifier?Users expect the name they uploaded when they download the file again, and support staff need it to match a record to what someone remembers sending. It belongs in a metadata field beside the generated identifier — bounded, escaped where rendered, and never used as a key.
- What should a handler do when two parts in one request carry the same filename?Nothing special, because identity should not depend on the name. Each part becomes its own stored object with its own generated identifier, and both original names are kept as attributes. A handler that keys a map by filename silently loses one of the two files.
saying these in an interview costs you the question
- Trusts the declared content type to decide what the bytes are
- Uses the supplied filename as the stored object's identity
- Believes a file extension proves the format of the content
- Assumes the supplied name is short, ASCII and non-empty
- Thinks a missing part content type means the content is plain text