When an HTTP server builds the request object it hands to a web framework, which parts are parsed already and which are left untouched?
answer
- head eager, body lazy
- server parsed what it needed itself
- raw query, parsed on demand
- connection facts are fields, not headers
- repeated header fields must survive
basics
~20 sThe start line and the whole header block are parsed and validated before hand-off; the body is usually handed over unread. Query strings and structured header values are often left raw, parsed on demand above the seam.
solid answer
~40 sA server will not invoke the framework until it has a **complete, valid head**: the method, the request target and the protocol version from the start line, plus every header field of the header block, exposed as a case-insensitive collection with repeated fields preserved. The **body is different** - it is normally presented as a stream that has not been consumed, because the server cannot know whether the application wants all of it, some of it, or none. Detail above that is usually deferred: the query string is often handed over raw and parsed lazily, and structured header values are parsed only when read. The object also carries facts that are not header fields at all, such as the peer address, whether the connection is secure and which protocol version was negotiated.
go deeper
Know that method, path and headers are ready to use the moment your code runs, while the body is something you choose to read. Do not write code that expects to parse the request line yourself.
Explain the eager head versus lazy body split and why it exists, and name what is deliberately deferred: query decoding, body interpretation and structured header values.
Show awareness of the failure modes deferral creates - a body that errors mid-handler, and two layers decoding the same raw value under different rules and disagreeing about the request.
Decide where in your stack a request is canonicalised, and make that one place. Ambiguity that survives to two independent parsers is the shape of a whole class of request-smuggling and authorization-bypass defects.
## Why the head is parsed and the body is not The server needs the head to do its own job. It cannot route bytes, enforce limits, decide on connection reuse or choose a framing strategy for the response without knowing the method, the target and the header fields. So the head is parsed eagerly and validated: a malformed start line or header block is refused before the framework is ever called. The body is the opposite case. It can be megabytes, it can be a long-lived upload, and the server has no idea whether the application will read it at all. Buffering it eagerly would turn every large request into memory pressure and would make streaming impossible. So the usual arrangement is: **the body crosses the seam as a stream that has not been read yet.** Small bodies are sometimes already buffered as an optimisation, and some execution models deliver the body as chunks pushed to the application rather than pulled from it - the contract differs, but the principle is the same, the body is not parsed content, it is bytes waiting for someone to decide what they are. ## What the request object typically carries | part | state at hand-off | who finishes the job | |---|---|---| | method | parsed, a plain token | nobody, it is final | | request target | split into path and raw query; percent-encoding may or may not be decoded | framework, when reading path variables or query values | | protocol version | negotiated and recorded | nobody | | header fields | fully parsed into name/value pairs, case-insensitive lookup, repeats preserved | framework, when interpreting structured values | | body | an unread stream (sometimes a small buffer) | framework, when it reads or decodes it | | peer address, secure flag | attached as fields, not headers | framework, if it cares | ## Facts that are not header fields A recurring source of confusion is expecting connection-level facts to appear as headers. They do not: the peer address, whether the connection was encrypted, and which protocol version was negotiated are properties of the connection, observed by the server and attached to the request object as separate fields. Any header field that appears to say something about the client came from the client or from an intermediary and is data, not observation. ## Deferred parsing, and why it matters Several things are deliberately **not** done at construction time: 1. **Query parsing.** The raw query string is a sequence of bytes with no single canonical interpretation - repeated keys, nested syntax and array notation all differ by convention. Servers usually hand it over raw and let the framework apply its own rules. 2. **Body interpretation.** Whether the bytes are a form submission, a serialized document or an upload is a content-type question the application answers. 3. **Structured header values.** Fields that carry lists, parameters and quality values are parsed when they are read, not up front, because most requests never read most headers. Deferral is a performance decision first - most of what arrives is never inspected - but it has a correctness consequence worth knowing: two layers that each parse the same raw value with different rules can disagree, and disagreement between layers about what a request says is a classic source of bugs. ## Details that surprise people - **Header names are matched without regard to case.** The collection is not a plain map keyed by the exact bytes that arrived; looking a header up by the casing you expect must work regardless of how the client wrote it. - **A header can appear more than once.** A request object has to preserve that - collapsing repeats to the first or last value loses information, and some fields legitimately repeat. - **The body may already be empty even though a length was announced**, because the server may have refused or truncated it under a limit before hand-off. - **Nothing about the request tells you which handler will run.** Selection happens above the seam; the server neither knows nor cares which part of the application will be given the object. ## The practical summary Think of the request object as *everything the server had to understand anyway*, plus a handle to everything it deliberately did not read. That division explains most of its shape: the head is complete because the server needed it, the body is a stream because the server did not, and the interpretation of anything ambiguous is left to the layer that actually has an opinion about what it means.
- Why does a server usually not decode the query string before hand-off?Because there is no single convention for repeated keys, nested structures or list syntax, and applying one at the server would impose it on every framework. Handing over the raw string lets the layer that has an opinion decode it, at the cost that two layers decoding the same string with different rules can disagree about what the request said.
- What is the consequence of the body arriving unread rather than buffered?The application controls when and whether bytes are pulled, which is what makes streaming uploads and early rejection possible without holding the payload in memory. It also means failures can surface late - a client that stops sending mid-body produces an error while the handler is already running, not before it started.
- Why must repeated header fields be preserved rather than collapsed?Several fields legitimately appear more than once, and collapsing them to a single value silently discards data an application may need. It also creates a security hazard: when two layers disagree on whether to take the first, the last, or the join of duplicates, they can be made to read the same request differently.
saying these in an interview costs you the question
- Thinks the framework receives raw bytes and parses the request itself
- Assumes the whole body is always read before the framework is called
- Expects header lookup to be case-sensitive on the exact bytes sent
- Expects the peer address to arrive as an ordinary header field
- Believes the server decides which handler will receive the request