skip to content

questions

3

HTTP fields are commonly grouped into categories such as control data, representation metadata, request context and response context. What does each group mean, and why is the distinction useful when you read or design a message?

level: juniorimportance: should knowfreq 44%

answer

  1. Control = steering the message
  2. Representation = describing the bytes
  3. Validators = which version (ETag, Last-Modified)
  4. Context = who asked / who answered
  5. Content-* survives on HEAD and 304

basics

~20 s

Control data steers how the message is handled (Host, Cache-Control, Range, Date). Representation metadata describes the bytes being carried (Content-Type, Content-Encoding, Content-Language) and validators identify their version (ETag, Last-Modified). Context fields describe the sender or the resource (User-Agent, Referer, Server, Allow). Grouping tells you what a field applies to: the message, the representation, or the party.

solid answer

~50 s

HTTP fields are not one flat bag; each one applies to something specific. - **Control data** directs processing of the message itself: `Host`, `Cache-Control`, `Range`, `Expect`, `Date`, `Location`, `Retry-After`. - **Representation metadata** describes the enclosed representation — what the bytes are and how they are encoded: `Content-Type`, `Content-Encoding`, `Content-Language`, `Content-Location`. - **Validators** identify which version of a representation you hold: `ETag`, `Last-Modified`. They feed conditional requests. - **Request context**: who is asking and from where — `User-Agent`, `Referer`, `From`. - **Response context**: facts about the resource or origin — `Server`, `Allow`. - **Authentication**: `Authorization`, `WWW-Authenticate`, `Proxy-Authenticate`. The practical payoff: representation metadata travels *with the bytes*, so it also appears on `HEAD` and `206 Partial Content` responses and describes what a `GET` would return, while control data affects only the message in hand. Confusing the two produces bugs such as caching on the wrong axis, or assuming Content-Type describes the resource rather than the representation returned.

code

http · 9 lines
http
HTTP/1.1 200 OK
Date: Tue, 12 Aug 2026 09:15:00 GMT
Server: nginx
Cache-Control: max-age=300
Vary: Accept-Encoding
ETag: "9f3c1"
Content-Type: application/json; charset=utf-8
Content-Encoding: gzip
Content-Length: 812

go deeper

for a junior

Name the groups with one or two examples each and say what each group applies to: the message, the bytes, or the sender.

for a middle

Add the representation-versus-resource point, why Content-* fields appear on HEAD and 206, and why Vary exists.

for a senior

Use the categories to reason about intermediary rewriting: who may legally change what, and which fields must be recomputed when a body is transformed.

for a principal

Treat the taxonomy as the contract boundary: which layer owns which field, what the edge may inject or strip, and how that shapes caching and negotiation strategy across the estate.

## Fields, not just headers HTTP calls them **fields**. A field appearing before the body is a *header field*; one appearing after a chunked body is a *trailer field*. Both use the same syntax, `name: value`, and the same registry of names. Thinking in terms of fields rather than headers is what makes trailers and HTTP/2 pseudo-headers fit into the same mental model. ## The categories **Control data** tells a recipient how to handle *this message*. Examples on requests: `Host` (which authority is addressed), `Range` (which part of the representation is wanted), `Expect`, `Max-Forwards`, `Cache-Control`, `If-*` preconditions. On responses: `Date`, `Location`, `Retry-After`, `Age`, `Vary`, `Cache-Control`. Control data is the steering wheel, and it must be visible before the body is processed — which is exactly why control data may never be sent as trailer fields. **Representation metadata** describes the bytes enclosed, or the bytes that *would* be enclosed. `Content-Type` says how to interpret them, `Content-Encoding` says what coding was applied on top (gzip, br), `Content-Language` names the natural language, `Content-Location` gives a direct URL for this particular representation. The key insight is the word *representation*: one resource can have many representations (JSON and HTML, English and German, gzipped and identity), selected by content negotiation. Representation metadata describes the one you got, not the resource in the abstract. That is why `Vary` exists — to tell caches which request fields selected this representation. **Validators** — `ETag` and `Last-Modified` — are a special slice of representation metadata: an opaque or timestamp identifier for the version you hold. They are what `If-None-Match` and `If-Match` compare against, so they are the hinge between representation metadata and conditional control data. **Request context** describes the client and the circumstances of the request rather than the payload: `User-Agent`, `Referer`, `From`. It is advisory; servers may log it or vary on it, but nothing in the protocol depends on it being accurate, and all of it is attacker-controlled. **Response context** describes the origin or the resource rather than the payload: `Server`, `Allow` (which methods the resource supports). **Authentication fields** form their own group because they follow a challenge/response pattern: `WWW-Authenticate` challenges, `Authorization` answers, and `Proxy-Authenticate` / `Proxy-Authorization` do the same for one hop only. ## Why the split earns its keep Two practical consequences show up constantly. First, **representation metadata rides with the representation, even when the body is absent**. A `HEAD` response carries `Content-Type` and `Content-Length` describing what a `GET` would return, though it has no body. A `304 Not Modified` carries validators and caching control data so the client can update its stored copy. A `206 Partial Content` describes the selected representation while `Content-Range` scopes the fragment. If you assume every `Content-*` field implies bytes present, you will misparse all three. Second, **the category tells you where a field may legitimately be modified**. An intermediary that decompresses a body is altering the representation, so it must fix `Content-Encoding` and the framing fields. An intermediary that adds `Age` or `Via` is adding control and context data about the hop. A gateway rewriting a `Location` header is touching control data with routing consequences. Knowing the category is knowing what breaks if you touch it. ## A note on Content-Length `Content-Length` looks like representation metadata because of its name, but it primarily frames the message. That dual nature is why it appears on `HEAD` responses as pure metadata and on bodied responses as framing. ## What interviewers look for Not memorised lists, but the ability to place an unfamiliar field: is it steering the message, describing the bytes, identifying a version, or describing a party? A candidate who can answer that will make correct decisions about caching, negotiation, trailers and proxy rewriting without looking anything up.

  • Why does a 304 Not Modified response carry ETag and Cache-Control but no Content-Type?
    Because 304 sends no representation: the client already has it. The fields present are the ones that update the client's stored entry, namely validators and caching control data. Representation metadata such as Content-Type is unnecessary, since the stored copy already carries it from the original 200.
  • Where does the Vary response header field fit, and what problem does it solve?
    Vary is control data aimed at caches. It names the request fields whose values selected this particular representation, such as Accept-Encoding or Accept-Language. Without it, a cache could serve a gzipped or German representation to a client that asked for neither, because it keys only on the URL.

saying these in an interview costs you the question

  • Treating Content-Type as a property of the resource rather than of the returned representation
  • Assuming any Content-* field means a body is present
  • Trusting request-context fields such as User-Agent or Referer for authorisation decisions
  • Believing header names are fixed by the protocol with no registry or extension model
  • Putting control data such as Cache-Control into trailer fields

context

open as a page

Are HTTP field names case-sensitive, and if the same field name appears on several lines of one message, how should a recipient interpret it? Give an example where the usual rule does not hold.

level: middleimportance: should knowfreq 40%

basics

~20 s

Field names are case-insensitive, so Content-Type and content-type are the same field (HTTP/2 and HTTP/3 require lowercase on the wire). Repeated lines of one field may be joined into a single comma-separated value, but only for fields defined as comma-separated lists. Set-Cookie is the exception: it must stay as separate lines.

open as a page

What is the difference between an end-to-end HTTP field and a hop-by-hop one, how does the Connection header field designate the hop-by-hop set, and what goes wrong if a proxy forwards them unchanged?

level: seniorimportance: should knowfreq 38%

basics

~20 s

End-to-end fields travel from origin sender to final recipient through every proxy. Hop-by-hop fields describe one connection only and must be consumed and removed by each hop. The Connection header field names them, plus a fixed set such as Keep-Alive, Transfer-Encoding, TE, Trailer, Upgrade and the Proxy-Authenticate pair. Forwarding them leaks one hop's connection state onto another and breaks framing or upgrades.

open as a page