Walk through how HTTP response compression is negotiated between a client and a server using the Accept-Encoding and Content-Encoding headers.
answer
- Accept-Encoding asks, Content-Encoding declares
- identity = no coding; just omit the header
- Content-Type unchanged; gzip is a wrapper
- Content-Length = compressed bytes
- Vary: Accept-Encoding always
basics
~20 sThe client sends Accept-Encoding listing codings it can decode, such as gzip, br, zstd. The server may compress the body with one of them and must say which in Content-Encoding. The client decompresses before parsing. Uncompressed is called identity.
solid answer
~50 sThe client advertises what it can decode: ``` Accept-Encoding: gzip, deflate, br, zstd ``` The server may pick one of those codings, compress the response body, and **must** declare it: ``` Content-Encoding: br Vary: Accept-Encoding ``` The client sees `Content-Encoding`, decompresses, and only then parses the payload. If the server compresses nothing, it simply omits `Content-Encoding` — the coding is `identity`, and sending `Content-Encoding: identity` is not the right way to say so. Key points: compression is optional and server-chosen, never mandatory. `Content-Length`, when present, is the length of the **compressed** bytes actually on the wire. `Vary: Accept-Encoding` is required so a shared cache never hands a Brotli body to a client that only speaks gzip. And the media type is unaffected: a gzipped JSON response is still `Content-Type: application/json` with `Content-Encoding: gzip` — the encoding is a wrapper, not a new format.
code
http · 9 linesGET /api/items HTTP/1.1
Host: example.com
Accept-Encoding: gzip, br, zstd
HTTP/1.1 200 OK
Content-Type: application/json
Content-Encoding: br
Vary: Accept-Encoding
Content-Length: 1042go deeper
Describe the two-header handshake and that the client decompresses before parsing. Know identity means uncompressed.
Add that Content-Type is unchanged, Content-Length counts compressed bytes, and Vary: Accept-Encoding is required.
Talk about where compression belongs in the stack, double-encoding bugs, ETag-per-encoding, minimum size thresholds, and not recompressing already-compressed media.
Frame it as a CPU-versus-bandwidth and latency policy decision across edge and origin, including cache-key normalisation and what the CDN should own versus the app.
## The exchange Compression in HTTP is proactive content negotiation on the *encoding* axis. **Step 1 — client advertises.** Nearly every browser and HTTP library sends: ``` Accept-Encoding: gzip, deflate, br, zstd ``` Each token is a **content coding** registered with IANA. `identity` means no transformation. An absent `Accept-Encoding` header means the client has no preference; in practice servers treat that conservatively and send identity. **Step 2 — server decides.** The server picks a coding it supports from that list, or picks none. Compression is always optional: a server may decline because the payload is tiny, already compressed, CPU is scarce, or policy forbids it. **Step 3 — server declares.** If it compressed, it must send: ``` Content-Encoding: gzip ``` This is not advisory — a body compressed without the header is undecodable garbage to the client. Conversely, claiming an encoding you did not apply produces the classic "incorrect header check"/decode error. **Step 4 — client decodes.** The client strips the coding, then interprets the result according to `Content-Type`. ## Content-Encoding is a property of the representation This is the mental model to carry: `Content-Encoding` says the resource's representation has been transformed and the recipient must reverse it to recover the payload. Consequences that follow directly: - **`Content-Type` does not change.** Gzipped JSON is `Content-Type: application/json` + `Content-Encoding: gzip`. It is *not* `application/gzip` — that type is for a file that genuinely *is* a gzip archive (downloading `backup.tar.gz` is `Content-Type: application/gzip` with no `Content-Encoding`, because the compression is the content, not a transfer optimisation). - **`Content-Length` counts compressed bytes.** It describes the octets on the wire. Middleware that compresses a response but forwards the original length breaks the connection framing — a very common proxy bug. - **Range requests operate on the encoded bytes.** A `Range` on a compressed response addresses the compressed octets, which is why partial fetches and compression interact awkwardly. - **`ETag` must differ per encoding.** The gzip and Brotli representations are different octet sequences, so they need different strong ETags; sharing one causes corrupt reassembly after a conditional request through a cache. ## Caching ``` Vary: Accept-Encoding ``` is mandatory on any response whose body may be compressed. Without it, a shared cache stores whatever the first requester got and replays it — Brotli to a gzip-only client, or a compressed body to a client that sent no Accept-Encoding at all. This is one of the most frequently seen real-world CDN misconfigurations, and its symptom is mojibake or hard parse failures for a subset of users. Accept-Encoding is one of the few headers CDNs *do* normalise safely, collapsing the many client variants into `gzip`, `br`, or nothing before keying the cache. ## What not to compress - **Already-compressed payloads:** JPEG, PNG, WebP, MP4, ZIP, and anything you serve as `application/gzip`. Recompressing spends CPU to add bytes. - **Tiny bodies.** Below roughly 1 KB the gzip/Brotli frame overhead can exceed the savings; servers normally have a minimum-size threshold. - **Streaming responses where latency matters**, unless the compressor flushes per chunk. ## Where it actually happens Compression is usually done by the reverse proxy, CDN, or web server (nginx `gzip on`/`brotli on`, a CDN toggle) rather than by application code. That placement matters: if both the app and the proxy compress, you can end up with a double-encoded body — legal in principle (`Content-Encoding: gzip, gzip`, applied in order) but a source of real bugs, since many clients and proxies handle multi-value encoding poorly. Pick one layer. ## Quick sanity check on the wire `curl -H 'Accept-Encoding: gzip' -i https://example.com/` shows you the header; adding `--compressed` makes curl advertise and transparently decode. If you see `Content-Encoding: gzip` and readable text in the same output, curl decoded it for you — that is not evidence the server sent plaintext.
- A gzipped JSON response — what are its Content-Type and Content-Encoding?`Content-Type: application/json` and `Content-Encoding: gzip`. The media type describes the payload after decoding; the encoding describes the transformation applied for transfer. `application/gzip` would mean the resource itself is a gzip archive, which is a different situation.
- What breaks if a server compresses a response but omits Vary: Accept-Encoding?Shared caches store the compressed body under a key that ignores Accept-Encoding and replay it to clients that cannot decode it, or store an identity body and starve everyone else of compression. Symptoms are garbled responses or parse errors for a subset of users, typically only behind a CDN.
- Which layer should perform compression — the application or the reverse proxy?Normally the reverse proxy or CDN: it is configured once, uses tuned native implementations, and keeps application code free of the concern. The important rule is to do it in exactly one place; if the app compresses and the proxy compresses again you get double encoding, which many intermediaries mishandle.
Content-Encoding is shrink-wrap on a parcel: the label describing the contents (Content-Type) is unchanged, but the recipient must unwrap before using it. application/gzip is when the parcel's actual contents are a roll of shrink-wrap.
saying these in an interview costs you the question
- Saying gzipped JSON is served as `Content-Type: application/gzip`.
- Claiming Content-Length is the uncompressed size.
- Sending `Content-Encoding: identity` instead of omitting the header.
- Believing the server must compress whenever Accept-Encoding lists gzip.
- Forgetting Vary: Accept-Encoding on compressible responses.