skip to content

questions

5

Walk through how HTTP response compression is negotiated between a client and a server using the Accept-Encoding and Content-Encoding headers.

level: juniorimportance: must knowfreq 58%

answer

  1. Accept-Encoding asks, Content-Encoding declares
  2. identity = no coding; just omit the header
  3. Content-Type unchanged; gzip is a wrapper
  4. Content-Length = compressed bytes
  5. Vary: Accept-Encoding always

basics

~20 s

The client sends Accept-Encoding listing codings it can decode, such as gzip, br, zstd. The server may compress the body with one of them and must say which in Content-Encoding. The client decompresses before parsing. Uncompressed is called identity.

solid answer

~50 s

The client advertises what it can decode: ``` Accept-Encoding: gzip, deflate, br, zstd ``` The server may pick one of those codings, compress the response body, and **must** declare it: ``` Content-Encoding: br Vary: Accept-Encoding ``` The client sees `Content-Encoding`, decompresses, and only then parses the payload. If the server compresses nothing, it simply omits `Content-Encoding` — the coding is `identity`, and sending `Content-Encoding: identity` is not the right way to say so. Key points: compression is optional and server-chosen, never mandatory. `Content-Length`, when present, is the length of the **compressed** bytes actually on the wire. `Vary: Accept-Encoding` is required so a shared cache never hands a Brotli body to a client that only speaks gzip. And the media type is unaffected: a gzipped JSON response is still `Content-Type: application/json` with `Content-Encoding: gzip` — the encoding is a wrapper, not a new format.

code

http · 9 lines
http
GET /api/items HTTP/1.1
Host: example.com
Accept-Encoding: gzip, br, zstd

HTTP/1.1 200 OK
Content-Type: application/json
Content-Encoding: br
Vary: Accept-Encoding
Content-Length: 1042

go deeper

for a junior

Describe the two-header handshake and that the client decompresses before parsing. Know identity means uncompressed.

for a middle

Add that Content-Type is unchanged, Content-Length counts compressed bytes, and Vary: Accept-Encoding is required.

for a senior

Talk about where compression belongs in the stack, double-encoding bugs, ETag-per-encoding, minimum size thresholds, and not recompressing already-compressed media.

for a principal

Frame it as a CPU-versus-bandwidth and latency policy decision across edge and origin, including cache-key normalisation and what the CDN should own versus the app.

## The exchange Compression in HTTP is proactive content negotiation on the *encoding* axis. **Step 1 — client advertises.** Nearly every browser and HTTP library sends: ``` Accept-Encoding: gzip, deflate, br, zstd ``` Each token is a **content coding** registered with IANA. `identity` means no transformation. An absent `Accept-Encoding` header means the client has no preference; in practice servers treat that conservatively and send identity. **Step 2 — server decides.** The server picks a coding it supports from that list, or picks none. Compression is always optional: a server may decline because the payload is tiny, already compressed, CPU is scarce, or policy forbids it. **Step 3 — server declares.** If it compressed, it must send: ``` Content-Encoding: gzip ``` This is not advisory — a body compressed without the header is undecodable garbage to the client. Conversely, claiming an encoding you did not apply produces the classic "incorrect header check"/decode error. **Step 4 — client decodes.** The client strips the coding, then interprets the result according to `Content-Type`. ## Content-Encoding is a property of the representation This is the mental model to carry: `Content-Encoding` says the resource's representation has been transformed and the recipient must reverse it to recover the payload. Consequences that follow directly: - **`Content-Type` does not change.** Gzipped JSON is `Content-Type: application/json` + `Content-Encoding: gzip`. It is *not* `application/gzip` — that type is for a file that genuinely *is* a gzip archive (downloading `backup.tar.gz` is `Content-Type: application/gzip` with no `Content-Encoding`, because the compression is the content, not a transfer optimisation). - **`Content-Length` counts compressed bytes.** It describes the octets on the wire. Middleware that compresses a response but forwards the original length breaks the connection framing — a very common proxy bug. - **Range requests operate on the encoded bytes.** A `Range` on a compressed response addresses the compressed octets, which is why partial fetches and compression interact awkwardly. - **`ETag` must differ per encoding.** The gzip and Brotli representations are different octet sequences, so they need different strong ETags; sharing one causes corrupt reassembly after a conditional request through a cache. ## Caching ``` Vary: Accept-Encoding ``` is mandatory on any response whose body may be compressed. Without it, a shared cache stores whatever the first requester got and replays it — Brotli to a gzip-only client, or a compressed body to a client that sent no Accept-Encoding at all. This is one of the most frequently seen real-world CDN misconfigurations, and its symptom is mojibake or hard parse failures for a subset of users. Accept-Encoding is one of the few headers CDNs *do* normalise safely, collapsing the many client variants into `gzip`, `br`, or nothing before keying the cache. ## What not to compress - **Already-compressed payloads:** JPEG, PNG, WebP, MP4, ZIP, and anything you serve as `application/gzip`. Recompressing spends CPU to add bytes. - **Tiny bodies.** Below roughly 1 KB the gzip/Brotli frame overhead can exceed the savings; servers normally have a minimum-size threshold. - **Streaming responses where latency matters**, unless the compressor flushes per chunk. ## Where it actually happens Compression is usually done by the reverse proxy, CDN, or web server (nginx `gzip on`/`brotli on`, a CDN toggle) rather than by application code. That placement matters: if both the app and the proxy compress, you can end up with a double-encoded body — legal in principle (`Content-Encoding: gzip, gzip`, applied in order) but a source of real bugs, since many clients and proxies handle multi-value encoding poorly. Pick one layer. ## Quick sanity check on the wire `curl -H 'Accept-Encoding: gzip' -i https://example.com/` shows you the header; adding `--compressed` makes curl advertise and transparently decode. If you see `Content-Encoding: gzip` and readable text in the same output, curl decoded it for you — that is not evidence the server sent plaintext.

  • A gzipped JSON response — what are its Content-Type and Content-Encoding?
    `Content-Type: application/json` and `Content-Encoding: gzip`. The media type describes the payload after decoding; the encoding describes the transformation applied for transfer. `application/gzip` would mean the resource itself is a gzip archive, which is a different situation.
  • What breaks if a server compresses a response but omits Vary: Accept-Encoding?
    Shared caches store the compressed body under a key that ignores Accept-Encoding and replay it to clients that cannot decode it, or store an identity body and starve everyone else of compression. Symptoms are garbled responses or parse errors for a subset of users, typically only behind a CDN.
  • Which layer should perform compression — the application or the reverse proxy?
    Normally the reverse proxy or CDN: it is configured once, uses tuned native implementations, and keeps application code free of the concern. The important rule is to do it in exactly one place; if the app compresses and the proxy compresses again you get double encoding, which many intermediaries mishandle.

Content-Encoding is shrink-wrap on a parcel: the label describing the contents (Content-Type) is unchanged, but the recipient must unwrap before using it. application/gzip is when the parcel's actual contents are a roll of shrink-wrap.

saying these in an interview costs you the question

  • Saying gzipped JSON is served as `Content-Type: application/gzip`.
  • Claiming Content-Length is the uncompressed size.
  • Sending `Content-Encoding: identity` instead of omitting the header.
  • Believing the server must compress whenever Accept-Encoding lists gzip.
  • Forgetting Vary: Accept-Encoding on compressible responses.

context

open as a page

Compare gzip, Brotli (br) and Zstandard (zstd) as HTTP content codings. When would you choose each?

level: middleimportance: should knowfreq 40%

basics

~20 s

gzip is the universal baseline — modest ratio, fast, supported everywhere. Brotli compresses text best, especially precompressed at high levels, but is slow to compress at high settings. zstd compresses near-Brotli quality far faster, making it good for dynamic responses; browser support is the newest.

open as a page

What is the difference between the HTTP Content-Encoding and Transfer-Encoding headers?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Content-Encoding is end-to-end: it transforms the representation itself, so only the final recipient decodes it and caches store it encoded. Transfer-Encoding is hop-by-hop framing for one connection — chunked is its only real use — and it does not exist in HTTP/2 or HTTP/3.

open as a page

How do you serve pre-compressed static assets over HTTP — for example .br and .gz files written at build time — and what must be correct for caches and proxies?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Build both a .br and .gz copy beside each asset. The server checks Accept-Encoding, serves the matching file with the original Content-Type, the right Content-Encoding, Vary: Accept-Encoding, a distinct ETag per encoding, and falls back to the uncompressed file.

open as a page

How would you set a compression policy for a service sitting behind a CDN — what gets compressed, at which layer, and at what cost?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Compress text-like types above a size threshold, never already-compressed media; precompress static assets at maximum level and use a fast coding for dynamic ones; do it in exactly one layer; and exclude responses that mix secrets with attacker-influenced input.

open as a page