A server compresses responses when the request carries Accept-Encoding: gzip. Why must it also send Vary: Accept-Encoding, and what concretely goes wrong in a shared cache if it does not?
answer
- Compression = negotiation on Accept-Encoding
- Missing Vary: gzip bytes to non-gzip client
- Or: silent loss of compression
- Put Vary on the identity variant too
- Normalize to gzip / br / none
basics
~20 sCompression makes one URL produce two different bodies. Without Vary: Accept-Encoding a cache stores one of them and serves it to everyone, so a client that cannot decompress may receive gzip bytes, or a gzip-capable client gets an uncompressed body.
solid answer
~50 sCompression is content negotiation: the bytes on the wire depend on the request's `Accept-Encoding`. `Vary: Accept-Encoding` tells caches to include that header in the key, so the gzip and identity variants are stored and served separately. Without it the outcomes are asymmetric. If a shared cache first stores the gzip response, a later client that did not advertise gzip receives a body with `Content-Encoding: gzip` it may not decode — mojibake, JSON parse errors, or a broken download. If the identity response is stored first, gzip-capable clients simply lose compression: correct output, wasted bandwidth. The first case is a real bug; the second is a performance regression. This is the canonical `Vary` case, and it is cheap: `Accept-Encoding` has low effective cardinality once normalized, so you usually pay at most two or three variants. Most reverse proxies and CDNs add it automatically when they compress, but an origin doing its own gzip must emit it itself.
code
bash · 2 linescurl -sSD - -o /dev/null -H 'Accept-Encoding: gzip' https://example.com/app.js
curl -sSD - -o /dev/null -H 'Accept-Encoding: identity' https://example.com/app.jsgo deeper
Say that gzip and non-gzip bodies are different responses for the same URL, so the cache must be told to distinguish them.
Name both failure modes, explain why browsers hide the bug, and note that the header belongs on the uncompressed variant too.
Add normalization for hit rate, ETag-per-variant correctness, range-request interaction, and where the header should be emitted (origin vs edge that does the compressing).
Decide where compression happens in the delivery chain and make the cache-key policy an explicit, tested part of the CDN configuration rather than an accident of origin defaults.
## Why compression is negotiation HTTP compression is opt-in by the client. A client advertises what it can decode with the request header `Accept-Encoding: gzip, br` and the server picks one, applies it, and marks the result with `Content-Encoding: gzip`. A client that sends no `Accept-Encoding`, or an explicit `Accept-Encoding: identity`, must get the raw bytes. So a single URL now has at least two legitimate representations that differ byte for byte. Anything that stores responses by URL alone — CDN edge, reverse proxy, corporate forward proxy — will conflate them. ## The two failure modes **Gzip stored first, non-gzip client second.** The cache returns the compressed body together with `Content-Encoding: gzip`. A browser handles that fine, because browsers always advertise gzip — which is exactly why this bug survives manual testing. The victims are the clients that do not: an old library, a curl invocation without `--compressed`, an embedded device, a corporate proxy that strips `Accept-Encoding` on the way out but forwards the response, a webhook consumer. Symptoms are binary garbage in logs, JSON parse errors on the first byte, or a downloaded file that will not open. **Identity stored first, gzip client second.** Everyone gets correct but uncompressed bytes. No errors — just an unexplained jump in egress and page weight. This one is usually found in a bandwidth graph, not a bug report. Both are non-deterministic: which failure you get depends on which client warmed that particular edge node. A CDN with hundreds of points of presence will show it in some regions and not others, which is a classic 'works for me' investigation. ## The fix Send `Vary: Accept-Encoding` on every response whose body was (or could have been) compressed. The cache then stores one variant per distinct `Accept-Encoding` value and matches on it. This is the canonical, near-universal use of `Vary` and it is the one case where the variant cost is negligible. It matters that the header is present even on the *uncompressed* variant. If only the gzip response carries `Vary`, the identity response is still stored under the naked URL and can be handed to a gzip-capable client. ## Normalization keeps the cost small Raw `Accept-Encoding` values are wildly diverse: `gzip, deflate`, `gzip, deflate, br`, `br;q=1.0, gzip;q=0.8`, `gzip;q=1.0, identity; q=0.5, *;q=0`. Keyed literally, each distinct string is a separate cache object, and you can end up with dozens of copies of the same resource. Caches and CDNs therefore normalize before keying — typically collapsing the header to one of a few tokens (`br`, `gzip`, or empty) based on what the origin can actually produce. Varnish deployments traditionally do this in `vcl_recv`; commercial CDNs do it by default. If you build your own caching layer, normalization is the difference between two variants and two hundred. ## Related traps - **Do not vary on `Accept-Encoding` for content you never compress** (already-compressed images, video). You gain nothing and split the cache. - **`Content-Encoding` is not `Transfer-Encoding`.** Compression applied as a transfer coding is hop-by-hop and does not need `Vary`; `Content-Encoding` is part of the representation and does. - **Range requests interact badly.** A byte range of a gzip variant is meaningless to a client expecting ranges of the identity variant; caches must keep ranges within a variant. - **ETags must differ between variants.** If the gzip and identity representations share one strong ETag, a conditional request can revalidate the wrong variant into place. Serve distinct ETags (many servers suffix the compressed one) or use a weak ETag. - **Compression plus secrets on one connection** is the ingredient of the BREACH class of attacks; that is a security consideration for what you compress, not a reason to skip `Vary`. ## Rule of thumb If `Content-Encoding` can ever appear on a response, `Vary: Accept-Encoding` belongs on that response — compressed or not.
- Which of the two failure modes is worse, and why is it the one that survives testing?Serving a gzip body to a client that never asked for it is worse, because it corrupts the response rather than just wasting bandwidth. It survives testing because every browser advertises gzip, so a developer clicking through the site never sees it; only non-browser clients — curl without --compressed, old SDKs, webhook receivers — hit it.
- How would you keep the variant count low if clients send dozens of distinct Accept-Encoding strings?Normalize the header at the cache before it becomes part of the key: inspect it, and rewrite it to one of the small set of encodings the origin can actually emit — for example br, gzip, or empty — discarding q-values and ordering. That turns an unbounded string space into two or three stored variants per URL.
saying these in an interview costs you the question
- Claiming Vary is unnecessary because 'all clients support gzip'
- Adding Vary only to the compressed response and not to the identity one
- Keying on the raw Accept-Encoding string and then complaining about hit rate
- Confusing Content-Encoding with Transfer-Encoding when reasoning about caching
- Giving both variants the same strong ETag