skip to content

How would you set a compression policy for a service sitting behind a CDN — what gets compressed, at which layer, and at what cost?

level: principalimportance: nice to knowfreq 26%

answer

  1. allow-list + ~1 KB floor
  2. one layer per response class
  3. level, not codec, drives latency
  4. normalise Accept-Encoding at the edge
  5. BREACH: secret + reflected input

basics

~20 s

Compress text-like types above a size threshold, never already-compressed media; precompress static assets at maximum level and use a fast coding for dynamic ones; do it in exactly one layer; and exclude responses that mix secrets with attacker-influenced input.

solid answer

~60 s

I would decide four things explicitly. **What.** An allow-list of text-like media types (HTML, CSS, JS, JSON, XML, SVG, plain text) and a minimum size around 1 KB. Never JPEG/PNG/WebP/MP4/WOFF2 — recompression costs CPU and adds bytes. **Where.** Exactly one layer. Precompressed static artifacts (Brotli 11 + gzip) served straight through the CDN with the CDN's own compression off; dynamic responses compressed at the edge or reverse proxy, not in application code. Two layers compressing means double encoding or your build-time work being silently re-encoded at lower quality. **How hard.** Level is the real latency knob: maximum for build-time artifacts, low levels (zstd fast, or Brotli 4–5) for per-request work. Watch time-to-first-byte on streamed responses, where a high-level compressor buffers. **What to exclude.** Responses that mix a secret with attacker-controlled input leak that secret through compressed size (the BREACH class). Mitigate by not reflecting input into secret-bearing responses, or by masking tokens per response, rather than by disabling compression everywhere. Then measure: egress saved, CPU used, p99 TTFB, and CDN hit ratio after the `Vary: Accept-Encoding` cache split.

code

http · 9 lines
http
GET /search?q=<attacker-controlled> HTTP/1.1
Accept-Encoding: gzip, br

HTTP/1.1 200 OK
Content-Type: text/html
Cache-Control: no-store
Vary: Accept-Encoding

<!-- reflects q AND embeds a CSRF token: compress off or mask the token -->

go deeper

for a junior

Know the basics that feed the policy: compress text, not images; compression costs CPU; the client must advertise support.

for a middle

Give the allow-list, the size threshold, and the static-versus-dynamic level split, and name Vary: Accept-Encoding as a caching consequence.

for a senior

Own the layering decision, the double-compression failure mode, edge normalisation of Accept-Encoding, and the concrete BREACH mitigations.

for a principal

Present it as an explicit cost model — egress versus CPU versus p99 latency versus cache cardinality — with a single owner per response class, named security exclusions, and the metrics that would make you reverse the decision.

## Framing the decision Compression trades **CPU and latency at the server** for **bytes and latency on the network**. On a fast fibre link, aggressive compression of a small response can be net-negative; on a congested mobile link it is a large win. A policy therefore has to be explicit about what, where, how hard, and what to exclude — rather than a global "gzip on". ## What to compress Run an **allow-list by media type**, not a deny-list: - Yes: `text/html`, `text/css`, `text/javascript`, `application/json`, `application/*+json`, `application/xml`, `image/svg+xml`, `text/plain`, `text/csv`. - No: JPEG, PNG, WebP, AVIF, GIF, MP4, ZIP, WOFF2 (Brotli-compressed internally), and anything already served as `application/gzip`. Add a **minimum size threshold** of roughly 1 KB. Below it, framing overhead can exceed the saving, and the CPU is spent for nothing. Note the `+json` suffix in the allow-list — vendor media types are the usual thing an allow-list accidentally omits, silently leaving your largest API payloads uncompressed. ## Which layer compresses The choices are application, reverse proxy/web server, CDN edge, or build step. The rule is: **exactly one layer per response**. - **Static assets** → build step, maximum level, served through the CDN as-is. You pay Brotli-11 once per deploy for millions of hits. Turn the CDN's own compression off for these paths so it does not re-encode at a lower quality or double-encode. - **Dynamic responses** → edge or reverse proxy, low level. The edge is closest to the user, has the CPU headroom, and keeps the concern out of application code. Compressing in the app and again at the proxy is the classic double-encoding incident. - **Internal service-to-service** → the client library, typically zstd at a fast level, where both ends are yours and you can even train a dictionary on your payload shapes. ## How hard to compress Compression **level** dominates the latency question far more than compressor choice. A useful decomposition of what compression costs a dynamic response: - Encode CPU, which lands directly in time-to-first-byte if the compressor buffers. - Buffering behaviour on streamed responses: a high-level compressor holds bytes back to find matches, which can destroy the point of streaming. Flush per chunk, or drop the level. - Tail latency under load: compression CPU competes with request handling, so p99 degrades before the mean does. So: maximum level where it is amortised, low level where it is per-request, and always look at p99 TTFB rather than average response size. ## Caching consequences Compressing means `Vary: Accept-Encoding`, which splits every cache entry across codings. Real-world `Accept-Encoding` values vary enough that a naive cache key fragments badly, so **normalise the header at the edge** into the small set you actually serve (`br`, `gzip`, none) before it enters the key. Track the cache hit ratio before and after enabling a new coding — adding zstd to a fleet can quietly triple the number of stored variants. ## Security exclusions Compression before encryption leaks information through **ciphertext length**. CRIME attacked compressed TLS/SPDY headers and was solved at the protocol level. **BREACH** is the still-live variant: if a response body contains both a **secret** (a CSRF token, part of a session identifier, a PII field) and **attacker-influenced reflected input**, the attacker can guess the secret byte by byte by observing compressed response sizes, because a correct guess compresses better. The right mitigations are surgical, not global: - Do not reflect user-controlled input into responses that also carry secrets. - Mask/randomise CSRF tokens per response so the secret's compressibility changes each time. - Add length randomisation, or disable compression for the specific endpoints that meet both conditions. - Keep robust CSRF protection so a leaked-then-guessed token is not the whole game. Globally disabling compression to "be safe" costs every user real bandwidth to defend a narrow condition — that tradeoff is exactly what a principal-level answer should articulate. ## Measuring the policy Define success before turning knobs: - **Egress bytes** by route (the money metric). - **CPU seconds per request** at the compressing layer. - **p50/p99 TTFB**, especially on streamed endpoints. - **CDN hit ratio** after each Vary split. - **Real-user page-load metrics** — the only number that says whether the smaller bytes actually helped. The answer that impresses is not "enable Brotli"; it is "here is the allow-list, here is the single owning layer per response class, here is the level per class, here are the endpoints excluded for BREACH, and here are the four metrics I watch."

  • Why not simply disable compression everywhere to eliminate BREACH?
    Because the vulnerable condition is narrow — a response must contain both a secret and attacker-influenced reflected input — while the cost of disabling compression is paid by every user on every text response. The proportionate answer is to mask tokens per response, avoid reflecting input alongside secrets, and disable compression only on the specific endpoints that meet both conditions.
  • How does enabling a third content coding affect your CDN?
    Every additional coding multiplies stored variants because responses vary on Accept-Encoding, which can lower hit ratio and raise origin load. Mitigate by normalising Accept-Encoding at the edge into the small set you actually serve, and by watching hit ratio and origin egress across the rollout rather than only response size.
  • Compression is on but p99 time-to-first-byte got worse. What do you look at?
    Level and buffering first. A high-level compressor holds bytes to find matches, which delays the first byte on streamed responses, and its CPU competes with request handling so the tail degrades before the mean. Drop the level on dynamic paths, flush per chunk for streams, and confirm you are not compressing tiny or already-compressed bodies.

saying these in an interview costs you the question

  • Treating 'enable Brotli everywhere at max level' as the answer.
  • Compressing in both the application and the proxy or CDN.
  • Ignoring the cache-key fragmentation that Vary: Accept-Encoding creates.
  • Dismissing BREACH entirely, or over-correcting by disabling compression globally.
  • Judging success by average response size alone, with no CPU or p99 latency metric.

context