An image CDN caches the derivatives it generates. What forms the cache key for a derivative, and how does a front-end team accidentally destroy that cache's hit rate?
answer
- hits require identical keys
- path plus params, plus varying headers
- computed widths mint endless variants
- a miss re-encodes the master
basics
~20 sThe cache key is the master's path plus the transformation parameters, and sometimes request headers the response varies on. Any change to that key — computed widths, extra tracking parameters, reordered or re-cased parameters — mints a new derivative, forcing an origin fetch and a re-encode.
solid answer
~60 sA derivative is cached under a key derived from the request: the path identifying the master, plus the transformation parameters, plus any header the response varies on. Two requests share a cached derivative only if they produce the same key, so hit rate is entirely a function of how few distinct keys your pages emit. Teams wreck it in predictable ways: computing a width from `element.clientWidth` at runtime, which produces hundreds of near-identical widths instead of a handful; appending analytics or cache-busting query strings to image URLs; letting parameter order or casing differ between templates; or embedding per-user or per-timestamp tokens in the URL. Each of those turns a would-be hit into a miss, and a miss costs an origin round trip plus a decode and re-encode of the master — often hundreds of milliseconds added to the byte that matters most, the hero image. The fixes are to normalize aggressively: snap requested widths to a small fixed set, generate every image URL through one helper so ordering and casing are identical, strip unrecognised parameters at the edge, and monitor miss rate rather than average latency.
go deeper
Know that the CDN stores each transformed image under its URL, so a different URL means a different stored copy and a fresh transformation.
Explain what goes into the key and walk through the miss path — origin fetch, decode, resize, encode — and name at least two habits that fragment the key.
Demonstrate diagnosis: read the cache-status header, separate miss latency from transfer time, and trace a hit-rate collapse back to the change that started minting new URLs.
Frame it as a constraint to enforce, not a habit to hope for — one URL-building helper, a fixed width ladder, parameter stripping at the edge, and a hit-ratio alarm that catches regressions.
## What is actually being cached An image CDN stores **derivatives** — the resized, re-encoded outputs it has already produced — so it does not have to redo that work. Every derivative sits under a **cache key**. Understanding what goes into that key is the whole game, because two requests reuse the same stored bytes if and only if they compute the same key. A typical key is built from: - **The path**, which identifies the master image. - **The transformation parameters**: width, height, fit/crop mode, quality, format, and anything else the vendor supports. - **Selected request headers**, when the response genuinely depends on them — most commonly the client's accepted formats, and sometimes a device-pixel-ratio hint. Everything in that key multiplies. Five widths times three formats is fifteen derivatives per master, each with its own miss to pay and its own storage. ## The cost of a miss On a hit, the edge streams bytes it already has. On a miss it must fetch the master from origin, decode a possibly very large image, resize it, encode the output — AVIF encoding in particular is CPU-hungry — and only then start responding. That is easily hundreds of milliseconds, and it lands on whichever unlucky visitor asked first. This matters disproportionately for the largest image on the page, because that image is usually what the browser paints last and what your loading metrics are measured against. "The hero is fast in staging" often just means staging's cache was warm. ## Four ways teams destroy hit rate **Runtime-computed dimensions.** Code that measures the container and requests exactly that many pixels feels precise and is a disaster for caching: ```js // every viewport width mints a distinct derivative const w = Math.round(el.clientWidth * devicePixelRatio); img.src = `/photos/hero.jpg?width=${w}`; ``` Across real devices this produces hundreds of near-identical widths — 731, 733, 736 — that look the same to a human and are entirely different to a cache. Snapping to a small fixed ladder collapses them back to a handful of keys with no visible difference. **Foreign query parameters.** Analytics tags, campaign parameters, or a deploy-stamped `?v=` appended to image URLs change the key even though they change nothing about the image. Some CDNs can be configured to ignore unknown parameters when keying; do not assume yours does. **Inconsistent formatting.** `?w=640&q=75` and `?q=75&w=640` may key differently, as may `AUTO` versus `auto`. Multiple templates written by multiple people drift apart quickly. **Per-request tokens.** Signed URLs whose signature includes a fresh timestamp, or a per-user identifier, make every request unique by construction. Signing is worth having — but the signature must be stable for a given transformation, or excluded from the key, or the cache is decorative. ## Making the key small on purpose The discipline is to treat the set of URLs your site can emit as a deliberately bounded set: - Generate every image URL through **one helper function** so parameter names, order, and casing are identical everywhere and no template can invent its own dialect. - **Quantize** any computed dimension to a fixed ladder before it reaches the URL. - **Strip** parameters the CDN does not understand, at the edge, so a stray tracking tag cannot fragment the cache. - Keep the master path **stable and versioned** — a new image gets a new path, rather than new bytes at an old path. ## Watching it in production The number to alarm on is **cache hit ratio for image requests**, split by whether the request was a miss. A healthy pipeline sits high and boring; a sudden fall almost always means new URLs are being minted — a shipped change that started computing widths, a new tracking parameter, or someone probing the endpoint. Pair that with **miss latency**, because hit ratio alone hides the impact. A 95% hit ratio sounds fine until you notice the 5% is concentrated on first views of new campaign pages, where the miss lands on the hero. One more subtlety: the edge cache is not the only cache. Repeat visitors may be served from the browser cache regardless of edge state, which is why field data can look better than a synthetic cold run — and why a cold-cache test is the honest one when you are judging the pipeline itself.
- How would you confirm that a slow hero image is caused by a cold derivative rather than a large file?Compare a first request with an immediate repeat. A cold derivative shows a long wait before the first byte and a cache-status header indicating a miss, then drops sharply on the repeat with identical byte count. A genuinely oversized file shows a short wait and a long transfer that does not improve on the second request. Vendors expose the status via a response header — check which one yours sends.
- Is it better to reduce the number of widths you request, or to keep them and accept the misses?Fewer widths, in almost every case. Each extra width buys a marginal reduction in bytes for some viewport but costs a whole additional derivative to warm and store. A small ladder concentrates traffic on a few keys that stay warm, which usually beats a perfectly-sized image that is cold when the visitor arrives.
- What can you do to avoid the very first visitor paying the miss on a launch page?Warm the cache before traffic arrives: after deploy, request the exact derivative URLs the critical pages will use, so the edge produces and stores them ahead of time. This works only if you can enumerate those URLs, which is another argument for generating them from one helper and from a bounded ladder of widths.
saying these in an interview costs you the question
- Thinks unknown query parameters are always ignored when keying
- Computes exact pixel widths at runtime for cache-friendliness
- Assumes a warm staging test proves production performance
- Believes more variants always means better performance