skip to content

A CDN dashboard reports a 55% cache hit ratio for a site's static assets. What does that number actually mean, and what are the usual reasons an edge cache misses on files that never change?

level: middleimportance: must knowfreq 52%

answer

  1. hits divided by total requests
  2. requests versus bytes tell different stories
  3. one object per exact URL
  4. each PoP caches on its own
  5. lifetime shorter than the request gap

basics

~20 s

Cache hit ratio is the share of requests an edge answered from its own cache instead of fetching from the origin. On never-changing files a low ratio usually means cache-key fragmentation from query strings, a lifetime shorter than the gap between requests, traffic spread thinly across independent PoPs, or responses the CDN was told not to store.

solid answer

~50 s

Hit ratio is simply hits divided by total requests over some window — the share of requests the edge answered itself. It is worth separating request hit ratio from byte hit ratio: a site can hit on most small requests and still pull most of its *bytes* from origin because the large media misses. For assets that genuinely never change, the usual causes of misses are: the cache key includes the full URL, so appended analytics parameters turn one file into many objects; the cached lifetime is shorter than the interval between requests for that object; traffic is spread across many PoPs that each cache independently, so a long-tail file is cold nearly everywhere; and responses the CDN refuses to store, typically because the origin sends `no-store` or attaches a `Set-Cookie` to asset responses. Deploys also guarantee a cold miss per new file per PoP.

code

bash · 2 lines
bash
curl -sI 'https://example.com/app.8f3a1c.js'
curl -sI 'https://example.com/app.8f3a1c.js?utm_source=news'

go deeper

for a junior

Know that a hit means the edge answered from its own copy and a miss means it had to ask your server, and that adding query parameters to a URL creates a different cached object.

for a middle

Explain the mechanics behind a low ratio: URL-based cache keys, lifetimes shorter than the interval between requests, per-PoP caches with their own eviction, and response headers that forbid storage.

for a senior

Show a diagnosis path — break the ratio down by path and by PoP, distinguish request from byte hit ratio, and pick the fix that matches the cause rather than raising lifetimes everywhere.

for a principal

Own the tradeoff: decide how much engineering the long tail is worth, where a shield tier or key normalisation pays for itself, and set the rule that stops correctness being traded away for a nicer dashboard number.

## What the number is Cache hit ratio is `hits / (hits + misses)` over a time window, at whatever scope the dashboard is aggregating — usually the whole property, sometimes per PoP. A hit means the edge server answered from its own stored copy without contacting your origin. A miss means it had to fetch. Two variants matter and they diverge: - **Request hit ratio** — the share of *requests* served from cache. - **Byte hit ratio** — the share of *bytes* served from cache. A site can serve 95% of requests from the edge and still push half its bytes from origin, because the misses are the big video and image objects. If you are optimising origin egress cost or origin bandwidth, byte hit ratio is the honest number; if you are optimising user-perceived latency, request hit ratio is closer to what users feel. A hit ratio also has a floor you cannot beat. Every distinct object must be fetched at least once per PoP, so a site with a huge catalogue of rarely-requested objects and a global footprint will never approach 100%, no matter how well it is configured. ## Why a file that never changes still misses **1. Cache-key fragmentation from the URL.** The cache key is built from the request URL, so `/app.8f3a1c.js` and `/app.8f3a1c.js?utm_source=newsletter` are two separate objects holding identical bytes. Campaign parameters, session identifiers spliced into links, cache-busting parameters added by well-meaning code, and third-party tracking suffixes all multiply one object into hundreds — each of which has to be fetched from origin the first time it is seen at each PoP. The fix is to normalise the key: strip or allow-list query parameters at the edge for asset paths, and sort the ones you keep, so equivalent requests collapse onto one object. **2. A lifetime shorter than the request interval.** If an object is cached for five minutes but a given PoP sees a request for it every eight minutes, the copy is expired on every arrival and every request is a miss. For genuinely immutable, content-hashed assets the answer is a long lifetime — the URL changes when the content does, so there is no reason to expire it early. **3. Independent caches at each PoP.** Edges do not share one global cache. Splitting worldwide traffic over dozens of PoPs means each one sees a fraction of the volume, and a moderately popular object may be requested too rarely at any single PoP to stay resident. Most CDNs offer a shield or upper-tier cache — a designated intermediate cache all PoPs miss into — which converts many origin fetches into one and lifts the effective hit ratio for exactly this long tail. **4. Eviction.** Edge storage is finite and managed roughly least-recently-used. A large catalogue competing with popular objects means older entries are evicted before their lifetime expires, so "cached for a year" is an upper bound, never a promise. **5. Responses the CDN will not store.** `Cache-Control: no-store` or `private` on an asset response, or a `Set-Cookie` attached to static files because the whole site shares one hostname and a framework sets a session cookie on every response, will stop most CDNs from caching. This is the one that most often surprises people, because the file looks perfectly static. **6. Deploys.** Every new content-hashed filename is, by definition, cold everywhere. A dip in hit ratio right after each release is expected and is not a defect. ## Diagnosing it rather than guessing Sample real URLs from logs and look at how many distinct keys map to what should be one object. Compare the same asset requested with and without a query string: ```bash curl -sI 'https://example.com/app.8f3a1c.js' curl -sI 'https://example.com/app.8f3a1c.js?utm_source=news' ``` Most CDNs report cache status in a response header (the name is a per-provider convention), and the standard `Age` header on a response tells you the copy was served from a cache rather than freshly fetched. Then break the ratio down by path prefix and by PoP: a low ratio concentrated on one path is usually a headers or key problem, while a low ratio spread evenly across small PoPs is usually traffic dilution, and the fixes are different. ## The judgment part A higher hit ratio is not automatically the goal. Raising it by caching things that should not be cached is a correctness bug, and the last few percent of a long tail may cost more engineering than the origin traffic it saves. What a strong answer shows is that you know what the number measures, which of the causes above your data points at, and which fix follows from that cause.

  • Your hit ratio drops sharply for an hour after every deploy and then recovers. Is that a problem?
    Normally no. New content-hashed filenames are cold at every PoP, so the first request for each one in each region must reach the origin. It is worth confirming the origin can absorb that burst, and a shield tier reduces it to roughly one fetch per object rather than one per PoP, but the shape of the graph itself is expected.
  • How would you decide whether to strip an unknown query parameter from the cache key or keep it?
    Ask whether the parameter changes the bytes the origin would return. If it does not — campaign tags, tracking suffixes — it is pure fragmentation and should be dropped from the key for those paths. If it does, keeping it is required for correctness. Allow-listing the parameters that matter is safer than blocklisting the ones you have seen.
  • Why can request hit ratio look excellent while origin bandwidth stays high?
    Because bytes are distributed very unevenly. Thousands of tiny hits on scripts and icons dominate the request count, while a handful of large media objects miss and account for most of the transferred volume. Byte hit ratio exposes this; request hit ratio hides it.

saying these in an interview costs you the question

  • Treating 100% hit ratio as the target
  • Assuming all PoPs share one global cache
  • Ignoring query strings as part of the cache key
  • Believing a static file is always cacheable regardless of headers
  • Reading one aggregate ratio without breaking it down

context