skip to content

Asset Caching Strategy

Long-lived caching only works if you can invalidate it, which is why hashed filenames exist. Expect to design the caching policy for HTML versus fingerprinted static assets.

on this pageshow

questions

4

A production build emits `app.9f2c1a4b.js` instead of `app.js`. Why is that hash computed from the file's contents rather than from the release number or build id, and what caching policy does it unlock?

level: juniorimportance: must knowfreq 62%

answer

  1. the URL is the cache key
  2. never change bytes behind a URL
  3. unchanged file keeps its name
  4. version bump renames everything at once
  5. HTML is the one uncached pointer

basics

~20 s

A content hash changes only when the file's bytes change, so unchanged files keep their URL across releases and stay cached, while changed files get a new URL nobody has cached. That is what makes a year-long cache lifetime safe.

solid answer

~50 s

The URL is the cache key, so the only way to cache an asset for a year and still ship changes is to change its URL. A content hash does that automatically and minimally: it is derived from the file's bytes, so a file that did not change keeps exactly the same name across releases and returning visitors keep using their cached copy, while a file that did change lands on a URL nobody has ever requested. A release number in the filename would be correct but wasteful — bumping the version renames every asset at once, so a one-line change forces a re-download of the whole build. Hashed files then get a very long `max-age`, and the one un-hashed thing, the HTML document, stays short-lived, because that document is what points at the new names.

go deeper

for a junior

Be ready to say plainly that the URL is what a cache stores things under, so shipping a change means shipping a new filename. Recognise app.9f2c1a4b.js on sight and explain that the hash comes from the file's contents.

for a middle

Explain the two-tier policy that fingerprinting enables — very long lifetimes for hashed assets, a short one for the HTML document — and why a build-number scheme invalidates far more than it needs to.

for a senior

Show that you have operated this: hashes compose bottom-up through the asset graph, previous builds must stay fetchable after a deploy, and the document's lifetime sets how long clients keep asking for old names.

for a principal

Frame it as an invalidation-surface decision: fingerprinting shrinks the set of mutable URLs to almost nothing, so purge tooling, rollback, and edge policy only have to be correct for the handful of addresses that stay stable.

## The cache key is the URL Every HTTP cache — the browser's disk cache, a proxy, a CDN edge — stores a response under the URL that produced it. Nothing in that URL says which release the file came from or whether it is still correct. When a cache holds a copy it considers fresh, it serves it without contacting the server at all. That property is the whole prize. On a repeat visit, a fresh cached bundle costs zero requests, zero bytes, and zero latency; it is the cheapest performance win a frontend has. It is also what makes long caching frightening: once a browser holds a year-long copy of `/app.js`, there is no reliable way to reach out and tell it that copy is now wrong. Caches are one-way — you can hand a client something, you cannot take it back. So the discipline reduces to one rule: **never change the bytes behind a URL; change the URL.** ## Fingerprinting: hash the contents, not the release A build tool implements that rule by naming each output file after a hash of its own contents: ``` # release 12 assets/app.9f2c1a4b.js assets/vendor.4d81ee02.js assets/styles.aa71b3c0.css # release 13 — only application source changed assets/app.5e0b77d1.js <- new URL, nobody has this cached assets/vendor.4d81ee02.js <- identical bytes, identical URL, still cached assets/styles.aa71b3c0.css <- unchanged too ``` Two things follow. First, **correctness**: a changed file is unreachable at its old address, so no cache anywhere can serve the old version of the new file. There is no invalidation step to forget, no purge to get wrong, no user stuck on a stale copy. Second, **economy**: only what actually changed is re-fetched. The dependency bundle, the icon sprite, the stylesheet you did not touch all keep their addresses, and the returning user downloads only the delta. ## Why a release number in the name is worse `app.v13.js` also solves correctness — every release is a fresh URL. But it renames *everything* every time, so a typo fix in one component costs every returning visitor a full re-download of the whole asset set. The same objection applies to a global query-string bump like `app.js?v=13`: browsers do include the query string in the cache key, so it busts the cache, but it busts it for every asset at once, and it makes the URL's stability depend on a number that has nothing to do with the file. Content hashing gets you the invalidation you need and nothing more. ## The policy that follows Once filenames are content-addressed, the caching policy is almost mechanical, and it has two tiers: - **Hashed static assets** get the longest lifetime you are willing to write — conventionally a year of `max-age`, commonly paired with `immutable`. The URL is a promise that these bytes will never change. - **The HTML document** gets a short lifetime or is revalidated on every navigation. It is the one address that must stay stable (users type it, link to it, bookmark it), so it cannot be fingerprinted, and it is the thing that carries the new asset names. The document is the invalidation point of the entire system. That asymmetry is the design. Everything expensive is immutable and cached forever; the one small, cheap file is uncached and re-fetched, and it redirects the browser to whatever the current build is. ## Hashes compose, deepest first Hashing is applied bottom-up because references travel upward. Swap a font file and its URL changes; the stylesheet that references that URL therefore has different bytes, so the stylesheet's hash changes; the document that references the stylesheet now names a different file. A single leaf change ripples up exactly one chain and leaves every unrelated file untouched. This is also why a build's asset graph has to be hashed in dependency order — you cannot know a parent's hash until every child's final URL is fixed. ## What hashing does not solve - **First visits.** A cold visitor downloads everything regardless; fingerprinting is a repeat-visit optimization. - **HTML freshness.** Hashed assets are safe, but if the document is over-cached, clients keep pointing at an old build long after deploy. - **Old files disappearing.** Because a client can hold an old document, the previous build's files must remain fetchable for a while after a deploy — deleting them is what turns "safe caching" into a post-deploy 404. - **Chunks that churn for no reason.** A hash only buys you stability if the bytes are actually stable between builds; embedded build metadata or renumbered internal identifiers can change a file that is semantically unchanged, silently costing every returning user a re-download.

  • If every asset is immutable, what still has to change on each deploy for a returning user to see the new build?
    The HTML document — or whatever small manifest lists the entry URLs. It is the one address that cannot be fingerprinted because users navigate to it directly, so it carries a short lifetime or is revalidated on every navigation. Everything else follows from what that document names, which is why over-caching the document quietly freezes users on an old release even though the assets are perfectly cacheable.
  • A CSS file references a hashed font URL. What happens to the CSS file's hash when the font changes?
    It changes too. The new font gets a new URL, that URL is written into the stylesheet, so the stylesheet's bytes differ and its hash differs, and the document referencing the stylesheet names a new file in turn. Hashing therefore runs bottom-up through the asset graph: a parent's hash cannot be computed until every child's final URL is fixed.
  • Does a content hash mean you never need to purge the CDN?
    For hashed assets, yes — a changed file arrives at a URL no cache has ever seen, so there is nothing to purge. Purging is only relevant for the addresses that stay stable across releases: the HTML document, and any un-fingerprinted file such as a manifest or a well-known path. That is a much smaller and much safer surface to operate.

A content hash names a file by its fingerprint rather than its edition number: two editions with identical text get the same name, so nobody re-fetches a chapter they already have.

saying these in an interview costs you the question

  • Thinks the hash is a version number that increments each release
  • Says a full CDN purge is required after every deploy
  • Wants to cache the HTML document for a year too
  • Claims browsers re-download every asset after any deploy
  • Believes appending ?v=13 to every asset is equivalent

context

open as a page

A build splits dependencies into a separate `vendor` chunk with a content-hashed filename, but that hash changes on nearly every deploy even when no dependency was upgraded. Why does that undermine the caching strategy, and what usually causes it?

level: middleimportance: should knowfreq 42%

basics

~20 s

A hash that changes without a real content change destroys the point of splitting out vendor code: returning users re-download it every release. Usual causes are renumbered internal module identifiers, embedded build metadata, or the chunk-URL map living inside the chunk.

open as a page

Minutes after a deploy, some users see a blank screen or a failed route, and their network log shows a 404 for `/assets/route-detail.3c19aa7e.js` — a filename that does not exist in the new build. What caching and deploy decisions produced that, and how would you stop it recurring?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An old HTML document, cached or sitting in a tab opened before the deploy, still names chunk URLs the deploy deleted. Fix it by publishing assets before the document, retaining previous builds, and keeping the HTML short-lived.

open as a page

Because content-hashed assets are never overwritten, every release leaves the previous build's files behind. How would you decide how long to keep old builds' assets available, and what does that number depend on?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Keep old assets at least as long as any document naming them can still be served: HTML cache lifetime plus the open-tab tail, plus your rollback window. Measure that tail with field data rather than guessing.

open as a page