A production build emits `app.9f2c1a4b.js` instead of `app.js`. Why is that hash computed from the file's contents rather than from the release number or build id, and what caching policy does it unlock?
answer
- the URL is the cache key
- never change bytes behind a URL
- unchanged file keeps its name
- version bump renames everything at once
- HTML is the one uncached pointer
basics
~20 sA content hash changes only when the file's bytes change, so unchanged files keep their URL across releases and stay cached, while changed files get a new URL nobody has cached. That is what makes a year-long cache lifetime safe.
solid answer
~50 sThe URL is the cache key, so the only way to cache an asset for a year and still ship changes is to change its URL. A content hash does that automatically and minimally: it is derived from the file's bytes, so a file that did not change keeps exactly the same name across releases and returning visitors keep using their cached copy, while a file that did change lands on a URL nobody has ever requested. A release number in the filename would be correct but wasteful — bumping the version renames every asset at once, so a one-line change forces a re-download of the whole build. Hashed files then get a very long `max-age`, and the one un-hashed thing, the HTML document, stays short-lived, because that document is what points at the new names.
go deeper
Be ready to say plainly that the URL is what a cache stores things under, so shipping a change means shipping a new filename. Recognise app.9f2c1a4b.js on sight and explain that the hash comes from the file's contents.
Explain the two-tier policy that fingerprinting enables — very long lifetimes for hashed assets, a short one for the HTML document — and why a build-number scheme invalidates far more than it needs to.
Show that you have operated this: hashes compose bottom-up through the asset graph, previous builds must stay fetchable after a deploy, and the document's lifetime sets how long clients keep asking for old names.
Frame it as an invalidation-surface decision: fingerprinting shrinks the set of mutable URLs to almost nothing, so purge tooling, rollback, and edge policy only have to be correct for the handful of addresses that stay stable.
## The cache key is the URL Every HTTP cache — the browser's disk cache, a proxy, a CDN edge — stores a response under the URL that produced it. Nothing in that URL says which release the file came from or whether it is still correct. When a cache holds a copy it considers fresh, it serves it without contacting the server at all. That property is the whole prize. On a repeat visit, a fresh cached bundle costs zero requests, zero bytes, and zero latency; it is the cheapest performance win a frontend has. It is also what makes long caching frightening: once a browser holds a year-long copy of `/app.js`, there is no reliable way to reach out and tell it that copy is now wrong. Caches are one-way — you can hand a client something, you cannot take it back. So the discipline reduces to one rule: **never change the bytes behind a URL; change the URL.** ## Fingerprinting: hash the contents, not the release A build tool implements that rule by naming each output file after a hash of its own contents: ``` # release 12 assets/app.9f2c1a4b.js assets/vendor.4d81ee02.js assets/styles.aa71b3c0.css # release 13 — only application source changed assets/app.5e0b77d1.js <- new URL, nobody has this cached assets/vendor.4d81ee02.js <- identical bytes, identical URL, still cached assets/styles.aa71b3c0.css <- unchanged too ``` Two things follow. First, **correctness**: a changed file is unreachable at its old address, so no cache anywhere can serve the old version of the new file. There is no invalidation step to forget, no purge to get wrong, no user stuck on a stale copy. Second, **economy**: only what actually changed is re-fetched. The dependency bundle, the icon sprite, the stylesheet you did not touch all keep their addresses, and the returning user downloads only the delta. ## Why a release number in the name is worse `app.v13.js` also solves correctness — every release is a fresh URL. But it renames *everything* every time, so a typo fix in one component costs every returning visitor a full re-download of the whole asset set. The same objection applies to a global query-string bump like `app.js?v=13`: browsers do include the query string in the cache key, so it busts the cache, but it busts it for every asset at once, and it makes the URL's stability depend on a number that has nothing to do with the file. Content hashing gets you the invalidation you need and nothing more. ## The policy that follows Once filenames are content-addressed, the caching policy is almost mechanical, and it has two tiers: - **Hashed static assets** get the longest lifetime you are willing to write — conventionally a year of `max-age`, commonly paired with `immutable`. The URL is a promise that these bytes will never change. - **The HTML document** gets a short lifetime or is revalidated on every navigation. It is the one address that must stay stable (users type it, link to it, bookmark it), so it cannot be fingerprinted, and it is the thing that carries the new asset names. The document is the invalidation point of the entire system. That asymmetry is the design. Everything expensive is immutable and cached forever; the one small, cheap file is uncached and re-fetched, and it redirects the browser to whatever the current build is. ## Hashes compose, deepest first Hashing is applied bottom-up because references travel upward. Swap a font file and its URL changes; the stylesheet that references that URL therefore has different bytes, so the stylesheet's hash changes; the document that references the stylesheet now names a different file. A single leaf change ripples up exactly one chain and leaves every unrelated file untouched. This is also why a build's asset graph has to be hashed in dependency order — you cannot know a parent's hash until every child's final URL is fixed. ## What hashing does not solve - **First visits.** A cold visitor downloads everything regardless; fingerprinting is a repeat-visit optimization. - **HTML freshness.** Hashed assets are safe, but if the document is over-cached, clients keep pointing at an old build long after deploy. - **Old files disappearing.** Because a client can hold an old document, the previous build's files must remain fetchable for a while after a deploy — deleting them is what turns "safe caching" into a post-deploy 404. - **Chunks that churn for no reason.** A hash only buys you stability if the bytes are actually stable between builds; embedded build metadata or renumbered internal identifiers can change a file that is semantically unchanged, silently costing every returning user a re-download.
- If every asset is immutable, what still has to change on each deploy for a returning user to see the new build?The HTML document — or whatever small manifest lists the entry URLs. It is the one address that cannot be fingerprinted because users navigate to it directly, so it carries a short lifetime or is revalidated on every navigation. Everything else follows from what that document names, which is why over-caching the document quietly freezes users on an old release even though the assets are perfectly cacheable.
- A CSS file references a hashed font URL. What happens to the CSS file's hash when the font changes?It changes too. The new font gets a new URL, that URL is written into the stylesheet, so the stylesheet's bytes differ and its hash differs, and the document referencing the stylesheet names a new file in turn. Hashing therefore runs bottom-up through the asset graph: a parent's hash cannot be computed until every child's final URL is fixed.
- Does a content hash mean you never need to purge the CDN?For hashed assets, yes — a changed file arrives at a URL no cache has ever seen, so there is nothing to purge. Purging is only relevant for the addresses that stay stable across releases: the HTML document, and any un-fingerprinted file such as a manifest or a well-known path. That is a much smaller and much safer surface to operate.
A content hash names a file by its fingerprint rather than its edition number: two editions with identical text get the same name, so nobody re-fetches a chapter they already have.
saying these in an interview costs you the question
- Thinks the hash is a version number that increments each release
- Says a full CDN purge is required after every deploy
- Wants to cache the HTML document for a year too
- Claims browsers re-download every asset after any deploy
- Believes appending ?v=13 to every asset is equivalent