skip to content

Minutes after a deploy, some users see a blank screen or a failed route, and their network log shows a 404 for `/assets/route-detail.3c19aa7e.js` — a filename that does not exist in the new build. What caching and deploy decisions produced that, and how would you stop it recurring?

level: seniorimportance: should knowfreq 48%

answer

  1. old document names deleted filenames
  2. the document is the only mutable URL
  3. open tabs ignore every cache header
  4. assets first, document last
  5. retain previous builds past the tail

basics

~20 s

An old HTML document, cached or sitting in a tab opened before the deploy, still names chunk URLs the deploy deleted. Fix it by publishing assets before the document, retaining previous builds, and keeping the HTML short-lived.

solid answer

~50 s

Those users are running the previous build's document. Either it is still being served from a browser or CDN cache, or their tab has been open since before the deploy, so the code asking for that chunk is old code that knows only old filenames — and the deploy replaced the asset directory, so the old name now 404s. Note that the hashed assets are behaving correctly; the failure is that the previous build was deleted while clients still referenced it. I would fix it in three places: publish all new assets before flipping the document, so there is never a moment when the document names files that are not there yet; keep the previous few builds' assets live for longer than any document referencing them can survive; and keep the document's cache lifetime short so that window is bounded. Then add a client-side backstop that reloads once when a dynamic chunk fails to load.

go deeper

for a junior

Know that the HTML document is what lists the current asset filenames, so a user on an old document asks for old files. Do not answer with "tell them to clear their cache".

for a middle

Explain the ordering rule — publish assets before swapping the document — and why hashed filenames let two builds coexist in one directory rather than replacing each other.

for a senior

Demonstrate the full diagnosis: identify which population holds the old document, check the lifetime actually served on it, and connect document lifetime to the retention window the previous build needs.

for a principal

Own it as a deploy contract: assets are append-only and retained past the longest client tail, the document is the single invalidation point, and rollback is only safe while both builds remain published.

## What the 404 is telling you A content-hashed URL is a promise that those exact bytes live at that exact address. A 404 on one means the promise was broken from the server side: something deleted a file that a live client still needs. The client is not confused and its cache is not corrupt — it is faithfully requesting a filename that the build it is running told it to request. The file it is running is the **document**. In a hashed-asset scheme the document is the only mutable address, and it is where the current build's filenames live. So any client holding an old document holds an old set of filenames, and it will keep asking for them until it gets a new document. ## The three populations that hold an old document 1. **Browser-cached documents.** If the document was served with any positive freshness lifetime, returning users navigate straight into the previous build with no request to the origin. 2. **Edge-cached documents.** A CDN holding the document for even a short TTL keeps handing the old build to everyone routed through that node until it expires or is invalidated. 3. **Open tabs.** The most stubborn population, and the one no cache header touches. A tab opened before the deploy is running the old build's JavaScript in memory. When the user finally clicks the link that triggers a lazily-loaded route, the request goes out for a filename that has not existed for hours. Every one of these is a normal, expected state. The system must tolerate them; the deploy design is what determines whether it does. ## The deploy shape that breaks it The breaking pattern is a deploy that treats the asset directory as a single mutable thing: wipe it, upload the new build, done. That is correct for the *new* clients and instantly fatal for the *old* ones. A close cousin is publishing the document first and the assets afterwards, which opens a window where the new document names files that have not finished uploading — the mirror-image failure, and worse because it hits fresh visitors. ## The four fixes, in order of leverage **Publish assets first, flip the document last.** Content-hashed names never collide, so the previous and the new build can coexist in the same directory. Upload every new asset, confirm they are all fetchable, and only then update the document. There is then no instant at which any document names a file that is not there. **Retain previous builds.** Keep the last several builds' assets in place. The retention window must exceed the longest time a client can still be holding an old document — the document's freshness lifetime plus edge TTL plus the tail of long-open tabs. Deleting a build is what turns a normal stale client into a 404. **Keep the document short-lived.** The document's cache lifetime is the dial that sets how big the stale population is and how long it persists. Short lifetimes with revalidation mean users converge on the new build within one navigation. This is also the reason the document is the one thing you never fingerprint and never cache aggressively. **Add a client-side backstop.** Even with all of the above, a tab open for days will eventually reach for something that is gone. Detect the failure of a dynamically loaded chunk and recover by reloading the page once, which fetches a fresh document and therefore fresh filenames. Guard it: record that a reload was already attempted, so a genuinely broken deploy does not put users in a reload loop, and consider showing an explicit "a new version is available" prompt instead of reloading under someone's cursor. ## Diagnosis when it is already happening Confirm the failing filename is absent from the current build's output — that distinguishes this from a partially propagated upload, where the file exists at the origin but not yet at every edge. Check the freshness lifetime actually being served on the document, since a proxy or platform default can quietly override what you thought you configured. Look at how far apart the failing users' document versions are from the current one; if they are many releases behind, the problem is document caching, and if they are one release behind, the problem is deploy ordering or retention. ## Fixes that look tempting and are wrong Shortening the assets' cache lifetime does nothing — the assets are not stale, they are gone. Purging the CDN does not help either, because the stale documents are largely in browsers, and it makes the deploy slower and more fragile. Fingerprinting the document is not an option, since users navigate to it by a stable address. And rolling back is only a partial escape: it restores the old assets and rescues the old clients, but unless the new build's assets are also retained, the clients that already took the new document break in exactly the same way.

  • Why does this failure get worse the longer the HTML document is cached?
    Because the document's freshness lifetime sets how long clients keep believing in old filenames. A one-minute lifetime means the stale population drains within a navigation or two; a one-day lifetime means a full day of users can still request last release's chunks. Retention has to cover that window, so a longer document lifetime directly forces a longer and more expensive asset retention policy.
  • A user's tab has been open for three days. How do you bound that tail?
    Cache headers cannot touch it, so you need an in-app mechanism: poll a small version endpoint or compare a build id embedded in the app against the current one, and prompt the user to reload when they differ. That converts an unbounded tail into a bounded one, which in turn lets you set a defensible retention window instead of keeping every build forever.
  • Is rolling back the deploy a fix for this symptom?
    Partly, and it can create the mirror problem. Rolling back restores the previous build's assets, so the stale-document clients recover — but every client that already received the new document now names files that the rollback removed. Rollback is only safe when both builds' assets remain published, which is another argument for retention rather than for treating the asset directory as a single mutable state.
  • How do you tell this apart from a CDN that has not finished replicating the new build?
    Check whether the failing filename exists in the current build's output at all. If it does not, the client is running an old document. If it does exist but only some edges 404, the upload is still propagating — a different failure with a different fix, namely publishing and verifying assets everywhere before flipping the document.

saying these in an interview costs you the question

  • Blames the browser cache and tells users to hard-refresh
  • Proposes shortening the assets' max-age to fix it
  • Wants to purge the CDN on every deploy as the fix
  • Thinks fingerprinting alone makes deploys safe
  • Ignores tabs that were open before the deploy

context