Why is shipping a service worker a different class of deploy risk from shipping ordinary static assets, and what do you do so that a fix can still reach users who already carry the bad one?
answer
- you can purge servers, not devices
- the client now holds the deciding code
- one narrow channel back in
- never long-cache the worker script
- fail open to the network
basics
~20 sA service worker puts code you can no longer revoke on the request path of every page in its scope. There is no server-side purge for a cache on a user's device, so the only way back in is the browser re-fetching the worker script — which is why that script must never be served with a long lifetime.
solid answer
~60 sEvery other layer of delivery converges on you: redeploy, purge the CDN, change a header, and clients come back to the current build. A service worker inverts that, because the code deciding what to fetch now lives on the device and can answer from its own cache without ever consulting you. If it caches the document aggressively and you push a fix, the affected users may never see it — they are pinned to a build you have already deleted, and stale JavaScript keeps calling an API that has moved on. The one channel back in is the browser periodically re-fetching the worker script when a page in scope loads, so protect that channel above everything: serve the worker script with a short or no-store lifetime, keep its logic small and boring, make it fall back to the network when its own code throws, and never let it be the only path to fresh HTML. Then rehearse the disable path before you need it, and report the active worker build with your real-user data so you can see how many people are still on the old one.
code
bash · 1 linecurl -sI https://example.com/sw.js | grep -i '^cache-control'go deeper
Remember the one-line version: a service worker is code on the user's device, so a server rollback does not reach it, and the worker script itself must never be cached for long.
Explain the recovery channel — the browser re-fetching the worker script on a later visit — and why serving that script short-lived from a stable, unhashed URL is what keeps the channel open.
Show you would treat it as an incident-management problem: fail open, keep the worker minimal, roll out gradually, measure how many sessions are still on the bad build, and rehearse the disable path before shipping.
Own the framing that this exports part of your delivery system onto devices you cannot reach, and set the bar — staged rollout, separate release train, telemetry on the active build, a practised recovery runbook — before any team ships one.
## What normally makes a bad deploy survivable On a conventional stack, every cache between you and the user is reachable. The origin is yours. The CDN has a purge API and obeys the lifetimes you set. The browser's own store is bounded by the headers you sent and, for the HTML entry point, by a deliberately short lifetime. Roll back and clients converge within minutes, because everything in the path is still asking you what to serve. A service worker breaks that chain. Once installed, it is code of yours running on the user's device, in front of every request from every page in its scope, and it may answer entirely from a store it controls. It does not consult your origin unless its own logic decides to. Nothing you do on the server — no purge, no header, no rollback — reaches it. This is the single fact the question is testing. ## The failure that follows The classic incident: a worker answers the document from its own cache. You ship a fix. Users load the site, the worker serves the old HTML, the old HTML names the old asset URLs, the worker serves those too, and the user sees the broken build indefinitely — including after a reload, which is the first thing they and your support team will try. Worse, this is not merely cosmetic staleness: old client code keeps calling a backend that has moved on, so you inherit a long tail of clients you cannot upgrade and must keep compatible with. The tail is measured in weeks, not minutes. Users who do not return do not get the fix, because the only opportunity to repair them is a visit. ## Protect the one channel back in When a page in scope is loaded, the browser checks for an updated worker script; if the bytes differ from the installed one, the new script is taken on and eventually replaces the old. That check is your entire recovery path, so treat it as production infrastructure. **Never serve the worker script with a long lifetime.** Modern evergreen browsers already refuse to trust a long `max-age` for it — they cap the value (24 hours) and by default bypass the HTTP cache for the script's update check — but do not lean on that. Serve it with a short lifetime or no-store, from a stable URL, and never with a content hash in its filename (a hashed worker URL is a worker nobody can update, because nothing points at the new one). ``` Cache-Control: no-cache ``` **Keep the worker's own code small, boring and rarely changed.** Its job is routing decisions; complexity belongs in the app, which you can update through the worker. A worker that changes on every release is a worker with a fresh chance to break on every release. **Fail open.** If your handler throws, the request should end up on the network rather than nowhere. A worker that returns nothing when its logic hits an unexpected state turns a small bug into a blank page for everyone who has it. **Do not make the worker the only route to fresh HTML.** Whatever policy you pick for the document, it must include a path where the network's copy wins reasonably promptly; otherwise you have designed the pinning failure in deliberately. ## Operate it like infrastructure - **Know your exposure.** Have the page report the active worker's build identifier along with your real-user telemetry, so you can answer "how many sessions are still on the bad build?" with a number instead of a guess. - **Rehearse the disable path.** Decide, and practise in a staging environment, exactly what you ship when the worker is the problem — including turning off registration so that new and repaired users stop acquiring it while you work. A recovery procedure first attempted during the incident is not a procedure. - **Ship worker changes on their own.** Bundling a risky worker change with a feature release means you cannot roll back one without the other. - **Roll out gradually.** A worker on 1% of sessions with a monitored error rate is a bad afternoon; the same worker on 100% is a multi-week tail. The underlying principle is worth saying explicitly in an interview: the moment you install a service worker you have moved a piece of your delivery system onto hardware you do not own and cannot reach. Everything above is about keeping one narrow, reliable door open through which you can still get in.
- What happens to users who install the bad worker and then do not return for three weeks?They stay broken until they visit again, because the update check only happens on a visit to a page in scope. That is why the blast radius is measured over weeks: your incident is over on the server long before it is over in the field, and your metrics should count sessions on the bad build rather than assume convergence.
- How do you find out how many users are still running the old worker?Stamp each build with an identifier and have the running app report the active worker's identifier alongside your real-user telemetry. Then the recovery becomes observable — you can watch the affected cohort shrink and know when the incident is genuinely closed rather than guessing from deploy time.
- Should the worker script itself be content-hashed like your other assets?No. A hashed filename gives a stable URL to nothing — the browser updates the worker by re-fetching the URL it registered, so that URL must stay constant and short-lived. Hash the assets the worker serves, never the worker's own entry point.
saying these in an interview costs you the question
- Thinks purging the CDN clears a client-side cache
- Serves the worker script with a long max-age
- Assumes users converge on the new build within minutes
- Lets a failure in the worker return nothing instead of hitting the network
- Bundles worker changes into ordinary feature releases