A broken service worker is live and returning visitors keep getting a broken app even after you redeploy. How does the browser decide it has a new worker script, and how do you actually kill a bad worker for users who already carry it?
answer
- updates are pull-based on navigation
- byte-for-byte script comparison
- cache the worker script and you are stuck
- unregister leaves open pages controlled
- ship a worker that removes itself
basics
~20 sThe browser re-fetches the worker script on navigations in scope and compares it byte for byte with the stored copy, so an unchanged or HTTP-cached script means no update. To kill a bad worker, ship a replacement sw.js that calls skipWaiting and clients.claim, then unregisters itself and reloads its clients.
solid answer
~50 sUpdates are pull-based: on a navigation in scope — and on functional events, throttled — the browser re-requests the worker script and does a byte-for-byte comparison with the stored version, including imported scripts. Anything that keeps the old bytes coming back blocks the fix, which is why `sw.js` must be served with a short or `no-cache` lifetime; current browsers also cap the worker script's HTTP cache age at 24 hours and, by default (`updateViaCache: 'imports'`), bypass the HTTP cache for the top-level script. Once the new script does download, `registration.unregister()` alone is weak — it stops future navigations from being controlled but leaves current documents with their controller. The reliable kill switch is a minimal `sw.js` that calls `self.skipWaiting()` in install, and in activate does `clients.claim()`, `self.registration.unregister()`, and then reloads its window clients. Deploy that, let it propagate, and only then ship a real worker again.
code
javascript · 13 lines// sw.js — kill switch: install it, and it removes itself and its predecessor.
self.addEventListener('install', () => self.skipWaiting());
self.addEventListener('activate', (event) => {
event.waitUntil(
(async () => {
await self.clients.claim();
await self.registration.unregister();
const windows = await self.clients.matchAll({ type: 'window' });
for (const client of windows) client.navigate(client.url);
})()
);
});go deeper
Know that the browser only replaces a worker when the script's bytes change, and that sw.js should not be cached aggressively.
Describe what triggers an update check, that the comparison is byte for byte including imported scripts, and what updateViaCache controls.
Be able to run the recovery: explain why unregister alone is insufficient, write the self-removing worker, and reason about how long to leave it deployed before shipping a real one again.
Treat the worker as production infrastructure — set the caching and review policy for the script, require a version signal you can observe, and decide in advance what your kill-switch procedure and its rollout window look like.
## Why a bad worker is uniquely painful Every other bad deploy is fixed by shipping a good one, because the browser asks the server for the page. A bad service worker is different: it sits between the user and the network for every request from every controlled page. If it is broken in a way that also breaks how the fix is delivered, the user has no path back to your server. That is the scenario worth being able to talk through. ## How the browser decides an update exists An update check is triggered by: - a **navigation** to an in-scope URL (the common case); - a functional event such as `push` or `fetch`, throttled so it happens at most about once a day; - an explicit `registration.update()` from the page; - the browser's own periodic check when the registration is older than 24 hours. The check fetches the worker script and compares it **byte for byte** with the stored copy. Not the ETag, not the modified date — the bytes. Scripts pulled in with `importScripts` are compared too. If nothing differs, no update happens; a worker whose only change is in files it fetches at runtime will never itself be replaced. ## The HTTP cache trap The classic incident: `sw.js` is served by the same static-asset rule as everything else, with `Cache-Control: max-age=31536000`. The browser's update check hits its own HTTP cache, gets the old bytes back, and concludes there is no update — indefinitely. Current browsers mitigate this in two ways. The worker script's cache lifetime is **capped at 24 hours** regardless of a longer `max-age`, and `register()`'s `updateViaCache` option defaults to `'imports'`, meaning the top-level worker script bypasses the HTTP cache entirely on update checks while imported scripts may still use it. You can force the strictest behaviour explicitly: ```js navigator.serviceWorker.register('/sw.js', { updateViaCache: 'none' }); ``` Even so, serve `sw.js` with `Cache-Control: no-cache` (or a very short max-age). Intermediaries you do not control — a CDN edge rule, a corporate proxy — do not necessarily implement the browser's cap, and 24 hours of a broken app is already an incident. ## Why unregister() alone is not a kill switch ```js const reg = await navigator.serviceWorker.getRegistration(); await reg.unregister(); ``` This removes the registration so **future** navigations are uncontrolled. It does not tear the controller out of documents that are already open — they keep their controller until they unload. And it only runs at all if a page of yours loads and executes that code, which requires the broken worker to have served a working enough page to run it. As a recovery mechanism it is a coin flip. ## The kill-switch worker The robust move is to ship a replacement worker whose whole job is to remove itself: ```js self.addEventListener('install', () => self.skipWaiting()); self.addEventListener('activate', (event) => { event.waitUntil((async () => { await self.clients.claim(); await self.registration.unregister(); const clients = await self.clients.matchAll({ type: 'window' }); for (const client of clients) client.navigate(client.url); })()); }); ``` What each line buys you: `skipWaiting()` refuses to sit in the waiting room behind the broken version; `clients.claim()` takes over the open pages immediately; `unregister()` removes the registration; the reload loop gets every open window back onto a network-served document. Note there is no `fetch` handler at all — with none registered, requests go to the network by default, so even before the reload the user is no longer being served by a broken interceptor. On a large site, leave the kill-switch worker deployed long enough for the update check to reach the long-tail users before you ship a real worker again; users who have not visited in the meantime are still carrying the broken version. ## Reducing the chance you ever need it - Keep the `fetch` handler defensive: if you do not call `event.respondWith()`, the browser performs its default network fetch, so an unhandled path degrades to normal browsing rather than to an error. - Never let the worker be the only way to load your HTML shell without a network fallback path. - Keep the worker script small and boring, and review it as production infrastructure, not as app code. - Ship a way to observe it: report the worker's version from the page so you can see how many users are still on an old one.
- Why does the kill-switch worker deliberately register no fetch handler?Because a worker with no `fetch` listener does not intercept anything — the browser performs its normal network fetch for every request. That makes the page behave as if there were no service worker at all, even in the window between activation and the reload, and it removes any chance that the recovery worker itself breaks a response.
- What Cache-Control would you put on sw.js, given browsers already cap it at 24 hours?`no-cache`, or a max-age of a few minutes at most. The browser's cap is a safety net, not a guarantee for the whole path: CDN edges and corporate proxies apply their own rules, and 24 hours of serving a broken worker is an incident on its own terms. Explicitly registering with `updateViaCache: 'none'` is a useful belt-and-braces addition.
- Your fix is deployed but metrics show a slice of users still on the old worker a week later. What explains it?They have not navigated. Update checks are pull-based, triggered mostly by an in-scope navigation, so a user who has not opened the app has no reason to notice the new script — and someone who keeps a tab open indefinitely may only hit the throttled periodic check. That is why the kill-switch worker must stay deployed until the long tail has drained.
saying these in an interview costs you the question
- Thinks unregister() immediately stops the worker in open pages
- Serves sw.js with a long max-age like other static assets
- Believes ETag or Last-Modified drives the update check
- Assumes redeploying the app updates the worker without a script change
- Expects all users to pick up the fix without visiting the site