Your team wants to add its first service worker to a high-traffic production site. What would you require in the rollout plan before approving it, and how do you bound the blast radius if it misbehaves?
answer
- treat it as infrastructure, not a feature
- prove the mechanism before the behaviour
- an off switch you can pull server-side
- split real-user metrics by controlled sessions
- rehearse the disable runbook first
basics
~20 sApprove it as infrastructure, not a feature: a first version that does almost nothing, registration behind a server-side switch, a staged rollout with real-user monitoring split by whether a worker was in control, a rehearsed disable procedure, a named owner, and a success criterion that permits removal.
solid answer
~60 sI would separate two risks that teams usually ship together: the risk of *having* a worker at all, and the risk of whatever caching behaviour they want it to do. So the first release installs a worker that adds no caching — it proves registration, update and removal work on real traffic, at a small percentage, with nothing to be stale about. Behaviour lands only after that. Around it I would require four things. A server-controlled switch on registration, so we can stop new installs without a code deploy. A staged rollout by cohort with real-user metrics dimensioned by whether the session was worker-controlled, plus the worker's build identifier on every beacon so exposure is countable. A disable-and-recover runbook that has actually been executed in staging, not just written. And a named owner with a success criterion agreed in advance — if the controlled cohort shows no field improvement in a month, we remove it. What may and may not be stored gets decided in the same review, in writing.
go deeper
Know that a service worker is not an ordinary reversible deploy, and that a rollout should start small with a way to stop it rather than going straight to all users.
Be able to describe the staged rollout concretely: a first version with no caching behaviour, registration behind a flag, and telemetry that says whether a session was worker-controlled.
Show you would run it as an operational change — soak periods, defined stop conditions, a rehearsed disable path, and worker changes shipped on their own release train.
Own the governance: who is accountable, what the worker may store, the exposure metric that closes an incident, and the pre-agreed criterion under which the worker gets removed again.
## Why this needs a plan at all Most frontend changes are reversible by redeploying. A service worker is not: it places code on users' devices, in front of every request in its scope, that you can only reach again when those users come back. That property — not the caching, not the offline page — is what makes it an infrastructure decision, and it is why "we'll add a service worker this sprint" deserves a review rather than a code review. The goal of the plan is to make the first weeks boring and every failure bounded. ## Separate the two risks Teams usually ship one change that both installs a worker and gives it a caching policy. That fuses two independent risks. The mechanism risk is whether registration, updating and removal behave correctly on your real traffic, across your real browser mix, behind your real CDN. The behaviour risk is whether the caching policy serves the right thing. Ship the mechanism first. Release a worker that adds no caching of its own — it exists, it updates, it can be removed. There is nothing for it to serve stale, so the worst case is a wasted release. Once you have watched it update itself across a full deploy cycle and confirmed you can remove it from a cohort on demand, you have earned the right to add behaviour, and any later incident has a known-good baseline to fall back to. ## Bound the blast radius **A switch you can pull without a deploy.** Registration should be conditional on a flag your servers control. Turning it off does not repair devices that already have a worker, but it immediately stops the population growing while you work — and on an incident clock, halting the bleed is worth more than the fix. **Staged exposure.** One percent, then a tenth, then everyone, with a defined soak at each step. Choose cohorts you can describe afterwards, because you will need to compare them. Publish, in advance, what would make you stop. **Exposure you can count.** Every real-user beacon should carry two extra dimensions: whether the session was controlled by a worker, and which build of the worker. Without them you cannot tell whether a metric moved because of the worker, and during an incident you cannot say how many people are still affected — you are reduced to guessing from deploy timestamps. **A rehearsed recovery.** Write down exactly what you ship when the worker is the problem, and then do it in staging with a device that carries the bad build. A procedure first attempted at 2am has not been tested. Include how the incident closes: not "we deployed the fix" but "the affected cohort fell below X". **Its own release train.** Worker changes ship separately from feature work, so a rollback of one never blocks the other, and so a risky change is never invisible inside a large diff. ## Decide the policy questions before the code Two things must be settled in writing at approval time. What the worker is allowed to store — the boundary between public assets and anything user-specific belongs in this review, with security and privacy in the room, not in a later pull request. And who owns it: a named team that reviews every change to the worker, keeps the runbook current, and is paged when worker-attributed errors spike. An unowned worker is the one that is still running two reorganisations later with nobody able to explain what it caches. ## Agree what success means, and be willing to remove it Approval should come with a number and a date: the field metric expected to move, for which cohort, by when. Real-user data comparing controlled and uncontrolled sessions decides it. If the worker cannot show its win within the agreed window, remove it — deliberately, through the same staged process, because removal is itself a change that has to reach devices. That last clause is the part teams resist and the part that matters most. A service-worker cache is a permanent liability sitting in front of every request your users make; it should have to keep earning that position, and the plan that installs it should already describe how it comes out.
- What is the smallest useful first service worker you would ship?One that registers, updates and can be removed, and caches nothing of its own. It proves the riskiest machinery on real traffic while there is nothing to serve stale. You learn how updates propagate through your actual browser mix and CDN before any user-visible behaviour depends on it.
- Turning off registration does not fix devices that already have the worker. So what is it for?Stopping the bleed. It caps the exposed population immediately and without a deploy, which buys time to build and validate the real fix. Recovery for devices that already carry the worker is a separate track, and the two run in parallel during an incident.
- How would you decide to remove a service worker that is already live?On the number agreed at approval. If field data from worker-controlled sessions shows no improvement over uncontrolled ones by the agreed date, remove it — through the same staged process, because removal is a change that also has to reach devices. Carrying an unjustified worker is a standing risk with no offsetting benefit.
- What would you monitor in the first week?Error rate and failed-request rate split by controlled versus uncontrolled sessions, the distribution of active worker builds, the field metrics you predicted would move, and support ticket themes mentioning stale content. The build distribution is the one teams forget, and it is what tells you whether updates are actually propagating.
saying these in an interview costs you the question
- Ships the worker and its caching behaviour in one release
- No way to stop registration without a deploy
- Rolls out to 100% of traffic immediately
- Cannot say how many sessions are running which worker build
- Writes a recovery runbook but never executes it