A team's incident plan says the mitigation for their recommendation service failing is a kill switch that disables recommendations, but that switch has never been flipped in production. Why is that not yet a control, and what would you do to make it one?
answer
- never flipped, never verified
- wiring drifts around the conditional
- is the degraded state serviceable?
- on-call runs it, not the author
- record the date next to the claim
basics
~20 sAn unexercised switch has unverified wiring, an unverified fallback experience and unverified access. Make it a control by drilling it on a schedule in production, measuring time from flip to effect, and confirming the degraded state against a service level indicator.
solid answer
~50 sA control is something you have evidence works, and a switch nobody has ever flipped has none. Three things are unknown: whether the flag still gates the code path that matters after months of refactoring, whether the degraded experience is actually serviceable rather than a page that errors with no recommendations, and whether the on-call engineer has access and knows the exact steps at 3am. You make it a control by drilling it — quarterly, in production, starting announced and at a low-traffic hour, later unannounced. Measure the interval from flip to observable effect, watch the user-facing indicator during the degraded window to prove the fallback is acceptable, and have the drill performed by whoever is on call rather than the feature's author. Record the result next to the mitigation step so the plan cites evidence with a date, not an assumption.
go deeper
Know that a kill switch is meant to disable a risky feature quickly, and that nobody can claim it works until someone has actually turned it off in production.
Explain what a drill measures: propagation time, whether the guarded path is fully covered by the conditional, and what the user sees while the feature is off.
Show the operating judgment — announced then unannounced cadence, on-call executes rather than the author, dated evidence recorded beside the mitigation, and re-verification after refactors.
Generalise it: every unexercised recovery mechanism decays. Argue for a drill budget across failovers, restores and switches, and for accepting a small deliberate degradation as the price of verified mitigations.
## Why an unexercised switch is not a control "We have a kill switch" describes an intention. A control is something whose behaviour you can state from evidence. Between the intention and the evidence sit at least four independent failure modes, each of which has taken down real mitigations: **The wiring drifted.** The flag was added around one call site. Six months of refactoring later, the expensive dependency is also reached from a second path — a batch job, a mobile endpoint, a prefetch — that the conditional never covered. Flipping the switch then reduces load by a fraction and the incident continues, with the responders now confident they have already applied the fix. This is the most common one and the hardest to spot by reading code. **The fallback is not serviceable.** "Recommendations off" is supposed to mean a page that renders without a recommendation strip. It can instead mean an empty carousel with a broken layout, a null-pointer error in a downstream template, or a client that treats an empty list as a fatal response. The degraded state is a product decision that somebody must have looked at while it was live. **Access and knowledge.** The switch may be flippable only by a role that the on-call engineer does not hold, or through a console behind an SSO flow that nobody has used from a phone at 3am. If reaching the lever requires waking the feature owner, the mitigation time is the owner's response time, not the flip time. **Timing.** Nobody knows how long the change takes to reach every instance, so nobody knows how long to wait before concluding it did not help. That uncertainty is what causes responders to stack a second change on top of the first. ## Turning it into a control: the drill The fix is deliberate practice: flip the switch on purpose, in production, when nothing is wrong. **Start announced, in a low-traffic window.** The first drill should be planned, communicated, and short. You are buying information at the price of a small, chosen degradation — perhaps two minutes of missing recommendations — instead of paying full price during an incident. **Measure the things you did not know.** Time from flip to first observable effect. Time to full propagation across instances. The shape of the user-facing indicator during the degraded window: does conversion drop, does latency improve, does error rate stay flat? The last question matters most, because the whole premise is that degraded is better than down. **Have the right person do it.** The drill should be run by the current on-call engineer following the written steps, not by the engineer who built the feature. If the steps are ambiguous or the access is missing, that is exactly the finding you wanted, and it costs nothing to fix on a Tuesday afternoon. **Restore deliberately and verify.** Turning the switch back on is half the exercise. Confirm the feature returns cleanly, that no state was corrupted while it was off, and that any queues or caches that drained refill without a thundering-herd effect. **Escalate the realism.** Once the announced drill is boring, run it unannounced during business hours, then during a broader exercise where responders discover it themselves. Each step tests something different: the mechanism, then the humans, then the discovery. ## Cadence and cost A sensible default is quarterly per critical switch, plus after any significant refactor of the guarded path. The cost is real and worth naming out loud: a small deliberate degradation, engineer time, and a customer-facing blip you have to be willing to accept. Weigh it against the alternative — discovering the switch is dead in the middle of an outage, when you have both the original failure and no mitigation. If a switch cannot be drilled in production at all — because flipping it corrupts data, or because the fallback is genuinely unacceptable for even a minute — that is a design finding. Either the guarded path needs restructuring so the off state is safe, or the team should stop calling it a mitigation and write a plan that is actually executable. ## Recording the evidence The drill's output belongs where the mitigation is claimed. "Kill switch verified 2026-05-14; propagation 8 seconds; conversion dropped 3 percent while off; owner: search team" is a control. "There is a kill switch" is a hope. When an interviewer asks how you know your mitigations work, the dated evidence line is the answer they are listening for. ## The general principle This generalises past flags: any recovery mechanism that is never exercised — a failover, a restore-from-backup, a manual runbook step, a standby region — decays to non-functional at a rate nobody is measuring. Feature flags just make it unusually cheap to keep the evidence fresh, which is why an SRE interview treats "do you drill it?" as a proxy for whether you have really operated the system.
- How would you find out whether the flag still gates every path to the expensive dependency?Do not rely on reading the code. Instrument the dependency's client with a counter and flip the switch in a drill: if calls do not fall to zero, something reaches it outside the conditional. Batch jobs, prefetchers, cache warmers and alternate API surfaces are the usual culprits, and only the measurement finds them reliably.
- What would make you decide a switch simply cannot be drilled in production?When the off state is genuinely unsafe rather than merely degraded — it drops writes, leaves records half-migrated, or breaks a contractual obligation. That is a design finding, not a reason to skip the drill: restructure the guarded path so the off state is safe, or stop calling it a mitigation and write a plan that can actually be executed.
- Announced or unannounced drills — which do you run, and in what order?Announced first. The goal of the first drill is to learn whether the mechanism works, and an announced run gets you that with minimal risk and a fast abort. Once it is boring, go unannounced during business hours to test discovery and the human path. Jumping straight to unannounced tends to produce a real incident and a team that refuses to drill again.
- Who should own a permanent kill switch once the feature ships?The team that owns the guarded dependency, with the on-call rotation holding the authority to use it. Ownership means keeping the fallback tested, running the drill on cadence, and re-verifying after refactors. A switch with no owner degrades exactly like a stale flag, except it is the one you were counting on.
saying these in an interview costs you the question
- The switch exists in the code, so it works
- Flipping it in production is too risky to ever practise
- The feature author should run the drill
- Degraded means the same thing as broken
- Testing it in staging proves it works in production