skip to content

The opt-out guardrail trips at the 25% step of a model ramp — what does halting the ramp leave running that a kill switch does not?

level: seniorimportance: should knowfreq 55%

answer

  1. halt stops growth, not exposure
  2. the exposed cohort stays exposed
  3. kill switch is a routing decision
  4. neither changes the pinned versions
  5. the queue outlives the switch

basics

~20 s

Halting freezes the percentage where it is, so the 25% already assigned keep being served by the candidate model version. It stops the ramp growing; it does not remove exposure. A kill switch routes everyone back to the incumbent immediately.

solid answer

~50 s

The three responses form a ladder, and they restore different things. **Halting** holds the ramp at its current percentage: no new users are assigned, but the exposed cohort stays exposed, which is useful when the harm is bounded and you still need that cohort to diagnose. **The kill switch** is a serving-path decision that sends every user back to the incumbent model version within seconds, without a deploy — but the candidate stays deployed, the feature-definition version stays where the candidate left it, and any work already queued under the candidate is untouched unless it is explicitly invalidated. **Rolling back** applies the previous pinned release and is the only durable one. In a queued sender the switch's effect is also delayed by the queue horizon, so a switch that only changes future planning runs looks instant on a graph and is not.

go deeper

for a junior

Recall the difference between stopping a ramp from growing and removing the candidate from traffic entirely; they are two different actions with two different effects.

for a middle

Explain why a kill switch takes effect in seconds while a version change does not, and what stays deployed after the switch is flipped.

for a senior

Choose between halt, switch and rollback from whether harm accrues per action, and name the queue horizon and the feature-definition version as what survives the switch.

for a principal

Decide organisationally which model rollouts must carry a tested switch and a tested drain, and how long a system may sit on a switch before a rollback is mandatory.

## Three responses, three different scopes When a guardrail moves during a ramp, an engineer has three distinct levers, and a design round usually wants to hear that they are not synonyms. | Action | What it changes | What it leaves in place | Time to effect | |---|---|---|---| | Halt the ramp | The percentage stops rising | The exposed cohort stays on the candidate; all versions unchanged | Immediate for new assignment only | | Kill switch | Routing: everyone served by the incumbent | Candidate deployed, feature-definition version as the candidate left it, queued work | Seconds — plus the queue horizon | | Roll back the release | The pinned versions: artifact, feature definitions, serving config | Nothing of the candidate remains authoritative | Minutes, plus re-materialisation and drain | ## What halting really buys Halting is a **hold**, not a retreat. Its value is that it caps the blast radius at today's number while preserving an exposed population you can still read. That matters because diagnosis usually needs the failure to keep happening: if you pull everyone back the moment a guardrail moves, you often lose the only cohort that could tell you whether the candidate is at fault or whether something else moved at the same time. So halting is the right response when: - the harm per exposed user is bounded and reversible; - the guardrail moved but the direction and the cause are not yet clear; - the exposed cohort is the evidence you need to decide between resume and revert. And it is the wrong response when the harm accrues with every additional action — every additional badly timed notification — because holding at 25% keeps producing exactly those actions. ## What the kill switch buys, and what it does not The kill switch is a **routing decision read at decision time**, which is why it can take effect without a build. Its whole reason to exist is that recovery time should not be a deploy time. But its scope is narrower than people assume: - **It does not change any version.** The candidate artifact is still deployed and the feature-definition version it introduced is still the one the pipeline is producing. If that definition replaced the previous one in place, the incumbent is now reading semantics it was not trained against — so a switch can hand traffic back to a model that is itself no longer correct. - **It does not undo work already committed.** In a system that plans sends ahead of delivery, the queue holds decisions the candidate already made. Unless the switch also invalidates or re-plans that queue, notifications keep going out at the candidate's hours for the full queue horizon while every dashboard shows 0% candidate traffic. - **It is not a fix.** Nothing about the switch is a state you can leave the system in for weeks: the candidate is still shipped, the config is still divergent, and the next deploy has to know about it. ## Choosing between them under pressure 1. **Is harm accruing per action?** If yes, cut exposure now — the switch — and diagnose from logged history rather than from live victims. 2. **Is the harm bounded and the cause unclear?** Halt, keep the cohort, and look. 3. **Is the candidate confirmed at fault?** Roll back the pinned release, drain or re-plan queued work, and only then take the system off the switch. ## The detail that separates a good answer The strong answer names the **queue horizon** and the **feature-definition version** as the two things that survive a kill switch. Both are the reason an engineer can flip the switch, watch the traffic graph go to zero candidate, and still get user complaints an hour later. A rollout design is only as fast as the slowest thing the candidate already wrote down.

  • When is halting the better response than the kill switch?
    When the harm per exposed user is bounded and the cause is still unclear. Halting caps the blast radius at today's percentage while keeping the exposed cohort producing the signal you need to tell whether the candidate is at fault or something else moved at the same time. Pulling everyone back immediately often destroys that evidence.
  • What has to be true for a kill switch to actually take effect in seconds?
    The incumbent model version must already be loaded and serving, the routing decision must be read at decision time rather than baked into a deployed artifact, and any work already queued under the candidate must be invalidated or re-planned. Without the last one the graph goes to zero while the candidate's decisions keep being delivered.

saying these in an interview costs you the question

  • Thinks halting the ramp removes the candidate from all traffic
  • Calls flipping the kill switch a rollback of the model version
  • Assumes the switch stops sends already queued under the candidate
  • Proposes redeploying the previous artifact as the fastest way to stop harm
  • Leaves the system on the kill switch as the permanent fix