skip to content

Before a staged rollout starts, what stopping rule and blast-radius limit must already be defined?

level: seniorimportance: must knowfreq 64%

answer

  1. Decide before anyone is invested
  2. Conditions, not adjectives
  3. The window must outlast the signal
  4. Consequence bounds exposure, not percentage
  5. Ambiguous means revert, not wait

basics

~20 s

Written-down abort criteria with an observation window per stage, a named decider, and a default of reverting when evidence is ambiguous; plus a bounded exposure - cohort, data, irreversible side effects - and a rehearsed revert path.

solid answer

~40 s

A staged rollout only verifies anything if the abort decision was made before the ramp began. That means predeclared criteria stated as conditions rather than adjectives, an observation window per stage at least as long as the slowest signal takes to report, a named decider, and an explicit default: ambiguous evidence means revert, not keep watching. Blast radius is the other half, and it is set by what the new version can do irreversibly rather than by the traffic percentage - a version at 2 percent that emits real payment instructions is unbounded. So you also need cohort ordering by consequence, version-tagged telemetry, a revert path rehearsed with a measured time-to-revert, and an expiry after which an unfinished rollout reverts itself. Measuring a metric difference is experiment design, a separate discipline.

code

pseudocode · 14 lines
pseudocode
rollout "gross-to-net-v2":
    stages       = [2%, 11%, 40%, 100%]
    cohort_order = [internal_staff, employers_under_50, all]
    observe_each = one_complete_pay_cycle
    baseline     = measure(before_exposure)

    abort_when:
        invariant_violations("gross-net-mismatch", version="v2") > 0
        p99(preview_latency, version="v2") > 1.4 * baseline.p99 for 25 minutes
        kill_switch == on
        signal_is_ambiguous == true            # default is revert, not wait

    on_abort: route_all(previous_version)      # 90 seconds, no deployment
    expires_after 3 weeks -> route_all(previous_version)

go deeper

for a junior

Know the basic separation: deploying the code and exposing it to users are different steps, and a staged rollout increases exposure gradually so a problem reaches few people before someone can switch it back.

for a middle

Be able to list what a stopping rule contains - measurable conditions, an observation window per stage, who decides, and what happens on ambiguous evidence - and explain why writing it before the ramp changes the decision.

for a senior

Show the operating judgement: matching the window to the slowest signal, ordering cohorts by consequence, spotting irreversible side effects and forward-incompatible writes, rehearsing the revert and knowing its measured time-to-revert.

for a principal

Own the policy across teams - which classes of change may ramp at all, mandatory gates before irreversible steps, automatic expiry so no rollout becomes a permanent fork, and a clear line between a coarse release decision and a measured experiment.

### Deploy is not release A staged rollout separates *deployed* from *exposed*. The new version is running, and a control - a flag, a routing rule, a cohort list - decides who reaches it. That separation is the precondition for everything else: if exposure cannot be changed without a deployment, there is no rollout, only a series of releases, and the response to bad news is measured in build minutes rather than seconds. ### The stopping rule, written before the ramp The reason to predeclare is human, not statistical. Once a change is at 40 percent and the team has spent three weeks on it, every ambiguous signal acquires a benign explanation. A stopping rule written before anyone is invested is the only version of the rule that is honest. It needs five parts. **Conditions, not adjectives.** *If it looks bad, we stop* is not a rule. *Any violation of the gross-equals-net-plus-deductions invariant, at all*; *preview p99 latency above 1.4 times the pre-ramp baseline sustained for 25 minutes*; *any manual correction raised against a payslip produced by the new version* - those are rules. Note the mixed shapes: money invariants are zero-tolerance counts, performance is a sustained relative comparison. **An observation window per stage.** The window must be at least as long as the slowest criterion takes to report. If the correction-rate signal for a payroll defect only becomes readable after a pay cycle completes, a stage that advances in two hours has observed nothing but latency. Teams that ramp to 100 percent inside a day and then discover the defect at month-end did have a stopping rule; they just never gave it time to fire. **A named decider.** Either an automated policy that halts and reverts on its own, or a specific role that owns the call. *The team will decide* means nobody decides at three in the morning. **A default for ambiguity.** The most valuable clause in the rule, because ambiguity is the common case. Ambiguous means roll back, investigate off the critical path, and ramp again. The alternative - hold and keep watching - is how a rollout stalls at 11 percent for a month, leaving two versions in production, two data shapes and doubled support load. **An expiry.** A stage that has not advanced or reverted within a set time reverts automatically. Without this, the rollout quietly becomes a permanent fork. ### Blast radius is about consequence, not percentage The most common mistake is to equate blast radius with the exposure share. Four dimensions actually bound it. **Who.** Cohort ordering from lowest to highest consequence: internal staff first, then a small low-risk customer segment, then general population. Never the largest payroll customer first, and be aware that a random percentage draws from every segment at once - random exposure is fine for a preview screen and reckless for a pay-run submission. **What data.** A version that writes a new field the old version cannot read has extended its blast radius into the data, and reverting the code no longer reverts the state. The safe shape is to make the schema change first, tolerated by both versions, and only then ramp the behaviour. **Which side effects are irreversible.** This dominates everything. A payment instruction sent, a filing submitted, a notification emailed to fourteen thousand employees - none of these can be rolled back. A version at 2 percent exposure that can disburse money has an effectively unbounded radius, and the mitigation is a hold-and-review gate on the irreversible step, not a smaller percentage. **For how long.** Exposure multiplied by duration is what determines how many records were affected before the rule fired. A short window at a small share is genuinely bounded; the same share running for a three-week release train is not. ### Preconditions that let a rollout be read at all Even a perfect rule is unusable without three things in place first. **Attribution**: every record and every telemetry event carries the version that produced it, or you can see a metric move without knowing which of the train's many changes moved it. **A verified revert path**: exercised before the ramp, with a measured time-to-revert - if reverting takes 40 minutes, your stopping rule is really *40 minutes of damage after the criterion fires*. **A pre-ramp baseline**: the same measurements taken before exposure, since every criterion here is comparative. ### A concrete shape For a gross-to-net rewrite on a three-week release train: stages at 2, 11, 40 and 100 percent; cohort order internal staff, then employers under 50 employees, then all; observation per stage of one complete pay cycle for anything touching money; abort on any money-invariant violation, on sustained p99 latency above 1.4 times baseline for 25 minutes, or on the kill switch; revert by routing to the previous version within 90 seconds with no deployment; automatic revert if a stage has not advanced within the train's three weeks. ### Where this stops The rollout decision described here is deliberately coarse: is this version harmful, yes or no. Deciding whether a metric difference between exposed and unexposed groups is real - randomisation unit, how much data is needed, the penalty for repeated looks - is experiment design, a different discipline with its own machinery. Conflating the two produces bad releases and bad measurements alike: a safety call delayed for statistical certainty, or a business result claimed from a ramp never designed to measure one.

  • Why is a rollout at 2% exposure not necessarily a small blast radius?
    Because exposure share bounds how many users are affected, not how badly or how reversibly. A version reaching 2 percent of pay runs that can emit a real payment instruction, submit a filing or email thousands of employees has caused effects nobody can undo. Radius is set by the irreversibility of the side effects, the data the version writes that the old one cannot read, and exposure multiplied by duration. The mitigation for an irreversible step is a hold-and-review gate, not a smaller percentage.
  • The signal at stage two is ambiguous - a small rise in corrections that might be seasonal. What is the default action?
    Revert, then investigate off the critical path. Ambiguity is the common case, and the reason the default must be written down before the ramp is that everyone's judgement degrades once the change is in flight and effort is sunk. Holding at the current stage while watching is the worst option: it leaves two versions and two data shapes in production indefinitely and doubles the support surface without resolving anything.
  • What makes a rollout unreadable no matter how good the stopping rule is?
    Missing attribution and a missing baseline. If records and telemetry events do not carry the version that produced them, a moving metric cannot be tied to this change rather than to the several others riding the same release train. Without measurements taken before exposure there is nothing to compare against, since every criterion is relative. An unrehearsed revert path is the third gap: an unmeasured time-to-revert is unbounded damage after the criterion fires.
  • Which part of judging a staged rollout is not a release decision at all?
    Measuring whether a metric difference between exposed and unexposed groups is real. Choosing a randomisation unit, deciding how much data is needed, handling repeated looks at the data and attaching confidence to an effect estimate belong to experiment design. The release decision is coarser and faster: is this version harming anyone, and can we still revert. Conflating the two either delays a safety call waiting for certainty or claims a business result from a ramp never designed to measure one.

saying these in an interview costs you the question

  • Decides the abort criteria while the ramp is running
  • Equates blast radius with the exposure percentage
  • Advances a stage faster than the signal can report
  • Holds at an ambiguous stage instead of reverting
  • Never rehearses the revert path or times it
  • Ramps a version that writes data the old one cannot read

context