skip to content

You set release policy for a service that runs in three regions with three availability zones each. After the initial canary passes, how would you order the rollout across those zones and regions, and how do you decide the total wall-clock time a rollout is allowed to take?

level: principalimportance: should knowfreq 40%

answer

  1. expand along a blast-radius ladder
  2. one zone, then one region, then serially
  3. keep a known-good region until last
  4. rollouts overlap when they run long
  5. define the express path before you need it

basics

~20 s

Expand along a blast-radius ladder: one zone in the least critical region, then that region, then the remaining regions one at a time, keeping a known-good region until last. Total duration is bounded by mixed-version tolerance, emergency-fix latency and how often you deploy.

solid answer

~1 min

The ordering principle is that each stage should be survivable and should teach you something the previous one could not. So: one zone in the least critical region, then the rest of that region, then the other regions serially — never two at once — with the largest or most revenue-critical region last. Keeping one region on the known-good version until the final stage is deliberate: it is the failover target if the new version misbehaves at scale. I also check the rollout against redundancy headroom — if the design tolerates the loss of one zone, a rollout must not degrade a second one at the same time. Duration is the harder question. It is bounded from above by how long the fleet can tolerate two versions serving simultaneously, and by how fast an emergency fix can reach production through the same pipeline, which means an express path defined in advance. It is bounded from below by the bakes needed to cross a diurnal cycle and any scheduled job. And there is a structural trap: if a rollout takes three days and you deploy daily, several rollouts overlap permanently, the fleet is never on one version, and attribution during an incident gets much harder.

go deeper

for a junior

Know that a change reaches zones and regions in stages rather than everywhere at once, and that the ordering exists to limit how many users a bad change can reach.

for a middle

Be able to describe the ladder — one zone, then one region, then the others serially — and explain why the largest or most critical region goes last.

for a senior

Show the operational reasoning: keeping a known-good region as a failover target, pausing the rollout on an ongoing zone impairment, and checking that the capacity out of service during a stage fits inside the redundancy headroom.

for a principal

Own the duration budget as a policy decision. Bound it by mixed-version tolerance, emergency-fix latency and deploy frequency; define the express path and who may authorise it; and delete stages that cannot say what they would catch that the previous one could not.

## The ladder, and what each rung is for A multi-region rollout is not one decision, it is a sequence, and each rung should add a *different* kind of exposure rather than just more of the same. 1. **A single zone in the least critical region.** The first rung adds real infrastructure diversity — its own network path, its own control-plane instances, its own hardware — that a traffic-percentage canary inside one zone never touched. Failure here is contained to a fraction of one region. 2. **The rest of that region.** Now the change is serving a full region's traffic mix, which for most services means every code path is exercised, and the region's shared dependencies (its database, its cache tier) are under the new version's full load pattern. 3. **The remaining regions, one at a time.** Serially, never in parallel. Two regions rolling simultaneously means a bad change can take both, and it also destroys the comparison — you have no untouched region to compare against. 4. **The largest, most revenue-critical, or most latency-sensitive region last.** This is the same reasoning as staging your biggest customer last: the region where a mistake costs the most gets the most evidence behind it. Order within that: some organisations start where the working day is quiet, others deliberately start where engineers are awake. The second is usually the better choice — a rollout that nobody is watching is a rollout whose abort is delayed by hours. ## Redundancy headroom is a hard constraint, not a preference If the deployment is designed to survive the loss of one zone (N+1) or two (N+2), the rollout consumes some of that same headroom while it runs: instances restarting are instances not serving. Rolling one zone while another is already impaired is how a routine release becomes an outage. Two rules follow. First, the rollout must check current health before advancing, not just the canary's metrics — an ongoing zone impairment should pause it. Second, the total capacity temporarily out of service during a stage must be smaller than the headroom the design claims. If it is not, either the rollout is too aggressive or the redundancy claim was fiction. ## Keeping a known-good region The most valuable property of serial regional rollout is that until the last stage, there is somewhere to send traffic that is running code known to work at production scale. That is qualitatively better than a rollback: failing over to an untouched region is a traffic decision measured in minutes, whereas rolling a version back across a fleet is a deployment operation measured in tens of minutes and carries its own risk. It is only true if the untouched region has, or can acquire, the capacity to absorb the shifted traffic — which is a capacity-planning commitment, and one that quietly evaporates when a cost review trims regional headroom. ## Deciding total duration The bakes set the floor. Three things set the ceiling. **Mixed-version tolerance.** For the whole rollout, two versions are live simultaneously and must interoperate — wire formats, shared cache entries, queue messages, database rows one version writes and the other reads. Every hour of rollout is an hour those constraints must hold, and a longer window means rarer cross-version interactions get exercised. If the change alters anything shared, the rollout should be *shorter*, not longer, or the change should be split so the shared part ships separately. **Emergency-fix latency.** How long until an urgent fix can reach every user? If the answer is "the length of a full rollout", then your worst-case time to mitigate an unrelated bug is that number. You need a documented express path — which stages may be skipped, who authorises it, what minimum verification still runs — and it must be exercised occasionally, because an untested route is not a control. **Deploy frequency.** If a rollout takes longer than the interval between deploys, rollouts overlap permanently. The consequences are concrete: the fleet is never uniform, so "which version is that host on" becomes a lookup rather than a fact; an incident has several changes mid-flight to attribute; and pausing one rollout does not stop the others. When rollout duration and deploy interval collide, something has to give — fewer stages, higher traffic shares with shorter bakes, batching changes per rollout (which raises per-rollout risk), or deploying less often (which raises per-change risk). ## The tradeoff to state out loud The pressure is always toward faster: engineers want their change out, and a slow pipeline is a visible cost while a prevented outage is invisible. The counter-argument is not "safety matters"; it is a specific claim about what each stage buys. If a stage cannot articulate what it would catch that the previous one could not, it is duration without value and should be deleted — that is how you buy speed honestly. And the numbers should be reviewed against evidence: how many rollouts were aborted at each rung, and at which rung the escapes were caught. A ladder whose middle rungs have never caught anything in a year is telling you something.

  • Why roll regions serially rather than pushing every region to 1% at the same time?
    Parallel regions expose every region's users at once and remove the untouched comparison group, so a subtle regression has no clean control and a control-plane-level mistake lands everywhere simultaneously. Serial rollout keeps at least one region on known-good code, which is a failover target you can use in minutes rather than a rollback you must execute across the fleet.
  • How does a rollout interact with the redundancy the service claims?
    Instances restarting during a rollout consume the same headroom that a zone failure would. If the design tolerates losing one zone, a stage that takes meaningful capacity out of a second one erases the margin. The rollout must therefore check live health before advancing and pause on an ongoing impairment, and each stage's out-of-service capacity must stay inside the claimed headroom.
  • Your rollout takes three days and the team deploys daily. What breaks?
    Rollouts overlap permanently. The fleet is never on a single version, so any incident has several in-flight changes to attribute and pausing one rollout does not stop the others. The resolutions are all trades: fewer stages, larger traffic shares with shorter bakes, batching more change per rollout, or deploying less often. Pick one deliberately rather than letting the overlap happen by default.
  • How do you decide a rollout stage is not earning its place?
    Ask what it would catch that the previous stage could not, and check the record: how many rollouts that stage has aborted, and whether escapes were caught earlier or later. A rung that has never fired in a year is adding duration without evidence, and deleting it is how you buy pipeline speed without pretending the risk changed.

saying these in an interview costs you the question

  • Rolls all regions in parallel to finish faster
  • Leaves no region on the known-good version
  • Ignores that rolling instances consume failover headroom
  • Has no defined express path for an urgent fix
  • Never notices that slow rollouts overlap each other

context