skip to content

Your rolling replacement of an order API completed within seconds and the service went dark, because every new copy reported ready the moment its process started — what happened?

level: seniorimportance: should knowfreq 38%

answer

  1. the gate is worth what its input is worth
  2. ready answered as if it meant alive
  3. every step passed with no wait
  4. success reported, capacity gone
  5. self-recovery points at the signal, not the spec

basics

~20 s

Every step's gate opened immediately, so old copies were removed at full speed and the whole set ended up on copies that could not yet serve. The batch structure survived; the property it protects did not, because the ready answer was true too early.

solid answer

~40 s

The rolling walk only protects availability because the removal of an old copy waits on a truthful statement that the replacement can serve. Here the answer was really an is-the-process-alive answer, which is true the instant the process exists. Each step therefore passed its gate with no wait at all, the removals chained back to back, and within seconds every serving copy had been traded for one still loading configuration and opening its pools. That is the same outage as replacing everything at once — reached through a rollout that reported success, which is worse, because there is no stalled step for anyone to notice and the only signal is the error rate. If the answer had been honest the rollout would have stalled after the first batch with the old copies still serving.

go deeper

for a junior

Take away the rule: a copy saying ready must mean it can answer a request. If it says ready when it has merely started, the rollout removes working copies for useless ones.

for a middle

Explain that the gate consumes an answer rather than inspecting the copy, and trace the serving count down to zero while the platform believes it is at the desired number.

for a senior

Distinguish this from a stall, use the unaided recovery as the diagnostic, and recognise that rolling back the spec leaves the real defect in place for the next change.

for a principal

Set the expectation organisation-wide: define what a ready answer must mean before automated rollouts are permitted, and treat a too-early answer as a fleet-wide latent outage rather than one service's incident.

## What the gate consumes A rolling replacement makes exactly one safety claim: *a copy that can serve is removed only after a copy that can serve replaces it.* The platform cannot verify the second half itself. It reads the answer the copy publishes and treats it as true. So the rollout's entire guarantee is inherited from that answer — the batch sizes, the surge allowance and the missing-copy budget are all arithmetic performed on top of it. ## Ready against alive The two questions a copy can be asked look similar and mean opposite things for a rollout: | the answer states | becomes true when | what a false answer causes | what it is worth to a rollout | |---|---|---|---| | the process is alive | the process exists | the copy is restarted | nothing — it is true before any start-up work is done | | the copy can serve a request | start-up work is genuinely finished | the copy is taken out of routing and left running | everything — it is the gate's only input | When the readiness answer is wired to the first row's fact, the label still says ready and the meaning is alive. That substitution passes every structural check: the spec is valid, the gate exists, the batches are small, the rollout reports success. ## Trace the outage With six desired copies, a batch of two and none allowed missing: 1. Two copies start on the new spec and immediately answer ready. Serving count reads eight; two of those eight cannot answer a request. 2. The guard passes, two old copies are removed. Real serving capacity is four; the platform believes it is six. 3. The next batch starts, answers ready instantly, two more old copies go. Real capacity two. 4. The final batch repeats it. Real capacity zero, believed six, rollout complete. 5. Requests fail until the new copies finish their start-up work — and if the failures make them crash or restart, the recovery stretches further. The structure did everything it was told. Each removal was preceded by a ready answer; the answers were simply worthless. ## Why this is worse than a stall A rollout stopped by an honest not-ready answer is the good failure: capacity intact, a clear signal, time lost and nothing else. This is the bad one, for three reasons: - **It reports success.** There is no outstanding step for anyone to look at, and the change has already left the deploy dashboard. - **It is discovered by error rate**, which means users found it first. - **It repeats.** Every service sharing that pattern for its ready answer has the same latent outage, and it will not fire until the next rollout. A quick tell in the aftermath: the service recovers on its own a minute or two later, with no intervention. That means the new spec was fine and the copies could serve once warm; what was wrong was the moment they were declared serviceable. Roll the spec back and you will find nothing wrong with it. ## The general lesson The rolling structure is a *consumer* of a signal, not a source of safety by itself. Any gate is worth exactly what its input is worth, and the failure mode of a gate fed a too-early truth is silent and complete. Whether a given copy's answer is a good one is decided where that check is defined, not in the rollout settings; what the rollout owes you is that it will not remove capacity without one. Two designs exist in the wild for the timing problem underneath this — some platforms offer a separate allowance that holds the other checks off during a long start-up, while others expect the readiness answer alone to carry it — and the concept to hold onto is that start-up time is a distinct thing from ongoing health. ## What interviewers listen for The strong answer separates the mechanism from its input immediately: the rollout is not broken, its signal was. It states the consequence of a too-early truth in counts, contrasts it with the stall, and notes that the self-recovery is the diagnostic. A weak answer blames the batch size, proposes slowing the rollout down, or concludes that rolling replacement does not prevent downtime after all.

  • How would the same rollout have behaved if the ready answer had been truthful?
    It would have paused after the first batch with every old copy still serving, and stayed there until the new copies finished their start-up work — then continued normally. The visible outcome is a slow rollout instead of an outage, which is the trade the gate exists to make.
  • The service recovered by itself ninety seconds later. What does that tell you?
    That the new spec was capable of serving all along and the defect is the moment it was declared serviceable. It also means rolling back would have hidden the bug: the previous spec would come up, look fine, and leave the same too-early answer in place for the next change.
  • Would a smaller batch size have prevented this?
    No, only slowed it. With every gate opening instantly, a batch of one still chains removals back to back and empties the set; it just takes marginally longer and reduces capacity in smaller increments. Nothing about the batch arithmetic fixes a signal that is true too early.

A handover where the outgoing worker leaves as soon as the replacement walks through the door, rather than once they have actually taken over the desk.

saying these in an interview costs you the question

  • Concludes that rolling replacement does not prevent downtime.
  • Blames the batch size rather than the signal feeding the gate.
  • Treats an alive answer and a can-serve answer as interchangeable.
  • Says the platform should have verified the new copies itself.
  • Rolls the spec back and calls the incident closed.