skip to content

In a rolling replacement, why is each step gated on the new copy reporting itself ready rather than on its process having started?

level: middleimportance: must knowfreq 63%

answer

  1. started is not serving
  2. the step trades away real capacity
  3. readiness decides routing membership
  4. an unready copy pays for no removal
  5. a timer is a guess, wrong both ways

basics

~20 s

A started process is not yet a serving one, and the step is trading away a copy that can serve. Gating on the readiness answer keeps the exchange honest; gating on process start removes serving capacity in exchange for capacity that does not exist yet.

solid answer

~50 s

Each step of a rolling replacement is an exchange: one old copy that is serving traffic for one new copy that is supposed to. The platform cannot see inside the new copy, so it uses the one signal the copy publishes about itself — whether it reports ready — to decide the exchange is safe. Process start is the wrong signal because the gap between `started` and `can answer a request` is exactly where config loading, connection pools and warm-up live, and that gap is often seconds to minutes. A fixed timer is the wrong signal too: it is a guess that is too short on a bad day and wasted time on a good one. The readiness answer is also what controls routing membership, so a copy that has not reported ready is receiving nothing — removing an old copy for it is a straight capacity loss.

code

pseudocode · 21 lines
pseudocode
desired      = 6   // copies that must be able to serve
batchSize    = 2   // copies replaced per step
extraAllowed = 2   // copies allowed above desired during the change
maxMissing   = 0   // copies of desired allowed to be absent

oldCopies = desired   // 6
newReady  = 0

while newReady < desired:
    start batchSize copies from the new spec
    // copies existing now: oldCopies + newReady + batchSize = 8,
    // which is desired + extraAllowed, so the start is permitted

    wait until every started copy answers readyCheck
    // while any of them answers false, the loop stays here
    // and no old copy is taken away
    newReady = newReady + batchSize        // 2, then 4, then 6

    if (oldCopies - batchSize) + newReady >= desired - maxMissing:
        stop batchSize old copies          // 8 existing -> 6 serving
        oldCopies = oldCopies - batchSize  // 4, then 2, then 0

go deeper

for a junior

Remember the order inside one step: start the new copy, wait for it to say it can serve, then remove an old one. Waiting for the process to exist is not enough.

for a middle

Explain the gap between started and serving — config load, pools, warm-up — and say why the readiness answer, not a timer, is what the removal guard counts.

for a senior

Show how the guard's arithmetic uses the serving count and the missing-copy budget, and say what happens to a rollout whose ready answer is true too early or never true at all.

for a principal

Treat the readiness answer as a contract every service owes the platform, and decide what you require of it before a team is allowed an automated rollout at all.

## The step is an exchange, not an addition A rolling replacement is a sequence of exchanges. Each step adds copies on the new spec and then removes the same number of old copies. The removal is the dangerous half: the copy being taken away is currently answering real requests, and the copy replacing it is only *supposed* to. Everything about the gate follows from wanting that exchange to be honest. ## Three signals the step could wait on | signal it could gate on | what it actually tells you | how it fails | |---|---|---| | the process has started | the platform successfully launched something | says nothing about whether a request would be answered; the start-up gap is invisible | | a fixed timer has expired | someone's estimate of start-up cost has elapsed | too short on a slow day, wasted time on a fast one, and never adapts to a changed spec | | the copy reports itself ready | the copy's own statement that it can serve | only as good as what that answer is derived from | The third is the only one that is a statement *about serving*, which is why rolling replacement is built on it, and why the quality of the answer becomes the whole risk. ## Why "started" is so misleading Between launch and first useful response, a copy typically does several things that take real time: - reads its configuration and refuses to serve if something required is missing; - opens connection pools to its dependencies and completes any handshake; - loads or rebuilds anything it keeps in memory — caches, indexes, compiled artefacts; - registers itself with whatever it needs to register with. During all of that the process exists. It has a state the platform can observe as running. It cannot answer a request. A step gated on process start therefore removes a serving copy in exchange for a copy that will start serving at some unknown later time — the same gap as replacing everything at once, just spread out in slices. ## The invariant the gate protects State the invariant in counts and the design stops being abstract. With a desired count `D` and a budget of `M` copies allowed to be missing, the platform will not remove a copy if doing so would leave fewer than `D - M` copies **able to serve**. The chain is: 1. Readiness membership decides which copies receive requests. 2. The serving count is the size of that membership, not the number of running processes. 3. A removal is permitted only while the serving count stays at or above the floor. 4. So a new copy that has not reported ready contributes nothing to the count, and cannot pay for a removal. That last line is the answer in one sentence: an unready copy is not currency the step is allowed to spend. ## Where the ready answer comes from The copy answers a check the platform periodically asks it, and the platform's only input is that answer. Designing the check — what it should test, how often, how many consecutive answers count — is a separate subject in its own right. What belongs here is the consequence: **the rollout is exactly as trustworthy as that answer is.** If the answer becomes true before the copy can serve, every step passes instantly and the rolling structure protects nothing. If the answer never becomes true, the rollout stops part-way with the old copies still serving. Both outcomes are the gate doing precisely what it was told. Platforms differ in one detail here: some stop waiting after a deadline and report the step as failed, while others will wait on a not-ready batch indefinitely. Do not assume either behaviour without checking the one in front of you. ## What interviewers listen for The strong answer connects three things: readiness controls routing membership, membership is what the removal guard counts, and therefore the gate is about capacity rather than about politeness. Naming the fixed-timer alternative and why it is worse is a good sign. The weak answer says "it waits for the container to be healthy" without distinguishing running from serving, or claims the platform inspects the new copy itself rather than reading the answer the copy publishes.

  • While a new copy still answers not-ready, what share of traffic is it taking?
    None. The readiness answer is what puts a copy into the routing membership, so an unready copy is addressed by nothing. That is precisely why it cannot pay for the removal of an old copy: the exchange would reduce the number of copies actually receiving requests.
  • Why not just wait a generous fixed period per step instead of reading an answer?
    Because the period encodes a guess about start-up cost and is wrong in both directions. Too short and the step removes an old copy before the replacement can serve; too long and every rollout pays the worst case on every step. It also silently goes stale when the spec's start-up work changes.
  • The new copy reports ready, the old one is removed, and errors appear anyway. Where do you look first?
    At what the ready answer is derived from. The gate behaved correctly given the answer it read, so either the answer became true before the copy could serve, or the copy could serve the check's path but not the traffic's. The rolling structure is intact; the signal feeding it is not.

saying these in an interview costs you the question

  • Treats a running process as a serving copy.
  • Says the platform inspects the new copy rather than reading its answer.
  • Thinks a fixed wait per step is equivalent to a readiness gate.
  • Believes an unready copy still takes a share of traffic.
  • Says the gate exists to be polite rather than to hold a capacity floor.