Your order API's rolling replacement started two new copies twenty minutes ago, has removed no old copy since, and both specs are still present — why?
answer
- count first, theorise second
- the removal guard is evaluating false
- unready copies buy no removals
- capacity intact, so users see nothing
- widening the gate turns a stall into an outage
basics
~20 sThe new copies are not reporting ready, so the guard that allows an old copy to be removed keeps evaluating false and the step never advances. The rollout is stuck on purpose, holding the old copies in service rather than trading them for copies that cannot serve.
solid answer
~50 sRead the counts before the theories. The desired copies are all still serving, two extra copies exist on the new spec, and no removal has happened — that is the signature of a readiness gate that has not opened. Each step may remove an old copy only while the number of copies able to serve stays at or above the floor, and copies answering not-ready count for nothing, so the guard is false and stays false. Nothing is wedged: the platform is refusing an exchange that would cost availability. Traffic is unaffected, because an unready copy receives none of it. The work is to find out why the new spec's copies answer not-ready — a missing configuration value, a dependency they cannot reach, a check whose target moved — and then either fix that or return to the previous spec, which is its own decision.
code
pseudocode · 17 linesdesired = 8 // copies that must be able to serve
maxMissing = 0 // copies of desired allowed to be absent
serving = count of copies answering readyCheck = true
// the guard evaluated before every single removal
canRemoveOne = (serving - 1) >= (desired - maxMissing)
// at the stall:
// 8 old copies answer ready, 2 new copies answer not-ready
// serving = 8
// canRemoveOne = (8 - 1) >= 8 -> 7 >= 8 -> false
// so no old copy is taken away, and the step stays where it is
// what would change it:
// a new copy starts answering ready -> serving = 9 -> 8 >= 8 -> true
// maxMissing is raised to 1 -> 7 >= 7 -> true, and capacity dropsgo deeper
Recognise the pattern: new copies exist, old copies are all still there, nothing has been removed. That means the rollout is waiting for the new copies to report ready.
Explain the guard — a removal is allowed only while enough copies can still serve — and say why copies answering not-ready contribute nothing to that count.
Diagnose from counts, keep capacity claims straight, and resist the loosened gate. Name the likely causes of a not-ready answer that is specific to the new spec while the old copies beside it serve normally.
Decide the policy: how long a stalled change may sit, who is allowed to widen a gate and under what evidence, and how stalls get noticed when no user-facing signal will ever fire.
## Read the counts before the theories The first move on a stalled rollout is arithmetic, not diagnosis. Three numbers settle what kind of stall this is: - **copies existing** — here, the desired count plus the two started; - **copies reporting ready** — here, only the old ones; - **copies removed so far** — here, none. Desired capacity intact, extra copies present, zero removals: that is a readiness gate that has not opened. It is a different picture from a platform that cannot place the new copies at all (in which case the extra copies would not exist) and from a rollout that has been given up on (in which case the extras would have been cleaned away). ## Why nothing moves Before each removal, the platform evaluates a guard of the form *would the number of copies able to serve still meet the floor if I took one away?* The floor is the desired count minus however many copies are permitted to be missing. Copies that answer not-ready are not in that number, however healthy their processes look. So with the two new copies silent, removing an old one would put the serving count below the floor, the guard evaluates false, and the step waits. It will keep waiting as long as the answer stays false — though platforms differ here, some abandoning the step after a deadline and reporting it failed, others waiting without limit. ## Who is actually serving This is where the description of a stall is most often stated backwards, so separate the two shapes: | stall shape | what is serving | what a caller sees | |---|---|---| | no batch has ever reported ready | only the old spec | one version's behaviour, unchanged | | earlier batches succeeded, a later one is stuck | both specs, in proportion to their ready copies | a mix, and any version-dependent behaviour with it | In the scenario above, the new copies have never reported ready, so despite two versions being *present* only the old one is *serving*. Saying "we are serving both versions" when nothing new is in the routing membership sends the next engineer looking for a mixed-version bug that cannot exist yet. ## The stall is the mechanism working The instinct is to treat a stuck rollout as a platform failure. Invert it: the platform was asked to guarantee a capacity floor while changing the spec, it was given copies that state they cannot serve, and it declined to trade working capacity for them. The alternative behaviour — carry on removing old copies regardless — is a self-inflicted outage delivered on a schedule. A stalled rollout is the cheapest possible outcome of a bad spec: full availability, a clear signal, and nothing lost but time. That also makes the tempting fix the dangerous one. Widening the gate — allowing more copies to be missing, or loosening what the check requires until the answers come back true — converts the stall into exactly the outage it prevented, because the underlying inability to serve has not changed. ## Working it 1. **Confirm the shape from counts** — serving count at the floor or above, extras present, removals zero. 2. **Ask the new copies' own answer** — the readiness answer is the platform's only input, so the question is what makes it false: configuration the new spec expects and does not get, a dependency it cannot reach from where it now runs, a credential, or a check whose target moved with the new spec. 3. **Check what is different about the new spec specifically**, not about the service generally — the old copies beside it are serving fine, which rules out most environment-wide causes for free. 4. **Decide fix-forward or go back.** Fixing the spec restarts the same gated walk; returning to the previous one is a decision with its own mechanics and its own risks. 5. **Do not relax the gate to make the rollout finish.** That is not a fix, it is a decision to remove the protection that is currently keeping the service up. ## What interviewers listen for The answer that lands says the guard is false, names the readiness answer as the input, states that capacity is intact, and calls the stall correct behaviour before proposing anything. A weaker answer reaches for the platform being broken, or proposes loosening the gate as step one. The strongest answers also notice that the stall is *silent* to users — there is no error-rate signal to alert on, so a rollout that stops moving has to be noticed by watching the rollout itself.
- What distinguishes this from a platform that is simply not acting on the change at all?The extra copies. A stalled gate has started the new copies and is holding at a stable count; a platform that never acted on the spec has created nothing, and the copy count equals the desired count exactly. The presence of unready extras is the evidence that the loop ran and then declined a removal.
- Does a stalled rollout page anyone?Not by itself, and that is the trap. Capacity is intact and error rates are normal, so user-facing alerting stays quiet. Noticing requires watching the change itself — how long a rollout has been outstanding, and whether its removal count has advanced — rather than waiting for a symptom that will not arrive.
- Someone proposes raising the number of copies allowed to be missing to get the rollout moving. What do you say?That it does not address why the new copies cannot serve, and it spends real capacity to keep replacing working copies with ones that answer not-ready. Each raise buys one more removal and moves the service one copy closer to having none serving. The stall is the diagnosis; the spec is the problem.
saying these in an interview costs you the question
- Calls the stall a platform bug rather than the gate holding.
- Loosens the readiness gate as the first remediation step.
- Claims users are being served by the stuck new copies.
- Assumes a stalled rollout shows up as a user-visible error rate.
- Says capacity has dropped, when the desired copies are all serving.