Your order API runs six copies behind one shared name; why does replacing all six at once drop requests, while replacing them a batch at a time does not?
answer
- capacity handover, not capacity gap
- the name is a directory, not a queue
- an empty membership list fails requests
- start the batch before removing the old
- the overlap in time is the whole trick
basics
~20 sStopping all six copies leaves the shared name with no copy able to serve until a replacement finishes starting, so every request in that window fails. Replacing a batch at a time keeps the rest serving throughout.
solid answer
~50 sThe copies sit behind a shared name, and the routing layer behind that name only hands requests to copies that report themselves ready. If you stop all six before starting any replacement, the membership list is empty for as long as the new copies take to become ready — not a restart, but the full start-up cost: process start, config load, connection pools, first-use caches. Every request arriving then fails. A rolling replacement instead starts a small batch on the new spec, waits until those copies report ready, and only then takes the matching old copies away, so at every instant some copies are serving. With two extra copies allowed above the desired six, the serving count never falls below six; if instead no extra copies are allowed and one may be missing, the floor is five — reduced capacity rather than an outage.
go deeper
Recall the shape: start the new copies, wait until they can serve, then remove the old ones. Downtime happens when nothing is in the routing membership list.
Explain that the gap equals the new spec's full start-up cost, not a process restart, and walk the serving count step by step for a stated copy count and batch size.
Show the cost side: extra copies for the duration or a lower capacity floor, plus a total rollout time equal to batches times start-up. Say which floor your service can actually survive.
Frame it as an availability budget: what capacity floor you require during any change, what you will pay in headroom to hold it, and which services are allowed the cheap all-at-once path.
## The set, and the name in front of it A service that has to stay available is not one process. It is a **set of copies** running the same spec, with a **shared name** in front of them. Callers address the name; a routing layer that the platform maintains keeps a **membership list** of the copies currently able to serve and hands each arriving request to one of them. Two properties of that arrangement carry the whole answer: - The name is a **directory, not a queue**. When the membership list is empty, requests are not parked until someone shows up — they fail or time out. - A copy joins the list only once it **reports itself ready**. A started process is not automatically a serving one. ## What "replace everything at once" really costs Stop all six copies, then start six copies on the new spec. Between the last stop and the first new copy reporting ready, the list is empty and the service is dark. Engineers underestimate that window because they price it as a restart. It is really the **full start-up cost of the new spec**, which usually includes: - process start and configuration load; - opening connection pools and any dependency handshake; - first-use caches, indexes or warm-up work the copy does before its first useful response; - whatever fixed delay sits between "the platform started it" and "it answers correctly". A second cost travels with it: the entire set moves to the new spec simultaneously, so if the new spec cannot serve, 100% of the service is already on it. ## What a rolling replacement changes One step of a rolling replacement is always the same shape, and the ordering is the point: 1. Start a **batch** of copies on the new spec — fewer than the whole set. 2. **Wait** until each of them reports itself ready. 3. Only then take away the matching number of old copies. 4. Repeat until every copy runs the new spec. Because step 3 is behind step 2, serving capacity is handed over rather than dropped and re-acquired. The change becomes a sequence of small overlaps instead of one gap. ## The arithmetic of a six-copy set Assume six desired copies, a batch of two, two extra copies allowed above the desired count while the change runs, and none of the six allowed to be missing: | moment | old copies | new copies ready | copies serving | |---|---|---|---| | before the change | 6 | 0 | 6 | | first batch started | 6 | 0 of 2 | 6 | | first batch reports ready | 6 | 2 | 8 | | two old copies stopped | 4 | 2 | 6 | | second batch ready, two old stopped | 2 | 4 | 6 | | third batch ready, two old stopped | 0 | 6 | 6 | The worst moment is six serving — the desired count — bought by briefly running eight copies. Change the setting so that no extra copies are allowed and one of the six may be missing, and the same walk floors at **five serving**: still no outage, but a sixth of the capacity is gone for the length of the rollout. Those two settings are the dial; what they trade against each other is a release-strategy subject of its own. ## What the rolling shape does not buy you - **It is not free.** You either pay for extra copies for the duration or serve on fewer than you sized for. - **It is slower.** Total wall-clock time is roughly the number of batches multiplied by the time a copy takes to become ready, so a small batch on a slow-starting service is a long change. - **Both versions exist for the duration**, which matters if they share anything that cares — a consequence owned by the mixed-version discussion, not by the mechanism here. - **It does not rescue a spec that cannot serve.** If the new copies never report ready, the rollout does not finish; it stops half-done with the old copies still serving. That is the design working, not failing. Platforms also differ in a detail worth knowing but not worth guessing at in an interview: some replace a copy on the host it was already running on, while others simply start the replacement wherever there is room and remove the old one afterwards. ## What interviewers listen for The strong answer names the membership list, says plainly that starting is not serving, and puts the removal of an old copy **after** the readiness of its replacement rather than beside it. The weak answer says "rolling updates mean zero downtime" with no mention of what the step waits for, and cannot say what the serving count is at the worst moment.
- With six desired copies, a batch of two and no extra copies allowed above six, how many are serving at the worst moment?Four. Two old copies must be stopped to make room for the batch before it can exist, so the set dips to four serving until the two replacements report ready. That setting is cheaper in resources and worse in capacity, and it only works if four copies can carry the traffic for the length of a step.
- Why does this argument need a shared name in front of the copies at all?Because the name is what lets a copy leave without the caller noticing. If clients held individual copy addresses, removing one would break the clients holding it, and no ordering of starts and stops would hide that. The rolling shape protects availability only for callers addressing the set rather than its members.
- Does a rolling replacement remove the disruption entirely, or just shrink it?It removes the whole-service gap. What remains is per-copy: requests already in flight on a copy that is being taken away still have to be finished or lost, which is what that copy's shutdown handling decides. The set stays available; individual in-flight work is a separate concern.
A shift handover where the outgoing worker leaves only once the replacement has actually taken over the desk, instead of everyone walking out and the next shift starting to get dressed.
saying these in an interview costs you the question
- Says a restart is quick, so nobody notices the gap.
- Assumes the shared name hides the gap so no request fails.
- Thinks the routing layer queues requests until a copy returns.
- Counts a copy as serving the moment its process starts.
- Calls a rolling replacement free, with no capacity or time cost.
- Believes a rolling change protects a spec that cannot serve at all.