A network fault makes twenty hosts stop reporting at once - what is the risk if the platform replaces everything they were running?
answer
- correlated, not independent
- the observer may be the fault
- no room for the whole fleet
- cold starts arrive together
- paced write-off, then a threshold
basics
~20 sMass replacement can turn an observation problem into a real outage: the surviving hosts cannot hold the whole fleet, every replacement pays a cold start at once, and if the hosts were never down, the cluster has just doubled its running workloads.
solid answer
~40 sReplacing one unreachable host's workloads is routine. Replacing twenty hosts' worth simultaneously is different in kind. The replacements compete for whatever unreserved capacity is left, so some of them stay unplaced while the workloads that did fit start together - a burst of image fetches, cold caches and connection storms aimed at dependencies that are also degraded. And the correlated silence is itself evidence: twenty hosts failing independently at the same instant is far less likely than one broken network path, or a problem on the observing side. Platforms therefore rate-limit how fast workloads on unreachable hosts are written off, and typically stop doing it altogether once the unreachable fraction of a failure domain crosses a threshold - treating widespread silence as a reason to doubt the observer rather than to act on it.
go deeper
The takeaway is that many hosts going quiet together usually means one shared cause, and that replacing all their work at once can be worse than waiting.
Explain the capacity arithmetic and the synchronised cold start, and why replacement of correlated losses is not just more of the single-host case.
Reason from the evidence: correlated silence points at the observing path, so name the paced write-off and the unreachable-fraction threshold and what each is protecting.
Decide how much of the cluster must survive a domain going silent, and what you are willing to pay in idle capacity for a mass failure to be absorbed rather than merely detected.
## Why correlated silence is a different problem The unreachable-node mechanism is designed around independent failures: one machine, one power supply, one disk. Its arithmetic works because a cluster sized to survive losing a host has somewhere to put that host's workloads. Correlated silence breaks both halves of that assumption at once. The capacity to absorb the loss may not exist, and the *cause* is probably not what the mechanism assumes it is. Twenty hosts going quiet in the same second has a short list of plausible explanations, and "twenty independent machine failures" is at the bottom of it. Far more likely: a network segment lost its path to the control plane; a switch or a routing change cut a rack, a row or an availability zone; or something on the observing side - the state store, the API layer, the component that processes reports - is degraded and is not recording reports that arrived normally. In most of those worlds the workloads are still running and still serving their clients. ## What mass replacement costs If the platform writes off all twenty hosts' workloads and asks for replacements: - **Placement pressure.** Every replacement carries the same reservations as the original. The remaining hosts have to fit all of them into their unreserved capacity, and whatever does not fit sits waiting - so recovery is partial and uneven rather than complete. - **A synchronised cold start.** Hundreds of instances start within seconds of each other. Each may fetch an image, open connections, and fill caches. The fetch traffic alone can saturate links that are already the suspect in this incident. - **A thundering herd on dependencies.** Cold instances hit shared dependencies - the data store, the credential service, the upstream the ingesters read from - all at once, at exactly the moment those dependencies may also be impaired. - **Doubling, not moving.** If the hosts were never down, the originals are still running. Now the cluster runs two of everything, with two sets of consumers, two sets of writers and twice the load on everything downstream. - **Eviction pressure on the survivors.** The hosts that absorbed the replacements are now much closer to their limits, so an ordinary spike can start reclaiming memory from workloads that were fine before the incident. The outcome is the failure mode worth naming: **an observation failure converted into a workload failure by the recovery mechanism itself.** ## What platforms do about it Two mitigations are near-universal in shape, if not in detail: 1. **Rate-limiting the write-off.** Rather than declaring every unreachable host's workloads gone the moment their timeouts expire, the platform paces it - so many hosts per interval. A genuinely dead rack still recovers, just more slowly; a transient fault often resolves before most of the hosts have been processed at all. 2. **A threshold on the unreachable fraction.** When the share of unreachable hosts within a failure domain, or across the cluster, exceeds some fraction, the platform slows or stops replacing altogether. The reasoning is explicit: at that scale the likeliest broken component is the path by which the platform observes, and acting on a broken observation is worse than not acting. Both are the same principle stated at different strengths: **the more correlated the evidence, the less it supports the conclusion the mechanism was built on.** A lone silent host is probably a dead host. A silent half of the cluster is probably not half a cluster of dead hosts. ## Operating on the other side of it This behaviour surprises people during an incident, because it looks like the platform has stopped helping. The right reading is the opposite: it has stopped guessing. Practically, that means a few things for whoever is holding the incident: - Check whether the hosts are actually unreachable *from the clients*, not just from the control plane. If clients are still being served, workload replacement is not the thing to be chasing. - Expect recovery to be paced when it does begin, and do not force it faster by hand while the cause is unknown - the pacing is what is keeping the survivors alive. - Remember that spreading copies across failure domains is what makes the threshold behaviour survivable in the first place; the domains have their own mechanisms and their own owners. And one thing does not change: every one of those replaced copies is a fresh instance from the same spec. If the twenty hosts really did die, whatever those workloads held locally died with them.
- Why would a platform ever stop replacing when many hosts are unreachable at once - is that not exactly when replacement matters?It is exactly when the evidence is least trustworthy. Widespread simultaneous silence points at the path by which the platform observes, not at the machines, and replacing on a false premise doubles the running fleet and floods dependencies. Stopping preserves the workloads that are almost certainly still serving.
- What does spreading copies across failure domains change about this scenario?It decides whether one segment going silent is survivable at all. With copies concentrated, a domain-wide fault takes the whole workload and mass replacement becomes the only option; spread, some copies keep serving throughout, which is what makes a paced or halted write-off tolerable. How that spreading is expressed is a placement concern.
- The twenty hosts genuinely died. How does paced replacement play out then?Recovery is complete but staged: hosts are written off in batches, so replacements start in waves rather than all at once, spreading image fetches and cold-start load over minutes. Total recovery takes longer than an unthrottled burst would in theory, and usually less time in practice, because nothing downstream is knocked over on the way.
saying these in an interview costs you the question
- Treats twenty simultaneously silent hosts as twenty independent failures
- Assumes the surviving hosts can absorb the whole fleet
- Says faster mass replacement is always the safer choice
- Ignores that the observing path, not the hosts, may be broken
- Forgets that every replacement pays its start-up cost at the same moment