skip to content

Every instance of a service fails its readiness check at once because a shared downstream store is slow — what does the platform do?

level: seniorimportance: should knowfreq 38%

answer

  1. everyone fails the same check together
  2. unrouting only helps a per-instance fault
  3. no healthy peer left to move traffic to
  4. a degradation became a total outage
  5. far worse when the restarting verdict agrees

basics

~20 s

It empties the routing set — every instance leaves at once, so even requests that never touch the slow store now fail, turning a partial degradation into a total outage. The processes keep running, unless the liveness check reads the same signal too.

solid answer

~40 s

The platform does exactly what it was told: each instance reports itself unable to serve, and each is removed from the routing set. Because the cause is shared, they all leave within one check interval and the routing set empties. Traffic that would have succeeded — anything not touching the slow store — now fails as well, so the blast radius is larger than the dependency's own failure. The instances themselves are untouched and rejoin as soon as the store recovers. The deeper point is that unrouting is a **reallocation** mechanism: it helps only when the fault is per-instance, because its whole effect is moving traffic to healthier peers. With a correlated fault there are no healthier peers, so the verdict has nowhere to send anything and only subtracts capacity.

go deeper

for a junior

Recall that a readiness failure removes an instance from the routing set and nothing more, and that if every instance fails at once there is no instance left to receive traffic.

for a middle

Explain why the blast radius grows: requests that never touch the slow dependency now fail too, because the address they arrive at has no instances behind it.

for a senior

Show the diagnosis and the rule you would take away: unrouting is reallocation and only works on per-instance faults, and the restarting verdict must never depend on anything remote.

for a principal

Frame the standard for a whole estate: decide which signals are allowed to influence which verdict, and make the correlated-failure case an explicit review question rather than a lesson each team learns during its own outage.

## What the platform actually does Nothing clever, and that is the point. Each instance is evaluated independently: 1. Each instance's readiness check fails, for its own reasons, which happen to be the same reason. 2. Each accumulates consecutive failures until its threshold is reached. 3. Each is removed from the routing set — the set of instances the service address is allowed to hand traffic to. 4. Within roughly one interval plus one threshold's worth of attempts, the set is empty. 5. No process is ended. Every instance keeps running, keeps its memory and keeps its slot on a node. 6. When the store recovers, each check passes again and each instance rejoins within an interval. ## Why the blast radius exceeds the dependency's own failure A slow downstream store degrades the requests that need it. An empty routing set fails **every** request, including: - endpoints that never touch the store at all; - cached responses the instances could still have served from memory; - the health and diagnostic surfaces operators reach through the same address; - anything that would have returned a useful partial answer. The platform has amplified a partial degradation into a complete one, and it did so by following instructions. This is the single most important thing to say in an interview: the outage was *caused by the health mechanism*, not merely revealed by it. ## Unrouting is a reallocation mechanism The routing verdict has exactly one effect: move traffic away from this instance toward the others. That is valuable when the fault is local — a bad node, a lost connection this instance alone holds, a queue only this instance has filled — because the traffic has somewhere better to go. When the fault is shared by every instance, nothing about that is true. The useful test before letting anything influence the routing verdict: > *If traffic moved to a different instance of this service, would it be served?* | fault | correlated? | does unrouting help? | |---|---|---| | this instance lost its connection to a partition it owns | no | yes — peers still serve | | this instance's local queue is saturated | no | yes — peers absorb it | | this instance is still loading at start-up | no | yes — peers are already serving | | a store every instance reads is slow | yes | no — there is nowhere to move traffic to | | a shared credential expired fleet-wide | yes | no — every peer fails identically | If the answer is no, the condition is not a routing verdict. It belongs to how individual requests are answered: fail that request, serve a degraded response, or serve from what is in memory — decisions made per request, by the instance, not by the platform about the instance. ## When the restarting verdict reads the same signal This is where an outage becomes an incident that does not end on its own. If the liveness check consults the same slow store, then every instance in the fleet is ended at nearly the same moment. The consequences compound: - every warm cache and loaded index in the fleet is destroyed simultaneously; - every instance restarts together and reconnects together, hitting the struggling store with a synchronised burst far heavier than the steady-state load it was already failing to serve; - if start-up itself needs the store, nothing completes start-up, and the fleet cannot come back until the dependency recovers on its own. The rule that follows is blunt and worth stating as a rule: **the restarting verdict should depend on nothing remote.** It should answer only whether this process is still capable of doing its own work. ## Designs differ on an empty routing set Platforms do not agree on what to do when a service has no healthy instances left. Some simply serve nothing, which is the behaviour described above. Others treat 'everything is unhealthy' as evidence that the health signal is wrong rather than the fleet, and fall back to sending traffic to all instances — on the theory that a degraded answer beats none. Knowing which behaviour you have is an operational fact worth establishing before an incident, not during one. ## Reading the incident afterwards The signature is distinctive: instances leave the routing set at nearly the same wall-clock moment, with no deployment in flight; each instance's own logs show it healthy and idle; and traffic that avoids the slow dependency was failing too. That last detail is the one that separates this from a genuine fleet-wide fault, and it is what points the review at the check rather than at the service.

  • What changes if the liveness check reads the same slow dependency?
    The fleet is ended, not just unrouted. Every instance restarts at once, every warm cache is lost, and they all reconnect simultaneously, hitting the struggling store harder than the load it was already failing. If start-up itself needs the store, nothing comes back until the store does.
  • When is it legitimate for a readiness check to depend on something downstream?
    When the dependency is genuinely per-instance — a connection this instance holds, a partition it owns, a lease it lost — because then moving traffic to a peer actually helps. The test is whether another instance of the same service would serve the request.
  • How would you tell this apart from a bad release rolling out across the fleet?
    By timing and by what is still healthy. A release moves through instances in batches over minutes and correlates with a change; this fails everything at nearly the same instant with nothing deployed, and each instance's own logs show it idle and well while requests that avoid the dependency were failing too.

saying these in an interview costs you the question

  • Says the platform will restart the instances to clear the condition
  • Thinks emptying the routing set is safe because instances come back
  • Says a correlated failure is solved by adding more replicas
  • Treats a slow shared dependency as evidence that each instance is unhealthy
  • Assumes traffic is shifted to the instances that are still healthy
  • Believes the routing verdict shields users from the dependency's latency