skip to content

During a traffic spike every instance in a backend pool slows down, health checks begin timing out, and the load balancer ejects instances one after another until almost nothing is left in rotation — turning a slowdown into a total outage. Explain the feedback loop, and what safeguards stop health checking from causing this.

level: seniorimportance: should knowfreq 47%

answer

  1. ejection concentrates load, it never removes it
  2. positive feedback: the fix feeds the cause
  3. failing on timeout, not on error
  4. common-mode failure, not independent failure
  5. fail open below a healthy floor

basics

~20 s

Overload makes checks time out, ejection moves that instance's traffic onto the survivors, the survivors slow further and are ejected in turn. The loop is closed because ejection increases the load that caused it. Safeguards: fail open below a healthy floor, cap the ejectable fraction, and never eject on latency alone.

solid answer

~60 s

The loop closes because ejection does not remove work, it concentrates it. Under load, response times cross the check timeout, so a probe that would have succeeded fails; the instance is ejected; its share of traffic redistributes onto instances that were already at the edge; they cross the timeout too, and the pool eats itself. The tell is that the checks are failing on *latency*, not on errors — a health check with a tight timeout is really a latency alarm wired to a capacity switch. Defences come in layers: fail open when the healthy fraction drops below a floor, because a pool that is entirely unhealthy almost always means the check is wrong rather than every instance dying at once; cap the fraction of the pool that may be ejected at any moment; give the probe generous slack over normal latency so it tests reachability rather than speed; prefer relative outlier detection, which ejects nobody when everyone degrades equally; and shed or throttle load at the edge, since the real problem is that demand exceeded capacity and no routing decision fixes that.

go deeper

for a junior

Know the basic mechanism: taking an instance out of rotation sends its requests to the remaining instances, so ejecting under load makes the survivors busier, not calmer.

for a middle

Explain why a probe that fails by timing out is a different signal from one that fails with an error, and why the timeout version couples health to load.

for a senior

Show the whole loop and name the safeguards you would have had in place — fail-open floor, ejection cap, generous probe timeout, relative outlier detection — and how you would confirm the cause afterwards.

for a principal

Own the position that under common-mode failure a load balancer should keep serving from unhealthy hosts, and set the platform defaults, headroom targets and shedding policy that make that stance survivable.

## The loop, stated precisely Health checking assumes that removing a bad instance improves things. That assumption holds when the instance is *individually* broken and fails silently. It inverts when the cause is *shared* load: 1. Demand rises; every instance queues and its response times climb. 2. The health probe's response time climbs with everything else and crosses the check timeout. A timed-out probe is a failed probe. 3. Consecutive failures accumulate and the instance is ejected. 4. Its traffic is redistributed across the remaining instances, so each survivor's load rises by roughly `1/(n−1)`. 5. Go to step 2, now faster — the increment compounds, so ejections accelerate. The system has positive feedback: the corrective action amplifies the disturbance. And it terminates in the worst possible state, an empty pool, where the load balancer has nothing to route to and returns errors for 100% of traffic — a strictly worse outcome than every instance being slow. ## The diagnostic tell When you review the incident, ask what the checks were failing *on*. Two very different answers: - **Connection refused, reset, or a 5xx from the check path.** The instance really is broken. Ejection was correct. - **Timeout.** The instance answered, just not fast enough. The check did not measure health; it measured queue depth. With a tight timeout, a health check is a latency alarm wired to a capacity switch, and you built an amplifier. That distinction is the heart of a good answer. Many proxies let you grade these differently, and a check timeout set just above the p50 makes a routine load spike indistinguishable from a dead host. ## Safeguards, roughly in order of value **Fail open below a healthy floor.** The single most important one. When the fraction of healthy hosts drops below a configured percentage, the proxy ignores health status entirely and balances across every host in the pool. The reasoning is Bayesian: it is far more likely that your check is wrong than that every instance died in the same second, and serving from a degraded backend beats serving nothing. Several proxies implement exactly this and it is on by default in at least one of them. **Cap the ejectable fraction.** Independently, limit how much of the pool passive ejection may remove at once — commonly a low double-digit percentage. This bounds the damage a mistaken verdict can do, and it breaks the loop at step 4 by refusing to concentrate load past a point. **Give the probe real slack.** The check timeout should be set against the failure you want to catch (a hung or unreachable instance) and not against normal latency. Seconds of slack, not milliseconds. Related: keep the probe path cheap, and do not let it queue behind expensive application work, or the probe becomes a load meter by construction. **Prefer relative outlier detection over absolute thresholds.** An outlier is a host much worse than its peers. Under a shared cause every host degrades together, so a relative rule correctly ejects nobody, while an absolute error-rate or latency threshold ejects everybody. This is the mechanism that distinguishes "this instance is bad" from "today is bad". **Do not chain checks to dependencies you share.** If the probe path calls the common database, a database blip fails every probe at once and you have wired a dependency outage directly into a pool-wide ejection. This is one of the sharpest self-inflicted versions of the loop. **Fix the actual problem: demand exceeded capacity.** Ejection is a routing decision, and no routing decision creates capacity. The instruments that do help are admission control and load shedding at the edge (reject or degrade a fraction of requests fast, so the rest are served well), concurrency limits per backend so queues stay bounded, and enough headroom that losing an instance does not push survivors over the line. Watch the retry interaction too: a client that retries on timeout multiplies demand at exactly the moment the pool is shrinking, which is the same loop running one layer up. ## What separates a senior answer Junior answers describe the ejection. Middle answers describe the redistribution. The senior move is to name the *inversion*: health checking assumes independent failures, the overload case is a common-mode failure, and every safeguard above is a way of detecting or tolerating common mode. Say that, then say the uncomfortable corollary — under a common-mode failure the correct behaviour of a load balancer is to keep sending traffic to instances it believes are unhealthy.

  • Why is failing open — routing to hosts marked unhealthy — the right behaviour when almost the whole pool fails its checks?
    Because the prior favours the check being wrong. Every instance failing simultaneously is far more consistent with a bad probe, a shared dependency, or a load-induced timeout than with genuinely simultaneous deaths. And the alternative is an empty pool, which fails 100% of requests — strictly worse than a degraded backend that still serves most of them.
  • How would you tell, after the incident, whether the ejections were justified?
    Look at how the probes failed. Connection refused, resets, or 5xx from the check path mean the instances were genuinely broken and ejection helped. Timeouts mean the instances answered too slowly, so the check was measuring queue depth. Correlate the ejection timestamps with request latency and CPU: if latency crossed the check timeout just before the first ejection, the check caused the outage.
  • A team suggests preventing this by health-checking a path that queries the shared database, so checks fail sooner. What is wrong with that?
    It converts a shared dependency into a synchronised pool-wide ejection. When the database hiccups, every instance fails its probe in the same second and the entire pool leaves rotation, even though the instances are fine and could still serve cached or degraded responses. Health checks should test the instance, not its dependencies; dependency failures belong in observability and in a relative, real-traffic signal.

It is like closing lanes on a jammed motorway because the traffic on them is moving slowly: every closure pushes the same cars into fewer lanes, so the next lane qualifies for closure sooner.

saying these in an interview costs you the question

  • Believes ejecting instances reduces the pool's load
  • Sets the check timeout just above normal response latency
  • Has no floor below which the proxy fails open
  • Health-checks a path that queries a shared dependency
  • Treats an entirely unhealthy pool as genuinely all-dead

context