skip to content

Health Checking & Outlier Ejection

You learn how a proxy decides a backend is usable: active probes versus passive observation of real traffic, how thresholds and intervals set your detection-versus-flapping trade, and what outlier ejection and slow start fix. Interviewers ask it as an incident story — a backend that half-fails, or a health check so aggressive it ejects the whole pool — because tuning these numbers badly is a common self-inflicted outage.

on this pageshow

questions

5

A reverse proxy load-balances across several backend instances and actively health-checks each one. What does the proxy do with a backend that starts failing those checks, and what has to happen before that backend receives traffic again?

level: juniorimportance: must knowfreq 72%

answer

  1. not one failed probe — a count
  2. consecutive failures out, successes back
  3. new requests stop, in-flight ones finish
  4. the proxy never restarts anything
  5. survivors absorb the ejected instance's load

basics

~20 s

After a configured number of consecutive failed checks the proxy marks the backend unhealthy and stops sending it new requests, while continuing to probe it. It rejoins the pool only after a configured number of consecutive successful checks.

solid answer

~40 s

A health check is the proxy's answer to one narrow question: should I send the next request here? A backend that fails a probe is not ejected immediately — proxies require N consecutive failures (an unhealthy threshold), because a single lost packet or a garbage-collection pause is not evidence of a broken instance. Once the threshold is crossed, the instance leaves the selection pool for new requests; requests already in flight are usually allowed to finish. The proxy keeps probing the ejected instance, and after a configured number of consecutive successes it goes back into rotation, often behind a ramp rather than at full share. Note what the proxy does *not* do: it never restarts the instance and never declares it globally broken — health here is per proxy, and another proxy may disagree.

go deeper

for a junior

Be ready to describe the loop out loud: repeated probes, several consecutive failures before ejection, several consecutive successes before return. Say clearly that the proxy stops routing but does not restart anything.

for a middle

Explain why the threshold exists at all and what it costs — every extra required failure adds one interval of detection delay while filtering out noise like a GC pause or a dropped probe packet.

for a senior

Show that you plan capacity for ejection: state how much headroom the pool needs so removing an instance does not push the survivors into the overload that ejects them too.

for a principal

Own the policy question of who defines check semantics across a fleet — what the platform mandates, what a service team may tune, and how you keep a hundred teams from each inventing a different meaning of healthy.

## The question the proxy is actually asking A reverse proxy or load balancer holds a list of backend instances it is allowed to send requests to. An active health check is how the proxy answers one narrow question on its own: *should I send the next request to this instance?* It is not a diagnosis of what is wrong, it is not a restart decision, and it is not a statement that the instance is broken for everyone. Almost every mistake in this area comes from forgetting that framing. ## The three numbers that define the check loop An active check runs on a timer, inside each proxy instance, against each backend. Three settings shape it: - **Interval** — how often a probe is sent (say, every 10 seconds). - **Timeout** — how long the proxy waits for the probe's response before calling that probe a failure. A probe that answers slowly is a failed probe, not a slow one. - **Thresholds** — how many *consecutive* failed probes flip a healthy backend to unhealthy, and how many *consecutive* successful probes flip it back. Every mainstream proxy exposes both; the names differ per product. The threshold is the part people forget. A single failed probe is a poor signal: probes are lost to dropped packets, a stop-the-world pause, a momentary scheduler delay, or a listener that was mid-reload. Requiring several consecutive failures buys a cheap, effective noise filter — at the cost of detection delay, which is the central trade of this whole topic. ``` probe ok -> failures = 0 probe fail -> failures += 1 ; if failures == unhealthyThreshold: mark DOWN (while DOWN) probe ok -> successes += 1 ; if successes == healthyThreshold: mark UP ``` ## What the transition actually changes Marking a backend unhealthy removes it from the pool the load-balancing algorithm chooses from **for new requests**. Requests already in flight to that instance normally run to completion; the proxy does not usually abort them, though it may stop reusing that instance's idle keep-alive connections. Probing continues — an ejected backend that stopped being probed could never come back. What does not happen matters just as much. The proxy has no authority over the process: it cannot restart it, redeploy it, or take it out of an orchestrator's rotation. Deciding to *restart* something unhealthy is a different system's job entirely. The proxy's power ends at "I will not route to you." ## Coming back The return path is deliberately asymmetric in most production configurations: quick to eject, slow to readmit. A common shape is two consecutive failures to eject and three to five consecutive successes to return, so a genuinely flapping instance cannot oscillate in and out of the pool every few seconds. Some proxies additionally lengthen the exclusion period each time the same instance is ejected again, and many ramp its traffic share up gradually rather than handing it a full share the instant it passes. ## The two consequences juniors miss **Ejecting a backend does not reduce the traffic.** The requests that instance was serving are now spread across the survivors. A pool of four instances running at 70% CPU cannot absorb one ejection — losing one puts the other three at roughly 93%, and the pool is one bad minute from ejecting itself into an outage. Health checking is only safe if the pool has headroom. **Health is per proxy, not global.** Each proxy instance runs its own check loop and keeps its own verdict. If you run six proxies, the backend receives six times the probe traffic, and during a partial network problem some proxies may consider a backend healthy while others do not. That is by design — each proxy routes based on reachability *from where it stands* — but it means "is instance 3 healthy?" is not a question with one answer, and a dashboard that shows one is aggregating. ## What a good answer sounds like Name the state machine (consecutive failures out, consecutive successes back), say plainly that in-flight requests usually finish while new ones stop, note that the proxy does not restart anything, and mention that the survivors absorb the load. That is the whole junior-level shape, and it sets up every harder question in this area.

  • Why is the number of successful checks needed to return often higher than the number of failures needed to eject?
    Because the two errors have different costs. Ejecting a healthy instance costs a little capacity; readmitting a broken one costs real user requests. Asymmetric thresholds also damp flapping: an instance that recovers for one probe and fails the next never accumulates enough consecutive successes to rejoin, so it stays out until it is genuinely stable.
  • If every backend in the pool fails its health check at once, what should the proxy do?
    Serving errors from an empty pool is strictly worse than trying a possibly-degraded backend, so many proxies fail open: when the healthy fraction falls below a configured floor, they ignore health status and balance across every host. It is a deliberate admission that a pool-wide failure is far more likely to mean the check is wrong than that every instance died simultaneously.
  • Does the proxy's health status have any effect on the backend process itself?
    None. The proxy only changes its own routing table. The process keeps running, keeps holding its connections, and keeps logging. Restarting or replacing an unhealthy instance is the job of whatever supervises it, and confusing the two leads people to expect a load balancer to fix a hung process, which it never will.

saying these in an interview costs you the question

  • Thinks one failed probe removes a backend immediately
  • Believes the proxy restarts the unhealthy backend
  • Assumes in-flight requests are killed on ejection
  • Treats health status as global rather than per proxy
  • Ignores that survivors inherit the ejected instance's load

context

open as a page

A load balancer can decide a backend is bad in two ways: by sending it probe requests of its own, or by watching how real client requests to that backend turn out. Compare the two — what does each detect that the other misses, and why do serious production setups run both?

level: middleimportance: must knowfreq 66%

basics

~20 s

Active probes are synthetic and cheap to run, so they catch a dead or unreachable instance before users do, but they only ever test the probe path. Passive detection judges an instance from real request outcomes, catching failures probes never touch — at the cost of real failed requests.

open as a page

A reverse proxy probes each backend every 10 seconds with a 2-second check timeout, and marks a backend unhealthy after 3 consecutive failed probes. Roughly how long can a dead backend keep receiving requests, and what goes wrong if you shrink those numbers to make detection nearly instant?

level: middleimportance: should knowfreq 54%

basics

~20 s

Roughly 22 to 32 seconds: up to one interval before the next probe, then three probes ten seconds apart, the last costing its two-second timeout. Tightening the numbers shortens that window but multiplies probe load and makes brief pauses eject healthy backends.

open as a page

During a traffic spike every instance in a backend pool slows down, health checks begin timing out, and the load balancer ejects instances one after another until almost nothing is left in rotation — turning a slowdown into a total outage. Explain the feedback loop, and what safeguards stop health checking from causing this.

level: seniorimportance: should knowfreq 47%

basics

~20 s

Overload makes checks time out, ejection moves that instance's traffic onto the survivors, the survivors slow further and are ejected in turn. The loop is closed because ejection increases the load that caused it. Safeguards: fail open below a healthy floor, cap the ejectable fraction, and never eject on latency alone.

open as a page

A backend instance that a load balancer had ejected is now passing its health checks again and is about to rejoin the pool. Why can handing it an immediate equal share of traffic break it a second time, and what does a slow-start or warm-up ramp do about that?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

A returning instance is cold: empty caches, no warm connection pools, unoptimised runtime. Worse, load-aware algorithms see an idle host as the most attractive target and send it a burst above its fair share. A slow-start ramp raises its weight gradually so it warms under partial load.

open as a page