A reverse proxy load-balances across several backend instances and actively health-checks each one. What does the proxy do with a backend that starts failing those checks, and what has to happen before that backend receives traffic again?
answer
- not one failed probe — a count
- consecutive failures out, successes back
- new requests stop, in-flight ones finish
- the proxy never restarts anything
- survivors absorb the ejected instance's load
basics
~20 sAfter a configured number of consecutive failed checks the proxy marks the backend unhealthy and stops sending it new requests, while continuing to probe it. It rejoins the pool only after a configured number of consecutive successful checks.
solid answer
~40 sA health check is the proxy's answer to one narrow question: should I send the next request here? A backend that fails a probe is not ejected immediately — proxies require N consecutive failures (an unhealthy threshold), because a single lost packet or a garbage-collection pause is not evidence of a broken instance. Once the threshold is crossed, the instance leaves the selection pool for new requests; requests already in flight are usually allowed to finish. The proxy keeps probing the ejected instance, and after a configured number of consecutive successes it goes back into rotation, often behind a ramp rather than at full share. Note what the proxy does *not* do: it never restarts the instance and never declares it globally broken — health here is per proxy, and another proxy may disagree.
go deeper
Be ready to describe the loop out loud: repeated probes, several consecutive failures before ejection, several consecutive successes before return. Say clearly that the proxy stops routing but does not restart anything.
Explain why the threshold exists at all and what it costs — every extra required failure adds one interval of detection delay while filtering out noise like a GC pause or a dropped probe packet.
Show that you plan capacity for ejection: state how much headroom the pool needs so removing an instance does not push the survivors into the overload that ejects them too.
Own the policy question of who defines check semantics across a fleet — what the platform mandates, what a service team may tune, and how you keep a hundred teams from each inventing a different meaning of healthy.
## The question the proxy is actually asking A reverse proxy or load balancer holds a list of backend instances it is allowed to send requests to. An active health check is how the proxy answers one narrow question on its own: *should I send the next request to this instance?* It is not a diagnosis of what is wrong, it is not a restart decision, and it is not a statement that the instance is broken for everyone. Almost every mistake in this area comes from forgetting that framing. ## The three numbers that define the check loop An active check runs on a timer, inside each proxy instance, against each backend. Three settings shape it: - **Interval** — how often a probe is sent (say, every 10 seconds). - **Timeout** — how long the proxy waits for the probe's response before calling that probe a failure. A probe that answers slowly is a failed probe, not a slow one. - **Thresholds** — how many *consecutive* failed probes flip a healthy backend to unhealthy, and how many *consecutive* successful probes flip it back. Every mainstream proxy exposes both; the names differ per product. The threshold is the part people forget. A single failed probe is a poor signal: probes are lost to dropped packets, a stop-the-world pause, a momentary scheduler delay, or a listener that was mid-reload. Requiring several consecutive failures buys a cheap, effective noise filter — at the cost of detection delay, which is the central trade of this whole topic. ``` probe ok -> failures = 0 probe fail -> failures += 1 ; if failures == unhealthyThreshold: mark DOWN (while DOWN) probe ok -> successes += 1 ; if successes == healthyThreshold: mark UP ``` ## What the transition actually changes Marking a backend unhealthy removes it from the pool the load-balancing algorithm chooses from **for new requests**. Requests already in flight to that instance normally run to completion; the proxy does not usually abort them, though it may stop reusing that instance's idle keep-alive connections. Probing continues — an ejected backend that stopped being probed could never come back. What does not happen matters just as much. The proxy has no authority over the process: it cannot restart it, redeploy it, or take it out of an orchestrator's rotation. Deciding to *restart* something unhealthy is a different system's job entirely. The proxy's power ends at "I will not route to you." ## Coming back The return path is deliberately asymmetric in most production configurations: quick to eject, slow to readmit. A common shape is two consecutive failures to eject and three to five consecutive successes to return, so a genuinely flapping instance cannot oscillate in and out of the pool every few seconds. Some proxies additionally lengthen the exclusion period each time the same instance is ejected again, and many ramp its traffic share up gradually rather than handing it a full share the instant it passes. ## The two consequences juniors miss **Ejecting a backend does not reduce the traffic.** The requests that instance was serving are now spread across the survivors. A pool of four instances running at 70% CPU cannot absorb one ejection — losing one puts the other three at roughly 93%, and the pool is one bad minute from ejecting itself into an outage. Health checking is only safe if the pool has headroom. **Health is per proxy, not global.** Each proxy instance runs its own check loop and keeps its own verdict. If you run six proxies, the backend receives six times the probe traffic, and during a partial network problem some proxies may consider a backend healthy while others do not. That is by design — each proxy routes based on reachability *from where it stands* — but it means "is instance 3 healthy?" is not a question with one answer, and a dashboard that shows one is aggregating. ## What a good answer sounds like Name the state machine (consecutive failures out, consecutive successes back), say plainly that in-flight requests usually finish while new ones stop, note that the proxy does not restart anything, and mention that the survivors absorb the load. That is the whole junior-level shape, and it sets up every harder question in this area.
- Why is the number of successful checks needed to return often higher than the number of failures needed to eject?Because the two errors have different costs. Ejecting a healthy instance costs a little capacity; readmitting a broken one costs real user requests. Asymmetric thresholds also damp flapping: an instance that recovers for one probe and fails the next never accumulates enough consecutive successes to rejoin, so it stays out until it is genuinely stable.
- If every backend in the pool fails its health check at once, what should the proxy do?Serving errors from an empty pool is strictly worse than trying a possibly-degraded backend, so many proxies fail open: when the healthy fraction falls below a configured floor, they ignore health status and balance across every host. It is a deliberate admission that a pool-wide failure is far more likely to mean the check is wrong than that every instance died simultaneously.
- Does the proxy's health status have any effect on the backend process itself?None. The proxy only changes its own routing table. The process keeps running, keeps holding its connections, and keeps logging. Restarting or replacing an unhealthy instance is the job of whatever supervises it, and confusing the two leads people to expect a load balancer to fix a hung process, which it never will.
saying these in an interview costs you the question
- Thinks one failed probe removes a backend immediately
- Believes the proxy restarts the unhealthy backend
- Assumes in-flight requests are killed on ejection
- Treats health status as global rather than per proxy
- Ignores that survivors inherit the ejected instance's load