A reverse proxy probes each backend every 10 seconds with a 2-second check timeout, and marks a backend unhealthy after 3 consecutive failed probes. Roughly how long can a dead backend keep receiving requests, and what goes wrong if you shrink those numbers to make detection nearly instant?
answer
- interval times threshold, plus the timeout
- up to one interval before anyone looks
- probe rate scales with proxies times backends
- one-failure threshold ejects on a GC pause
- retries and passive ejection close the gap
basics
~20 sRoughly 22 to 32 seconds: up to one interval before the next probe, then three probes ten seconds apart, the last costing its two-second timeout. Tightening the numbers shortens that window but multiplies probe load and makes brief pauses eject healthy backends.
solid answer
~60 sWork it out from the timer. The backend can die just after a successful probe, so you wait up to one full interval for the next one; then you need three consecutive failures at interval spacing, and each failing probe burns its timeout before it counts. That is roughly `interval × failures + timeout`, so about 32 seconds worst case and about 22 in the best case — and every request routed to that backend in the window fails. The instinct is to drop the interval to one second and the threshold to one, but that buys detection by paying twice: probe traffic scales as proxies × backends ÷ interval, and a single-failure threshold means any garbage-collection pause, brief scheduling delay, or dropped packet ejects a healthy instance. The right fix for the gap is usually not a faster probe but retrying the failed request onto another backend and letting passive ejection react to real traffic, both of which respond on the first failure rather than on the next probe cycle.
go deeper
Know that detection is not instant: a dead backend keeps getting requests until enough probes have failed, and that delay comes straight from the interval and the failure count.
Be able to compute the window on the spot as interval times threshold plus timeout, and explain what each of the three numbers costs if you shrink it.
Argue that the gap is better closed with retries and real-traffic ejection than with a faster timer, and show that you have priced the probe load your tuning creates across every proxy.
Set the platform default and the ceiling teams may not exceed, so that no single service's aggressive check tuning becomes a synthetic load problem or an alert-fatigue source for everyone.
## Do the arithmetic out loud Interviewers ask this to see whether you can turn three config numbers into a user-visible outage duration. The chain is: 1. The backend dies at some point in the probe cycle. In the worst case it dies immediately after a probe succeeded, so nothing is learned for a full **interval**. 2. The next probe fails. If it fails by timing out — which is the common case for a hung process, a full accept queue, or a black-holing host — that failure is only recorded after the **timeout** elapses. If it fails by connection refused, it is recorded immediately. 3. You still need the remaining failures to reach the **unhealthy threshold**, arriving one interval apart. So the window is approximately `interval × threshold + timeout`, and the best case (death just before a probe) is approximately `interval × (threshold − 1) + timeout`. With 10 s / 2 s / 3 that is roughly **22–32 seconds**. Half a minute of routing requests into a dead instance. In a four-instance pool that is a quarter of traffic failing, and no amount of retry configuration makes those numbers smaller — retries hide the failures, they do not shorten the window. One more term matters if the timeout is larger than the interval: probes then overlap or queue, and depending on the implementation the effective detection can be *slower* than the formula suggests. Keep the timeout comfortably below the interval. ## Why not just make it 1 s / 1 s / 1? Three costs, all real. **Probe volume.** The rate is `proxyInstances × backends ÷ interval`. Six proxies in front of a 200-instance pool at a 10-second interval is 120 probes per second; at a 1-second interval it is 1,200. That is a nontrivial synthetic workload for the backends, for the proxies' own event loops, and for any logging or metrics pipeline that records it. **False positives.** A threshold of one means every transient becomes an ejection: a garbage-collection pause, a CPU-scheduling delay on a noisy host, one dropped UDP or TCP packet, a config reload, a momentary listener backlog. Real instances hit these constantly. The consecutive-failure threshold exists because a single sample is a weak measurement, and setting it to one throws away the only noise filter you have. **Flapping, which is worse than either.** An instance that oscillates in and out of the pool every few seconds is not merely a false positive; it is actively destructive. Each transition reshuffles load across the survivors, invalidates whatever locality or affinity the balancer had built, discards and rebuilds connection pools, sends the instance a traffic burst the moment it returns (a returning host with no open connections looks attractive to a least-request balancer), and fills alerting with noise until the signal is ignored. Aggressive tuning turns a marginal instance into a pool-wide perturbation. ## Tune the knobs against each other Useful shapes, roughly in order of how often they help: - **Asymmetric thresholds.** Two failures out, four or five successes back in. Quick to protect users, slow to trust a recovering instance. Detection stays fast without the flapping. - **A timeout matched to the failure you want to catch, not to normal latency.** A probe timeout barely above the p50 turns a load spike into an ejection storm. Give it real slack — the check answers "is this reachable", not "is this fast". - **Jitter.** Without it, N proxies started together probe in lockstep and every backend sees a synchronised burst each interval. Most implementations add jitter by default; if yours does not, it matters at scale. - **A dedicated, cheap check path.** The probe should not queue behind expensive work, or you are measuring load rather than liveness. ## The better answer to "detect faster" The detection window is a property of a *sampling* scheme, and there is a strictly better instrument for the same job: the real request. Two mechanisms respond on the first failure rather than the next probe cycle: - **Retry onto a different backend.** When a request to a dead instance fails to connect or is refused, the proxy can immediately try another instance. The user never sees the failure, and the response time is measured in milliseconds rather than in probe intervals. Note the constraint: this is only safe where re-sending the request is safe, so the retry policy has to be chosen per route, not globally. - **Passive ejection.** Real-traffic errors accumulate at the rate of real traffic, not at the probe rate, so a dead instance under load is convicted almost immediately. The mature position, then, is: set active checks slow enough to be quiet and stable, and rely on retries plus real-traffic observation to cover the seconds in between. A one-second probe interval is usually a sign that someone tried to solve a routing problem with a timer.
- Which of the three settings would you change first to halve the detection window, and why that one?The interval, because it multiplies the threshold in the formula while the timeout only adds once. Halving 10 s to 5 s takes the worst case from about 32 s to about 17 s and keeps the noise filter intact. Cutting the threshold instead removes exactly the protection against transient failures, and cutting the timeout starts converting slow responses into ejections.
- What happens if the check timeout is set larger than the check interval?Probes stop being independent: a probe is still outstanding when the next one is due, so depending on the implementation they overlap, queue, or are skipped. Failure counting no longer advances one interval at a time and detection can be slower than the arithmetic predicts, while the backend sees several concurrent probes. Keep the timeout well under the interval.
- Why do proxies jitter their health-check timers?Without jitter, many proxy instances configured identically and started together fire probes in lockstep, so each backend sees a synchronised burst once per interval instead of a smooth trickle. At scale that burst competes with real traffic and can itself cause the latency that fails the checks. Jitter spreads the same total probe volume evenly across the interval.
saying these in an interview costs you the question
- Assumes detection is immediate when a backend dies
- Forgets the timeout is spent before a probe counts as failed
- Sets the unhealthy threshold to one for faster detection
- Ignores that probe volume scales with the proxy count
- Sets the probe timeout just above normal response latency