A load balancer can decide a backend is bad in two ways: by sending it probe requests of its own, or by watching how real client requests to that backend turn out. Compare the two — what does each detect that the other misses, and why do serious production setups run both?
answer
- synthetic traffic versus real traffic
- probes are proactive but sample one path
- passive detection pays in failed requests
- an ejected host gets no real traffic
- relative outlier, not absolute threshold
basics
~20 sActive probes are synthetic and cheap to run, so they catch a dead or unreachable instance before users do, but they only ever test the probe path. Passive detection judges an instance from real request outcomes, catching failures probes never touch — at the cost of real failed requests.
solid answer
~50 sActive checking means the proxy periodically sends its own request to each backend and grades the response. It is proactive — it can find a dead listener before a single user hits it — and it is the only way to tell whether an already-ejected instance has recovered, since an ejected instance sees no real traffic. Its blind spot is that it tests exactly one synthetic path: an instance that answers the probe fine while failing one route, one dependency, or 5% of requests looks perfectly healthy. Passive detection, usually called outlier ejection, judges each backend from the outcomes of real traffic — connection errors, timeouts, 5xx rates — so it sees precisely the failures users are experiencing, including partial ones. Its cost is that the signal *is* user harm: something has to fail before it reacts. So you run active checks for reachability and for the recovery decision, and passive detection for the half-failing instance the probe cannot see.
go deeper
Know the two words and one sentence each: active means the proxy sends its own probe on a timer, passive means it grades the backend from the results of real user requests.
Be able to name the blind spot of each without prompting — a probe samples one synthetic path at low rate, and passive detection needs real traffic and only reacts after requests have already failed.
Show the design instinct: active for reachability and for the readmission decision, passive with a relative threshold for the half-failing instance, and an explanation of why deepening the probe is the dangerous fix.
Own the trade at fleet scale — probe cost across every proxy-to-backend pair, the failure semantics you standardise on, and how much user harm you are willing to spend to detect a partial failure.
## Two different sources of evidence Every proxy needs evidence that a backend is usable. There are only two places to get it: traffic the proxy generates itself, or traffic the users generated. Those are active and passive health checking, and they fail in opposite directions — which is why the mature answer is almost never "pick one". ## Active checking: synthetic, proactive, narrow An active check is a request the proxy manufactures on a timer — a TCP connect, or an HTTP request to a nominated path whose response code (and sometimes body) is graded. Its strengths follow from being synthetic: - **It works with zero traffic.** A backend that nobody has called in ten minutes is still known-good or known-bad. Passive detection is blind here — no requests, no signal. - **It is the only recovery signal.** Once an instance is out of the pool it receives no real traffic, so its real-traffic error rate is undefined forever. Something must poke it. Even a purely passive ejection scheme has to fall back to a timed probation period, which is really just a very crude active probe. - **It is proactive.** Detection can happen before a user is harmed, which is the entire point of running it. And its weaknesses follow from the same fact: - **It tests one path only.** If the probe path returns a static 200 without touching the code, the config, or the dependency that is actually broken, a comprehensively broken instance is graded healthy. This is the classic half-failing backend: probes green, users getting errors. - **It has no statistical power.** A probe every 10 seconds is six samples a minute. A backend failing 3% of real requests will pass essentially every probe. Rare-but-real failure rates are invisible to a low-rate sampler. - **It costs the backend something.** Probe load is (number of proxy instances × number of backends) ÷ interval, and on a large mesh that is not nothing. ## Passive checking: real evidence, reactive, needs volume Passive detection — outlier ejection, in the vocabulary most proxies use — records the outcome of every real request the proxy forwards, and ejects a backend whose outcomes are bad: connection refused or reset, request timeouts, a run of consecutive 5xx responses, or a 5xx rate that stands out from the rest of the pool. Its strengths are the mirror image: - **It measures what users experience**, on the paths users actually call, at production traffic volume. The 3% failure rate the probe missed shows up in hundreds of samples per minute. - **It needs no cooperation from the application.** No health endpoint to write, no risk of that endpoint drifting away from what the service really does. Its costs are equally structural: - **The signal is harm.** By the time a backend is ejected passively, real requests have already failed. Passive detection limits the blast radius; it never prevents the first failures. This is why it pairs with retries: the retry hides the individual failure while the ejection stops the bleeding. - **It needs traffic.** A backend receiving three requests a minute cannot be judged; ejection on tiny samples is superstition, which is why implementations require a minimum request volume before evaluating a host. - **It can be fooled by the request mix.** If a hashing or affinity rule sends all requests for one broken tenant to one instance, that instance looks like the outlier when the data is at fault. This is also why the better implementations judge *relatively* — a host is an outlier when it is much worse than its peers — so that a pool-wide failure ejects nobody, which is exactly the desired behaviour. ## Why both, concretely Put them side by side and each covers the other's blind spot: | Failure | Active probe | Passive observation | |---|---|---| | Process dead, port closed | Detected in one interval | Detected on the first real request | | Instance idle, no traffic | Detected | Invisible | | One route broken, probe path fine | Invisible | Detected | | 3% error rate | Invisible | Detected | | Recovery of an ejected host | Detected | Impossible to observe | The standard production shape is therefore: an active check that is cheap and stable, deciding reachability and the return-to-service decision; plus passive ejection with a relative threshold and a bounded ejection duration, catching the partial failures the probe cannot see. Neither is a substitute for the other, and an interviewer asking this is usually listening for whether you know that the recovery decision *cannot* be passive. ## The trap in the follow-up When people discover the probe's blind spot, the instinct is to make the probe deeper until it exercises everything. That instinct has a sharp edge: a probe that touches the shared database now fails on every instance at once when that database hiccups, converting a dependency blip into a pool-wide ejection. Passive detection is generally the safer way to catch a partial failure, precisely because it grades an instance against its peers rather than against an absolute.
- If passive ejection catches the failures users actually hit, why not rely on it alone?Because it cannot see recovery. An ejected backend receives no real requests, so its real-traffic error rate never updates and it could never be readmitted on evidence. Purely passive schemes fall back to a timed probation, which is a blunt active probe. Passive detection is also blind on a low-traffic or idle backend, where there are simply not enough samples to judge.
- Why do outlier ejection implementations compare a backend against the rest of the pool instead of a fixed error rate?A fixed threshold cannot distinguish a broken instance from a broken dependency. If a shared database degrades, every instance's error rate rises together and an absolute rule ejects the entire pool. A relative rule asks whether this host is much worse than its peers — which is the actual definition of an outlier — so a pool-wide problem ejects nobody and stays a partial failure instead of an outage.
- What limits should bound passive ejection so it cannot empty the pool?Three: a minimum request volume per host before any verdict, a cap on the fraction of the pool that may be ejected at once, and a bounded ejection duration after which the host is retried. Together they mean a mistaken verdict costs a slice of capacity for a short window instead of the whole service.
saying these in an interview costs you the question
- Claims an active health check proves the service works
- Thinks passive ejection can also detect recovery
- Judges a backend from a handful of real requests
- Uses an absolute error threshold, ejecting the whole pool
- Fixes probe blind spots by probing every dependency