skip to content

How do active and passive health checks work in a load balancer, and what production failure modes can occur if health-check thresholds are configured too aggressively or too loosely?

level: seniorimportance: must knowfreq 75%

answer

  1. active = synthetic probe on interval
  2. passive = watch real traffic outcomes
  3. N-consecutive-failure debounce
  4. flapping cascades load onto survivors
  5. asymmetric thresholds: slow out, fast back in

basics

~20 s

Active health checks are the load balancer regularly pinging each server with a test request to see if it's alive. Passive health checks watch real traffic and mark a server unhealthy if its actual responses start failing. Bad thresholds can either kick out healthy servers too fast or take too long to notice a broken one.

solid answer

~1 min

Active health checks have the load balancer periodically send a synthetic probe (often an HTTP GET to a dedicated /health or /healthz endpoint) to each backend on a fixed interval, and mark it unhealthy after N consecutive failures (the unhealthy threshold) and healthy again after M consecutive successes (the healthy threshold). Passive health checks instead observe real client traffic and infer health from it - marking a backend unhealthy if it returns a run of 5xx errors, times out, or resets connections. Active checks catch a dead backend even with zero live traffic and let you check deeper (DB connectivity, dependency status), but add constant background load and can pass a shallow check while the app itself is broken. Passive checks reflect real behavior with no extra load but need actual traffic to detect a problem and can be slower to react. The failure-mode risk is a tuning problem: too-aggressive thresholds (short interval, low failure count) cause flapping - a backend under a momentary GC pause or brief blip gets pulled from rotation, concentrating its load onto the remaining servers and potentially cascading; too-loose thresholds delay detection, so a genuinely broken backend keeps receiving and failing real user requests for longer than necessary.

go deeper

for a junior

Should know a load balancer periodically checks if servers are alive and stops sending traffic to a dead one.

for a middle

Should distinguish active vs passive checks and explain the consecutive-failure threshold concept.

for a senior

Should explain the flapping/cascading-failure risk from aggressive thresholds and the shallow-vs-deep health endpoint trade-off.

for a principal

Should design asymmetric threshold policy, combine active and passive/outlier detection deliberately, and reason about health-check design as part of overall fleet resilience, including its own cost (probe load, expensive deep checks).

## Why health checking exists A load balancer can only route around a broken backend if it knows the backend is broken, and health checking is the mechanism that supplies that knowledge. There are two broad families. | Family | Where the signal comes from | |---|---| | **active checks** | the load balancer itself originates synthetic probe traffic | | **passive checks** | the load balancer infers health by observing the outcome of real client requests it is already forwarding | ## Active checks: synthetic probes on an interval An active health check works by the load balancer opening a connection (for L4) or sending an HTTP request (for L7, typically a GET to a dedicated endpoint like `/health` or `/healthz`) to each backend on a configured interval - say every 5 or 10 seconds. The response is judged against expected criteria: - **for TCP checks**, simply whether the connection succeeds; - **for HTTP checks**, whether the status code falls in an expected range (usually 200) within a timeout, and sometimes whether the response body matches an expected pattern. Crucially, the load balancer does not act on a single failed probe - it applies: - an **unhealthy threshold**, a count of consecutive failures (commonly 2-3) required before the backend is pulled out of rotation; - a **healthy threshold**, a count of consecutive successes required before it is added back. This debouncing exists because a single missed probe is common and often meaningless - a momentary network blip, a garbage-collection pause, or a probe that happened to race with a deploy - and pulling a backend out over one failed check would cause unnecessary churn. The health endpoint itself is a design decision with real consequences: - a **shallow check** that just returns 200 from the web server tells you the process is alive but nothing about whether it can reach its database or downstream dependencies; - a **deep check** that verifies DB connectivity and dependency health gives a much more accurate picture of true readiness - at the cost of the health endpoint itself becoming a potential bottleneck or single point of failure if it, say, does an expensive query on every probe. ## Passive checks: riding on real traffic A passive health check instead rides on real traffic: the load balancer watches the actual responses it gets back from each backend as it forwards genuine client requests, and tracks a rolling count or rate of failures - HTTP 5xx statuses, connection resets, timeouts. If that rate crosses a configured threshold within a time window (for example, more than 50% of the last 10 requests failed, or 5 consecutive timeouts), the backend is marked unhealthy and removed from rotation, usually for a cooldown period after which it is retried. This is sometimes called an 'outlier detection' or 'circuit breaker' pattern at the load-balancer level, and it has the advantage of adding zero synthetic load and reflecting the exact failure modes real users are experiencing, including ones a shallow active probe would never surface (e.g. a backend that answers health checks fine but fails specifically on a particular heavy query pattern). Its weakness is that it requires actual traffic to detect a problem at all - a backend with no incoming requests yet, or one that is broken for a code path nobody has hit recently, will look healthy - and by definition some real user requests must fail before the system reacts, unlike an active check which can catch a dead process before any user ever reaches it. ## Thresholds too aggressive: flapping and cascade The failure modes show up at both extremes of threshold tuning. Configuring checks too aggressively - a short interval combined with a low unhealthy threshold, such as checking every 2 seconds and pulling a backend after a single failure - makes the system hypersensitive to transient noise: a backend experiencing a normal JVM garbage-collection pause of a few hundred milliseconds, or briefly saturated during a deploy, gets yanked from the pool. This is called **flapping**, and it's actively dangerous under load, because removing a backend concentrates its share of traffic onto the remaining, already-busy instances, which can push them past their own tipping point and cause them to fail their own health checks in turn - a cascading failure that can take an entire healthy fleet down from what started as one noisy probe. ## Thresholds too loose: slow detection Configuring checks too loosely - a long interval and a high failure threshold, such as checking every 60 seconds and requiring 10 consecutive failures - means a genuinely broken backend can sit in rotation minutes, continuing to serve errors or timeouts to real users, degrading the overall error rate and latency the whole fleet reports, before it is finally pulled. ## What production tuning looks like Production systems tune these deliberately asymmetric: it's common to - require several consecutive failures to mark unhealthy (avoiding flapping on transient blips) but only one or two successes to mark healthy again (so a recovered backend rejoins quickly); - pair active checks (for catching total outages fast, even with no traffic) with passive/outlier detection (for catching subtler, traffic-dependent breakage) rather than relying on either alone. AWS ALB target groups, Kubernetes readiness probes, and Envoy's outlier-detection filter are all concrete, widely used implementations of exactly this active-plus-passive combination.

  • Why is it common to require more consecutive failures to mark a backend unhealthy than successes to mark it healthy again?
    Because the cost of the two mistakes is asymmetric - prematurely pulling a healthy backend (a false positive on 'unhealthy') concentrates load onto fewer servers and risks cascading failure, while re-adding a backend slightly too eagerly just risks a brief return of a still-recovering instance, which is a much cheaper mistake. Requiring a larger streak of failures before removal filters out transient noise, while requiring only one or two successes gets capacity back online quickly.
  • What is the risk of using a shallow health check endpoint that only checks 'is the web server process running'?
    It can report healthy while the backend is actually unable to serve real requests correctly - for example if its database connection pool is exhausted or a critical downstream dependency is down - so the load balancer keeps sending it real traffic that fails. A deeper check that verifies key dependencies gives a more accurate readiness signal, at the cost of a heavier, potentially slower probe.
  • How can a health-check-triggered cascading failure happen, concretely?
    If thresholds are too tight, a burst of latency (a GC pause, a brief network blip) causes several backends to fail their checks around the same time and get pulled from rotation; the surviving backends absorb their share of traffic, which pushes their own latency and resource usage up, which can cause them to start failing their own checks too, in a feedback loop that can empty the whole pool.

Active health checks are like a manager walking the floor every few minutes and asking each worker 'you okay?'. Passive health checks are like noticing that a particular worker's completed tasks keep coming back wrong or late, without ever having to ask them directly.

saying these in an interview costs you the question

  • Thinks a single failed probe should immediately pull a backend
  • Doesn't distinguish active from passive checks
  • Assumes deeper health checks have no downside
  • Can't explain how aggressive thresholds cause cascading failure
  • Believes health checks and load-balancing algorithms are the same mechanism

context