A circuit breaker uses a sliding window of recent call outcomes to decide whether to trip. What's the difference between a failure-rate threshold and a slow-call-rate threshold, and how does the sliding window shape when those thresholds actually get evaluated?
answer
- failure rate = errors / total in window
- slow-call rate = successes-but-too-slow / total
- count-based vs time-based window
- minimum-number-of-calls floor before evaluating
- slow calls can trip breaker with 0% errors
basics
~20 sThe breaker looks at the last batch of calls. The failure-rate threshold trips it if too many of those calls errored out. The slow-call threshold trips it separately if too many calls succeeded but took too long. Both are percentages over a recent batch, not a single call.
solid answer
~30 sFailure-rate threshold counts calls that threw an exception or returned a defined-failure result against the total calls recorded in the sliding window; crossing a percentage (e.g. 50%) trips the breaker. Slow-call-rate threshold is evaluated independently: calls that completed successfully but exceeded a configured duration are counted as 'slow', and if their percentage also crosses a limit, the breaker trips even though nothing technically errored — this catches degraded-but-not-failing dependencies. The window can be count-based (last N calls) or time-based (last N seconds), and a minimum-number-of-calls floor prevents evaluating either threshold against a statistically meaningless sample.
go deeper
Should know thresholds are percentages over recent calls, not single-call triggers.
Should clearly separate failure-rate from slow-call-rate as two independent signals and explain why slowness matters even without errors.
Should discuss count-based vs time-based windows and the minimum-number-of-calls floor, and reason about tuning trade-offs for a specific dependency's latency profile.
Should connect threshold tuning to broader system behavior — how misconfigured thresholds interact with cascading failure across a service graph and how you'd derive thresholds from observed SLOs rather than guesswork.
## The two statistics that matter A circuit breaker doesn't make trip decisions off a single bad call; it maintains a **sliding window** of recent outcomes and evaluates aggregate statistics against configured thresholds. Two of those statistics matter most: the **failure rate** and the **slow-call rate**, and they exist for genuinely different reasons even though they trigger the same open transition. ## The failure-rate threshold The failure-rate threshold answers 'how often are calls to this dependency erroring out?' Every completed call is classified as success or failure — failure usually means one of: - an exception was thrown, - a configured failure predicate matched (e.g. a 5xx HTTP status, a specific exception type), - or the call timed out outright. The breaker divides failed calls by total calls in the window and compares that percentage to a configured threshold, commonly 50%. This is the classic 'the dependency is visibly broken' signal. ## The slow-call-rate threshold The slow-call-rate threshold answers a different and easily overlooked question: 'how often are calls succeeding but taking too long?' A dependency can be in a degraded state where it never throws an exception — every call eventually returns 200 OK — but response times balloon from 50ms to 8 seconds under load, saturating connection pools and thread capacity on the caller side exactly as badly as outright failures would, just without tripping any error-based signal. To catch this, a **slow-call duration threshold** is configured (e.g. 2 seconds), calls exceeding it are marked 'slow' regardless of their eventual outcome, and if the percentage of slow calls in the window crosses its own threshold, the breaker trips — independently of whatever the failure rate is doing. A dependency can trip the breaker purely on slowness with a 0% error rate. ## What the sliding window changes The sliding window determines what population these percentages are computed over, and its type changes the breaker's behavior under different traffic patterns. | Window type | Strength | Weakness | |---|---|---| | A count-based window ('the last N calls') | gives statistically stable percentages regardless of traffic volume | but under low traffic it can take a long time to fill, delaying detection | | A time-based window ('the last N seconds') | reacts on a predictable wall-clock cadence | but under low traffic can evaluate thresholds against a tiny, noisy sample, and under bursty traffic a single burst can dominate the window | Because a fresh service, a fresh window, or a low-traffic period can otherwise produce misleadingly extreme percentages (2 calls, 1 failed = 50%), implementations enforce a **minimum-number-of-calls floor**: the breaker will not evaluate either threshold — won't even consider tripping — until the window has accumulated at least that many calls. ## Tuning: precision versus responsiveness The trade-off in tuning these two thresholds independently is precision versus responsiveness. - Setting the slow-call duration threshold **too aggressively** (say, 200ms for a dependency whose normal p99 is 300ms) causes the breaker to classify healthy traffic as slow and trip on noise; setting it **too loose** means real degradation goes undetected until it eventually produces outright failures, by which point the caller-side resource damage (thread/connection exhaustion) may already be done. - Similarly, a failure-rate threshold set **too low** trips on transient blips (a GC pause causing two timeouts in a row); set too high, it tolerates a genuinely struggling dependency for too long. - A **wider** sliding window smooths out noise but reacts to real outages more slowly; a **narrower** window reacts fast but is more susceptible to short-lived spikes causing unnecessary trips. ## A concrete production failure mode A concrete production failure mode this design guards against: a payment gateway degrades under load such that 95% of calls still return success, just slowly — 4 to 6 seconds instead of 200ms. A breaker configured only with a failure-rate threshold would never trip, because almost nothing is technically failing, while every caller thread waiting on those calls slowly exhausts the checkout service's own capacity. A slow-call-rate threshold configured at, say, '50% of calls slower than 2 seconds trips the breaker' catches this scenario and fails fast well before the caller-side pool is exhausted, routing to a fallback instead. This is precisely why libraries like Resilience4j expose both `failureRateThreshold` and `slowCallRateThreshold/slowCallDurationThreshold` as separate, independently tunable configuration values on the same breaker instance, alongside `slidingWindowType` (count-based or time-based) and `minimumNumberOfCalls`.
- Why classify a slow call as 'slow' even if it eventually succeeds, rather than just letting the failure-rate threshold handle genuinely broken calls?Because a slow-but-successful call still consumes exactly the same caller-side resources — a thread, a connection-pool slot, request budget — for its entire duration as a failed call does, so from the caller's resource-exhaustion perspective it's just as dangerous. Ignoring slowness and only counting hard failures means a breaker can stay closed right up until the caller's own capacity collapses under successful-but-glacial responses.
- What's the risk of setting the minimum-number-of-calls floor too low, say 2?With a floor of 2, a single failed call out of two gives a 50% failure rate, which is enough to trip most default thresholds off essentially no statistical evidence — a couple of unlucky calls during a deploy or a brief network blip can trip the breaker even though the dependency is fine. A reasonable floor (commonly 10-20+) ensures the percentage reflects a real pattern rather than noise.
- How would you decide the slow-call duration threshold for a specific dependency rather than picking an arbitrary number?You'd look at that dependency's normal latency distribution — typically its p95 or p99 under healthy conditions — and set the threshold somewhat above that, so ordinary tail latency doesn't get misclassified as 'slow' but a genuine shift in the distribution (degradation) does. Setting it independent of the dependency's actual baseline is a common misconfiguration that causes either constant false trips or a threshold so loose it never fires.
It's like a restaurant health inspector who flags a kitchen not only for food that makes people sick (failure rate) but separately for orders that take absurdly long to come out even when the food is fine (slow-call rate) — both are signs something's wrong even though only one involves an actual failure.
saying these in an interview costs you the question
- Thinks a single failed call trips the breaker
- Doesn't know slow-but-successful calls can trip a breaker at all
- Can't explain what a sliding window is evaluating against
- Assumes the window always resets to empty right after a trip with no minimum-calls consideration
- Conflates the slow-call duration threshold with the overall call timeout