skip to content

In a microservice that calls a downstream payment service over HTTP, what problem does wrapping that call in a circuit breaker solve, and how does the breaker decide when to stop letting calls through?

level: juniorimportance: must knowfreq 85%

answer

  1. CLOSED/OPEN/HALF_OPEN states
  2. fail fast, not fail slow
  3. rolling window failure rate
  4. half-open trial calls
  5. protects caller not just callee

basics

~20 s

A circuit breaker watches for repeated failures calling another service and, once it sees too many, stops sending new requests for a while so the failing service isn't overloaded and your own service doesn't get stuck waiting.

solid answer

~40 s

Circuit breaker wraps outbound calls to a dependency and tracks success/failure (and slow-call) rate over a rolling window. In CLOSED state calls pass through normally. Once the failure ratio crosses a threshold, breaker trips to OPEN: calls fail fast immediately without hitting the network, protecting the caller's threads/connections and giving the downstream service room to recover instead of being hammered by retries. After a configured wait, breaker moves to HALF_OPEN and allows a small number of trial calls through; if they succeed it closes again, if they fail it reopens. This prevents cascading failure - a slow/broken dependency causing thread pool exhaustion and timeouts to ripple upstream through the whole call chain. Key trade-off: threshold tuning - too sensitive trips on transient blips, too lax lets pressure build before failing fast.

go deeper

for a junior

Should know a circuit breaker exists to stop hammering a failing dependency and roughly the closed/open idea; doesn't need state-machine precision.

for a middle

Should describe all three states, name a library (resilience4j/Hystrix), and know it needs a fallback path for the open state.

for a senior

Should discuss threshold tuning (count vs time window, minimum call volume), which failures should/shouldn't count, and how breaker interacts with retries and timeouts in the same call path.

for a principal

Should discuss breaker granularity (per-instance vs per-endpoint), thundering herd on half-open across a fleet, and where infra-level circuit breaking (service mesh outlier detection) may be preferable to app-level libraries.

## What the breaker wraps and what it counts A **circuit breaker** is a stateful wrapper placed around every outbound call a service makes to a particular dependency (a database, another microservice's HTTP endpoint, a third-party API). Concretely, it holds a **rolling window** of recent call outcomes - counted either by a fixed number of calls (e.g., the last 20) or by a time window (e.g., the last 10 seconds) - and classifies each outcome as: - **success**; - **failure** - an exception, connection error, or configured HTTP status like 5xx; - or, in more advanced implementations, **"slow call"** - a call that succeeded but exceeded a latency threshold. ## The three states The breaker itself is a three-state machine. 1. In the `CLOSED` state, everything behaves normally: calls pass straight through to the dependency and their outcomes feed the rolling window. Once the failure rate (or slow-call rate) in that window crosses a configured threshold - and only after a minimum number of calls have been observed, so a handful of early failures on a cold window doesn't trip it prematurely - the breaker transitions to `OPEN`. 2. In `OPEN` state, calls are rejected immediately, in-process, without ever touching the network: the caller gets an exception or a fallback value in microseconds instead of waiting out a timeout. 3. After a configured wait duration, the breaker moves to `HALF_OPEN`, where it allows a small, limited number of "trial" calls through to test whether the dependency has recovered; if enough of those succeed it transitions back to `CLOSED`, and if they fail it reopens for another wait period. ## Why the pattern exists The reason this pattern exists is to stop a single struggling dependency from taking down the caller, and by extension everything upstream of the caller. In a synchronous microservices architecture, a slow or hanging downstream call ties up the caller's own resources - threads, connections, memory buffering the in-flight request - for as long as the caller is willing to wait. If that dependency is degraded rather than fully down, every retry and every default socket timeout compounds the problem, because the caller keeps issuing new calls that also hang, exhausting its own thread pool. Once a service's thread pool is exhausted handling calls to one bad dependency, it can no longer serve requests that don't even touch that dependency, and its own callers start timing out waiting on it - and the failure cascades up the call graph. The circuit breaker breaks this chain by making the caller **fail fast** once it has enough evidence the dependency is unhealthy, freeing its own resources immediately instead of spending them waiting. ## The trade-off The trade-off is real: - An **aggressive threshold** protects the caller quickly but risks tripping on transient blips - a single slow GC pause or a momentary network hiccup - and then serving degraded behavior to users who would have been fine with a real call. - A **lax threshold** avoids false positives but lets pressure build for longer before the breaker actually helps, during which the caller's own resource exhaustion may already be underway. Threshold tuning (failure-rate cutoff, window size, minimum call count, wait duration) has to be calibrated per dependency based on its normal error rate and latency profile - there's no universal setting. A second, often-missed trade-off is that a circuit breaker requires you to decide what happens when it's open: some **fallback behavior** (cached/stale data, a degraded feature, a queued-for-later write, or simply a fast, clear error) has to exist, or the breaker just converts a slow failure into a fast one without actually improving the user's experience. ## Production failure modes Production failure modes cluster around a few patterns. 1. **First, misclassifying faults**: if the breaker's classifier counts expected client-side responses (like a 404 for a genuinely missing resource) as failures, it can trip for reasons that have nothing to do with the dependency's actual health - typically only 5xx responses, timeouts, and connection errors should count. 2. **Second, "flapping,"** where the dependency is right on the threshold boundary and the breaker cycles between `OPEN` and `HALF_OPEN` repeatedly, adding instability and confusing alerts. 3. **Third, granularity mismatches**: a breaker configured per-service when the service exposes many endpoints of very different reliability can trip broadly for a problem confined to one endpoint. 4. **Fourth**, at fleet scale, if every instance's breaker opens around the same time and their half-open cooldowns are synchronized, all instances can send trial calls simultaneously when the wait expires, re-tripping the dependency right as it was starting to recover - a thundering-herd problem usually fixed with jitter on the cooldown timer. ## Where you meet it Netflix popularized this pattern at scale with Hystrix; the more current standard in the JVM ecosystem is `resilience4j` (commonly used via Spring Boot's resilience4j integration), and infrastructure-level equivalents exist too - Envoy/Istio's "outlier detection" implements similar logic at the service-mesh layer, ejecting unhealthy endpoints from the load-balancing pool without any application code.

  • What should a service do when the circuit is OPEN and a request comes in - just return an error?
    It depends on business criticality: options include returning cached/stale data, a sensible default value, a degraded version of the feature, queuing the write for later processing, or a fast, clear error if none of those are safe. The point of the breaker is only to fail fast - the fallback strategy is a separate design decision that has to exist, or users just get a faster failure instead of a better experience.
  • How is a circuit breaker different from a simple timeout?
    A timeout bounds how long a single call is allowed to wait before giving up, but each call still attempts the network round trip and pays that cost. A circuit breaker tracks the aggregate history across many calls and, once the trend is clearly bad, stops attempting the call at all - avoiding even the cost of trying and timing out repeatedly.
  • Should a circuit breaker trip on a 404 Not Found response from the downstream service?
    Generally no. The breaker should typically only count faults that indicate infrastructure distress - 5xx responses, timeouts, connection refused - not expected client-side outcomes like a legitimate 404 for a missing resource. Counting 404s as failures trips the breaker for something that isn't actually a downstream health problem.
  • What happens if all your service instances open their breaker to the same dependency around the same time and all move to half-open simultaneously?
    You get a thundering-herd effect: trial requests from every instance spike together, which can be enough load to re-trip the dependency's health right as it was starting to recover. This is usually mitigated by adding jitter to the half-open cooldown timer so instances don't retest in lockstep.

Like a household circuit breaker that trips to cut power when there's a short circuit, protecting the wiring from burning out - instead of letting the fault keep drawing current until something catches fire.

saying these in an interview costs you the question

  • thinks circuit breaker retries automatically
  • conflates with load balancer failover
  • doesn't know about half-open trial state
  • trips breaker on 4xx client errors
  • no fallback defined for open state
  • assumes breaker fixes the downstream outage

context