skip to content

In the circuit breaker pattern used to protect a caller from a failing downstream dependency, what are the three states a breaker cycles through, and what makes it flip from closed to open?

level: juniorimportance: must knowfreq 82%

answer

  1. closed→open→half-open→closed
  2. trip on threshold breach
  3. fast-fail, no real call in open
  4. probe calls in half-open
  5. electrical fuse analogy

basics

~20 s

A circuit breaker watches calls to another service. Normally it's 'closed' and lets calls through. If too many fail, it 'opens' and stops calling that service for a while, failing fast instead. After a timeout it lets a few test calls through ('half-open') to check if the service recovered.

solid answer

~30 s

Closed is the normal state: every call passes through to the real dependency and the outcome (success/failure/slow) is recorded in a rolling window. Once failures — or slow calls — cross a configured threshold within that window, the breaker trips to open: calls are short-circuited immediately, without touching the dependency, typically invoking a fallback. After a wait duration the breaker moves to half-open and lets a small number of trial calls through; if those succeed it closes, if they fail it reopens.

go deeper

for a junior

Should recall there are three states and describe the closed→open transition in plain language: too many failures trips it.

for a middle

Should describe all three states correctly including that half-open exists specifically to test recovery with limited traffic, not to resume full traffic immediately.

for a senior

Should articulate why fast-fail protects caller-side resources (threads, connection pools) and connect the pattern to preventing cascading failure across a service graph.

for a principal

Should discuss the trade-offs in wait-duration and threshold tuning at a systems level, including flapping risk and how it interacts with overall system resilience strategy, not just the mechanics of one breaker instance.

## What the breaker is, and what CLOSED does A circuit breaker is a **stateful proxy** sitting between a caller and a remote dependency — an HTTP call to another microservice, a database query, a third-party API. It has three states: `CLOSED`, `OPEN`, and `HALF_OPEN`. In `CLOSED`, every call executes for real against the dependency, and each outcome (success, exception, timeout) is recorded into a rolling record of recent calls. The breaker continuously evaluates that record against configured thresholds — most commonly a **failure-rate percentage** — and the moment the threshold is breached it transitions to `OPEN`. ## Why the pattern exists The reason this pattern exists is to stop **cascading failure**. When a downstream dependency starts erroring or hanging, callers that keep hitting it as if nothing is wrong pile up blocked threads, exhausted connection pools, and growing latency, and that pressure propagates upstream: - a slow database can take down the API layer, - which takes down the UI, - which takes down everything depending on the UI. It also hurts the struggling dependency itself: a flood of retrying callers is exactly the load it can least handle while trying to recover. The circuit breaker is named after its electrical namesake: it trips to interrupt the circuit before the damage spreads, rather than letting current (traffic) keep flowing into a fault. ## What happens once it trips Once `OPEN`, calls are short-circuited immediately: no network call is attempted at all, and the caller gets an instant rejection (an exception, or ideally a fallback response) instead of waiting out a slow timeout. The breaker starts a **wait-duration timer** at this point. When that timer elapses, the breaker moves to `HALF_OPEN`, a probationary state where it permits a small, configured number of real calls through as probes while everything else is still rejected. 1. If enough of those probes succeed, the breaker closes and resets its counters to a clean slate. 2. If they fail, it reopens and restarts the wait timer, often with some backoff so it doesn't hammer a still-broken dependency every few seconds. ## The core trade-off The core trade-off is **availability versus protection**. Failing fast during `OPEN` reduces load and latency pressure on both the caller and the dependency, but it also rejects calls that might have succeeded if the dependency had already quietly recovered — you're trading some false rejections for a guarantee against pile-up. Tuning the wait duration is itself a balancing act, and threshold tuning has the same shape: | Knob | Set too tight | Set too loose | |---|---|---| | Wait duration | too short and the breaker flaps between open and half-open under sustained partial failure, hammering a dependency that hasn't fully recovered | too long and you needlessly extend an outage after the dependency is healthy again | | Trip threshold | too sensitive and transient blips (a couple of timeouts during a GC pause) trip the breaker unnecessarily | too lax and a real outage takes too long to trigger protection | ## Failure modes in production In production: - A common failure mode is **only counting hard exceptions** toward the threshold and ignoring calls that succeed but are pathologically slow — a hung downstream that eventually returns 200 after 30 seconds looks 'healthy' to a naive breaker while it quietly exhausts every caller thread. This is why real implementations pair a failure-rate threshold with a separate **slow-call-rate threshold**. - Another common issue is **evaluating the threshold against too small a sample** — three calls at startup, two of which fail, trips a breaker that hasn't seen enough traffic to justify the decision, which is why implementations enforce a **minimum-number-of-calls floor** before they'll even evaluate. ## A concrete example A concrete example: an e-commerce checkout service wraps its call to a payment gateway in a circuit breaker. Under normal operation the breaker stays closed and every checkout hits the gateway directly. If the gateway starts timing out during an incident, the breaker trips open after a handful of failures, and every subsequent checkout attempt immediately gets a 'payments are temporarily unavailable, please retry' message instead of hanging for 30 seconds waiting on a doomed call — protecting the checkout service's own thread pool from being drained by calls to a dependency that isn't going to answer.

  • Why is failing fast in the open state considered better than letting the call proceed and just time out normally?
    A normal timeout still ties up a thread, a connection-pool slot, and often a downstream resource for the full timeout duration, and it does this for every single caller simultaneously during an outage. Failing fast in open returns immediately with near-zero cost, freeing those resources for other work and preventing the caller's own capacity from being drained by calls that are statistically very likely to fail anyway.
  • What happens if a caller doesn't check the breaker's state and calls the dependency directly, bypassing the breaker?
    The breaker only protects calls that are routed through it; a direct call bypasses the recorded outcome tracking and the fast-fail behavior entirely, so that caller experiences the full failure/timeout cost and its outcomes never contribute to the breaker's threshold evaluation, effectively creating an unprotected hole in the protection strategy.
  • Should the wait duration in the open state be fixed or does it make sense to make it adaptive?
    A fixed wait duration is simplest and is what most libraries default to, but some implementations support exponential backoff on repeated trip-reopen cycles so a persistently broken dependency isn't probed as frequently, reducing wasted probe traffic during a long outage while still recovering promptly for short-lived blips.

It's like an electrical circuit breaker in a house: normal current flows fine (closed), but if it detects a dangerous overload it trips (open) and cuts power immediately rather than letting the wiring overheat, and someone has to reset it (half-open test) before power flows normally again.

saying these in an interview costs you the question

  • Says the breaker retries the same call several times before giving up (that's retry, not circuit breaker)
  • Can't explain what happens to calls during the open state
  • Thinks half-open lets all traffic through instead of a limited probe set
  • Believes the breaker state is per-call rather than shared across calls to the same dependency
  • Confuses circuit breaker with a rate limiter or a load balancer health check

context