skip to content

questions

6

In the circuit breaker pattern used to protect a caller from a failing downstream dependency, what are the three states a breaker cycles through, and what makes it flip from closed to open?

level: juniorimportance: must knowfreq 82%

answer

  1. closed→open→half-open→closed
  2. trip on threshold breach
  3. fast-fail, no real call in open
  4. probe calls in half-open
  5. electrical fuse analogy

basics

~20 s

A circuit breaker watches calls to another service. Normally it's 'closed' and lets calls through. If too many fail, it 'opens' and stops calling that service for a while, failing fast instead. After a timeout it lets a few test calls through ('half-open') to check if the service recovered.

solid answer

~30 s

Closed is the normal state: every call passes through to the real dependency and the outcome (success/failure/slow) is recorded in a rolling window. Once failures — or slow calls — cross a configured threshold within that window, the breaker trips to open: calls are short-circuited immediately, without touching the dependency, typically invoking a fallback. After a wait duration the breaker moves to half-open and lets a small number of trial calls through; if those succeed it closes, if they fail it reopens.

go deeper

for a junior

Should recall there are three states and describe the closed→open transition in plain language: too many failures trips it.

for a middle

Should describe all three states correctly including that half-open exists specifically to test recovery with limited traffic, not to resume full traffic immediately.

for a senior

Should articulate why fast-fail protects caller-side resources (threads, connection pools) and connect the pattern to preventing cascading failure across a service graph.

for a principal

Should discuss the trade-offs in wait-duration and threshold tuning at a systems level, including flapping risk and how it interacts with overall system resilience strategy, not just the mechanics of one breaker instance.

## What the breaker is, and what CLOSED does A circuit breaker is a **stateful proxy** sitting between a caller and a remote dependency — an HTTP call to another microservice, a database query, a third-party API. It has three states: `CLOSED`, `OPEN`, and `HALF_OPEN`. In `CLOSED`, every call executes for real against the dependency, and each outcome (success, exception, timeout) is recorded into a rolling record of recent calls. The breaker continuously evaluates that record against configured thresholds — most commonly a **failure-rate percentage** — and the moment the threshold is breached it transitions to `OPEN`. ## Why the pattern exists The reason this pattern exists is to stop **cascading failure**. When a downstream dependency starts erroring or hanging, callers that keep hitting it as if nothing is wrong pile up blocked threads, exhausted connection pools, and growing latency, and that pressure propagates upstream: - a slow database can take down the API layer, - which takes down the UI, - which takes down everything depending on the UI. It also hurts the struggling dependency itself: a flood of retrying callers is exactly the load it can least handle while trying to recover. The circuit breaker is named after its electrical namesake: it trips to interrupt the circuit before the damage spreads, rather than letting current (traffic) keep flowing into a fault. ## What happens once it trips Once `OPEN`, calls are short-circuited immediately: no network call is attempted at all, and the caller gets an instant rejection (an exception, or ideally a fallback response) instead of waiting out a slow timeout. The breaker starts a **wait-duration timer** at this point. When that timer elapses, the breaker moves to `HALF_OPEN`, a probationary state where it permits a small, configured number of real calls through as probes while everything else is still rejected. 1. If enough of those probes succeed, the breaker closes and resets its counters to a clean slate. 2. If they fail, it reopens and restarts the wait timer, often with some backoff so it doesn't hammer a still-broken dependency every few seconds. ## The core trade-off The core trade-off is **availability versus protection**. Failing fast during `OPEN` reduces load and latency pressure on both the caller and the dependency, but it also rejects calls that might have succeeded if the dependency had already quietly recovered — you're trading some false rejections for a guarantee against pile-up. Tuning the wait duration is itself a balancing act, and threshold tuning has the same shape: | Knob | Set too tight | Set too loose | |---|---|---| | Wait duration | too short and the breaker flaps between open and half-open under sustained partial failure, hammering a dependency that hasn't fully recovered | too long and you needlessly extend an outage after the dependency is healthy again | | Trip threshold | too sensitive and transient blips (a couple of timeouts during a GC pause) trip the breaker unnecessarily | too lax and a real outage takes too long to trigger protection | ## Failure modes in production In production: - A common failure mode is **only counting hard exceptions** toward the threshold and ignoring calls that succeed but are pathologically slow — a hung downstream that eventually returns 200 after 30 seconds looks 'healthy' to a naive breaker while it quietly exhausts every caller thread. This is why real implementations pair a failure-rate threshold with a separate **slow-call-rate threshold**. - Another common issue is **evaluating the threshold against too small a sample** — three calls at startup, two of which fail, trips a breaker that hasn't seen enough traffic to justify the decision, which is why implementations enforce a **minimum-number-of-calls floor** before they'll even evaluate. ## A concrete example A concrete example: an e-commerce checkout service wraps its call to a payment gateway in a circuit breaker. Under normal operation the breaker stays closed and every checkout hits the gateway directly. If the gateway starts timing out during an incident, the breaker trips open after a handful of failures, and every subsequent checkout attempt immediately gets a 'payments are temporarily unavailable, please retry' message instead of hanging for 30 seconds waiting on a doomed call — protecting the checkout service's own thread pool from being drained by calls to a dependency that isn't going to answer.

  • Why is failing fast in the open state considered better than letting the call proceed and just time out normally?
    A normal timeout still ties up a thread, a connection-pool slot, and often a downstream resource for the full timeout duration, and it does this for every single caller simultaneously during an outage. Failing fast in open returns immediately with near-zero cost, freeing those resources for other work and preventing the caller's own capacity from being drained by calls that are statistically very likely to fail anyway.
  • What happens if a caller doesn't check the breaker's state and calls the dependency directly, bypassing the breaker?
    The breaker only protects calls that are routed through it; a direct call bypasses the recorded outcome tracking and the fast-fail behavior entirely, so that caller experiences the full failure/timeout cost and its outcomes never contribute to the breaker's threshold evaluation, effectively creating an unprotected hole in the protection strategy.
  • Should the wait duration in the open state be fixed or does it make sense to make it adaptive?
    A fixed wait duration is simplest and is what most libraries default to, but some implementations support exponential backoff on repeated trip-reopen cycles so a persistently broken dependency isn't probed as frequently, reducing wasted probe traffic during a long outage while still recovering promptly for short-lived blips.

It's like an electrical circuit breaker in a house: normal current flows fine (closed), but if it detects a dangerous overload it trips (open) and cuts power immediately rather than letting the wiring overheat, and someone has to reset it (half-open test) before power flows normally again.

saying these in an interview costs you the question

  • Says the breaker retries the same call several times before giving up (that's retry, not circuit breaker)
  • Can't explain what happens to calls during the open state
  • Thinks half-open lets all traffic through instead of a limited probe set
  • Believes the breaker state is per-call rather than shared across calls to the same dependency
  • Confuses circuit breaker with a rate limiter or a load balancer health check

context

open as a page

A circuit breaker uses a sliding window of recent call outcomes to decide whether to trip. What's the difference between a failure-rate threshold and a slow-call-rate threshold, and how does the sliding window shape when those thresholds actually get evaluated?

level: middleimportance: must knowfreq 76%

basics

~20 s

The breaker looks at the last batch of calls. The failure-rate threshold trips it if too many of those calls errored out. The slow-call threshold trips it separately if too many calls succeeded but took too long. Both are percentages over a recent batch, not a single call.

open as a page

Once a circuit breaker's wait duration expires and it enters half-open, how does it decide whether to fully close again or trip back to open, and why does it deliberately limit how many calls are allowed through during that phase instead of just resuming full traffic?

level: middleimportance: must knowfreq 70%

basics

~20 s

Half-open lets only a small number of test calls through, not all traffic. If enough of those test calls succeed, the breaker fully reopens for business (closes). If too many fail, it goes back to blocking everything (reopens). Limiting the number of test calls avoids overwhelming a dependency that might still be fragile.

open as a page

When wiring a fallback for an open circuit breaker — say, returning a cached value or a default response instead of calling the failing dependency — what should that fallback avoid doing, and what can go wrong if it's implemented carelessly?

level: seniorimportance: should knowfreq 62%

basics

~20 s

A fallback is the backup answer given when the breaker blocks a call. It should be fast, safe, and not depend on the thing that just failed. Done badly, the fallback can itself be slow, call another struggling service, or quietly hide a real problem from the people who need to know about it.

open as a page

How do circuit breaker implementations differ across Resilience4j (Java), Hystrix (Netflix's now-legacy library), and Polly (.NET) in how you configure and wire a breaker into a call, and why did the industry largely move away from Hystrix's approach?

level: seniorimportance: should knowfreq 52%

basics

~30 s

All three let you wrap a risky call so it can fail fast and recover automatically. Hystrix was the older, heavier one from Netflix, built around wrapping calls in special 'command' objects. Resilience4j is a newer, lighter Java library that wraps a plain function call instead. Polly is the .NET equivalent, using a similar lightweight wrapping style. Hystrix is no longer actively developed, so most new Java projects use Resilience4j instead.

open as a page

In a production system running many replica instances of the same service, each with its own in-process circuit breaker guarding calls to a shared downstream dependency, what operational pitfalls can emerge from that per-instance breaker design, and how would you mitigate them?

level: principalimportance: nice to knowfreq 38%

basics

~20 s

If every copy of your service has its own separate circuit breaker watching the same downstream dependency, they can trip and recover at different, uncoordinated times, or all probe recovery at once and re-overload the dependency. Fixes include adding randomness to timing, or moving the breaker logic to a shared layer like a service mesh so it's coordinated.

open as a page