Once a circuit breaker's wait duration expires and it enters half-open, how does it decide whether to fully close again or trip back to open, and why does it deliberately limit how many calls are allowed through during that phase instead of just resuming full traffic?
answer
- small trial-call count, not full traffic
- trial outcomes evaluated against same-style threshold
- close on success, reopen on failure
- avoids re-overloading a fragile dependency
- thundering herd risk across many breaker instances
basics
~20 sHalf-open lets only a small number of test calls through, not all traffic. If enough of those test calls succeed, the breaker fully reopens for business (closes). If too many fail, it goes back to blocking everything (reopens). Limiting the number of test calls avoids overwhelming a dependency that might still be fragile.
solid answer
~30 sIn half-open, the breaker permits a configured number of trial calls — e.g. 10 — and evaluates their outcomes against the same kind of failure/slow-call-rate threshold used in closed state, just over that small trial population. Meeting the threshold on the trial calls closes the breaker and resets its window; failing it reopens the breaker and restarts the wait-duration timer. The call count is capped deliberately because resuming full production traffic immediately risks re-triggering the exact overload that tripped the breaker in the first place, especially if the dependency is only partially recovered.
go deeper
Should know half-open means only a few test calls go through, and success closes the breaker while failure reopens it.
Should explain why the trial count is capped — protecting a possibly-fragile dependency from a full traffic resumption — not just that it is.
Should discuss the trade-off between trial-count size (statistical reliability vs blast radius) and describe backoff strategies on repeated failed recovery.
Should identify the thundering-herd-on-recovery problem across many replica instances and propose mitigations like jitter or centralizing breaker state at a mesh/gateway layer.
## A deliberately constrained state Half-open is a deliberately constrained **probationary state**, and understanding why it's constrained the way it is matters more than memorizing that it exists. When the open-state wait-duration timer expires, the breaker doesn't simply flip back to closed and let all traffic resume — it transitions to half-open and permits only a small, explicitly configured number of calls through (a 'permitted number of calls in half-open state', commonly single digits to low tens) while every other concurrent request during this window continues to be short-circuited exactly as in open state. ## How the trial probes are judged Those permitted calls are **trial probes**: real calls to the real dependency, with real outcomes recorded. The breaker evaluates those outcomes against essentially the same kind of threshold logic used in closed state — a failure-rate and/or slow-call-rate percentage — just computed over this small trial population instead of the full sliding window. - If the trial calls meet the success bar, the breaker transitions to **closed** and resets its statistics to a clean slate, resuming full unrestricted traffic. - If the trial calls fail to meet the bar, the breaker transitions back to **open**, restarting the wait-duration timer (sometimes with backoff, so repeated failed recovery attempts wait progressively longer between tries). ## Why the trial count is capped The reason the number of trial calls is capped rather than just resuming full traffic is **risk containment**. A dependency can be in one of several states when the wait duration expires: 1. fully recovered, 2. partially recovered (able to handle some load but not full production volume yet), 3. or still down. If the breaker dumped its entire backlog of accumulated, pent-up production traffic onto the dependency the instant the timer expired, a partially-recovered dependency could be knocked straight back into failure by the very traffic surge the breaker was protecting it from — the breaker would effectively cause the exact re-trip it exists to avoid, and worse, do so with a burst rather than a controlled trickle. Limiting the trial to a handful of calls means the breaker tests recovery cheaply: if the dependency is still broken, only those few calls pay the cost of finding out, not the entire caller population. ## The trade-off in sizing the probe The trade-off is **speed of detection versus safety of the probe**. - A very small trial count (say, 1-2 calls) is the gentlest possible test but is statistically noisy — a single unlucky timeout on an otherwise-recovered dependency sends the breaker straight back to open, extending the outage from the caller's perspective for another full wait-duration cycle even though the dependency was fine. - A larger trial count gives a more reliable signal about true recovery but exposes more of the (possibly still-fragile) dependency to load during the probe, and takes longer for all trial calls to complete before a decision can be made, especially if the dependency is still slow. Tuning this number is the same kind of statistical-noise-versus-responsiveness trade-off seen in threshold tuning generally. ## The thundering-herd-on-recovery failure mode A production failure mode specific to half-open is what's sometimes informally called the **thundering-herd-on-recovery** problem: if a service is deployed as many replicas, each running its own in-process breaker independently against the same shared downstream dependency, those breakers tend to trip around the same time under a shared incident and their wait-duration timers, started at roughly the same moment, expire at roughly the same moment too. Dozens or hundreds of replicas can then simultaneously enter half-open and simultaneously send their permitted trial calls at the exact same instant, effectively reconstituting the very load spike that caused the outage, just synchronized. Mitigations include: - **jittering the wait duration** per instance so timers don't expire in lockstep, - or **moving circuit breaking to a shared layer** (like a service mesh sidecar) that can coordinate state across replicas rather than each instance probing independently. ## A concrete scenario A concrete scenario: a search service's breaker trips against a recommendations backend that's been overloaded and returning timeouts. After a 30-second wait duration, the breaker allows 5 trial calls through. If all 5 succeed within acceptable latency, the breaker closes and the search service resumes calling recommendations normally for all subsequent requests. If 3 of the 5 still time out, the breaker reopens, waits another 30 seconds (or a backed-off longer interval), and tries again — protecting the recommendations backend from a full traffic surge while it's still digging itself out, and protecting the search service's own threads from being re-exhausted by a premature full resumption.
- Would it ever make sense to let the permitted-calls-in-half-open count be dynamic rather than a fixed number?Yes — some implementations or custom wiring scale the trial count based on how confident you want to be, or based on the criticality of the dependency; a payment gateway might warrant a larger, more statistically reliable trial batch before fully trusting it again, while a low-stakes recommendations service might tolerate a smaller, faster trial. It's a deliberate design knob, not a fixed law of the pattern.
- What should happen to requests that arrive during half-open but aren't among the permitted trial calls?They're treated the same as requests during open state: short-circuited immediately, typically routed to the fallback, without consuming one of the limited trial slots. Only a bounded number of concurrent or sequential calls actually reach the dependency during half-open; everything else continues to fail fast until the breaker resolves to closed or open.
- How does backoff on repeated open→half-open→open cycles change the recovery behavior compared to a fixed wait duration?With backoff, each failed recovery attempt increases the wait duration before the next probe (e.g. doubling it, up to a cap), which reduces wasted probe traffic against a dependency that's down for an extended period and avoids the breaker uselessly hammering it every few seconds; the cost is that a dependency which recovers quickly after several failed early probes will take longer to be detected as healthy again since the wait window has grown.
It's like reopening a bridge after repairs by first sending a single test truck across at low speed instead of immediately letting the full rush-hour traffic jam back onto it — if the test truck makes it fine, you open the lanes; if it doesn't, you close the bridge again before a thousand cars pile onto it.
saying these in an interview costs you the question
- Thinks half-open resumes all traffic immediately
- Can't explain what determines close-vs-reopen from half-open
- Doesn't recognize the risk of overwhelming a partially recovered dependency
- Unaware that concurrent breaker instances can synchronize their half-open probes into a load spike
- Assumes half-open is a permanent steady state rather than transient