skip to content

Resilience & Observability

Applying circuit breakers, bulkheads and retries with backoff so one failing dependency does not cascade, and instrumenting everything so you can see it happen. Trace context and correlated logs are what make a cross-service incident debuggable at all.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a microservice that calls a downstream payment service over HTTP, what problem does wrapping that call in a circuit breaker solve, and how does the breaker decide when to stop letting calls through?

level: juniorimportance: must knowfreq 85%

answer

  1. CLOSED/OPEN/HALF_OPEN states
  2. fail fast, not fail slow
  3. rolling window failure rate
  4. half-open trial calls
  5. protects caller not just callee

basics

~20 s

A circuit breaker watches for repeated failures calling another service and, once it sees too many, stops sending new requests for a while so the failing service isn't overloaded and your own service doesn't get stuck waiting.

solid answer

~40 s

Circuit breaker wraps outbound calls to a dependency and tracks success/failure (and slow-call) rate over a rolling window. In CLOSED state calls pass through normally. Once the failure ratio crosses a threshold, breaker trips to OPEN: calls fail fast immediately without hitting the network, protecting the caller's threads/connections and giving the downstream service room to recover instead of being hammered by retries. After a configured wait, breaker moves to HALF_OPEN and allows a small number of trial calls through; if they succeed it closes again, if they fail it reopens. This prevents cascading failure - a slow/broken dependency causing thread pool exhaustion and timeouts to ripple upstream through the whole call chain. Key trade-off: threshold tuning - too sensitive trips on transient blips, too lax lets pressure build before failing fast.

go deeper

for a junior

Should know a circuit breaker exists to stop hammering a failing dependency and roughly the closed/open idea; doesn't need state-machine precision.

for a middle

Should describe all three states, name a library (resilience4j/Hystrix), and know it needs a fallback path for the open state.

for a senior

Should discuss threshold tuning (count vs time window, minimum call volume), which failures should/shouldn't count, and how breaker interacts with retries and timeouts in the same call path.

for a principal

Should discuss breaker granularity (per-instance vs per-endpoint), thundering herd on half-open across a fleet, and where infra-level circuit breaking (service mesh outlier detection) may be preferable to app-level libraries.

## What the breaker wraps and what it counts A **circuit breaker** is a stateful wrapper placed around every outbound call a service makes to a particular dependency (a database, another microservice's HTTP endpoint, a third-party API). Concretely, it holds a **rolling window** of recent call outcomes - counted either by a fixed number of calls (e.g., the last 20) or by a time window (e.g., the last 10 seconds) - and classifies each outcome as: - **success**; - **failure** - an exception, connection error, or configured HTTP status like 5xx; - or, in more advanced implementations, **"slow call"** - a call that succeeded but exceeded a latency threshold. ## The three states The breaker itself is a three-state machine. 1. In the `CLOSED` state, everything behaves normally: calls pass straight through to the dependency and their outcomes feed the rolling window. Once the failure rate (or slow-call rate) in that window crosses a configured threshold - and only after a minimum number of calls have been observed, so a handful of early failures on a cold window doesn't trip it prematurely - the breaker transitions to `OPEN`. 2. In `OPEN` state, calls are rejected immediately, in-process, without ever touching the network: the caller gets an exception or a fallback value in microseconds instead of waiting out a timeout. 3. After a configured wait duration, the breaker moves to `HALF_OPEN`, where it allows a small, limited number of "trial" calls through to test whether the dependency has recovered; if enough of those succeed it transitions back to `CLOSED`, and if they fail it reopens for another wait period. ## Why the pattern exists The reason this pattern exists is to stop a single struggling dependency from taking down the caller, and by extension everything upstream of the caller. In a synchronous microservices architecture, a slow or hanging downstream call ties up the caller's own resources - threads, connections, memory buffering the in-flight request - for as long as the caller is willing to wait. If that dependency is degraded rather than fully down, every retry and every default socket timeout compounds the problem, because the caller keeps issuing new calls that also hang, exhausting its own thread pool. Once a service's thread pool is exhausted handling calls to one bad dependency, it can no longer serve requests that don't even touch that dependency, and its own callers start timing out waiting on it - and the failure cascades up the call graph. The circuit breaker breaks this chain by making the caller **fail fast** once it has enough evidence the dependency is unhealthy, freeing its own resources immediately instead of spending them waiting. ## The trade-off The trade-off is real: - An **aggressive threshold** protects the caller quickly but risks tripping on transient blips - a single slow GC pause or a momentary network hiccup - and then serving degraded behavior to users who would have been fine with a real call. - A **lax threshold** avoids false positives but lets pressure build for longer before the breaker actually helps, during which the caller's own resource exhaustion may already be underway. Threshold tuning (failure-rate cutoff, window size, minimum call count, wait duration) has to be calibrated per dependency based on its normal error rate and latency profile - there's no universal setting. A second, often-missed trade-off is that a circuit breaker requires you to decide what happens when it's open: some **fallback behavior** (cached/stale data, a degraded feature, a queued-for-later write, or simply a fast, clear error) has to exist, or the breaker just converts a slow failure into a fast one without actually improving the user's experience. ## Production failure modes Production failure modes cluster around a few patterns. 1. **First, misclassifying faults**: if the breaker's classifier counts expected client-side responses (like a 404 for a genuinely missing resource) as failures, it can trip for reasons that have nothing to do with the dependency's actual health - typically only 5xx responses, timeouts, and connection errors should count. 2. **Second, "flapping,"** where the dependency is right on the threshold boundary and the breaker cycles between `OPEN` and `HALF_OPEN` repeatedly, adding instability and confusing alerts. 3. **Third, granularity mismatches**: a breaker configured per-service when the service exposes many endpoints of very different reliability can trip broadly for a problem confined to one endpoint. 4. **Fourth**, at fleet scale, if every instance's breaker opens around the same time and their half-open cooldowns are synchronized, all instances can send trial calls simultaneously when the wait expires, re-tripping the dependency right as it was starting to recover - a thundering-herd problem usually fixed with jitter on the cooldown timer. ## Where you meet it Netflix popularized this pattern at scale with Hystrix; the more current standard in the JVM ecosystem is `resilience4j` (commonly used via Spring Boot's resilience4j integration), and infrastructure-level equivalents exist too - Envoy/Istio's "outlier detection" implements similar logic at the service-mesh layer, ejecting unhealthy endpoints from the load-balancing pool without any application code.

  • What should a service do when the circuit is OPEN and a request comes in - just return an error?
    It depends on business criticality: options include returning cached/stale data, a sensible default value, a degraded version of the feature, queuing the write for later processing, or a fast, clear error if none of those are safe. The point of the breaker is only to fail fast - the fallback strategy is a separate design decision that has to exist, or users just get a faster failure instead of a better experience.
  • How is a circuit breaker different from a simple timeout?
    A timeout bounds how long a single call is allowed to wait before giving up, but each call still attempts the network round trip and pays that cost. A circuit breaker tracks the aggregate history across many calls and, once the trend is clearly bad, stops attempting the call at all - avoiding even the cost of trying and timing out repeatedly.
  • Should a circuit breaker trip on a 404 Not Found response from the downstream service?
    Generally no. The breaker should typically only count faults that indicate infrastructure distress - 5xx responses, timeouts, connection refused - not expected client-side outcomes like a legitimate 404 for a missing resource. Counting 404s as failures trips the breaker for something that isn't actually a downstream health problem.
  • What happens if all your service instances open their breaker to the same dependency around the same time and all move to half-open simultaneously?
    You get a thundering-herd effect: trial requests from every instance spike together, which can be enough load to re-trip the dependency's health right as it was starting to recover. This is usually mitigated by adding jitter to the half-open cooldown timer so instances don't retest in lockstep.

Like a household circuit breaker that trips to cut power when there's a short circuit, protecting the wiring from burning out - instead of letting the fault keep drawing current until something catches fire.

saying these in an interview costs you the question

  • thinks circuit breaker retries automatically
  • conflates with load balancer failover
  • doesn't know about half-open trial state
  • trips breaker on 4xx client errors
  • no fallback defined for open state
  • assumes breaker fixes the downstream outage

context

open as a page

A service calls three downstream dependencies (inventory, pricing, and recommendations) using a shared thread pool for all outbound HTTP calls. Recommendations starts responding slowly. Why can that alone stall inventory and pricing calls too, and what does the bulkhead pattern do about it?

level: middleimportance: must knowfreq 65%

basics

~20 s

If all outgoing calls share one pool of worker threads, a slow dependency can hog every thread, leaving none free for calls to healthy dependencies too. A bulkhead gives each dependency its own separate, limited pool so one slow dependency can't starve the rest.

open as a page

A client calls a downstream service and gets a connection timeout. It retries three times with exponential backoff and jitter before giving up. Why is the jitter necessary in addition to the exponential growth, and what can go wrong if retries are added without a cap on total attempts or a retry budget?

level: middleimportance: must knowfreq 80%

basics

~10 s

Adding random jitter to retry delays stops many clients from all retrying at exactly the same moment and slamming the recovering service again; without limits, retries can multiply traffic and make an outage worse.

open as a page

When a single user-facing request fans out across five microservices, what mechanism lets a tool like Jaeger reconstruct the full call tree and show which of those five services caused the added latency, and what has to happen at each service boundary for that to work?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Each call in the chain gets tagged with a shared trace ID, and each step gets its own ID linked back to whoever called it. Tools like Jaeger stitch these into one timeline showing where time actually went.

open as a page

Three microservices each write their own log lines to separate log streams while handling one incoming HTTP request. What has to be threaded through that request for an engineer to later pull up every log line from all three services that belongs to that one request, and how does this relate to (but differ from) distributed tracing?

level: seniorimportance: should knowfreq 60%

basics

~20 s

A shared ID gets attached to a request and passed to every service it touches. If each service logs that ID on every line, you can search for the ID and pull all related lines from every service, in order.

open as a page

A Kubernetes-deployed microservice exposes both a liveness probe and a readiness probe, plus a /metrics endpoint scraped by Prometheus. What's the functional difference between the two probes, and what happens to the service's traffic and pod status if only one of them is misconfigured to always report healthy?

level: principalimportance: should knowfreq 55%

basics

~20 s

Liveness checks whether the app should be restarted because it's stuck; readiness checks whether it should currently receive traffic. If readiness is broken and always says "yes," a struggling instance keeps getting new requests it can't handle while it's failing.

open as a page