Service A calls Service B synchronously over REST, and B in turn calls Service C synchronously. C starts responding slowly under load. What failure mode does this create for A, and what are the standard mitigations?
answer
- thread pool exhaustion backward through the chain
- circuit breaker fails fast
- bulkhead isolates per-dependency pools
- retry storm from naive retries
- timeouts must shrink going up the chain
basics
~20 sC being slow makes B slow, which makes A slow too, and if nothing stops it, all three can eventually run out of capacity and fail together, like one traffic jam backing up onto every road behind it. Fixes include setting time limits, giving up early instead of piling up more requests, and having a backup plan when the answer doesn't come.
solid answer
~60 sThis is a cascading failure caused by synchronous chaining: A's threads or connections waiting on B are themselves waiting on B's threads waiting on C, so C's slowness propagates backward and can exhaust connection pools, thread pools, or memory at every hop, turning one slow dependency into an outage across the whole chain, especially if A also serves other callers who now queue up behind the ones stuck waiting on C. The standard mitigations are: aggressive, tuned timeouts at every hop so no caller waits indefinitely; circuit breakers that stop sending requests to a dependency once its failure/latency rate crosses a threshold, failing fast instead of piling on; bulkheads that isolate resource pools per dependency so a slow C can't starve threads needed for calls to healthy dependencies; and load shedding or backpressure so A can reject excess work rather than queue it unboundedly. Retries need care too, since naive retries on a slow dependency add more load to an already struggling C and can make things worse, so they should be bounded, use backoff, and ideally only retry idempotent operations.
go deeper
Should recognize that a slow downstream service can make upstream services slow too, and name 'timeout' as a basic mitigation.
Should explain the resource-exhaustion mechanism (threads/connections held waiting) and name circuit breakers and bounded retries as mitigations beyond just timeouts.
Should explain bulkheading, retry storms, and how timeouts need to be tuned per hop, and connect the pattern to a concrete framework or real incident (e.g., Hystrix, a documented cascading-failure outage).
Should reason about this at a systems level: how service mesh or platform-level resilience defaults (rather than per-service library code) reduce this risk org-wide, how to design SLOs and timeout budgets across a whole call graph, and how to build organizational practices (chaos testing, dependency health dashboards) that catch this class of risk before it causes an incident.
## How latency propagates backward When Service A calls B synchronously, and B calls C synchronously, the three services form a chain where each caller's completion time is lower-bounded by the sum of the calls beneath it. If C starts responding slowly, say its p99 latency climbs from 50ms to 5 seconds under load: - B's calls to C now take up to 5 seconds each, which means B's own response time to A climbs correspondingly; - A's response time to whatever is calling A climbs on top of that. This is the mechanical root of a **cascading failure**: latency doesn't stay contained at the layer where it originated, it propagates backward through every synchronous hop in the chain. ## Why it is dangerous: resource exhaustion The reason this is dangerous rather than just "slightly slower" is **resource exhaustion**. Most services handle concurrent requests using a bounded pool of threads, connections, or both. 1. If B is waiting 5 seconds per call to C instead of 50ms, and B is receiving a steady stream of requests from A, B's thread pool fills up with threads all blocked waiting on C, much faster than it would drain normally, because each thread is now occupied 100x longer per request. 2. Once B's pool is exhausted, B stops accepting new requests, or starts queuing them, which makes B itself look slow or unresponsive to A, even though B's own code has no bug at all, it's simply out of capacity because of C. 3. The same thing then happens one layer up: A's pool fills up waiting on now-slow-or-unresponsive B, and if A serves other callers or other endpoints unrelated to the B-and-C path, those unrelated callers can get starved too, because they're competing for the same exhausted thread pool. This is how one slow dependency deep in a call graph can take down services that have no direct relationship to it at all, which is exactly what happened in several well-documented AWS and other cloud-provider outages where a single overloaded internal dependency cascaded into a much broader service disruption. ## The standard mitigations The standard mitigations each address a different part of this mechanism. - **Timeouts** bound how long any single hop will wait, which limits (but doesn't eliminate) the damage; the key subtlety is that timeouts need to be tuned per hop, with each caller's timeout shorter than what it can tolerate given its own callers' timeouts, otherwise you just get failures further up the chain instead of contained ones. - **Circuit breakers** go a step further: rather than waiting out the full timeout on every request to a dependency that's clearly unhealthy, a circuit breaker tracks recent failure or latency rates and, once a threshold is crossed, "opens" and fails fast (often returning a fallback or a cached response) without even attempting the call, which protects both the caller's own resources and the struggling dependency from additional load while it's already in trouble. - **Bulkheads** borrow the naval engineering idea of watertight compartments: instead of one shared thread or connection pool for all outbound calls, each dependency gets its own isolated pool, so C exhausting its allocated pool doesn't starve the threads needed to call an unrelated, healthy dependency D. - **Load shedding and backpressure** address the front door: rather than accepting unlimited incoming requests and letting them all queue up waiting for scarce resources, a service can start actively rejecting or degrading requests once it's near capacity, which keeps the requests it can serve fast rather than letting everything degrade together. ## Why retries can make it worse Retries deserve special mention because they're an intuitive fix that can actively worsen this failure mode. If B automatically retries a failed or slow call to C, and C is slow because it's overloaded, that retry adds more load to C, which can push its latency even higher, which triggers more retries from more callers, in a feedback loop sometimes called a **retry storm**. The mitigation is to make retries: - **bounded** — a small fixed number of attempts; - **backoff** — use exponential backoff with jitter so retries from many concurrent callers don't all land on C at the same moment; - **safe to repeat** — only retry operations that are idempotent, meaning calling twice has the same effect as calling once, since retrying a non-idempotent operation like "charge the card" can cause duplicate side effects on top of the latency problem. ## Where the defenses were hardened A concrete, well-documented pattern: Netflix's Hystrix library, one of the earliest widely-adopted circuit-breaker implementations for microservices, was built specifically because Netflix's own service-to-service call graphs were experiencing exactly this cascading-failure pattern at scale, and it popularized bulkheading (per-dependency thread pools) and circuit breaking as a combined, standard defensive pattern that's since become a normal expectation, whether implemented via a library, a service mesh's built-in resilience features (like Istio or Linkerd), or an API gateway layer, in any production system with more than one or two hops of synchronous service-to-service calls.
- Why isn't simply setting a long, generous timeout at every hop a sufficient fix on its own?A generous timeout still lets each hop hold its thread or connection for a long time while waiting, which is exactly what exhausts the pool; a long timeout delays the failure and makes the resource exhaustion window bigger, not smaller. Timeouts need to be tight enough that a struggling dependency's threads get freed quickly, combined with circuit breakers so repeated failed attempts don't keep re-triggering the same long wait.
- How does a circuit breaker decide when to 'open' and stop sending requests to a dependency, and what happens while it's open?It typically tracks a rolling window of recent call outcomes (failures, timeouts, or elevated latency) and opens once the failure or slow-response rate crosses a configured threshold within that window. While open, it fails fast for a cooldown period, often returning a fallback value or cached response instead of calling the dependency at all, then periodically allows a small number of trial requests through (a 'half-open' state) to check whether the dependency has recovered before fully closing again.
- If C's slowness is caused by C itself being overwhelmed by too much traffic, how do bulkheads on B's side actually help C recover, if at all?Bulkheads on B's side don't directly help C recover; their job is to contain the blast radius on B so that C's problem doesn't also take down B's calls to other, healthy dependencies. What actually helps C recover is B's circuit breaker and load shedding reducing the volume of requests reaching C once it's clearly struggling, combined with C having its own capacity-based backpressure or autoscaling to shed or absorb load.
It's a traffic jam on a single lane feeding a highway on-ramp: one stalled car (slow C) backs up traffic onto the on-ramp (B), which then backs up onto the surface street feeding it (A), even though the surface street itself has no problem at all.
saying these in an interview costs you the question
- Suggests only 'add more retries' as the fix without mentioning backoff or idempotency
- Doesn't connect the failure to thread/connection pool exhaustion, treating it as vague 'slowness'
- Thinks a timeout alone fully prevents cascading failure
- Can't explain why retries can make an overloaded dependency worse
- Doesn't distinguish bulkheading (resource isolation) from circuit breaking (fail-fast on unhealthy dependency)