How do you stop a single failing dependency from cascading and taking down the whole service? Which Resilience4j patterns combine to achieve failure isolation?
answer
- cascade = threads block, pool exhausts, spreads up
- Bulkhead isolates concurrency/threads
- TimeLimiter bounds the wait (CompletableFuture)
- CircuitBreaker short-circuits a dead dep
- retry storms cause cascades — backoff+jitter
basics
~20 sIsolate the failing call so it can't exhaust shared resources: a circuit breaker stops hammering a dead dependency, a bulkhead caps how many concurrent calls (and threads) it can consume, a time limiter bounds waiting, and a fallback returns a degraded response instead of an error.
solid answer
~40 sCascading failure happens when a slow/dead dependency makes callers block, threads pile up, and the outage spreads upstream. Resilience4j isolates it with layered patterns: **Bulkhead** limits concurrent calls to that dependency so it can't consume the whole thread pool (`SemaphoreBulkhead` caps in-flight calls; `ThreadPoolBulkhead` gives it a dedicated pool). **TimeLimiter** caps how long a call may block, converting a hang into a fast failure. **CircuitBreaker** trips OPEN after a failure-rate/slow-call threshold, short-circuiting further calls so you stop wasting resources on a known-dead dependency, then probes via HALF_OPEN. **Fallback** turns each rejected/failed call into a degraded but valid response (graceful degradation). Order in Resilience4j's aspect (outer→inner): Retry → CircuitBreaker → RateLimiter → TimeLimiter → Bulkhead, with the fallback as the outermost recovery. Together they contain the blast radius to that one dependency.
code
java · 21 lines@Service
public class PricingGateway {
private final PricingClient client; // Feign or RestClient wrapper
public PricingGateway(PricingClient client) { this.client = client; }
// Full isolation stack for a critical, high-fan-in dependency.
// ThreadPoolBulkhead requires a CompletableFuture return type; TimeLimiter caps the wait.
@CircuitBreaker(name = "pricing", fallbackMethod = "pricingFallback")
@Bulkhead(name = "pricing", type = Bulkhead.Type.THREADPOOL)
@TimeLimiter(name = "pricing")
public CompletableFuture<Price> quote(String sku) {
return CompletableFuture.supplyAsync(() -> client.quote(sku));
}
// Matches return type + trailing Throwable; handles OPEN, timeout, bulkhead-full alike.
private CompletableFuture<Price> pricingFallback(String sku, Throwable t) {
return CompletableFuture.completedFuture(Price.lastKnownOrDefault(sku));
}
}go deeper
Know the circuit breaker + fallback idea at a high level.
Name bulkhead, time limiter, circuit breaker, fallback and what each protects against.
Explain the aspect composition order, retry-storm danger, and semaphore vs thread-pool bulkhead trade-offs.
Design the isolation policy per dependency, budget timeouts across the call chain, and reason about system-wide blast-radius containment and load-shedding.
## What 'cascading failure' means Service A calls dependency B. B gets slow (GC pause, DB lock, network). Each request to A now holds a thread waiting on B. Under load, A's thread pool fills with threads blocked on B; A can no longer serve *any* request — even ones that don't need B. A's callers then block on A, and the failure **cascades** across the system. The goal of failure isolation is to **contain the blast radius** to the one broken dependency. ## The Resilience4j toolkit (each addresses a different failure mode) ### 1. Bulkhead — resource isolation Named after ship compartments: a leak in one doesn't sink the ship. Two implementations: - **`SemaphoreBulkhead`** (`@Bulkhead(type = SEMAPHORE)`, default): caps the number of **concurrent** calls; excess calls are rejected immediately with `BulkheadFullException`. Cheap, runs on the caller thread. - **`ThreadPoolBulkhead`** (`@Bulkhead(type = THREADPOOL)`): runs calls on a **dedicated bounded thread pool + queue**, so a slow dependency can never borrow the main request threads. Returns a `CompletableFuture`. Bulkheads are the primary defence against thread-pool exhaustion — the core cascade mechanism. ### 2. TimeLimiter — bound the wait `@TimeLimiter(name=...)` caps execution time and cancels the running future, throwing `TimeoutException`. Requires the method to return a `CompletableFuture` (it works with `ThreadPoolBulkhead`). Prevents unbounded hangs — a hang is worse than an error because it holds resources indefinitely. (Also always set sane HTTP client connect/read timeouts as a baseline.) ### 3. CircuitBreaker — stop calling a dead dependency Tracks a sliding window of outcomes. When the **failure rate** or **slow-call rate** exceeds a threshold, it transitions CLOSED → OPEN and **short-circuits** all calls (throws `CallNotPermittedException`) for `waitDurationInOpenState`. Then HALF_OPEN lets a few trial calls through: success → CLOSED, failure → OPEN again. This stops you from wasting resources retrying a known-dead dependency and gives it room to recover. ### 4. Fallback — graceful degradation The recovery path (see the fallback question): return cached/default/partial data or a controlled 503 so callers get a usable response even while B is isolated. ### 5. Retry (use with care) `@Retry` re-attempts transient failures. **Danger:** retries *amplify* load on a struggling dependency and can *cause* the cascade. Only retry idempotent ops, with backoff + jitter, and keep the circuit breaker outside the retry so repeated failures still trip it. ## How they compose — aspect order With the Spring Boot starter, when multiple annotations are on one method, the Resilience4j aspects nest in a fixed order (outer → inner): `Retry( CircuitBreaker( RateLimiter( TimeLimiter( Bulkhead( call )))))` and the **fallback** wraps the whole thing as the last-resort recovery. Rationale: the bulkhead/time-limiter constrain the actual call; the circuit breaker observes aggregate health; retry sits outside so a retried call still respects the breaker; fallback catches whatever escapes. ## Putting it together ```java @CircuitBreaker(name = "pricing", fallbackMethod = "fallback") @Bulkhead(name = "pricing", type = Bulkhead.Type.THREADPOOL) @TimeLimiter(name = "pricing") public CompletableFuture<Price> quote(String sku) { ... } ``` Now a pricing outage: bulkhead confines it to pricing's own pool, time-limiter caps waits, breaker opens after the failure rate spikes, and the fallback serves last-known prices. The checkout endpoint keeps working. ## Gotchas - **Retry storms**: aggressive retry without backoff is the #1 self-inflicted cascade. - **Shared thread pools**: a `SemaphoreBulkhead` limits concurrency but the call still runs on the caller thread — for true thread isolation you need `ThreadPoolBulkhead` (or reactive/virtual threads). - **Timeouts must be shorter than the caller's**: an inner timeout longer than the upstream deadline is useless. - **Don't fall back into another flaky call** unprotected. - **Tune per dependency** — one global config rarely fits both a fast cache and a slow report service. ## When to use Every synchronous cross-service/cross-process call in a distributed system. Critical, high-fan-in dependencies get the full stack (bulkhead + timeout + breaker + fallback); trivial internal calls may need only a timeout.
- Why can a naive `@Retry` make cascading failure worse instead of better?Retries multiply the request rate against a dependency that's already overloaded, pushing it further down and holding caller threads longer per logical request. Without backoff + jitter and idempotency checks, a synchronized retry storm turns a partial outage into a full one. Keep the circuit breaker outside retry so exhausted retries still trip it, and only retry idempotent operations.
- What's the practical difference between a SemaphoreBulkhead and a ThreadPoolBulkhead for isolation?A `SemaphoreBulkhead` only counts concurrent permits and runs the call on the *caller's* thread — it caps concurrency but doesn't give the dependency its own threads, so a slow call still ties up a request thread. A `ThreadPoolBulkhead` executes on a dedicated bounded pool with a queue, so a slow dependency can never consume the main request threads — true thread isolation, at the cost of async/CompletableFuture semantics and context-propagation care.
saying these in an interview costs you the question
- Believing a circuit breaker alone prevents thread-pool exhaustion (it needs a bulkhead/timeout too).
- Adding retries without backoff/jitter, amplifying the outage.
- Thinking SemaphoreBulkhead gives the dependency its own threads.
- Setting an inner timeout longer than the caller's deadline.