How do you add a circuit breaker around a WebClient call in a reactive Spring app, and why?
answer
- CLOSED -> OPEN -> HALF_OPEN state machine
- Resilience4j (no native WebFlux breaker)
- transformDeferred(CircuitBreakerOperator.of(cb))
- breaker OUTSIDE retries; pair with timeout
- CallNotPermittedException + fallback
basics
~20 sSpring doesn't have a built-in circuit breaker; use Resilience4j (via spring-cloud-circuitbreaker or its reactor operators). Wrap the WebClient Mono with a CircuitBreakerOperator or transformDeferred(cb) so that after too many failures it 'opens' and fast-fails instead of calling the dying upstream.
solid answer
~40 sA circuit breaker stops you from hammering a failing dependency: after the failure rate crosses a threshold it trips **OPEN** and calls fail fast (no network hit); after a wait it goes **HALF_OPEN**, letting a few probe calls through; success returns it to **CLOSED**. WebFlux has no native breaker — you use **Resilience4j**. Reactively you wrap the Mono: `mono.transformDeferred(CircuitBreakerOperator.of(circuitBreaker))`, or use Spring Cloud CircuitBreaker's `ReactiveCircuitBreakerFactory` / `ReactiveCircuitBreaker.run(mono, fallback)`. Provide a **fallback** (cached value, default, or a fast error) via `onErrorResume` or the factory's fallback arg. Order matters: put the breaker **outside** retries so a burst of retries counts as failures and can trip it, and pair with a bulkhead/timeout. The breaker protects *you* (fail fast, free threads) and the upstream (stops the pile-on).
code
java · 25 linesimport io.github.resilience4j.circuitbreaker.CircuitBreaker;
import io.github.resilience4j.reactor.circuitbreaker.operator.CircuitBreakerOperator;
import io.github.resilience4j.circuitbreaker.CallNotPermittedException;
CircuitBreaker cb = CircuitBreaker.of("quotesService",
CircuitBreakerConfig.custom()
.slidingWindowType(SlidingWindowType.COUNT_BASED)
.slidingWindowSize(20)
.minimumNumberOfCalls(10)
.failureRateThreshold(50f)
.waitDurationInOpenState(Duration.ofSeconds(10))
.permittedNumberOfCallsInHalfOpenState(3)
.build());
Mono<Quote> quote = webClient.get()
.uri("/quotes/{s}", symbol)
.retrieve()
.bodyToMono(Quote.class)
.timeout(Duration.ofSeconds(2))
.retryWhen(Retry.backoff(2, Duration.ofMillis(200)))
// breaker OUTSIDE the retry: it records the final outcome
.transformDeferred(CircuitBreakerOperator.of(cb))
// graceful fallback when the breaker is OPEN (fast-fail) or on error
.onErrorResume(CallNotPermittedException.class,
ex -> Mono.just(Quote.cachedFallback(symbol)));go deeper
Know a circuit breaker fast-fails after repeated failures and that Resilience4j provides it (not built into WebFlux).
Explain CLOSED/OPEN/HALF_OPEN and wiring via transformDeferred(CircuitBreakerOperator.of(cb)) plus a fallback.
Discuss breaker-vs-retry ordering, pairing with TimeLimiter/Bulkhead, per-dependency scoping, and config thresholds.
Define resilience patterns org-wide, budget failure thresholds against SLOs, avoid cascading failure, and standardize fallbacks/observability across clients.
**The problem a circuit breaker solves:** when a dependency is down or overloaded, every call to it wastes time (waiting for timeouts), ties up connections/threads, and adds load to the struggling upstream — potentially cascading the failure back into your service. A **circuit breaker** detects sustained failure and *stops calling* for a while, failing fast instead. **State machine (Resilience4j):** - **CLOSED** — normal; calls pass through. The breaker records outcomes in a sliding window (count-based or time-based). - **OPEN** — once the **failure rate** (or slow-call rate) in the window exceeds a threshold, it trips. All calls fail *immediately* with `CallNotPermittedException` — no network attempt. It stays open for `waitDurationInOpenState`. - **HALF_OPEN** — after the wait, a limited number of *probe* calls are allowed. If enough succeed, it closes; if they fail, it re-opens. This tests recovery without a full flood. Key config: `failureRateThreshold`, `slowCallRateThreshold` + `slowCallDurationThreshold`, `slidingWindowType` (COUNT_BASED/TIME_BASED), `slidingWindowSize`, `minimumNumberOfCalls`, `waitDurationInOpenState`, `permittedNumberOfCallsInHalfOpenState`. **Why not built-in?** Spring WebFlux/WebClient has no native circuit breaker. The standard solution is **Resilience4j** (Hystrix is retired). Two integration styles: 1. **Resilience4j reactor operators directly:** `mono.transformDeferred(CircuitBreakerOperator.of(circuitBreaker))`. `transformDeferred` (not `transform`) ensures the breaker is applied per-subscription. Add `TimeLimiterOperator`, `BulkheadOperator`, `RetryOperator` similarly. 2. **Spring Cloud CircuitBreaker abstraction:** inject `ReactiveCircuitBreakerFactory`, create a breaker, and call `breaker.run(webClientMono, throwable -> fallbackMono)`. This is a vendor-neutral facade (backed by Resilience4j). **Fallbacks:** a breaker is most useful with a graceful fallback — a cached/last-known value, a sensible default, an empty result, or a fast, clear error. With raw operators you add `.onErrorResume(CallNotPermittedException.class, ex -> fallback)`; with Spring Cloud you pass the fallback lambda to `run`. **Ordering with retries/timeouts (critical):** - Put the **circuit breaker outside the retry** so that repeated retry failures are recorded as failures and can trip the breaker — otherwise retries mask the failures from the breaker. - Put a **timeout / TimeLimiter** so slow calls are counted (as failures or slow-calls) rather than hanging. - Add a **bulkhead** to cap concurrent in-flight calls so one slow dependency can't consume all resources. Resilience4j documents a recommended order: Bulkhead → TimeLimiter → RateLimiter → CircuitBreaker → Retry (retry outermost) for some setups, but the practical rule for breaker-vs-retry is: **breaker sees the outcome of the retries**, i.e. retry is inside the breaker's guarded call, OR the breaker wraps the retried mono. Confirm the intended semantics for your case. **Gotchas:** - Using `transform` instead of `transformDeferred` shares state incorrectly across subscriptions. - A breaker without a timeout may never trip on *hangs* (only on errors) — pair them. - `minimumNumberOfCalls` must be reached before the rate is evaluated; low traffic services may not trip as expected. - Sharing one breaker instance across unrelated endpoints couples their health; scope breakers per dependency/endpoint. **When to use:** any call to a remote dependency that can fail or slow down, especially on hot paths. Combine breaker + timeout + bounded retries + fallback for a resilient client.
- What are the three circuit breaker states and the transitions between them?CLOSED (calls pass, outcomes recorded); when failure/slow-call rate exceeds threshold it trips to OPEN (calls fail fast with CallNotPermittedException, no network); after waitDurationInOpenState it moves to HALF_OPEN and allows a few probe calls — enough successes return it to CLOSED, failures send it back to OPEN.
- Should the circuit breaker wrap the retries or the other way around?The breaker should observe the outcome of the retries — put it outside/around the retried Mono (breaker records the final failure after retries exhaust). If retries were outside the breaker, they'd mask failures and the breaker might never trip.
saying these in an interview costs you the question
- Claiming Spring WebFlux has a built-in @CircuitBreaker for WebClient (it doesn't — you use Resilience4j / Spring Cloud CircuitBreaker).
- Suggesting Hystrix as the current recommendation (it's retired/maintenance-only).
- Putting retries outside the breaker so failures are hidden from it.
- Using transform instead of transformDeferred, breaking per-subscription semantics.
- A breaker with no timeout that never trips on hung (non-erroring) calls.