skip to content

How do you compose timeouts, retries, and a circuit breaker on a WebClient call correctly, and what ordering and idempotency concerns govern the design?

level: principalimportance: should knowfreq 35%

answer

  1. inside->out: timeout, retry, breaker, bulkhead, fallback
  2. timeout inside retry (per-attempt); breaker outside retry
  3. idempotent-only / idempotency keys for POST
  4. budget = attempts x (timeout+backoff) < caller deadline
  5. retry amplification down the chain; one breaker per dependency

basics

~20 s

Layer them: a per-attempt timeout inside bounded, filtered, jittered retries; a circuit breaker outside the retries so it sees final failures; a fallback outermost. Only retry idempotent calls, keep the total time budget under upstream deadlines, and scope one breaker per dependency.

solid answer

~40 s

Build a defense-in-depth stack per call. Innermost: a **per-attempt timeout** (`responseTimeout` on the connector, plus a Reactor `.timeout()` bounding each try). Around it: **retries** — `Retry.backoff` with `.maxBackoff`, `.jitter`, a `.filter` for transient-only (5xx/timeouts, never 4xx), and `.onRetryExhaustedThrow`. Around that: a **circuit breaker** (`transformDeferred(CircuitBreakerOperator.of(cb))`) so a burst of retry failures trips it and future calls fast-fail. Optionally a **bulkhead** to cap concurrency. Outermost: a **fallback** via `onErrorResume`. Governing constraints: only retry **idempotent** verbs (or use idempotency keys); the *total* worst-case latency (attempts × (timeout + backoff)) must stay under the deadline your callers give you — otherwise deadline inversion; scope a breaker per dependency, not globally; and emit metrics on retries/breaker state. Prefer failing fast over unbounded waiting.

code

java · 22 lines
java
// Reusable resilient call: timeout (per attempt) -> retry -> breaker -> fallback
Mono<Quote> resilientQuote(String symbol) {
    return webClient.get()
        .uri("/quotes/{s}", symbol)
        .retrieve()
        .onStatus(HttpStatusCode::is5xxServerError,
            r -> Mono.error(new UpstreamServerException()))
        .bodyToMono(Quote.class)
        // (1) per-attempt bound (in addition to connector responseTimeout)
        .timeout(Duration.ofSeconds(2))
        // (2) bounded, filtered, jittered retries — GET is idempotent
        .retryWhen(Retry.backoff(2, Duration.ofMillis(200))
            .maxBackoff(Duration.ofSeconds(1))
            .jitter(0.5)
            .filter(ex -> ex instanceof java.util.concurrent.TimeoutException
                || ex instanceof UpstreamServerException)
            .onRetryExhaustedThrow((spec, sig) -> sig.failure()))
        // (3) breaker OUTSIDE retries: sees the final outcome, trips if hard-down
        .transformDeferred(CircuitBreakerOperator.of(quotesBreaker))
        // (4) graceful fallback outermost
        .onErrorResume(ex -> Mono.just(Quote.cachedFallback(symbol)));
}

go deeper

for a junior

Understand the pieces exist (timeout, retry, breaker, fallback) even if not the exact ordering.

for a middle

Wire timeout-inside-retry and a fallback; know idempotency matters for retries.

for a senior

Reason about breaker-outside-retry, per-attempt vs overall timeout placement, and transient-only filtering.

for a principal

Own system-wide policy: latency/deadline budgets, retry amplification limits, per-dependency breaker scoping, idempotency for writes, centralized reusable client config, and resilience observability/alerting.

**Goal:** make one remote call robust without amplifying failures. Each mechanism guards a different failure mode; combined incorrectly they fight each other. **The recommended layering (inside → out):** 1. **Per-attempt timeout.** Set transport `responseTimeout` on the Reactor Netty `HttpClient` and/or a Reactor `.timeout(D)` *below* `retryWhen` so each individual attempt is bounded. This ensures a single slow attempt can't consume the whole budget and lets a retry actually happen. 2. **Retry (bounded, filtered, backed off).** `retryWhen(Retry.backoff(maxAttempts, minBackoff).maxBackoff(...).jitter(0.5).filter(transientOnly).onRetryExhaustedThrow((s,sig)->sig.failure()))`. Retry only transient, idempotent failures. 3. **Circuit breaker.** `transformDeferred(CircuitBreakerOperator.of(cb))` placed *outside* the retry so the breaker records the *final* outcome after retries. If the upstream is hard-down, the breaker trips OPEN and short-circuits — which also stops the retry machinery from piling on. 4. **Bulkhead / concurrency limit** (optional but valuable): `BulkheadOperator` caps simultaneous in-flight calls so one slow dependency can't exhaust the connection pool / event loop. 5. **Fallback (outermost):** `onErrorResume(...)` returns a cached value, default, empty, or fast domain error — so callers get a graceful degrade rather than a raw exception. **Ordering rationale:** - **Timeout inside retry** → bounds each attempt. Timeout *outside* retry bounds the whole sequence — use that only as an overall deadline cap, and be aware it can cut off legitimate retries. - **Breaker outside retry** → the breaker must *see* failures. If retries sat outside the breaker, they'd absorb the failures and the breaker would never trip. Conversely with the breaker outside, once OPEN it prevents retries entirely. - Resilience4j's documented composition order for its operators is Bulkhead → TimeLimiter → RateLimiter → CircuitBreaker → Retry (Retry outermost in *their* decorator model), but note their 'outermost' decorator is the *first to be entered*; the practical invariant is unchanged: the breaker's recorded outcome should reflect the retried result. Validate the exact semantics for your library version. **Idempotency — the governing safety rule:** retries and even timeouts create ambiguity. After a *response* timeout the server may have already processed the request. So: - Retry only **idempotent** operations (GET/PUT/DELETE/HEAD). - For writes (POST), require an **idempotency key** so the server de-duplicates, or don't retry. - Treat 'sent-but-no-response' as *maybe-succeeded*, not *definitely-failed*. **Latency budget / deadline propagation:** worst-case time ≈ `maxAttempts × (attemptTimeout) + sum(backoffs)`. This must be **shorter** than the deadline your own callers grant you; otherwise you'll be killed by *their* timeout while still retrying — **deadline inversion**. Propagate deadlines: shrink inner budgets as the remaining deadline shrinks. **Retry amplification across a system:** if service A retries B, and B retries C, retries multiply geometrically down the chain, turning a small blip into a load spike. Mitigate with: low retry counts (2–3), retry budgets (cap the % of traffic that is retries), circuit breakers to cut chains, and jitter. **Scoping & observability:** one breaker per remote dependency (or per critical endpoint) — sharing couples unrelated health. Emit metrics/events: retry counts, breaker state transitions, fallback rate, timeout rate. Alert on OPEN breakers and rising fallback rates. **Config location:** centralize this in a reusable `WebClient` builder customizer / a resilience wrapper so every client gets consistent timeouts, and per-dependency retry/breaker policy — rather than sprinkling ad-hoc operators. **When to simplify:** not every call needs all five layers. A read from a fast internal service might just need a timeout + fallback. Reserve the full stack for critical, flaky, or expensive dependencies.

  • Why must the per-attempt timeout sit inside retryWhen rather than outside?
    Inside, it bounds each individual attempt so a single slow try is cut short and a retry can proceed. Outside retryWhen, one .timeout() bounds the entire retry sequence — a first slow attempt could consume the whole budget, and the timeout would also abort the whole retry chain, defeating the point.
  • What is deadline inversion and how do you avoid it here?
    Deadline inversion is when your total worst-case latency (attempts × timeout + backoffs) exceeds the deadline your callers granted you, so they time out and kill the request while you're still retrying — wasting work. Avoid it by budgeting: keep the aggregate under the caller's deadline and propagate/shrink inner timeouts as remaining time decreases.
  • How can retries in a service chain amplify a small outage?
    If A retries B and B retries C, the retry factors multiply (e.g. 3×3 = 9× load on C), turning a brief blip into a load spike that can knock the upstream over. Mitigate with low retry counts, retry budgets (cap % of retried traffic), jitter, and circuit breakers to cut the chain.

saying these in an interview costs you the question

  • Putting retries outside the circuit breaker so failures never trip it.
  • One .timeout() outside retryWhen assumed to bound each attempt (it bounds the whole sequence).
  • Retrying non-idempotent writes without idempotency keys.
  • Ignoring the total latency budget vs the caller's deadline (deadline inversion).
  • A single global breaker shared across unrelated dependencies.
  • Adding all five layers to every trivial internal call regardless of need.

context