How do you add retries with exponential backoff to a WebClient call, and how do you retry only the right failures?
answer
- retryWhen(Retry.backoff(n, minBackoff))
- .maxBackoff + .jitter (thundering herd)
- .filter -> only 5xx/timeouts, never 4xx
- .onRetryExhaustedThrow -> original error
- idempotent only (POST double-executes)
basics
~10 sUse retryWhen(Retry.backoff(maxAttempts, minBackoff)) on the Mono/Flux. Add a .filter(...) predicate so you only retry transient failures (like 5xx or timeouts), not 4xx client errors, and set .jitter(...) and .maxBackoff(...).
solid answer
~30 sReactor's `retryWhen(Retry retrySpec)` drives retries reactively. `Retry.backoff(maxAttempts, minBackoff)` gives exponential backoff; chain `.maxBackoff(Duration)` to cap growth and `.jitter(0.5)` to de-synchronize retrying clients (thundering herd). Crucially add `.filter(Predicate<Throwable>)` so you retry only *transient* errors — connection failures, timeouts, 5xx (e.g. `WebClientResponseException` where `getStatusCode().is5xxServerError()`) — and never non-idempotent-unsafe or 4xx client errors, which won't get better. After exhausting attempts, `Retry.backoff` wraps the last error; use `.onRetryExhaustedThrow((spec, signal) -> signal.failure())` to propagate the original. Only retry **idempotent** operations (GET, PUT, DELETE); retrying a POST can double-execute. Watch retry+timeout interaction: place `.timeout()` per-attempt (below retryWhen) vs overall (above) intentionally.
code
java · 20 linesimport reactor.util.retry.Retry;
import org.springframework.web.reactive.function.client.WebClientResponseException;
Mono<Quote> quote = webClient.get()
.uri("/quotes/{sym}", symbol)
.retrieve()
.bodyToMono(Quote.class)
// bound EACH attempt (below retryWhen)
.timeout(Duration.ofSeconds(2))
.retryWhen(Retry.backoff(3, Duration.ofMillis(200))
.maxBackoff(Duration.ofSeconds(2))
.jitter(0.5)
// retry only transient failures, never 4xx client errors
.filter(ex ->
ex instanceof java.util.concurrent.TimeoutException
|| (ex instanceof WebClientResponseException w
&& w.getStatusCode().is5xxServerError()))
.doBeforeRetry(sig -> log.warn("retry #{}", sig.totalRetries() + 1))
// after exhaustion, throw the ORIGINAL error, not RetryExhaustedException
.onRetryExhaustedThrow((spec, sig) -> sig.failure()));go deeper
Know retryWhen(Retry.backoff(...)) exists and that you shouldn't retry everything.
Add .filter for transient-only, .maxBackoff and .jitter, and know idempotency matters.
Discuss exhaustion handling, per-attempt vs overall timeout placement, and retry/circuit-breaker interplay.
Own retry budgets across services, prevent retry amplification, mandate idempotency keys for writes, and standardize the transient-error predicate.
**Reactor retry model:** Reactor exposes `retryWhen(Retry retrySpec)` on `Mono` and `Flux`. The `Retry` companion (from `reactor.util.retry.Retry`) is a strategy object. The two main factory methods: - `Retry.max(long maxAttempts)` — immediate retries up to N times. - `Retry.backoff(long maxAttempts, Duration minBackoff)` — exponential backoff: waits roughly `minBackoff`, then ~2×, ~4×… between attempts, with built-in jitter by default. - `Retry.fixedDelay(long maxAttempts, Duration fixedDelay)` — constant delay. **Backoff tuning (chained on the spec):** - `.maxBackoff(Duration)` — caps the exponential growth so waits don't explode. - `.jitter(double factor)` — randomizes each delay by up to ±factor (0..1). Jitter is essential: if 100 clients all fail at once and retry after exactly the same delay, they hammer the recovering upstream simultaneously (**thundering herd / retry storm**). Jitter spreads them out. `Retry.backoff` applies 0.5 jitter by default. - `.filter(Predicate<Throwable>)` — **only** retry errors matching the predicate. Retry transient faults (timeouts, connection resets, HTTP 5xx, 429 with care), never HTTP 4xx client errors (400/401/403/404) — those are deterministic and will just waste attempts. **Reading the status in the filter:** on the `retrieve()` path a failed response is a `WebClientResponseException`, so: ```java .filter(ex -> ex instanceof WebClientResponseException w && w.getStatusCode().is5xxServerError()) ``` You can also retry on `IOException`/`java.util.concurrent.TimeoutException` for connection/timeout faults. **Exhaustion behavior:** when attempts run out, `Retry.backoff` throws a `RetryExhaustedException` wrapping the last failure — often not what callers expect. Use `.onRetryExhaustedThrow((retrySpec, retrySignal) -> retrySignal.failure())` to re-throw the **original** underlying exception instead. **Idempotency — the big gotcha:** a retry re-executes the request. That's safe only for **idempotent** operations (GET, PUT, DELETE, and reads). Retrying a **POST** (create) can create duplicates. Mitigations: only retry idempotent verbs; or use an **idempotency key** so the server de-duplicates; or don't retry on responses that may have partially succeeded (e.g. after the request was sent, a *response* timeout is ambiguous — the server may have processed it). **Retry × timeout interaction:** placement matters. - `.timeout(D)` *below* `retryWhen` (closer to the call) bounds *each attempt*. - `.timeout(D)` *above* `retryWhen` bounds the *whole* retry sequence. Get this wrong and either each attempt is unbounded, or the first slow attempt eats the whole budget leaving no room to retry. **Observability & side effects:** use `.doBeforeRetry(signal -> log...)` to log attempt counts, and emit metrics. Excessive retries hide upstream problems and amplify load. **When to use:** transient, idempotent calls to flaky upstreams. Combine with per-attempt timeouts and, above a certain failure rate, a **circuit breaker** (next question) so you stop retrying a hard-down dependency entirely.
- Why add jitter to backoff?Without jitter, many clients that failed at the same instant retry after the same delay, hitting the recovering upstream in synchronized waves (retry storm / thundering herd). Jitter randomizes each delay so load is spread out.
- Which HTTP methods are safe to retry and why?Idempotent ones — GET, PUT, DELETE, HEAD — because re-executing them yields the same end state. POST is generally unsafe (can create duplicates); to retry it safely, use a server-side idempotency key so duplicates are de-duplicated.
- After exhausting retries, callers see RetryExhaustedException instead of the real cause. How do you fix it?Add .onRetryExhaustedThrow((spec, signal) -> signal.failure()) to the Retry spec so it re-throws the original underlying exception.
saying these in an interview costs you the question
- Retrying all errors including 4xx — 400/401/404 are deterministic and won't succeed on retry.
- Retrying non-idempotent POSTs without an idempotency key.
- Fixed-delay retries with no jitter causing synchronized retry storms.
- Forgetting onRetryExhaustedThrow, so callers get RetryExhaustedException instead of the real cause.
- Placing .timeout() carelessly so it bounds the wrong scope relative to retryWhen.