How does retryWhen with Retry.backoff work, and how does it differ from a plain retry()? What happens when attempts are exhausted?
answer
- retry() = immediate resubscribe, no delay
- Retry.backoff(maxAttempts, minBackoff) = exponential + jitter
- retry == re-run upstream side effects
- exhaustion -> RetryExhaustedException (cause = last error)
- .filter transient, .maxBackoff cap, .transientErrors reset
basics
~10 sretry() re-subscribes to the source immediately on error, up to N times. retryWhen(Retry.backoff(max, minDelay)) re-subscribes with exponential backoff plus jitter. When retries run out it fails with a RetryExhaustedException wrapping the last error.
solid answer
~50 s`retry(n)` simply re-subscribes to the upstream immediately when it errors, up to n times — no delay, which can hammer a struggling dependency. `retryWhen(Retry)` is the configurable form: you pass a `Retry` strategy, most commonly `Retry.backoff(maxAttempts, minBackoff)`, which retries with **exponential backoff and jitter** (default 50% jitter) so retries spread out and avoid thundering-herd. You can refine it: `.filter(predicate)` to retry only transient errors, `.maxBackoff(Duration)` to cap the delay, `.jitter(double)`, `.transientErrors(true)`, and `.onRetryExhaustedThrow(...)`. Retrying means **re-subscribing**, so the whole upstream (and its side effects) runs again — fine for idempotent reads, dangerous for non-idempotent writes. Crucially, retry only triggers on `onError`; a completed or empty stream is never retried. When `maxAttempts` is used up, by default Reactor emits `onError` with a `RetryExhaustedException` (via `Exceptions.retryExhausted`) whose cause is the last upstream error; you typically follow it with `onErrorResume` for a final fallback.
code
java · 17 linesimport reactor.util.retry.Retry;
import java.time.Duration;
Mono<Quote> fetchQuote(String symbol) {
return quoteClient.get(symbol) // cold: re-runs on retry
.retryWhen(
Retry.backoff(3, Duration.ofMillis(200)) // 3 attempts, exp backoff
.maxBackoff(Duration.ofSeconds(2))
.jitter(0.5)
.filter(ex -> ex instanceof java.io.IOException) // only transient
.doBeforeRetry(sig ->
log.warn("retry #{} after {}", sig.totalRetries(), sig.failure().toString()))
// propagate the ORIGINAL error instead of RetryExhaustedException:
.onRetryExhaustedThrow((spec, sig) -> sig.failure())
)
.onErrorResume(ex -> Mono.just(Quote.stale(symbol))); // final fallback
}go deeper
Aware retry() re-tries on failure; not expected to know backoff internals.
Knows Retry.backoff gives exponential delay and that exhaustion errors out.
Must explain resubscribe semantics, filtering transient errors, jitter, and exhaustion wrapping.
Designs retry policy with idempotency, jitter to avoid thundering herd, scoping, and fallback strategy across services.
## retry() — the blunt instrument - `retry()` — resubscribe indefinitely on every `onError`. - `retry(long n)` — resubscribe up to `n` times, then propagate the last error. Resubscription means Reactor calls `subscribe()` on the upstream **again from scratch**. For a cold publisher (e.g., a fresh `WebClient` call) that re-executes the request. There is **no delay**, so on a persistent failure you fire n rapid-fire retries — a classic way to DDoS your own dependency. ## retryWhen(Retry) — the configurable form `retryWhen(Retry retrySpec)` drives retries from a `Retry` strategy (the modern `reactor.util.retry.Retry` builder API; the older `Function<Flux<Throwable>, Publisher>` form is deprecated). Common factories: - `Retry.max(long)` — up to N immediate retries. - `Retry.fixedDelay(long maxAttempts, Duration delay)` — constant delay between attempts. - `Retry.backoff(long maxAttempts, Duration minBackoff)` — **exponential backoff with jitter**. Delay roughly doubles each attempt starting from `minBackoff`, and a default **0.5 jitter** randomizes it to de-correlate retriers. ### Tuning `Retry.backoff` ``` Retry.backoff(3, Duration.ofMillis(200)) .maxBackoff(Duration.ofSeconds(2)) // cap the growth .jitter(0.5) // 0.0..1.0 randomization .filter(ex -> ex instanceof IOException) // only transient errors .transientErrors(true) // reset counter after a value .onRetryExhaustedThrow((spec, signal) -> signal.failure()); // custom terminal ``` - **`.filter` / (predicate)** — retry only errors you consider transient (timeouts, 503). Non-matching errors propagate immediately — don't retry a 400 or a bug. - **`.transientErrors(true)`** — the attempt counter **resets** after each successfully emitted value; good for long-lived Flux where bursts of errors shouldn't accumulate toward the cap. - **`.maxBackoff`** — ceiling so exponential growth doesn't reach minutes. - **`.doBeforeRetry` / `.doAfterRetry`** — side-effects like logging/metrics per attempt. ## Exhaustion behavior When `maxAttempts` is reached, `Retry.backoff`/`max`/`fixedDelay` **by default** emit `onError` with an exception created by `Exceptions.retryExhausted(...)` — a `RetryExhaustedException`-type wrapper whose **cause is the last source error**. Override with `.onRetryExhaustedThrow((retrySpec, retrySignal) -> retrySignal.failure())` to propagate the original exception unwrapped, or any custom exception. You almost always chain a final `onErrorResume`/`onErrorReturn` after `retryWhen` to provide a fallback once retries are spent. ## Key semantics & gotchas - **Retry == resubscribe.** All upstream side effects re-run. Safe for idempotent GETs; for POST/PUT ensure idempotency keys or don't retry. - **Only onError triggers retry.** An empty completion is *not* an error; retry won't fire. Use `switchIfEmpty(Mono.error(...))` if empty should be retried. - **Placement matters.** Put `retryWhen` around just the flaky operation (e.g., wrap the `WebClient` call), not the whole pipeline, so unrelated errors aren't retried and downstream work isn't duplicated. - **Backoff runs on a scheduler.** By default `Schedulers.parallel()`; delays are non-blocking (no thread is tied up sleeping). - **Don't retry non-transient errors.** Always add a `.filter` so validation/auth failures fail fast. - **Blocking retry loop myth.** Nothing blocks; the delay is a timer, and the subscriber sees values only after a successful attempt.
- What exception does the subscriber see if all backoff attempts fail and you don't override the behavior?A RetryExhaustedException produced by Exceptions.retryExhausted, whose cause() is the last upstream error. Use .onRetryExhaustedThrow((spec, sig) -> sig.failure()) to surface the original exception instead.
- Why is retrying a non-idempotent POST dangerous, and how do you guard it?Retry re-subscribes and re-executes the upstream, so a POST could run multiple times causing duplicate side effects. Guard with idempotency keys, retry only safe/read operations, or filter to errors that are known not to have committed.
saying these in an interview costs you the question
- Saying retry() waits between attempts (it doesn't — it's immediate)
- Retrying all errors, including 4xx/validation/bugs
- Assuming retry resumes mid-stream instead of re-subscribing from scratch
- Thinking an empty Mono is retried
- Blocking a thread during backoff (it's a non-blocking timer)