skip to content

How does retryWhen with Retry.backoff work, and how does it differ from a plain retry()? What happens when attempts are exhausted?

level: seniorimportance: should knowfreq 65%

answer

  1. retry() = immediate resubscribe, no delay
  2. Retry.backoff(maxAttempts, minBackoff) = exponential + jitter
  3. retry == re-run upstream side effects
  4. exhaustion -> RetryExhaustedException (cause = last error)
  5. .filter transient, .maxBackoff cap, .transientErrors reset

basics

~10 s

retry() re-subscribes to the source immediately on error, up to N times. retryWhen(Retry.backoff(max, minDelay)) re-subscribes with exponential backoff plus jitter. When retries run out it fails with a RetryExhaustedException wrapping the last error.

solid answer

~50 s

`retry(n)` simply re-subscribes to the upstream immediately when it errors, up to n times — no delay, which can hammer a struggling dependency. `retryWhen(Retry)` is the configurable form: you pass a `Retry` strategy, most commonly `Retry.backoff(maxAttempts, minBackoff)`, which retries with **exponential backoff and jitter** (default 50% jitter) so retries spread out and avoid thundering-herd. You can refine it: `.filter(predicate)` to retry only transient errors, `.maxBackoff(Duration)` to cap the delay, `.jitter(double)`, `.transientErrors(true)`, and `.onRetryExhaustedThrow(...)`. Retrying means **re-subscribing**, so the whole upstream (and its side effects) runs again — fine for idempotent reads, dangerous for non-idempotent writes. Crucially, retry only triggers on `onError`; a completed or empty stream is never retried. When `maxAttempts` is used up, by default Reactor emits `onError` with a `RetryExhaustedException` (via `Exceptions.retryExhausted`) whose cause is the last upstream error; you typically follow it with `onErrorResume` for a final fallback.

code

java · 17 lines
java
import reactor.util.retry.Retry;
import java.time.Duration;

Mono<Quote> fetchQuote(String symbol) {
    return quoteClient.get(symbol)                       // cold: re-runs on retry
        .retryWhen(
            Retry.backoff(3, Duration.ofMillis(200))     // 3 attempts, exp backoff
                 .maxBackoff(Duration.ofSeconds(2))
                 .jitter(0.5)
                 .filter(ex -> ex instanceof java.io.IOException) // only transient
                 .doBeforeRetry(sig ->
                     log.warn("retry #{} after {}", sig.totalRetries(), sig.failure().toString()))
                 // propagate the ORIGINAL error instead of RetryExhaustedException:
                 .onRetryExhaustedThrow((spec, sig) -> sig.failure())
        )
        .onErrorResume(ex -> Mono.just(Quote.stale(symbol))); // final fallback
}

go deeper

for a junior

Aware retry() re-tries on failure; not expected to know backoff internals.

for a middle

Knows Retry.backoff gives exponential delay and that exhaustion errors out.

for a senior

Must explain resubscribe semantics, filtering transient errors, jitter, and exhaustion wrapping.

for a principal

Designs retry policy with idempotency, jitter to avoid thundering herd, scoping, and fallback strategy across services.

## retry() — the blunt instrument - `retry()` — resubscribe indefinitely on every `onError`. - `retry(long n)` — resubscribe up to `n` times, then propagate the last error. Resubscription means Reactor calls `subscribe()` on the upstream **again from scratch**. For a cold publisher (e.g., a fresh `WebClient` call) that re-executes the request. There is **no delay**, so on a persistent failure you fire n rapid-fire retries — a classic way to DDoS your own dependency. ## retryWhen(Retry) — the configurable form `retryWhen(Retry retrySpec)` drives retries from a `Retry` strategy (the modern `reactor.util.retry.Retry` builder API; the older `Function<Flux<Throwable>, Publisher>` form is deprecated). Common factories: - `Retry.max(long)` — up to N immediate retries. - `Retry.fixedDelay(long maxAttempts, Duration delay)` — constant delay between attempts. - `Retry.backoff(long maxAttempts, Duration minBackoff)` — **exponential backoff with jitter**. Delay roughly doubles each attempt starting from `minBackoff`, and a default **0.5 jitter** randomizes it to de-correlate retriers. ### Tuning `Retry.backoff` ``` Retry.backoff(3, Duration.ofMillis(200)) .maxBackoff(Duration.ofSeconds(2)) // cap the growth .jitter(0.5) // 0.0..1.0 randomization .filter(ex -> ex instanceof IOException) // only transient errors .transientErrors(true) // reset counter after a value .onRetryExhaustedThrow((spec, signal) -> signal.failure()); // custom terminal ``` - **`.filter` / (predicate)** — retry only errors you consider transient (timeouts, 503). Non-matching errors propagate immediately — don't retry a 400 or a bug. - **`.transientErrors(true)`** — the attempt counter **resets** after each successfully emitted value; good for long-lived Flux where bursts of errors shouldn't accumulate toward the cap. - **`.maxBackoff`** — ceiling so exponential growth doesn't reach minutes. - **`.doBeforeRetry` / `.doAfterRetry`** — side-effects like logging/metrics per attempt. ## Exhaustion behavior When `maxAttempts` is reached, `Retry.backoff`/`max`/`fixedDelay` **by default** emit `onError` with an exception created by `Exceptions.retryExhausted(...)` — a `RetryExhaustedException`-type wrapper whose **cause is the last source error**. Override with `.onRetryExhaustedThrow((retrySpec, retrySignal) -> retrySignal.failure())` to propagate the original exception unwrapped, or any custom exception. You almost always chain a final `onErrorResume`/`onErrorReturn` after `retryWhen` to provide a fallback once retries are spent. ## Key semantics & gotchas - **Retry == resubscribe.** All upstream side effects re-run. Safe for idempotent GETs; for POST/PUT ensure idempotency keys or don't retry. - **Only onError triggers retry.** An empty completion is *not* an error; retry won't fire. Use `switchIfEmpty(Mono.error(...))` if empty should be retried. - **Placement matters.** Put `retryWhen` around just the flaky operation (e.g., wrap the `WebClient` call), not the whole pipeline, so unrelated errors aren't retried and downstream work isn't duplicated. - **Backoff runs on a scheduler.** By default `Schedulers.parallel()`; delays are non-blocking (no thread is tied up sleeping). - **Don't retry non-transient errors.** Always add a `.filter` so validation/auth failures fail fast. - **Blocking retry loop myth.** Nothing blocks; the delay is a timer, and the subscriber sees values only after a successful attempt.

  • What exception does the subscriber see if all backoff attempts fail and you don't override the behavior?
    A RetryExhaustedException produced by Exceptions.retryExhausted, whose cause() is the last upstream error. Use .onRetryExhaustedThrow((spec, sig) -> sig.failure()) to surface the original exception instead.
  • Why is retrying a non-idempotent POST dangerous, and how do you guard it?
    Retry re-subscribes and re-executes the upstream, so a POST could run multiple times causing duplicate side effects. Guard with idempotency keys, retry only safe/read operations, or filter to errors that are known not to have committed.

saying these in an interview costs you the question

  • Saying retry() waits between attempts (it doesn't — it's immediate)
  • Retrying all errors, including 4xx/validation/bugs
  • Assuming retry resumes mid-stream instead of re-subscribing from scratch
  • Thinking an empty Mono is retried
  • Blocking a thread during backoff (it's a non-blocking timer)

context