skip to content

Reactor Error Handling

onErrorResume, onErrorReturn and onErrorMap recover or translate, retryWhen with backoff retries, and an error is a terminal signal that ends the stream. Interviewers probe onErrorContinue because its behaviour surprises people.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

questions

5

In Project Reactor, what happens to a Flux or Mono when an error occurs, and why can't you just use a try/catch around it?

level: juniorimportance: must knowfreq 70%

answer

  1. onNext / onComplete / onError = the three signals
  2. onError is terminal — stream stops
  3. assembly-time vs subscription-time
  4. try/catch guards building, not running
  5. recover with an operator, not a catch

basics

~20 s

An error is a terminal signal: the stream emits onError, stops, and sends nothing more. try/catch doesn't work because the data flows asynchronously later, not on the line that builds the pipeline. You recover with operators like onErrorResume.

solid answer

~40 s

A reactive sequence emits three kinds of signals: onNext (values), onComplete (success end), and onError (failure end). onError is terminal — once it fires, no further onNext or onComplete follows and the subscription is done. A plain try/catch only guards the assembly code that builds the pipeline, not the asynchronous execution that happens when someone subscribes, so the exception surfaces inside the stream, not on your catch line. To handle it you insert an error operator into the chain: onErrorReturn (fallback value), onErrorResume (fallback publisher), onErrorMap (translate the exception), or doOnError (side-effect only). Anything placed after the point of failure is skipped; the error short-circuits straight to the subscriber unless an operator intercepts it first.

code

java · 17 lines
java
// try/catch does NOT catch the async error:
try {
    Flux<Integer> f = Flux.just(1, 2, 0)
        .map(n -> 10 / n); // divide-by-zero fires at subscription, not here
    f.subscribe(System.out::println,
                err -> System.out.println("onError: " + err));
} catch (ArithmeticException e) {
    System.out.println("never reached"); // assembly didn't throw
}
// Output: 10, 5, then onError: ArithmeticException — 0 is never emitted,
// and the terminal error goes to the subscriber's error consumer.

// Correct recovery with an operator in the chain:
Flux.just(1, 2, 0)
    .map(n -> 10 / n)
    .onErrorReturn(-1)      // emit fallback, then complete
    .subscribe(System.out::println); // 10, 5, -1

go deeper

for a junior

Must know onError is terminal and that recovery uses operators, not try/catch.

for a middle

Should distinguish doOnError (side-effect) from real recovery operators and understand short-circuiting.

for a senior

Explains assembly vs subscription time and operator placement precisely.

for a principal

Frames it in terms of the Reactive Streams contract and signal semantics across module boundaries.

## The Reactive Streams signal model Every Reactor `Flux`/`Mono` communicates with its subscriber through exactly three signal types defined by the Reactive Streams spec: - **onNext(value)** — zero or more data items. - **onComplete()** — the sequence finished *successfully*; terminal. - **onError(Throwable)** — the sequence finished *with a failure*; terminal. "Terminal" means final: after `onComplete` **or** `onError`, the publisher will never call anything again on that subscription. They are mutually exclusive — you get one or the other, never both. So an error is not an exception you throw and catch; it is a **value-like signal that travels down the stream** to the subscriber. ## Why try/catch does not work Reactor is **assembly-time vs. subscription-time**. Writing `Flux.just(...).map(...)` only *builds* a pipeline description; nothing runs yet. Execution happens later, when something calls `subscribe()`, often on a different thread. A `try { flux.map(...) } catch (Exception e)` block only wraps the *building* of the chain, which rarely throws. The real work — and thus the real exception — happens asynchronously at subscription, so it escapes as an `onError` signal instead of hitting your catch. (There is one nuance: an exception thrown *while assembling* the chain can be caught, but that's a bug in your builder code, not the normal data-flow error.) ## Error short-circuits the chain When an operator's function throws, or an upstream emits `onError`, Reactor immediately routes that error **downstream to the next operator**, skipping any remaining `onNext` processing. Operators placed *after* the failure point simply pass the error along untouched — unless they are error-handling operators. So the fix is: put an error-handling operator in the chain, before the subscriber. ## The core recovery operators (this leaf) - **`doOnError(Consumer<Throwable>)`** — a *side-effect* hook (log, metric). It does **not** recover; the error keeps propagating. - **`onErrorReturn(fallbackValue)`** — emit a single static fallback value, then `onComplete`. - **`onErrorResume(Function<Throwable, Publisher>)`** — switch to an alternate `Mono`/`Flux` (e.g., a cache lookup or default stream). - **`onErrorMap(Function<Throwable, Throwable>)`** — translate one exception type into another; the stream *still* ends in error, just a different one. ## Gotchas - Placement matters: `onErrorResume` only catches errors from operators **upstream** of it. An error produced downstream is not caught by an earlier handler. - `onComplete` is *not* an error, so error operators never fire on normal completion. - An empty `Mono` (`onComplete` with no value) is a success, not an error — use `switchIfEmpty` for that, not `onErrorResume`. - Nothing happens until you subscribe; forgetting to subscribe means the error never even occurs.

  • Does doOnError recover from the error?
    No. doOnError is a side-effect-only hook (logging, metrics). The error still propagates downstream and remains terminal; only operators like onErrorResume/onErrorReturn actually recover.
  • If an error happens, do later map() operators still run?
    No. The error short-circuits past all downstream onNext processing straight to the subscriber (or the first error-handling operator). Non-error operators just pass it along.

saying these in an interview costs you the question

  • Thinking a surrounding try/catch will catch async stream errors
  • Believing doOnError handles/swallows the error
  • Assuming operators after the failure still process values
  • Confusing an empty Mono (onComplete) with an error signal

context

open as a page

Compare onErrorReturn, onErrorResume, and onErrorMap. When would you choose each?

level: middleimportance: must knowfreq 75%

basics

~10 s

onErrorReturn emits a fixed fallback value then completes. onErrorResume switches to another Mono/Flux (e.g., a cache or default stream). onErrorMap translates the exception into a different one but the stream still ends in error.

open as a page

What does onErrorContinue do, why is it considered controversial, and how does it differ from onErrorResume inside a flatMap?

level: seniorimportance: should knowfreq 45%

basics

~20 s

onErrorContinue drops the element that caused an error and keeps the stream going instead of terminating. It's controversial because it breaks the 'error is terminal' model and only works with operators that explicitly support it, so behavior is surprising. Prefer onErrorResume inside flatMap.

open as a page

How does retryWhen with Retry.backoff work, and how does it differ from a plain retry()? What happens when attempts are exhausted?

level: seniorimportance: should knowfreq 65%

basics

~10 s

retry() re-subscribes to the source immediately on error, up to N times. retryWhen(Retry.backoff(max, minDelay)) re-subscribes with exponential backoff plus jitter. When retries run out it fails with a RetryExhaustedException wrapping the last error.

open as a page

Design the error-handling layer for a WebFlux endpoint that aggregates two downstream services. Discuss ordering of doOnError/retry/onErrorMap/onErrorResume, and how error signals interact with context propagation.

level: principalimportance: should knowfreq 35%

basics

~20 s

Scope handling per call: retry transient errors with backoff around each downstream, log with doOnError, translate infra exceptions with onErrorMap at the boundary, then provide a fallback with onErrorResume. Keep operator order deliberate because each catches only upstream errors, and rely on the Reactor Context to carry request/trace data through the error path.

open as a page