Design the error-handling layer for a WebFlux endpoint that aggregates two downstream services. Discuss ordering of doOnError/retry/onErrorMap/onErrorResume, and how error signals interact with context propagation.
answer
- handle per-dependency, not once at top
- order: call -> retry(filter transient) -> doOnError -> onErrorMap -> onErrorResume
- each operator catches only upstream errors
- Mono.zip short-circuits — give optional leg its own fallback
- Reactor Context carries trace/MDC on the error path (not thread-locals)
basics
~20 sScope handling per call: retry transient errors with backoff around each downstream, log with doOnError, translate infra exceptions with onErrorMap at the boundary, then provide a fallback with onErrorResume. Keep operator order deliberate because each catches only upstream errors, and rely on the Reactor Context to carry request/trace data through the error path.
solid answer
~50 sI'd handle errors close to each dependency, not once at the top. Around each service call: `retryWhen(Retry.backoff(...).filter(transient))` so only retryable failures re-run; `doOnError` for logging/metrics (side-effect, non-recovering); `onErrorMap` at the adapter boundary to turn `WebClientResponseException`/timeouts into domain exceptions so the aggregator doesn't couple to transport types. Then, when composing the two results (e.g., `Mono.zip`), decide per-service whether a failure is fatal or degradable: use `onErrorResume(e -> Mono.just(defaultPart))` for the optional service and let the critical one propagate. Order matters because each operator only intercepts **upstream** errors — retry must sit closest to the flaky call, mapping after retries are exhausted, fallback last. Error signals carry through the Reactor `Context` (immutable, subscription-scoped, flows upstream), so request/trace/MDC data placed via `contextWrite` remains available to `doOnError` and fallback logic even on the error path. Finally, `Mono.zip` short-circuits on the first error, so wrap optional legs with their own fallback before zipping.
code
java · 27 linesMono<Dashboard> aggregate(String userId) {
Mono<Profile> profile = profileClient.get(userId) // CRITICAL
.retryWhen(Retry.backoff(3, Duration.ofMillis(200))
.maxBackoff(Duration.ofSeconds(2))
.filter(this::isTransient))
.doOnError(e -> log.error("profile failed for {}", userId, e))
.onErrorMap(WebClientResponseException.class,
e -> new ProfileUnavailable(userId, e)); // fail fast (no resume)
Mono<Prefs> prefs = prefsClient.get(userId) // OPTIONAL/degradable
.retryWhen(Retry.backoff(2, Duration.ofMillis(100))
.filter(this::isTransient))
.onErrorResume(e -> { // own fallback BEFORE zip
log.warn("prefs degraded for {}", userId, e);
return Mono.just(Prefs.defaults());
});
return Mono.zip(profile, prefs) // A error propagates; B never errors
.map(t -> new Dashboard(t.getT1(), t.getT2()))
// read trace id from Context on both success and error paths:
.contextWrite(ctx -> ctx.put("traceId", tracer.currentTraceId()));
}
private boolean isTransient(Throwable ex) {
return ex instanceof java.io.IOException
|| (ex instanceof WebClientResponseException w && w.getStatusCode().is5xxServerError());
}go deeper
Not expected; this is an architecture-level design question.
Can place retry/onErrorResume around a single call but may miss composition and context concerns.
Reasons about operator ordering, transient filtering, and per-leg fallback with zip.
Designs the whole resilience/observability policy: selective degradation, exception translation at boundaries, Context-based propagation, idempotency, and retry-storm avoidance.
## Goal A controller method aggregates **Service A (critical)** and **Service B (optional/degradable)** via `WebClient`, then combines them. We want: retry transient faults, translate exceptions at boundaries, degrade gracefully where allowed, fail cleanly where not, and preserve observability (trace/MDC) across the error path — all non-blocking. ## Principle 1: handle per-dependency, not once at the top Each downstream call gets its **own** error stack so failures are contained and policies differ per service. A single top-level handler can't distinguish "B is optional" from "A is fatal". ## Principle 2: operator ordering (each catches only *upstream* errors) Recommended order **around one call**, from source outward: 1. **the call** (`webClient...retrieve().bodyToMono(...)`) 2. **`retryWhen(Retry.backoff(maxAttempts, min).filter(::isTransient))`** — closest to the source so it re-subscribes just the flaky call; filter so only transient errors (timeouts, 503) retry, not 4xx/bugs. 3. **`doOnError(...)`** — log/metric each *final* error (after retries) as a side effect; it does not recover, so put it where you want to observe. 4. **`onErrorMap(TransportEx.class, e -> new DomainEx(...))`** — translate infrastructure exceptions to domain exceptions *after* retries are exhausted, at the module boundary, so upstream logic depends on domain types. 5. **`onErrorResume(...)`** — the recovery/fallback, applied **last** for the parts that may degrade. For the critical service, you may omit this and let it propagate. Why this order: `retry` must be innermost or it would re-run downstream operators too; `onErrorMap`/`doOnError` see the post-retry error; `onErrorResume` is the outermost decision ("given a final domain error, do we substitute?"). ## Principle 3: composition & short-circuit semantics `Mono.zip(a, b)` (or `zipWith`) **short-circuits on the first error** — if B fails, the whole zip errors even though B is optional. So give B its **own** `onErrorResume(e -> Mono.just(Part.empty()))` *before* zipping. A's failure is allowed to propagate out of the zip and become the endpoint error (mapped to a domain exception, later turned into an HTTP status by the infrastructure `@ExceptionHandler` — out of scope here). ## Principle 4: error signals and the Reactor Context The Reactor **`Context`** is an immutable, per-subscription key/value map that flows **from the subscriber upstream** (opposite to data). You populate it with `contextWrite(...)` (e.g., trace id, tenant, MDC). Because the error path is just another signal traveling to the subscriber, operators on that path — `doOnError`, `onErrorResume`, `onErrorMap` — can still read the Context via `Mono.deferContextual` / `Flux.deferContextual`. That's how you keep correlation ids and MDC in your *error* logs without thread-locals (which don't survive thread hops in WebFlux). Libraries like Micrometer context-propagation bridge this Context to MDC around logging. Key point for a principal: **thread-locals are unreliable across schedulers; the Reactor Context is the durable carrier, and it remains readable on the error branch.** ## Principle 5: don't mask, don't over-retry - Filter retries to transient errors; never retry validation/auth failures or NPEs. - Scope `onErrorResume`/`onErrorReturn` by exception type so genuine bugs surface. - Ensure retried calls are idempotent (safe for GET; guard writes). - Use `.maxBackoff` and jitter to avoid synchronized retry storms across instances. ## Putting it together (see code) The critical leg fails fast with a mapped domain error; the optional leg degrades to a default; the Context carries trace data through both success and error paths. HTTP status mapping is intentionally left to the infrastructure layer (`@ExceptionHandler`/`ResponseStatusException`), which is a sibling concern. ## Gotchas - Putting one big `onErrorResume` around the whole `zip` collapses A and B into indistinguishable failures — lose the ability to degrade selectively. - `doOnError` after `onErrorResume` never fires (the error was already recovered upstream) — order it before the recovery if you want to see the original error. - `retryWhen` around the *combined* pipeline re-runs *both* services on any failure — usually wrong; scope retry to each call. - Blocking calls or thread-locals inside the reactive chain break both non-blocking guarantees and context propagation.
- Why not wrap the whole Mono.zip in a single onErrorResume?Because zip short-circuits on the first error and one blanket handler can't tell whether the critical or the optional service failed. You lose selective degradation and might substitute a default when the critical service actually failed. Handle each leg's failure locally before composing.
- In WebFlux, why can't you rely on thread-locals (e.g., SLF4J MDC) for correlation ids in error logs, and what's the fix?Reactive operators hop threads across schedulers, so thread-locals set on one thread aren't visible on another. The Reactor Context (subscription-scoped, immutable, flows upstream and is readable on the error path) is the durable carrier; Micrometer context-propagation bridges it to MDC around logging boundaries.
saying these in an interview costs you the question
- One top-level onErrorResume around the whole pipeline, losing per-service policy
- Retrying the combined pipeline so both services re-run on any failure
- Assuming Mono.zip continues when one leg errors (it short-circuits)
- Relying on thread-locals/MDC instead of Reactor Context across thread hops
- Ordering doOnError after onErrorResume and expecting it to log the original error