In a coroutine-based WebFlux endpoint, how are threading, cancellation, and blocking calls handled, and what must you get right to keep the event loop healthy?
answer
- handlers run on event loop — never block
- withContext(Dispatchers.IO) for blocking work
- disconnect/timeout -> subscription cancel -> CancellationException
- cooperative cancel: finally + NonCancellable, rethrow
- MDC/thread-local context NOT auto-propagated
basics
~20 sSuspend handlers run on the non-blocking Reactor event-loop threads, so never call blocking code directly — offload it with withContext(Dispatchers.IO). Client disconnects/timeouts cancel the Reactor subscription, which cancels the coroutine (a CancellationException); write cancellation-cooperative code.
solid answer
~40 sA coroutine WebFlux handler is bridged to a `Mono` that, when subscribed, runs the coroutine on whatever thread the reactive pipeline is on — typically a Reactor/Netty event-loop thread. There is a small, fixed number of these, so any blocking call (JDBC, `Thread.sleep`, `.block()`, blocking HTTP/file I/O) stalls unrelated in-flight requests and collapses throughput; move such work to `withContext(Dispatchers.IO) { ... }` (a bounded, blocking-appropriate pool). Cancellation is unified: when the client disconnects or a `withTimeout` fires, Reactor cancels the subscription and structured concurrency cancels the coroutine, surfacing as `CancellationException` — so your code must be cancellation-cooperative (suspend points check for cancellation; use `NonCancellable`/`finally` for cleanup). Reactor context is bridged via the `ReactorContext` element, but thread-local-based context (MDC logging, some security) does not automatically follow coroutine dispatch and needs explicit propagation.
code
kotlin · 16 lines@RestController
class ReportController(private val legacyJdbc: LegacyDao, private val webClient: WebClient) {
@GetMapping("/report/{id}")
suspend fun report(@PathVariable id: Long): Report = withTimeout(3_000) {
// Reactive call: safe on the event loop
val meta = webClient.get().uri("/meta/{id}", id).retrieve().awaitBody<Meta>()
// Blocking JDBC MUST be offloaded so it never runs on the event loop
val rows = withContext(Dispatchers.IO) { legacyJdbc.query(id) }
Report(meta, rows)
}
// If the client disconnects or 3s elapses, the coroutine is cancelled:
// CancellationException unwinds; finally-blocks run for cleanup.
}go deeper
Know handlers run on non-blocking threads and you must not block them.
Use withContext(Dispatchers.IO) for blocking work and know client disconnect cancels the request.
Explain the shared cancellation model (CancellationException, cooperative cancel, finally/NonCancellable) and event-loop sizing implications.
Reason end-to-end about capacity from non-blocking discipline, context propagation gaps (MDC/security/tracing), BlockHound verification, structured-scope choices for fire-and-forget, and treating cancellation as a first-class outcome in metrics/error handling.
**Threading model.** WebFlux on Netty uses a small pool of event-loop threads (roughly one per CPU core). All non-blocking request processing shares them. A coroutine handler is adapted to a `Mono` via `mono { }`; when subscribed, the coroutine executes using the `CoroutineContext`/dispatcher present at that point, which for a plain suspend handler is effectively the event-loop thread (there is no automatic dispatch to a separate pool). Consequence: **the event loop must never block.** **The blocking hazard.** Blocking operations include JDBC (`spring-boot-starter-data-jpa`), `Thread.sleep`, `Object.wait`, synchronous file I/O, `Mono.block()`/`Flux.blockFirst()`, and blocking client libraries. One blocked event-loop thread means every request queued behind it stalls; a handful of blocked threads freezes the server. Remedies: - `withContext(Dispatchers.IO) { blockingCall() }` — runs the blocking work on a large, elastic pool designed for it, suspending the handler (not blocking the loop) until it returns. - Prefer reactive/coroutine-native drivers (R2DBC, reactive `WebClient`) so nothing blocks in the first place. - Optionally a custom dispatcher for isolation/limits. BlockHound can be used in tests to detect accidental blocking on non-blocking threads. **Cancellation.** WebFlux and coroutines share a cancellation story: - If the HTTP client disconnects, Netty/Reactor cancels the response subscription; the bridge cancels the coroutine, which throws `CancellationException` at the next suspension point and unwinds via structured concurrency (child coroutines are cancelled too). - `withTimeout(...) { }` cancels the block on deadline. - Cancellation-cooperative code: long CPU loops should call `yield()`/check `isActive`; cleanup that must run during cancellation goes in `finally`, wrapping suspend cleanup in `withContext(NonCancellable)` if it must complete. Don't swallow `CancellationException` — rethrow it. - Because coroutines are structured, launching background work with the handler's own scope means it is cancelled when the request ends; use an application-scoped `CoroutineScope` for fire-and-forget that must outlive the request. **Context propagation.** Two context worlds meet: - Reactor `ContextView` ↔ coroutine `ReactorContext` element: bridged, so Reactor context (e.g. reactive security's `ReactiveSecurityContextHolder`) is reachable. - Thread-local context (SLF4J MDC, thread-bound security) does NOT automatically follow coroutine dispatch or event-loop hops. To carry MDC/trace ids you need explicit propagation — e.g. a `ThreadContextElement` (`MDCContext` from `kotlinx-coroutines-slf4j`), Micrometer context-propagation, or manual copying. This is a classic production gotcha: logs lose their correlation id after a `withContext` switch unless propagated. **Backpressure.** A `Flow<T>` handler adapts to `Flux<T>`, so transport-level backpressure (Netty demand) flows back into the `Flow`. Emit lazily; avoid buffering an unbounded `Flow` into a list. `Dispatchers`/`buffer`/`flowOn` change where upstream Flow work runs without breaking backpressure. **Observability & error handling.** Errors thrown in the coroutine become `Mono.onError`, handled by `@ExceptionHandler`/`ErrorWebExceptionHandler`; ensure `CancellationException` is not treated as a server error. Metrics/timers must account for cancellation as a distinct outcome. **When it matters.** These concerns dominate real production WebFlux+coroutine systems: capacity is set by keeping the event loop non-blocking; correctness under load depends on cancellation cooperation; and debuggability depends on context propagation. Getting them right is the difference between the reactive stack's throughput benefits and a subtly worse system than blocking MVC.
- A trace/correlation id logged at the start of a handler disappears from logs after a withContext(Dispatchers.IO) switch. Why, and how do you fix it?MDC is thread-local; switching dispatchers moves execution to another thread that has no MDC. Propagate it explicitly — e.g. add `MDCContext()` (kotlinx-coroutines-slf4j) to the coroutine context, or use Micrometer context-propagation so the id is restored on the new thread.
- How should cleanup code behave when a request coroutine is cancelled mid-flight?Put cleanup in `finally`. Since the coroutine is already cancelling, suspend cleanup would immediately throw; wrap must-complete suspend cleanup in `withContext(NonCancellable) { ... }`. Never catch-and-swallow `CancellationException`; rethrow it so structured cancellation completes.
saying these in an interview costs you the question
- Assuming coroutines make blocking JDBC safe on WebFlux threads.
- Believing dispatch to a background pool happens automatically without withContext.
- Thinking MDC/thread-local context follows coroutine dispatcher switches automatically.
- Catching and swallowing CancellationException.
- Ignoring that client disconnect cancels the handler coroutine.