Why must you never run blocking code on a reactor-http-nio thread, and what do you do when you have an unavoidable blocking call?
answer
- few loops, many connections -> block one = stall many
- symptom: high latency, low CPU
- boundedElastic() for blocking I/O
- R2DBC / WebClient = never block
- BlockHound detects it; .block() throws on the loop
basics
~10 sEvent-loop threads are few and shared across many connections. Blocking one freezes every request it serves. Offload blocking work to a separate scheduler with .subscribeOn(Schedulers.boundedElastic()).
solid answer
~40 sReactor Netty runs on a small set of event-loop threads (reactor-http-nio-*, roughly one per CPU core). Each loop multiplexes many connections, so it must return quickly to service the next ready socket. If you call blocking JDBC, a synchronous HTTP client, Thread.sleep, or .block() on that thread, the loop is stuck and every connection it owns stalls — throughput collapses and latency spikes even though CPU is idle. For unavoidable blocking work, move it off the loop with .subscribeOn(Schedulers.boundedElastic()), which uses a growable pool of worker threads sized for blocking tasks. Prefer fully reactive drivers (R2DBC, WebClient) so nothing blocks at all. In tests/dev you can enable BlockHound to detect accidental blocking calls on non-blocking threads at runtime.
code
java · 16 lines@RestController
class ReportController {
private final LegacyBlockingJdbcDao dao; // synchronous, blocks
ReportController(LegacyBlockingJdbcDao dao) { this.dao = dao; }
@GetMapping("/report")
Mono<Report> report(@RequestParam long id) {
// Wrap the blocking call and push it OFF the event loop.
return Mono.fromCallable(() -> dao.loadReport(id))
.subscribeOn(Schedulers.boundedElastic());
// Without subscribeOn, dao.loadReport() would block
// a reactor-http-nio thread and stall other requests.
}
}go deeper
Knows the rule 'don't block the event loop' even if fuzzy on why.
Must explain the starvation mechanism and use subscribeOn(boundedElastic()) correctly; distinguishes reactive drivers from blocking ones.
Distinguishes subscribeOn/publishOn, picks the right scheduler per workload, and uses BlockHound.
Reasons about capacity: bounded-elastic sizing, backpressure interplay, and when a partially-blocking service should just stay on MVC instead.
## The constraint Reactor Netty serves all HTTP traffic on a **small, fixed** pool of **event-loop** threads (`reactor-http-nio-N`, default `max(4, CPU cores)`). Each thread is responsible for **many** connections via **I/O multiplexing** — it watches many sockets and runs the handler for whichever is ready, then moves on. The model only works if every handler yields the thread quickly. ## Why blocking is catastrophic here If your handler makes a **blocking** call — synchronous JDBC, a blocking `RestTemplate`/`HttpURLConnection`, `Thread.sleep`, `Mono.block()`, a synchronous file read, or an unbounded synchronized lock — the event-loop thread parks. While parked it cannot service **any** of the other connections assigned to it. With only a few loops, a handful of concurrent blocking calls can freeze the whole server. The symptom is nasty: **high latency / dropped throughput with low CPU usage**, because threads are waiting, not working. This is the single most common WebFlux production failure. ## The fix: offload with a scheduler Reactor provides **`Schedulers`** to move work onto other threads: - **`Schedulers.boundedElastic()`** — the right choice for blocking I/O. A pool that grows on demand up to a cap (default 10 × CPU cores) and reuses idle threads. Wrap the blocking call and `.subscribeOn(Schedulers.boundedElastic())` so it executes off the event loop. - `Schedulers.parallel()` — for CPU-bound, **non-blocking** work; do not put blocking calls here (same starvation risk). - Never invent your own unbounded pool without thought — bound it. ## Better: don't block at all The cleanest answer is to use **non-blocking drivers** end to end: **R2DBC** for SQL, reactive Mongo/Redis drivers, and **`WebClient`** (not `RestTemplate`) for HTTP. Then nothing ever blocks and no offloading is needed. ## Detecting accidental blocking **BlockHound** is an agent that instruments the JVM and throws when a known blocking call executes on a thread marked non-blocking (Reactor's event loops and `parallel()` scheduler are marked). Enable it in tests to catch regressions early. Reactor also emits schedule hooks you can use. ## `publishOn` vs `subscribeOn` - **`subscribeOn`** decides which scheduler the *subscription/source* runs on — affects the whole chain upstream regardless of position. - **`publishOn`** switches the thread for operators *downstream* of it. Use `publishOn` to hop back onto an appropriate scheduler after a blocking region, or to isolate a segment. ## Gotcha: `.block()` on an event loop throws Calling `Mono.block()` on a `reactor-http-nio` thread now throws `IllegalStateException: block()/blockFirst()/blockLast() are blocking, which is not supported in thread reactor-http-nio-*` — Reactor guards against it. Non-guarded blocking (raw JDBC) still silently stalls, hence BlockHound. ## When to use what - Reactive driver exists -> use it, no offload. - Only a blocking library exists -> wrap in `Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic())`. - CPU-heavy pure computation -> `Schedulers.parallel()`.
- Why boundedElastic() rather than a fixed thread pool you create yourself?boundedElastic() is purpose-built for blocking I/O: it caps threads (default 10 × cores), reuses idle ones, evicts them after a TTL, and queues excess work — so it survives bursts without unbounded thread growth. A hand-rolled pool needs you to get all of that right yourself.
- What is the difference between subscribeOn and publishOn?subscribeOn sets the scheduler for the subscription/source and influences the whole upstream chain no matter where it is placed. publishOn switches threads only for operators downstream of it, so you can hop threads mid-pipeline (e.g., back onto parallel after a blocking region).
- How would you catch accidental blocking calls before production?Enable BlockHound in tests. It instruments the JVM and throws when a blocking method runs on a thread Reactor marks as non-blocking (event loops, parallel scheduler), surfacing the exact stack trace of the offending call.
saying these in an interview costs you the question
- Suggesting you just add more event-loop threads to fix blocking
- Using Schedulers.parallel() for blocking I/O
- Calling .block() inside a controller on the event loop
- Using RestTemplate/JDBC in WebFlux and assuming it scales
- Thinking high latency with low CPU means you need a bigger machine