skip to content

Why must you never run blocking code on a reactor-http-nio thread, and what do you do when you have an unavoidable blocking call?

level: middleimportance: must knowfreq 75%

answer

  1. few loops, many connections -> block one = stall many
  2. symptom: high latency, low CPU
  3. boundedElastic() for blocking I/O
  4. R2DBC / WebClient = never block
  5. BlockHound detects it; .block() throws on the loop

basics

~10 s

Event-loop threads are few and shared across many connections. Blocking one freezes every request it serves. Offload blocking work to a separate scheduler with .subscribeOn(Schedulers.boundedElastic()).

solid answer

~40 s

Reactor Netty runs on a small set of event-loop threads (reactor-http-nio-*, roughly one per CPU core). Each loop multiplexes many connections, so it must return quickly to service the next ready socket. If you call blocking JDBC, a synchronous HTTP client, Thread.sleep, or .block() on that thread, the loop is stuck and every connection it owns stalls — throughput collapses and latency spikes even though CPU is idle. For unavoidable blocking work, move it off the loop with .subscribeOn(Schedulers.boundedElastic()), which uses a growable pool of worker threads sized for blocking tasks. Prefer fully reactive drivers (R2DBC, WebClient) so nothing blocks at all. In tests/dev you can enable BlockHound to detect accidental blocking calls on non-blocking threads at runtime.

code

java · 16 lines
java
@RestController
class ReportController {

    private final LegacyBlockingJdbcDao dao; // synchronous, blocks

    ReportController(LegacyBlockingJdbcDao dao) { this.dao = dao; }

    @GetMapping("/report")
    Mono<Report> report(@RequestParam long id) {
        // Wrap the blocking call and push it OFF the event loop.
        return Mono.fromCallable(() -> dao.loadReport(id))
                   .subscribeOn(Schedulers.boundedElastic());
        // Without subscribeOn, dao.loadReport() would block
        // a reactor-http-nio thread and stall other requests.
    }
}

go deeper

for a junior

Knows the rule 'don't block the event loop' even if fuzzy on why.

for a middle

Must explain the starvation mechanism and use subscribeOn(boundedElastic()) correctly; distinguishes reactive drivers from blocking ones.

for a senior

Distinguishes subscribeOn/publishOn, picks the right scheduler per workload, and uses BlockHound.

for a principal

Reasons about capacity: bounded-elastic sizing, backpressure interplay, and when a partially-blocking service should just stay on MVC instead.

## The constraint Reactor Netty serves all HTTP traffic on a **small, fixed** pool of **event-loop** threads (`reactor-http-nio-N`, default `max(4, CPU cores)`). Each thread is responsible for **many** connections via **I/O multiplexing** — it watches many sockets and runs the handler for whichever is ready, then moves on. The model only works if every handler yields the thread quickly. ## Why blocking is catastrophic here If your handler makes a **blocking** call — synchronous JDBC, a blocking `RestTemplate`/`HttpURLConnection`, `Thread.sleep`, `Mono.block()`, a synchronous file read, or an unbounded synchronized lock — the event-loop thread parks. While parked it cannot service **any** of the other connections assigned to it. With only a few loops, a handful of concurrent blocking calls can freeze the whole server. The symptom is nasty: **high latency / dropped throughput with low CPU usage**, because threads are waiting, not working. This is the single most common WebFlux production failure. ## The fix: offload with a scheduler Reactor provides **`Schedulers`** to move work onto other threads: - **`Schedulers.boundedElastic()`** — the right choice for blocking I/O. A pool that grows on demand up to a cap (default 10 × CPU cores) and reuses idle threads. Wrap the blocking call and `.subscribeOn(Schedulers.boundedElastic())` so it executes off the event loop. - `Schedulers.parallel()` — for CPU-bound, **non-blocking** work; do not put blocking calls here (same starvation risk). - Never invent your own unbounded pool without thought — bound it. ## Better: don't block at all The cleanest answer is to use **non-blocking drivers** end to end: **R2DBC** for SQL, reactive Mongo/Redis drivers, and **`WebClient`** (not `RestTemplate`) for HTTP. Then nothing ever blocks and no offloading is needed. ## Detecting accidental blocking **BlockHound** is an agent that instruments the JVM and throws when a known blocking call executes on a thread marked non-blocking (Reactor's event loops and `parallel()` scheduler are marked). Enable it in tests to catch regressions early. Reactor also emits schedule hooks you can use. ## `publishOn` vs `subscribeOn` - **`subscribeOn`** decides which scheduler the *subscription/source* runs on — affects the whole chain upstream regardless of position. - **`publishOn`** switches the thread for operators *downstream* of it. Use `publishOn` to hop back onto an appropriate scheduler after a blocking region, or to isolate a segment. ## Gotcha: `.block()` on an event loop throws Calling `Mono.block()` on a `reactor-http-nio` thread now throws `IllegalStateException: block()/blockFirst()/blockLast() are blocking, which is not supported in thread reactor-http-nio-*` — Reactor guards against it. Non-guarded blocking (raw JDBC) still silently stalls, hence BlockHound. ## When to use what - Reactive driver exists -> use it, no offload. - Only a blocking library exists -> wrap in `Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic())`. - CPU-heavy pure computation -> `Schedulers.parallel()`.

  • Why boundedElastic() rather than a fixed thread pool you create yourself?
    boundedElastic() is purpose-built for blocking I/O: it caps threads (default 10 × cores), reuses idle ones, evicts them after a TTL, and queues excess work — so it survives bursts without unbounded thread growth. A hand-rolled pool needs you to get all of that right yourself.
  • What is the difference between subscribeOn and publishOn?
    subscribeOn sets the scheduler for the subscription/source and influences the whole upstream chain no matter where it is placed. publishOn switches threads only for operators downstream of it, so you can hop threads mid-pipeline (e.g., back onto parallel after a blocking region).
  • How would you catch accidental blocking calls before production?
    Enable BlockHound in tests. It instruments the JVM and throws when a blocking method runs on a thread Reactor marks as non-blocking (event loops, parallel scheduler), surfacing the exact stack trace of the offending call.

saying these in an interview costs you the question

  • Suggesting you just add more event-loop threads to fix blocking
  • Using Schedulers.parallel() for blocking I/O
  • Calling .block() inside a controller on the event loop
  • Using RestTemplate/JDBC in WebFlux and assuming it scales
  • Thinking high latency with low CPU means you need a bigger machine

context