skip to content

Why does the event-loop model scale better than thread-per-request under high concurrency with slow I/O, and where does thread-per-request hit its wall?

level: middleimportance: must knowfreq 75%

answer

  1. Throughput ≈ threads / latency (200/0.5s ≈ 400 rps)
  2. Blocked thread = held hostage, CPU idle
  3. Event loop returns thread during wait
  4. Only pays off if end-to-end non-blocking (R2DBC/WebClient)
  5. Backpressure + BlockHound

basics

~20 s

In thread-per-request, each slow call ties up a whole thread doing nothing but waiting, so once all threads are waiting no new requests can be served. Event-loop threads release themselves during waits, so a few threads serve thousands of waiting requests.

solid answer

~50 s

Thread-per-request couples one thread to each in-flight request for its entire duration, including idle wait time. When downstream I/O is slow (a 500ms DB or remote call), each thread spends most of its life parked. Concurrency is then capped by pool size: with a 200-thread pool and 500ms calls you top out near 400 req/s, and beyond that requests queue while CPU sits idle. You can raise the pool, but thousands of threads cost ~1MB stack each plus context-switch overhead, so it doesn't scale linearly. The event loop decouples thread from request: on a would-block I/O op the pipeline suspends and the thread returns to the loop to serve others, resuming via callback when data arrives. So a per-core pool multiplexes tens of thousands of *mostly-waiting* connections. The catch: it only pays off if the *entire* chain is non-blocking (WebClient, R2DBC) — one blocking call on the loop erases the benefit.

code

java · 24 lines
java
// The scaling difference in outbound-call form.

// MVC: RestTemplate blocks the worker thread for the full 500ms round-trip.
@GetMapping("/mvc/aggregate")
Report aggregateBlocking() {
    var a = restTemplate.getForObject("/svc-a", A.class); // thread parked ~500ms
    var b = restTemplate.getForObject("/svc-b", B.class); // parked again
    return new Report(a, b);
}

// WebFlux: both calls fly on WebClient; the event-loop thread is free during
// the waits, and the two calls run concurrently via zip.
@GetMapping("/flux/aggregate")
Mono<Report> aggregateReactive() {
    Mono<A> a = webClient.get().uri("/svc-a").retrieve().bodyToMono(A.class);
    Mono<B> b = webClient.get().uri("/svc-b").retrieve().bodyToMono(B.class);
    return Mono.zip(a, b, Report::new); // no thread held during I/O
}

// If you MUST call blocking code, keep it off the event loop:
Mono<String> safeBlocking() {
    return Mono.fromCallable(legacyDao::blockingLookup)
              .subscribeOn(Schedulers.boundedElastic());
}

go deeper

for a junior

Grasp that a blocked thread wastes capacity because it sits idle but unavailable, and that MVC has a fixed thread ceiling.

for a middle

Compute the throughput = threads/latency ceiling and explain why the event loop lifts it by releasing threads during waits.

for a senior

Stress the end-to-end non-blocking requirement (R2DBC/WebClient), boundedElastic offloading, and backpressure as a scaling lever.

for a principal

Weigh connection/FD limits, downstream DB connection ceilings, and operational complexity when deciding whether reactive is worth it at your concurrency.

## The core insight: thread ≠ work when you're waiting Most web requests spend the majority of their wall-clock time **waiting** on something else — a database, a cache, another microservice. During that wait, no CPU work happens. The two models differ in what they do with a thread during that idle wait. ### Thread-per-request: the thread is *held hostage* by the wait In Spring MVC on a servlet container, the worker thread that started the request stays assigned to it until the response is written. A synchronous JDBC call or `RestTemplate` call **blocks** that thread — the OS parks it, and it is neither doing work nor available for another request. **The wall.** Concurrency = pool size. With Tomcat's default `server.tomcat.threads.max=200` and each request making a 500 ms downstream call: - Steady-state throughput ≈ threads / latency = 200 / 0.5s = **~400 requests/sec**, no matter how idle the CPU is. - Request #201 (concurrently) waits in the accept/connection queue; if the queue fills, clients get connection timeouts. - Naively raising `threads.max` to, say, 5,000 costs ~5 GB of thread stacks and heavy context-switching, and databases usually can't accept 5,000 connections anyway. So you can't just add threads forever. ### Event-loop: the thread is *returned* during the wait WebFlux on Reactor Netty runs a fixed per-core pool. When a reactive pipeline reaches a would-block point — say a `WebClient` HTTP call — Reactor registers a callback with the non-blocking network layer and the event-loop thread **immediately goes back to the loop** to process other ready events. When the response bytes arrive, an event-loop thread picks the continuation back up. **Result:** the number of *concurrent connections* is bounded by memory and file descriptors, not thread count. A 4-core box with 4–8 event-loop threads can hold tens of thousands of concurrent, mostly-idle connections. This is exactly the workload where thread-per-request falls over: **high concurrency + slow, I/O-bound waiting**. ## The load-bearing condition: end-to-end non-blocking The event loop only scales if nothing blocks it. If a WebFlux handler calls blocking JDBC/JPA or `RestTemplate` directly on the loop, that event-loop thread parks — and because there are only a few of them, you stall *many* multiplexed requests at once, which is *worse* than MVC. Requirements to actually get the benefit: - Use the reactive `WebClient` (not `RestTemplate`) for outbound HTTP. - Use **R2DBC** (or reactive Mongo/Redis) instead of JDBC/JPA. - If a library is unavoidably blocking, wrap it: `Mono.fromCallable(this::blockingCall).subscribeOn(Schedulers.boundedElastic())` so the block happens on a separate elastic pool, not the loop. ## Edge cases & gotchas - **CPU-bound work gets no benefit.** If requests are compute-heavy rather than I/O-waiting, the event loop is busy anyway; you're limited by cores in both models, and MVC is simpler. - **Backpressure.** Reactive Streams lets a slow consumer signal a fast producer to slow down (`request(n)`), preventing unbounded buffering — a scaling tool MVC lacks natively. - **BlockHound.** A test/diagnostic agent that detects accidental blocking calls on non-blocking threads — worth wiring in to enforce the golden rule. - **Latency isn't magically lower.** A single request against an idle server has similar latency in both models. The difference shows up as *throughput and stability under concurrency*. ## When to reach for which - Thread-per-request: blocking data stack (JPA), moderate concurrency, team familiarity, linear debuggability. - Event-loop: API gateways, fan-out aggregators, streaming (SSE/WebSocket), very high concurrent connection counts with slow backends.

  • You have a 200-thread Tomcat pool and each request makes one 250ms downstream call. Roughly what's your concurrency ceiling and throughput?
    About 200 concurrent requests; steady-state throughput ≈ 200 / 0.25s = ~800 req/s. Beyond that requests queue while CPU stays idle, because threads are parked in the wait.
  • Does switching to WebFlux help if your service is CPU-bound rather than I/O-bound?
    No. With no idle waiting, event-loop threads stay busy computing; you're core-bound in both models. WebFlux only helps when threads would otherwise be parked waiting on I/O.

saying these in an interview costs you the question

  • Thinking event-loop scaling comes from creating more threads under load
  • Claiming WebFlux helps CPU-bound workloads
  • Forgetting that a single blocking call on the loop erases the benefit
  • Believing raising Tomcat's max-threads to thousands scales cleanly

context