What does it actually mean for a thread to "block" on an operation, and why is one blocked thread far more damaging inside an event-driven runtime than inside a server that dedicates a thread to each request?
answer
- blocked = holds a thread, makes no progress
- cost = how many tasks that thread carries
- 1 thread : 1 request vs 1 thread : all requests
- CPU-heavy handler == blocking, to everyone else
- async isn't faster, waiting is just thread-free
basics
~20 sBlocking means the thread stops running and cannot do anything else until the operation finishes. In a thread-per-request server that costs one request. In an event-driven runtime a handful of threads carry every request, so one blocked thread freezes many unrelated ones.
solid answer
~60 sTo block is to hold a thread that makes no progress: the thread is parked (or spinning) until some external event — bytes arriving, a lock releasing, a disk seek finishing — lets it continue. The thread's stack, its slot in the pool, and its scheduling entity stay reserved the whole time. The cost depends entirely on how many requests that thread represents. In a thread-per-request server the ratio is 1:1, so blocking is contained: one slow call delays one client. Pools are sized generously precisely because most threads are expected to be blocked. An event-driven runtime inverts the ratio. A small number of threads — often one per core — carry *all* connections by interleaving short handlers. Blocking one of them stops every connection assigned to it, so a single 200 ms call can add 200 ms to hundreds of unrelated requests. With four loops, four concurrent blocking calls halt the entire service. That asymmetry — not the blocking call itself — is why async runtimes forbid blocking in handlers.
go deeper
Define blocking as the thread being held and unable to do other work, and say the harm scales with how many requests that thread was serving.
Contrast 1:1 thread-per-request containment with 1:many multiplexing, and include CPU-heavy handlers as equivalent to blocking.
Name the full inventory of blocking operations including hidden ones, and give the detection signals — event-loop latency and blocking-call instrumentation.
Frame it as blast radius: which threads carry shared work, what isolation boundary protects them, and how the service degrades when the offload path saturates.
## What blocking is A thread is *blocked* when it cannot make progress and is waiting for something external. Usually the operating system parks it: the thread leaves the CPU, is placed on a wait queue, and becomes runnable again when the event it waits for occurs. Sometimes it spins instead, burning CPU while polling. In both cases one property holds — **the thread is unavailable for other work, and everything it owns stays reserved**: its stack, its entry in whatever pool handed it out, and any locks it holds. The common blocking operations are worth naming, because people only remember the first one: - Network reads and writes on a blocking handle. - Regular-file reads and writes, including implicit ones like loading a class, reading a config file, or touching a memory-mapped page that is not resident. - Acquiring a contended lock or semaphore. - Waiting on a queue, a latch, a future, or a child process. - Name resolution — often forgotten, frequently slow, frequently synchronous even in "async" libraries. - Long CPU-bound computation. It does not park the thread, but it is *indistinguishable from blocking* to everyone else waiting for that thread. Cryptography, compression, big JSON serialisation and regular-expression backtracking all qualify. ## Why the same call has two very different costs The damage from blocking equals *how many logical tasks that thread was carrying*. **Thread-per-request.** Each request owns a thread for its lifetime. Blocking is the normal state — a pool of 200 threads exists on the assumption most are parked waiting on I/O. One slow dependency delays the requests that touch it; unrelated requests keep flowing on other threads. The system is wasteful in memory but *contained* in blast radius. It fails only when the pool is exhausted, which is a resource limit you can watch, size and alarm on. **Event-driven.** A few threads run an endless loop: take the next ready event, run a short handler, take the next. Every connection is multiplexed over those threads, so a thread is not *a* request — it is *all of them, in turn*. Handlers are expected to be short and to yield at every wait point, returning the thread to the loop. If a handler blocks for 200 ms: - Every event already queued for that loop waits 200 ms longer. - Nothing is preempted out — the loop is a cooperative scheduler; it cannot take the thread back. - The symptom is latency on endpoints that have nothing to do with the slow call, which is deeply confusing during an incident. - With N loop threads, N concurrent blocking calls stop the service completely, and N is typically the core count, not a big number. The same reasoning applies to lightweight-thread runtimes, one level down: a task that blocks in a way the runtime cannot intercept *pins its carrier thread*, and there are only a few carriers. ## The rule and its consequence The rule is simple: **never block, and never run long, on a thread that carries other people's work.** The consequence is that async systems need a place to put unavoidable blocking work — a separate, bounded pool of ordinary threads whose only job is to absorb it, so the event loops stay free. Async code then waits for that pool's result *asynchronously*, which keeps the loop yielding. ## Blocking is not the same as being slow An asynchronous call to a slow service still takes a long time; the difference is that no thread is held while it waits. Latency is a property of the dependency; thread occupancy is a property of your code. Async I/O does not make anything faster — it makes waiting free in thread terms, so a fixed number of threads can hold a much larger number of in-flight operations. If a candidate says "we went async and it got faster", the honest version is "we stopped running out of threads". ## How to notice Two signals catch this reliably. First, **event-loop (or scheduler) latency**: the delay between an event becoming ready and its handler starting. It rises for *all* work when any handler blocks, which distinguishes hidden blocking from a genuinely slow dependency. Second, **blocking-call detectors**: instrumentation or agents that flag known blocking operations executing on a loop thread. Both belong in a service that mixes blocking libraries with async plumbing.
- Is a long CPU-bound computation inside an event-loop handler the same problem as a blocking I/O call?From the loop's point of view, yes — the thread is unavailable either way, and nothing preempts a running handler. It is arguably worse, because a blocked thread at least releases the CPU for other loops, while a CPU-heavy handler holds both the loop and the core. The fix differs, though: CPU work goes to a pool sized near the core count, whereas blocking I/O goes to a larger pool sized by wait time.
- Does making a call asynchronous make it finish sooner?No. The dependency takes exactly as long; asynchrony only means no thread is parked while you wait. The benefit is capacity — a fixed, small number of threads can hold far more operations in flight — not per-request latency. Latency usually improves only indirectly, because you stop queueing behind an exhausted thread pool.
A thread-per-request server is one till per shopper: a slow shopper delays only that queue. An event loop is one cashier serving many shoppers a step at a time — the moment that cashier stands still for a phone call, everybody's shopping stops.
saying these in an interview costs you the question
- Saying asynchronous code is faster than blocking code for the same operation.
- Assuming only network calls block — missing file reads, name resolution, lock acquisition and class loading.
- Thinking the runtime will preempt a long-running handler and let other work through.
- Believing a bigger event-loop pool fixes blocking rather than just delaying collapse.
- Treating CPU-bound work in a handler as harmless because it is not "waiting".