skip to content

What exactly does "suspend without blocking a thread" mean, and how does it improve scalability over thread-per-request blocking?

level: middleimportance: should knowfreq 60%

answer

  1. Suspend -> thread returned to pool; state on the heap
  2. Blocking -> idle thread reserved per request
  3. Threads ~1MB stack each; coroutines are heap objects
  4. Cooperative: only suspends at suspension points
  5. CPU-bound loops still hog a thread

basics

~20 s

Suspending pauses a coroutine but hands its thread back to do other work, so one thread serves many waiting tasks. Blocking keeps the thread idle and busy, so you need one thread per waiting task.

solid answer

~40 s

"Non-blocking suspension" means when a coroutine hits a suspension point and must wait, it releases the underlying thread rather than holding it idle. The thread is free to run other coroutines; when the awaited result is ready, the coroutine's continuation is rescheduled and resumes. Contrast thread-per-request blocking: each in-flight request owns a thread even while it sits idle on IO, so concurrency is capped by thread count, and threads are expensive (~1MB stack each, context-switch overhead). With coroutines, thousands of concurrent suspended operations can share a tiny pool (e.g. CPU-count threads on `Dispatchers.Default`, a larger pool on `Dispatchers.IO`). The result: far higher concurrency for IO-bound workloads at much lower memory and scheduling cost. Crucially, suspension is cooperative — it only happens at suspension points; CPU-bound code that never suspends still ties up a thread.

go deeper

for a junior

Grasps that suspending frees the thread so one thread serves many waiting tasks.

for a middle

Contrasts thread-per-request blocking with suspension and cites memory/throughput benefits.

for a senior

Explains cooperative scheduling, the CPU-bound caveat, and dispatcher sizing rationale.

for a principal

Reasons about system-level capacity planning, when coroutines do/don't help, and chooses concurrency models accordingly.

## "Blocking" vs "suspending" precisely - A **thread** is an OS-scheduled execution resource with its own stack (~1MB on the JVM). Threads are limited and costly. - **Blocking** a thread means it sits in a wait state (e.g. `Thread.sleep`, a synchronous socket read) doing nothing but unavailable to anyone else. - **Suspending** a coroutine means the coroutine pauses at a **suspension point** (a `suspend` call like `delay` or a suspending IO read) and **returns its thread to the dispatcher**. The coroutine's state lives on the heap (in its continuation), not on a parked thread. ## Why this scales for IO-bound work Thread-per-request servers map one thread to each in-flight request. If a request spends most of its time waiting on a database or remote API, that thread is idle but reserved. To serve 10,000 concurrent slow requests you'd need ~10,000 threads — gigabytes of stack and heavy context-switching. With coroutines, a suspended request consumes only a small heap object. Ten thousand suspended coroutines can be served by a handful of threads, because at any instant only the few that are actually running need a thread. ```kotlin // Conceptually: many suspended coroutines, few threads coroutineScope { repeat(10_000) { launch { // 10k coroutines delay(1000) // all suspended at once -> threads are free } } } ``` ## The cooperative catch Suspension is **cooperative**, not preemptive. The runtime cannot forcibly pause a coroutine; it only suspends at suspension points. So a tight CPU loop that never calls a suspend function will **hold its thread** the entire time and can starve others — coroutines don't magically parallelize CPU work. For CPU-bound work you still need real parallelism (multiple threads via `Dispatchers.Default`), and you may insert `yield()` to be cooperative. ## Dispatchers and thread pools - `Dispatchers.Default` — CPU-bound; pool sized to core count. - `Dispatchers.IO` — blocking-IO-friendly; larger elastic pool (default cap 64). - `Dispatchers.Main` — UI thread (Android/JS/Swing). The dispatcher decides which thread resumes a coroutine after suspension; suspension itself is dispatcher-agnostic. ## Bottom line Non-blocking suspension turns waiting from a per-thread cost into a per-heap-object cost, which is what makes coroutines scale for IO-bound concurrency.

  • If coroutines are so scalable, why doesn't a CPU-heavy computation benefit?
    Because suspension is cooperative — a tight CPU loop never reaches a suspension point, so it holds its thread. CPU work needs real parallelism across threads, not suspension.
  • Why is Dispatchers.IO allowed a much larger pool than Dispatchers.Default?
    IO threads spend most time blocked on IO, so more of them improve throughput without overloading CPUs; Default is sized to cores because it runs CPU-bound work.

Blocking is reserving a whole phone line per customer even while they're on hold; suspending is a callback system — you hang up, the line serves others, and you're rung back when it's your turn.

saying these in an interview costs you the question

  • Claiming coroutines make CPU-bound code faster automatically
  • Saying suspension is preemptive / the runtime can pause anywhere
  • Not distinguishing IO-bound from CPU-bound scalability
  • Believing each coroutine owns a dedicated thread
  • Ignoring that blocking code inside coroutines breaks the model

context