Why are coroutines described as 'cheap', and what does that buy you compared to a thread-per-task model?
answer
- Coroutine = small heap object, not an OS stack
- Suspended coroutine holds zero threads
- Few threads multiplex many coroutines
- Concurrency cheap; parallelism still = thread count
- Blocking calls pin a pool thread and break it
basics
~10 sEach coroutine is just a small object, not an OS thread, so creating one barely costs memory. That lets you have huge numbers of waiting tasks at once without exhausting threads.
solid answer
~50 sCoroutines are cheap because they are **user-space objects**, not OS threads. Creating one allocates a small continuation/state-machine object on the heap rather than reserving an OS thread with its ~0.5–1 MB stack and kernel scheduling overhead. Crucially, a **suspended** coroutine consumes no thread at all — `delay`, suspending I/O, and channel waits free the thread back to the dispatcher's pool, so a handful of threads can serve tens of thousands of in-flight coroutines. Compared to thread-per-task, this removes two limits: memory (you can't hold 50k native stacks) and context-switch cost (the kernel doesn't preempt-switch every coroutine). The trade-off: coroutines help with **concurrency of waiting work** (I/O-bound, lots of pending tasks). They do **not** add CPU parallelism by themselves — actual parallelism still comes from the threads in the dispatcher (e.g. `Dispatchers.Default` sized to CPU cores).
go deeper
Knows coroutines are lighter than threads and you can make a lot of them.
Explains the memory/scheduling savings and that suspended coroutines free their thread, enabling many over a small pool.
Distinguishes concurrency vs parallelism precisely, ties parallelism to dispatcher thread count, and flags blocking-call pitfalls.
Reasons about sizing dispatchers, backpressure, and when thread-per-task or virtual threads might be the better fit instead.
## What 'cheap' means concretely Two costs dominate a thread-per-task design: 1. **Memory**: every OS thread reserves a stack (commonly ~0.5–1 MB). Ten thousand threads is gigabytes of stacks. A coroutine instead is a small heap object (a generated `Continuation`/state machine holding its label and spilled locals) — on the order of bytes-to-kilobytes. 2. **Scheduling**: the kernel preemptively context-switches threads, which has overhead and contention. Coroutines switch **cooperatively** at suspension points within user space — much lighter. Because of this, you can have a very large number of coroutines that are mostly **waiting**, multiplexed over a small thread pool. ## Why suspension is the key The cheapness isn't only about creation — it's that a **suspended coroutine holds no thread**. When a coroutine calls a suspending function like `delay` or a suspending HTTP client, it returns its thread to the dispatcher, which immediately uses that thread to run another coroutine. So the number of threads is bounded (sized to cores or I/O capacity) while the number of coroutines can be huge. ```kotlin import kotlinx.coroutines.* fun main() = runBlocking(Dispatchers.Default) { val results = (1..50_000).map { i -> async { delay(100) // all suspend; threads are reused i * 2 } }.awaitAll() println(results.sum()) // runs fine on just a few threads } ``` ## What it buys you vs thread-per-task - **Scale of concurrent waiting work**: handle 50k open requests/timers/connections without 50k threads. - **Lower memory and switch overhead**. - **Structured lifetimes**: coroutines are organized under a scope/`Job` tree, so cancellation and completion propagate — something raw threads don't give you for free. ## The important caveat Coroutines give **concurrency**, not automatic **parallelism**. Parallelism (running code on multiple cores at once) comes from the **dispatcher's threads**. A single-threaded dispatcher runs coroutines concurrently but never in parallel. For **CPU-bound** work, you are still limited by core count — coroutines don't make computation faster; they make *waiting* cheap. Blocking calls inside a coroutine (e.g. `Thread.sleep`, blocking JDBC) defeat the model by pinning a pool thread. ## APIs to name `launch`, `async`/`awaitAll`, `delay`, `Dispatchers.Default` (CPU-sized), `Dispatchers.IO` (elastic for blocking I/O), `Job`, `CoroutineScope`.
- Do coroutines make CPU-bound work faster?No. CPU parallelism comes from the dispatcher's threads (e.g. Dispatchers.Default sized to cores). Coroutines make waiting cheap; they don't add cores or speed up computation.
- What happens if you call a blocking function (like Thread.sleep or blocking JDBC) inside a coroutine?It pins the underlying pool thread for the whole duration, so other coroutines can't use it — defeating the cheapness. Use suspending equivalents or Dispatchers.IO for unavoidable blocking.
A restaurant with 3 waiters (threads) serving 200 tables (coroutines): when a table is just thinking (suspended), the waiter helps someone else instead of standing idle.
saying these in an interview costs you the question
- Claims coroutines give automatic multi-core parallelism
- Says a suspended coroutine still occupies a thread
- Thinks more coroutines always means more speed
- Ignores that blocking calls pin pool threads
- Confuses concurrency with parallelism