skip to content

Async and Event-Driven Models

Concurrency without a thread per task: event loops and callbacks, futures and promises, async/await coroutines, and the reactor and proactor patterns. Expect to explain how cooperative scheduling differs from preemptive threading and what one blocking call does to an event loop.

part ofComputer science fundamentalsoverview, primer and where to startread it →
on this pageshow

questions

30

What is a coroutine, and what actually happens when one suspends, compared with an operating-system thread that blocks waiting for I/O?

level: juniorimportance: must knowfreq 58%

answer

  1. pause and resume, not stop and restart
  2. suspend frees the carrier thread; block does not
  3. state in a heap record, not a reserved stack
  4. cooperative: control moves only at suspension points
  5. concurrency unbounded, parallelism = carrier threads

basics

~20 s

A coroutine is a function that can pause partway through and resume later. Suspending saves its state, hands control back to a scheduler, and frees the underlying thread for other work. A blocked thread keeps its whole stack and OS slot while doing nothing.

solid answer

~50 s

A **coroutine** is a computation that can pause at defined points and later resume exactly where it stopped, carrying its local variables with it. When a coroutine **suspends**, the runtime captures its resume point and live locals, registers interest in whatever it is waiting for, and returns control to the scheduler. The carrier thread is then free to run other coroutines. When the awaited event completes, the scheduler resumes the coroutine, possibly on a different carrier thread. When an OS thread **blocks**, the kernel deschedules it, but the thread still exists: its stack (hundreds of kilobytes to megabytes of reserved address space), its kernel task structure, and its scheduler entry stay reserved. Nothing else can use it. The difference is who pays for the wait. Blocking costs one whole thread per outstanding operation; suspension costs one small heap object. That is why a process can hold hundreds of thousands of suspended coroutines but not hundreds of thousands of blocked threads.

go deeper

for a junior

Define a coroutine as a function that can pause and resume, and state the core contrast: suspending releases the thread, blocking holds it hostage.

for a middle

Add the mechanics — the continuation record, the carrier-thread pool, and the cost comparison of a heap object versus a reserved stack plus kernel context switch.

for a senior

Emphasise the cooperative-scheduling consequence: no suspension point means no yield, so CPU-bound work starves the pool. Note that resumption can move threads, breaking thread-affine assumptions.

for a principal

Frame it as a resource-model choice: coroutines decouple in-flight concurrency from thread count so capacity scales with outstanding requests instead of cores, at the cost of needing every I/O path in the stack to be non-blocking.

## The problem coroutines solve A server handling many slow operations must remember, for each one in flight, "where was I, and what do I do when the answer arrives?" The traditional place to store that memory is a thread: the call stack *is* the record of where you were. It reads beautifully — you write straight-line code — but it is expensive. Each thread reserves a contiguous stack sized for the deepest call it might ever make, plus kernel bookkeeping. Ten thousand concurrent waits means ten thousand stacks sitting idle. A coroutine keeps the same straight-line readability but stores "where was I" in a much smaller, heap-allocated record instead of a reserved stack. ## Definition A coroutine is a routine with more than one entry point: it can *yield* control back to its caller or scheduler and later be re-entered at the point it left off, with its locals intact. Ordinary functions have exactly one entry (the top) and give up their state when they return. Coroutines generalise that. ## Suspend versus block - **Block**: the calling thread makes a request the kernel cannot satisfy immediately. The kernel marks the thread not-runnable and switches to another thread. The blocked thread's stack, registers, and task structure remain allocated until the wait ends. Cost: one thread per concurrent wait, plus a kernel context switch (typically a few microseconds) on each transition. - **Suspend**: the coroutine reaches a suspension point, the runtime stores its continuation (resume point plus live locals) into a small object, arranges to be notified when the awaited result is ready, and *returns* to the scheduler. Cost: one small allocation and a normal function return; commonly tens to hundreds of nanoseconds. No kernel involvement at all in the common case. Crucially, the thread that was running the coroutine is not waiting for anything. It goes back to the scheduler and picks up the next ready coroutine. ## Cooperative, not preemptive OS threads are **preemptive**: a timer interrupt can take the CPU away from a thread at essentially any machine instruction. Coroutines are **cooperative**: control only changes hands at a suspension point the code chose. That has two consequences. The good one: interleaving is bounded and visible. Between two suspension points, a coroutine runs without another coroutine on the same loop interleaving with it, which simplifies reasoning about invariants. The bad one: a coroutine that never suspends never yields. A long CPU-bound computation with no suspension point starves every other coroutine on that carrier thread. The scheduler has no way to intervene — that is what "cooperative" means. ## Carrier threads and multiplexing Coroutines do not execute themselves; they run on some pool of real threads, often called carrier or worker threads. A common shape is N carrier threads for N cores, multiplexing many thousands of coroutines. Concurrency (many things in progress) is unbounded; parallelism (many things executing at the same instant) is capped by the carrier pool. A coroutine that is suspended occupies no thread at all. Because a coroutine may resume on a different carrier thread than it suspended on, thread-affine state — thread-local variables, thread-owned locks, thread identity used as an ownership token — becomes unsafe in a way it is not with plain threads. ## Why it matters in practice The payoff is memory and switching cost per concurrent operation. A workload that is 99% waiting (network calls, database round trips) needs concurrency proportional to the number of outstanding requests, not to CPU count. Coroutines make that number cheap, so a single process can hold far more in-flight work than its thread budget would allow, while still being written as ordinary sequential-looking code rather than a chain of callbacks.

  • If a coroutine runs a long computation and never reaches a suspension point, what does the scheduler do about it?
    Nothing — that is the defining limitation of cooperative scheduling. The carrier thread is occupied until the coroutine either suspends or finishes, so every other coroutine assigned to that thread is delayed. The fix is in the code: insert explicit yield points, chunk the work, or move it to a pool dedicated to CPU-bound tasks.
  • Do coroutines make a program parallel?
    No. Coroutines give you cheap concurrency — many operations in progress. Parallelism comes from the number of carrier threads and cores underneath. A million coroutines on one carrier thread execute strictly one at a time; the win is that waiting ones consume almost no resources.

A blocked thread is a checkout lane closed because the cashier is on the phone with a supplier — the lane is unusable. A suspended coroutine is the cashier setting your basket aside, serving the next customer, and picking your basket back up when the supplier calls back.

saying these in an interview costs you the question

  • Saying coroutines are scheduled by the OS kernel — they are scheduled in user space by the runtime
  • Claiming suspension is the same as sleeping, or that it blocks the thread briefly
  • Claiming coroutines make code faster or parallel by themselves
  • Assuming a coroutine resumes on the same thread it suspended on, so thread-local state is safe
  • Describing suspension as preemptive — the runtime cannot interrupt a coroutine that will not yield

context

open as a page

What is an event loop, and what does run-to-completion mean for the callbacks it dispatches?

level: juniorimportance: must knowfreq 60%

basics

~20 s

An event loop is a single thread cycling forever: wait for ready events, take the next task from a queue, run its handler to completion, repeat. No handler is ever interrupted mid-way by another handler, so loop-owned state needs no locks — but a slow handler delays everything behind it.

open as a page

In asynchronous programming, what is the distinction between a future and a promise, and why do many libraries hand out two separate objects for a single asynchronous result?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A future is the read side of a result that is not ready yet: you await it or attach a continuation. A promise is the write side: the producer completes it once, with a value or an error. Splitting them stops consumers from completing results.

open as a page

When you build a pipeline from steps that each return a future, what is the difference between transforming a future's value with a plain mapping function and chaining a step that itself returns a future — and what breaks if you confuse the two?

level: middleimportance: must knowfreq 56%

basics

~20 s

Mapping applies a synchronous function to the value. Chaining applies a function that returns another future and flattens it, so the result completes when the inner work finishes. Mapping with an async function yields a future of a future: the outer completes early, before the real work is done.

open as a page

How do failures travel through a chain of asynchronous steps built from futures, and why does wrapping the call in a try/catch at the call site usually fail to catch them?

level: middleimportance: must knowfreq 52%

basics

~20 s

A future completes with either a value or an error, and an error short-circuits the downstream transforms until a handler that accepts errors is reached. A try/catch at the call site only sees synchronous failures, because the function returns a pending handle long before the failure exists.

open as a page

In a reactive-streams style pipeline the producer pushes items to the consumer, yet the consumer sends demand signals upstream — messages that say request N more items. If delivery is push-based, why is that demand channel needed at all?

level: middleimportance: must knowfreq 45%

basics

~20 s

Pure push has no rate control: a fast producer and slow consumer means unbounded queue growth and eventual memory exhaustion. Demand signalling lets the consumer declare capacity, and the producer may never emit more than the outstanding demand, so items in flight stay bounded without blocking any thread.

open as a page

A server that dedicates one operating-system thread to each open connection behaves fine at a few hundred clients but collapses somewhere in the tens of thousands. Concretely, what runs out, and why did that problem drive servers toward event-driven designs?

level: middleimportance: must knowfreq 55%

basics

~20 s

Each thread costs reserved stack memory, a kernel scheduling entity, and cache footprint. At 10k+ mostly idle connections you pay gigabytes of stacks and constant context switching for almost no work, so throughput falls while connections rise. Event loops decouple connection count from thread count.

open as a page

A service built on a small pool of event-loop threads starts showing high p99 latency on many unrelated endpoints whenever one particular feature is used. How does a single blocking or CPU-heavy call inside a handler produce that pattern, and how would you locate it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Loop threads are shared by all connections, so a handler that blocks or computes long stops every queued event behind it — latency appears everywhere, not on the guilty endpoint. Find it by measuring event-loop lag, instrumenting handler durations, and profiling loop threads.

open as a page

If tasks in a runtime switch only at explicit suspension points and share one carrier thread, can you still get race conditions? Explain what is and is not guaranteed.

level: seniorimportance: must knowfreq 42%

basics

~20 s

Yes. Single-threaded cooperative scheduling removes low-level data races — no two tasks touch memory at the same instant — but not logical races. Any suspension point is a place where arbitrary other tasks ran, so check-then-act sequences spanning an await can still be violated.

open as a page

A service built on a single event loop has one request handler that spends 400 milliseconds doing CPU work. What happens to the rest of the system, and how would you detect it in production?

level: seniorimportance: must knowfreq 54%

basics

~20 s

Everything queued behind it waits: all in-flight requests take up to 400 ms longer, timers fire late, heartbeats and health checks miss, and client timeouts fire in a burst that triggers retries. Detect it by measuring loop lag — schedule a known-interval timer and record its overshoot — plus per-handler duration histograms.

open as a page

Asynchronous work can be expressed by passing a callback that is invoked on completion, or by returning a future object that the caller composes. What does the future style buy you over raw callbacks, and what does it cost?

level: juniorimportance: should knowfreq 55%

basics

~20 s

A future is a first-class value: you can store, return, and combine it, and it has one defined outcome — value or error. Callbacks invert control, nest as they compose, and each one invents its own error and completion convention. Futures cost allocation and clarity of stack traces.

open as a page

You already have futures and promises for asynchronous results. When is a future the wrong abstraction and an asynchronous stream the right one?

level: juniorimportance: should knowfreq 38%

basics

~20 s

A future carries exactly one outcome, once. A stream carries zero to many items over time plus a terminal completion or failure. Use a stream when results arrive incrementally, when there may be many or infinitely many, or when the producer's rate could outpace the consumer.

open as a page

What does it mean for a network socket read to be non-blocking, and how is that different from asynchronous (completion-based) I/O where the operating system performs the whole read and then tells you it finished?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Non-blocking: the read returns instantly with whatever bytes already exist or a "not ready" answer, so the thread never sleeps but you must retry when told the handle is ready. Asynchronous: you hand the OS a buffer, it does the read, then signals completion.

open as a page

What does it actually mean for a thread to "block" on an operation, and why is one blocked thread far more damaging inside an event-driven runtime than inside a server that dedicates a thread to each request?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Blocking means the thread stops running and cannot do anything else until the operation finishes. In a thread-per-request server that costs one request. In an event-driven runtime a handful of threads carry every request, so one blocked thread freezes many unrelated ones.

open as a page

Compare stackful and stackless coroutines: how does each remember where it was, and what can one do that the other cannot?

level: middleimportance: should knowfreq 44%

basics

~20 s

A stackful coroutine owns a real stack, so it can suspend from any depth, including inside a function that knows nothing about coroutines. A stackless one is compiled into a state machine holding only its own locals, so it can suspend only at explicit points in itself, but costs far less memory.

open as a page

Some event-loop runtimes keep two queues: a task queue (macrotasks) and a microtask/continuation queue drained after each task. What is the ordering rule, why does the second queue exist, and what hazard does it create?

level: middleimportance: should knowfreq 46%

basics

~20 s

Rule: run one macrotask to completion, then drain the entire microtask queue — including microtasks that microtasks add — before the next macrotask. Microtasks exist so continuations of already-settled operations finish before new external events. Hazard: an endlessly self-enqueuing microtask starves all I/O and timers.

open as a page

In an event-loop runtime, a callback scheduled for 50 milliseconds from now often fires later than that and essentially never earlier. What guarantee do such timers actually give, and how do loops implement them?

level: middleimportance: should knowfreq 38%

basics

~20 s

A timer promises a lower bound, not a deadline: the callback becomes eligible at the deadline and runs when the loop next gets to it. Delays come from a long-running handler, coarse OS wakeup granularity, and clamping. Loops keep deadlines in a min-heap or timer wheel and set their poll timeout from the nearest one.

open as a page

Early readiness interfaces required the program to pass its full set of file descriptors on every call and received the full set back (the select/poll style). Later ones (epoll, kqueue) register interest once and return only what changed. Explain the difference in complexity terms and why it mattered.

level: middleimportance: should knowfreq 38%

basics

~20 s

Scan-based interfaces are O(n) per call in both copying and kernel work, where n is all watched handles — even if one is ready. Registration-based ones pay O(1) amortised per registration and return O(k) ready events, so cost tracks activity, not the size of the watch set.

open as a page

How does a compiler turn a sequential-looking function containing await points into something that can pause and resume? Explain the continuation-passing / state-machine transform.

level: seniorimportance: should knowfreq 34%

basics

~20 s

It splits the body at each await into separate blocks and rewrites it as a state machine: an object holds a state number plus every local that outlives an await. Resuming calls a resume function that switches on the state and jumps to the right block. The continuation is the callback you never wrote.

open as a page

A client abandons its request while three asynchronous calls it triggered are still pending. What does cancelling a future actually mean, and what has to be true for the pending work to genuinely stop?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Cancelling usually only completes the read side with a cancelled outcome so the consumer stops waiting. The producer keeps working unless it cooperates — polling a cancellation signal, being interrupted, or having its socket closed. Real cancellation needs a channel back to the producer plus resource cleanup.

open as a page

Compare three ways for a consumer to receive a sequence of items from a producer: a blocking pull loop where the consumer asks for the next item, unrestrained push where the producer invokes a consumer callback, and demand-driven push where the consumer sends batched requests upstream. What does each cost, and when does the third win?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Blocking pull gives free rate control but parks a thread per stream. Unrestrained push parks nothing but hands rate control to the producer, so surplus items pile up. Demand-driven push combines them: batched permits bound memory, push delivery keeps threads free. It wins with many concurrent, remote, rate-mismatched sources.

open as a page

Readiness notification can be level-triggered — you are told repeatedly while the condition holds — or edge-triggered, where you are told only when it changes. Explain the difference and the classic stall bug that edge-triggered code hits.

level: seniorimportance: should knowfreq 28%

basics

~20 s

Level-triggered reports "data is available" on every wait until you drain it; edge-triggered reports only the transition from empty to non-empty. The stall bug: read once, leave bytes in the buffer, and no further edge ever arrives, so the connection hangs forever with data waiting.

open as a page

Compare the reactor pattern — the operating system tells you a handle is ready and you then perform the I/O yourself — with the proactor pattern, where you submit the I/O and are told when it completed. What changes for buffer ownership, error handling, and portability?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Reactor: demultiplex readiness, then your loop does the transfer with your buffer, allocated on demand; errors surface on your call. Proactor: submit operation plus buffer, kernel transfers, completion event carries bytes and error. Proactor pins buffers per in-flight operation and batches better; reactor is more portable.

open as a page

Your asynchronous service must call a library that only offers blocking operations, so you plan to run it on a separate worker pool and await the result. How do you decide that pool's size and queue bound, and what problem does offloading not solve?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Size it from the queueing relation: concurrency needed = arrival rate x service time. Bound the queue so overload rejects fast instead of growing latency, and give each dependency its own pool. Offloading protects the event loop; it does not make the dependency faster or add capacity.

open as a page

Explain how synchronously waiting for the result of an asynchronous operation can deadlock a service even though no mutual-exclusion locks are involved. What conditions have to hold for it to happen?

level: seniorimportance: should knowfreq 42%

basics

~20 s

If the thread that must run the async operation's continuation is the very thread now sitting blocked waiting for its result, the work can never finish. It needs a bounded execution resource — one loop thread, a single-threaded context, or an exhausted pool — plus a blocking wait held by a member of it.

open as a page

In some runtimes, calling an asynchronous function forces the caller to be asynchronous too, splitting an API into two parallel worlds. What causes that split, what problems does it create, and which designs avoid it?

level: principalimportance: should knowfreq 30%

basics

~20 s

It comes from the stackless transform: only rewritten frames can pause, so the ability to pause must propagate to every caller. That duplicates APIs, blocks incremental adoption, and tempts unsafe sync-over-async bridging. Stackful coroutines, user-mode threads, and effect systems remove the split.

open as a page

A single event loop saturates one CPU core. Compare running one independent loop per core with shared-nothing state against a single loop feeding a shared worker pool, and say when you would choose each.

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Loop-per-core shards connections and state across cores with no shared mutable data, so it scales near-linearly and avoids locks and cache-line contention — but it cannot rebalance a hot shard. One loop plus a worker pool balances load automatically but pays handoff latency, cache misses, and synchronisation on shared state.

open as a page

A team is deciding whether to build a new service's asynchronous layer around demand-driven reactive streams, around suspending sequential code (coroutine or async-await style), or around plain futures dispatched on an event loop. How would you reason about that choice?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Choose by problem shape. One result per request means futures or suspending code. Many items over time with a real rate mismatch and a bounded-memory requirement justifies streams. Weigh debuggability, the viral spread of the model through signatures, and whether the whole path is genuinely non-blocking.

open as a page

Lightweight user-space threads have made blocking-style code cheap again on several platforms. Given that, how would you decide whether a new high-connection network server should be written as an explicit event loop or as one lightweight thread per connection?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Decide on workload shape and team cost, not fashion. Lightweight threads keep sequential code and stack traces and are the default. Choose an explicit loop when you need per-core sharding, tight buffer control, predictable tail latency, or the runtime cannot hide blocking calls you depend on.

open as a page

A library maintainer wants every operation available in both a blocking form and a non-blocking form, implementing one as a thin wrapper over the other. What goes wrong with each wrapping direction, and what would you recommend instead?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Sync wrapping async ships a hidden blocking wait that can deadlock the caller's scheduler. Async wrapping sync ships a fake — it hides a thread the caller cannot size or bound. Best: implement each natively over a shared core, or pick one and let callers bridge explicitly at their boundary.

open as a page