skip to content

Threaded vs Event-Loop Models

Thread-per-request, event-loop and coroutine frameworks: which thread runs the handler, what a blocking call stalls, how workers are sized. Asked because one slow handler can sink a whole server.

on this pageshow

questions

5

In server-side web frameworks, which thread runs your handler under thread-per-request, event-loop, and coroutine execution models?

level: juniorimportance: must knowfreq 70%

answer

  1. who owns a thread while waiting
  2. one pooled worker, whole request
  3. few loop threads, short turns
  4. suspension releases the carrier thread
  5. cheaper to hold, not faster

basics

~20 s

Thread-per-request frameworks give each request a pooled worker thread for the whole exchange. Event-loop frameworks run handlers as short turns on a few shared loop threads. Coroutine frameworks suspend the handler at await points and release the carrier thread meanwhile.

solid answer

~50 s

Three shapes dominate server-side frameworks. In **thread-per-request**, the server borrows a worker from a pool, runs your handler on it start to finish, and returns it when the response is written; while the handler waits on a database or an HTTP call, that worker is idle but still held. In **event-loop** frameworks, a small set of loop threads (often one per core) run handlers as short turns between I/O readiness events; the handler is expected to return quickly and finish later through a callback or completion, so the thread is never parked on a wait. **Coroutine** frameworks keep the sequential look of the first style with the thread economics of the second: at each suspension point the handler releases its carrier thread and is resumed later, possibly on a different thread. The practical question in all three is the same: *who owns a thread while the request is waiting?*

go deeper

for a junior

Be able to name the three shapes and say, for each, whether a thread is held while the request waits on a database or another service.

for a middle

Explain why the loop thread count tracks cores rather than concurrent requests, and why a handler under a loop is a chain of short turns instead of one continuous run.

for a senior

Relate the model to observed behaviour: what saturation looks like in each, and why per-request state must travel with the request once one thread serves many requests.

for a principal

Frame the model as a stack-wide commitment with ecosystem and debuggability costs, not a free throughput win, and insist the bottleneck be measured before it drives a rewrite.

## The question behind the question Every server-side web framework has to answer one thing before it answers anything else: when a request arrives, **what unit of execution carries it from the first byte in to the last byte out**, and what happens to that unit while the request is waiting on something slow. Almost all the differences people attribute to "async frameworks" fall out of that single answer, so it is worth being able to describe the three common shapes precisely. ## Thread-per-request The server keeps a pool of worker threads. When a request is ready to be dispatched, one worker is borrowed, your handler runs **on that worker, top to bottom**, and the worker is returned once the response is written. - The code is ordinary sequential code: a call to a database or another service simply parks the worker until the answer comes back. - Everything the handler touches — the call stack, per-request values held against the running thread, the debugger's view — lines up with one request. - The waiting is the cost. A handler that spends 200 ms waiting on a downstream service holds a worker for 200 ms without using any processor time. Note the common misreading: "thread-per-request" almost never means a fresh operating-system thread is created per request. Threads are pooled and reused; creating one per request is what the pool exists to avoid. ## Event-loop A small number of loop threads — typically sized to the number of processor cores, not to the number of requests — sit in a readiness loop. Each turn of the loop picks up work that is ready (bytes arrived on a socket, a timer fired, a downstream response completed) and runs the corresponding piece of handler code **to completion of that turn**, then moves on. - A handler is therefore not one continuous run: it is a chain of short turns, stitched together by callbacks or completion values. - Nothing waits on a loop thread. The handler registers interest in a result and returns; the loop serves thousands of other sockets in the meantime. - Concurrency stops being bounded by thread count. Idle or slow connections cost memory and bookkeeping, not a thread each. ## Coroutine / suspending A coroutine framework runs handlers as resumable units of work on a small pool of carrier threads. The handler *looks* sequential — you write the downstream call as if you were waiting for it — but at each suspension point the runtime saves the handler's position, releases the carrier thread for other work, and resumes the handler when the result arrives. - You get thread economics close to the event-loop model with control flow close to the thread-per-request model. - The resumption may land on a **different** carrier thread, and that is the detail most people miss, because anything tied to the identity of the running thread stops being reliable across a suspension point. ## Side by side | | Thread-per-request | Event-loop | Coroutine | |---|---|---|---| | Runs the handler | one pooled worker, whole request | a shared loop thread, one turn at a time | a carrier thread, between suspensions | | While waiting on I/O | the worker is held, idle | nothing is held; the loop serves others | the carrier thread is released | | Handler code shape | plain sequential | callbacks / completions | sequential with suspension points | | Natural in-flight ceiling | the worker count | memory and file-handle limits | task count, not thread count | | Cost of a slow blocking call | one worker parked | one loop thread unavailable for every socket it owns | one carrier thread parked | ## What actually follows from the choice 1. **None of the models makes a single request faster.** They change how *cheaply the server can hold many requests at once*, which is a throughput and memory story, not a latency one. A handler that needs 200 ms of downstream time still takes 200 ms. 2. **The cost of blocking moves.** In the first model a blocked handler costs one worker out of many; in the other two it occupies a thread that was meant to serve a large share of the server's sockets, which is why those frameworks are strict about what may run inline. 3. **Thread identity stops being a reliable key.** Under the second and third models, one thread interleaves many requests and a single request may touch several threads, so per-request state must travel with the request rather than with the thread. 4. **The model is largely inherited, not chosen per endpoint.** A framework commits its whole surface — its client libraries, its middleware chain, its testing story — to one of these shapes, so this is a property of the framework you adopt rather than a switch you flip later. Being able to say which shape you are working in, and therefore what a wait costs, is the entry-level version of every capacity conversation that follows.

  • Does a suspending handler resume on the same thread it started on?
    Usually not guaranteed. The runtime picks a free carrier thread when the result arrives, so the code after a suspension point may run on a different thread than the code before it. Some dispatchers do pin work to a single thread; unless you have configured that, assume the thread can change.
  • If handlers never block, why run more than one event-loop thread?
    To use more than one core. A single loop saturates one core, so servers typically run several loops and spread accepted connections across them. Each connection is then served by one loop for its lifetime, which is also why a stall on one loop hurts a fixed share of clients rather than all of them.
  • Is thread-per-request simply the outdated choice?
    No. It is the simplest model to write, profile and debug, and with lightweight user-space threads a blocked handler no longer pins a costly operating-system thread, which removes much of its old penalty. It remains a poor fit when a server must hold very many mostly-idle connections.

Thread-per-request is a bank where a teller stays with one customer until they leave, even while the customer fills in a form. An event loop is one clerk who only ever touches paperwork that is ready, and never stands at a desk waiting.

saying these in an interview costs you the question

  • Says asynchronous frameworks make each individual request faster
  • Thinks thread-per-request creates a brand-new thread for every request
  • Believes an event loop gives each handler its own thread
  • Assumes a handler always resumes on the thread it started on
  • Claims blocking calls are harmless because the framework is asynchronous
open as a page

In a thread-per-request web framework, what caps the number of requests in flight, and what happens beyond that cap?

level: middleimportance: must knowfreq 62%

basics

~20 s

The handler worker count is the ceiling: one request occupies one worker for its whole exchange, waiting included. Beyond it requests sit in the server's connection or dispatch queue, so clients see queue time, then connection refusals or timeouts, while handler timings still look healthy.

open as a page

Why does per-request context keyed to the running thread break in event-loop or coroutine web frameworks?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Thread-keyed storage assumes one thread serves one request end to end. Under loop or coroutine execution a thread interleaves many requests and a request touches several threads, so context goes missing after a resumption, or belongs to somebody else.

open as a page

When one event-loop thread in a web server stalls, which server duties stop, and what limits the blast radius?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A stalled loop stops every duty it owns: reading and writing its sockets, completing handshakes, firing its own timers, often accepting connections. Sockets stay with the loop that accepted them, so the damage covers that loop's share of clients.

open as a page

How would you choose a thread-per-request, event-loop or coroutine web framework for a new service, and what constrains that choice?

level: principalimportance: should knowfreq 40%

basics

~20 s

Choose from the workload and the ecosystem, not from fashion: connection count and waiting time argue for loop or coroutine models, processor-bound work does not, and blocking client libraries can cancel the benefit entirely. Weigh it as a stack-wide, hard-to-reverse commitment.

open as a page