skip to content

Lightweight user-space threads have made blocking-style code cheap again on several platforms. Given that, how would you decide whether a new high-connection network server should be written as an explicit event loop or as one lightweight thread per connection?

level: principalimportance: nice to knowfreq 30%

answer

  1. who writes the state machine — runtime or you
  2. sequential code + stack traces vs sharding + buffer control
  3. memory floor at a million connections
  4. both die on unsuspendable work
  5. measure scheduler/loop latency

basics

~20 s

Decide on workload shape and team cost, not fashion. Lightweight threads keep sequential code and stack traces and are the default. Choose an explicit loop when you need per-core sharding, tight buffer control, predictable tail latency, or the runtime cannot hide blocking calls you depend on.

solid answer

~60 s

Both designs multiplex non-blocking handles underneath; the question is who writes the state machine — the runtime or you. **Default to lightweight threads per connection.** You keep sequential control flow, real stack traces, ordinary error propagation and structured cancellation, which is most of the maintenance cost of a server. Per-connection cost falls to a small growable stack. **Reach for an explicit event loop when** you need per-core sharded state with no synchronisation and no work-stealing migration; when tail latency must be predictable and you want explicit control over batching, buffer reuse and fairness; when memory per connection must be measured in hundreds of bytes; or when the protocol is naturally a state machine anyway (proxies, brokers). **Check the escape hatches either way.** The failure mode of both is the same: a blocking call the runtime cannot see — a native library, a file read, a long CPU burst — pins a carrier thread or stalls a loop. Decide up front where such work is offloaded and how that pool is bounded and observed.

go deeper

for a junior

Recognise that both approaches use non-blocking I/O underneath and that lightweight threads let you keep ordinary sequential code.

for a middle

Compare readability and diagnostics against memory per connection, and know that blocking calls must be offloaded in both models.

for a senior

Drive the decision from connection scale, latency contract and protocol shape, and specify the offload pool and the scheduler-latency metric.

for a principal

Own it as an architecture and staffing decision: lifetime maintenance cost, runtime maturity, hybrid layering, and what evidence would reverse the choice.

## Reframing the question For twenty years the answer was forced: kernel threads were too expensive to spend one per connection, so anything connection-heavy became an event loop, and developers paid for that in callback-shaped code. Lightweight user-space threads (green threads, fibers, goroutine-style schedulers, virtual threads) removed the forcing function. They give each connection a small stack that grows on demand and multiplex many of them onto a few carrier threads, suspending in user space when I/O would block. Underneath, the runtime registers non-blocking handles with a readiness or completion mechanism and runs a scheduler loop — the same machinery as an explicit event loop, hidden. So the decision is no longer "can we afford a thread per connection" but "who should own the state machine, the runtime or us?" ## What lightweight threads buy - **Sequential code.** Read, parse, call downstream, write, in one function, with ordinary control flow, loops, and try/finally-style cleanup. - **Diagnostics.** A stack trace shows the whole logical operation. In callback-style code the stack shows the loop and one handler, and correlating it back to a request needs explicit context propagation. - **Composability with cancellation and timeouts.** Scoped, tree-shaped cancellation is far easier to express over a stack than over a chain of callbacks. - **Team throughput.** Most engineers write correct blocking-style code faster than correct state-machine code, and reviewers catch more. ## What an explicit loop buys - **Sharding and locality.** One loop pinned per core, owning its connections and its data, needs no locks and no cross-core migration. Work-stealing schedulers can move a task between cores mid-request, which costs cache locality — usually invisible, occasionally the whole story. - **Buffer and allocation control.** Loop-local scratch buffers, arena reuse, vectored I/O and explicit batching are natural when you own the loop; a runtime abstraction may allocate per operation. - **Tail latency predictability.** You control fairness — how many bytes or messages a connection may consume per visit — instead of trusting a general-purpose scheduler. - **Memory floor.** Per-connection state can be a few hundred bytes; even a small growable stack is typically kilobytes. At a million connections that difference is the design. - **Protocols that are state machines anyway.** Proxies, message brokers, and multiplexed protocols with per-stream flow control map onto a loop naturally; forcing them into blocking style can be more code, not less. ## The shared failure mode Both models die the same way: work that the scheduler cannot suspend. A native library call, a regular-file read on a readiness-based substrate, a page fault on memory-mapped data, or a 40 ms CPU-bound computation will pin a carrier thread or stall a loop iteration. With a handful of carriers or loops, a few such calls in flight can idle the whole server. Whichever model you pick, you must: 1. Inventory blocking and CPU-heavy calls, including inside dependencies. 2. Route them to a bounded, separately sized pool with its own queue limit, so exhaustion degrades one dependency rather than the process. 3. Instrument loop or scheduler latency — the delay between a task becoming runnable and running — and alarm on it. That single metric detects hidden blocking in either architecture. ## Deciding, in order 1. **Connection scale and idleness.** Millions of mostly-idle connections push toward an explicit loop for memory reasons; tens of thousands do not. 2. **Latency contract.** A hard tail-latency budget favours explicit control of fairness and batching. 3. **Protocol shape.** Naturally state-machine protocols favour a loop; request/response business logic with several downstream calls favours sequential code. 4. **Runtime maturity.** Does the platform's lightweight-thread scheduler actually make your I/O layer, your drivers, and your locks suspend rather than pin? An immature stack turns "cheap threads" into carrier starvation. 5. **Team and lifespan.** A long-lived service maintained by a rotating team pays the callback tax every year. Weigh that as a real cost, not a soft one. 6. **Measure the workload you have.** Prototype the hot path both ways if the decision is expensive to reverse; the difference is often smaller than the architectural discussion suggests, and the operational simplicity of sequential code frequently wins. ## The hybrid answer Many strong systems are both: an explicit, sharded loop owning the socket layer and protocol framing, handing complete messages to lightweight threads that run application logic sequentially, with a separate bounded pool for genuinely blocking dependencies. That keeps the memory and fairness benefits where connection count lives and keeps the readability benefits where business logic lives. If forced to give one answer: start with lightweight threads plus a strict offload policy, and buy an explicit loop only where measurement says you need it.

  • What single metric tells you a lightweight-thread runtime is being starved by hidden blocking?
    Scheduler (or event-loop) latency: the delay between a task becoming runnable and actually running. It rises the moment carriers are pinned by unsuspendable calls, and it rises across all endpoints at once, which distinguishes it from a slow downstream dependency. Alarm on its high percentiles and record which task was on the carrier when the delay occurred.
  • Where does an explicit event loop still clearly beat one lightweight thread per connection?
    At extreme connection counts where even a small growable stack times millions of connections dominates memory, and in systems needing per-core sharded ownership so no lock or cross-core migration is involved. It also wins where you need explicit fairness — capping bytes served per connection per visit — and tight control over buffer reuse and batching for predictable tail latency.

saying these in an interview costs you the question

  • Treating lightweight threads as removing the need to think about blocking calls at all.
  • Choosing an event loop for "performance" without a connection-count, memory or latency number behind it.
  • Assuming lightweight threads eliminate races — shared mutable state is still shared.
  • Forgetting that regular-file I/O and native calls can pin carrier threads.
  • Presenting it as a binary choice when sharded loop plus sequential handlers is a common hybrid.

context