skip to content

When one event-loop thread in a web server stalls, which server duties stop, and what limits the blast radius?

level: seniorimportance: should knowfreq 50%

answer

  1. a loop owns sockets, not a request
  2. reads, writes, handshakes, timers stop
  3. its own timeouts cannot fire
  4. connection pinned to one loop
  5. victims share a loop, not a route

basics

~20 s

A stalled loop stops every duty it owns: reading and writing its sockets, completing handshakes, firing its own timers, often accepting connections. Sockets stay with the loop that accepted them, so the damage covers that loop's share of clients.

solid answer

~40 s

A loop thread is not dedicated to one request; it owns a set of connections and every duty attached to them. While it is stuck, those sockets are not read or written, queued completions are not delivered, and its timers do not fire, so the deadlines that should have cut the work short never run. If the accept duty sits on that loop, new connections stop being picked up too. Because a connection is normally bound to one loop for its lifetime, the victims are that loop's connections rather than all traffic: a **partial** outage, where a share of users hangs while health checks answered by another loop still report the service healthy. Frameworks bound the damage with several loops, handler execution kept off the I/O loops, and loop-lag monitoring.

go deeper

for a junior

Remember that a loop thread serves many connections, so occupying it affects other clients rather than only the request you are handling.

for a middle

List the duties that stop — socket reads and writes, handshakes, completions, timers, sometimes accepting — and explain why connections stay with the loop that accepted them.

for a senior

Recognise the partial-outage signature in an incident: green health checks, no server errors, victims spread across unrelated endpoints, and timers that never fired.

for a principal

Set the structural defence: how many loops, whether handlers run off the I/O threads, where deadlines are enforced, and what per-loop signal must exist before this failure can be alerted on.

## A loop thread owns duties, not a request In a thread-per-request server the unit of damage is obvious: a stuck handler costs one worker. In an event-loop server the thread that runs your handler is also the thread that runs the server's own machinery for a **set of connections**. While it is inside your code, none of those duties happen. Concretely, the stalled loop is not: - reading request bytes off any socket it owns, so requests already on the wire are not even parsed; - writing response bytes, so responses already produced are not delivered and buffers back up; - completing transport handshakes for connections assigned to it; - delivering completions for downstream calls whose results have already arrived; - running its timer wheel, so idle timeouts, request deadlines and keep-alive expiry do not fire; - accepting new connections, in designs where the accept duty lives on a loop rather than on a dedicated thread. That last pair deserves emphasis. The timeouts you configured are themselves scheduled work on the loop. A stalled loop cannot enforce its own deadlines, so the safety net you were relying on is disabled by the same event it was meant to catch. ## Why the outage is partial, and why that is confusing Servers typically run several loops, and a connection is assigned to one loop when it is accepted and stays there for its lifetime. Two consequences follow: 1. **The affected population is the loop's connections.** With four loops, a single stalled loop hangs roughly a quarter of connected clients, not everyone. 2. **Which requests those are is essentially arbitrary** — it depends on which connection landed on which loop, not on which endpoint was called. A user whose connection is pinned to the stalled loop hangs on *every* endpoint, including trivial ones. That produces a signature that is easy to misread: | Observation | Why it happens | |---|---| | A fraction of users report total unresponsiveness | Their connections are pinned to the stalled loop | | Health checks stay green | The probe's connection landed on a healthy loop | | Error rates look normal | Hung requests produce no server-side error; the client times out | | The slow endpoint is not the one users complain about | Victims share a loop with the offender, not a route | Because the victims are selected by connection rather than by feature, the usual instinct — look at the endpoint people complain about — points away from the cause. ## What actually limits the blast radius - **Several loops with connections spread across them.** This is the structural bound: the share of traffic one stall can affect is roughly one over the loop count. It caps the damage; it does not prevent it. - **Separating I/O loops from handler execution.** Many frameworks run transport duties on one small set of threads and application handlers on another, so application code that misbehaves cannot stop the server from reading sockets and firing timers. This converts a hang into a queue, which is a far better failure. - **Connection-level limits.** Bounded per-connection buffers and request limits stop one stalled loop from turning into unbounded memory growth while its writes back up. - **Timeouts enforced outside the affected loop** — at a proxy or load balancer — because in-process deadlines scheduled on the stalled thread cannot fire. - **Loop-lag detection.** A watchdog that schedules a tick on every loop and measures how late it arrives turns an invisible stall into a metric. Alerting on the maximum lag across loops, not the mean, is what makes a single-loop stall visible. ## Investigating one 1. **Check per-loop lag or per-thread state first.** One thread busy while the others idle, with modest overall processor usage, is the fingerprint. 2. **Take stacks of the loop threads,** not of the pool at large. A loop thread parked inside application code — a synchronous call, a lock, a long computation — names the offender directly. 3. **Correlate victims by connection, not by route.** If the affected requests span unrelated endpoints but cluster in time and by client connection, you are looking at a shared-loop problem rather than a slow feature. 4. **Confirm timers stopped.** Deadlines that should have fired and did not is strong evidence that the loop itself, and not a downstream dependency, was the stuck component. The framework-level lesson is that a loop thread is shared infrastructure. Anything that occupies it for long is not slowing a request; it is taking a slice of the server offline, including the parts that were supposed to notice.

  • Why can health checks stay green while a share of users hangs?
    The probe opens its own connection, which is likely assigned to a healthy loop, and it answers normally. Health checking one connection proves one loop is alive. Per-loop lag metrics, or a check that exercises several connections, is what makes the partial failure visible.
  • Why do in-process request timeouts fail to rescue a stalled loop?
    Those deadlines are timer entries on the loop itself. A thread stuck in application code is not running its timer wheel, so nothing fires until it returns. Enforcement has to come from outside the loop — a proxy, a load balancer or a separate watchdog thread.
  • What does running handlers off the I/O loops buy you?
    It changes the failure from a hang into a queue. Transport duties keep running, sockets are still read and written and timers still fire, so slow application code shows up as growing queue depth and enforced deadlines rather than as connections that silently stop responding.

saying these in an interview costs you the question

  • Thinks a stalled loop only delays that one request
  • Expects in-process request timeouts to fire during the stall
  • Trusts a green health check to prove no loop is stuck
  • Looks only at the complained-about endpoint for the cause
  • Assumes more loop threads eliminate rather than bound the damage
  • Reads low overall processor usage as evidence nothing is stuck