A service built on a single event loop has one request handler that spends 400 milliseconds doing CPU work. What happens to the rest of the system, and how would you detect it in production?
answer
- run-to-completion means no preemption: everyone waits
- tails explode; mean barely moves
- late timers, missed heartbeats, accept backlog overflow
- timeouts expire during the stall, then retries pile on
- loop lag = scheduled interval vs actual; watch p99/max
basics
~20 sEverything queued behind it waits: all in-flight requests take up to 400 ms longer, timers fire late, heartbeats and health checks miss, and client timeouts fire in a burst that triggers retries. Detect it by measuring loop lag — schedule a known-interval timer and record its overshoot — plus per-handler duration histograms.
solid answer
~60 s**Effects.** Run-to-completion means no preemption, so the 400 ms is added to the wait of every ready event. Symptoms cascade: - Latency for *all* concurrent requests spikes, not just the offending one — classic head-of-line blocking; p99 degrades far more than average. - Timers, heartbeats, and health probes fire late. Long enough, and the instance is marked unhealthy, or drops out of a cluster/leadership. - New connections queue in the kernel accept backlog; if it overflows, connections are refused or dropped. - Client timeouts expire *during* the stall, so when the loop resumes it does work whose answers nobody wants, then serves a wave of retries. Under sustained load that feedback loop is congestion collapse. **Detection.** The canonical metric is **event-loop lag**: schedule a timer for a fixed interval, measure actual minus expected, and track its high percentiles — a timer can only be late if the loop is busy. Add per-handler duration histograms to find the culprit, a watchdog on a separate thread that samples the loop thread's stack when lag exceeds a threshold, and correlate spikes with payload sizes and garbage-collection pauses.
code
text · 7 linesP = 100ms
expected = monotonicNow() + P
repeat every P:
actual = monotonicNow()
lag = actual - expected # >= 0; large means loop was busy
record(lag) # histogram, alert on p99 / max
expected = actual + Pgo deeper
Say that nothing can interrupt a running handler, so every other pending event is delayed by the full duration, and timers fire late.
Add the queueing view — tails degrade far more than the mean — and describe measuring loop lag with a known-interval timer.
Walk the cascade: late heartbeats, health-check removal, accept-backlog overflow, timeout expiry during the stall and retry amplification; then detection via lag percentiles, per-handler histograms, and an off-loop watchdog that samples the loop's stack.
Frame it as bounding per-event service time as a system invariant: cap input sizes, chunk with yields, propagate deadlines so dead work is dropped, apply retry budgets, and shed load rather than growing queues.
## Why one slow handler hurts everything An event loop dispatches handlers **run to completion**: nothing preempts the running handler. A 400 ms CPU burst therefore contributes 400 ms of queueing delay to every event that is ready or becomes ready during it. With threads, the kernel would timeslice and spread the pain; here the pain is concentrated in a single ordered queue. The latency shape is distinctive: mean response time rises modestly (the stall is intermittent), but tail percentiles explode, because during each stall a whole cohort of requests takes the full residual. If arrivals continue during the stall, the queue is longer when the loop resumes, so the excess persists past the stall itself — a queueing effect, not just a one-off delay. ## The cascade 1. **Head-of-line blocking.** Every ready socket, timer, and posted task waits. 2. **Late timers.** Anything the loop schedules — retry timers, idle timeouts, metric flushes, heartbeats to a coordinator or leader-election service — fires late. A 400 ms stall may be survivable; a few seconds can cost a lease or leadership. 3. **Missed health checks.** Probes are just more queued work. Enough missed probes and a load balancer removes the instance, shifting its traffic to peers, whose load rises — and if the stall is caused by traffic characteristics (large payloads), peers now stall too. That is a correlated failure, not an isolated one. 4. **Accept backlog growth.** New connections accumulate in the kernel queue while the loop is not accepting. Overflow means resets or silently dropped connection attempts, which clients experience as connection failures rather than slow responses. 5. **Timeout and retry storm.** Callers whose deadline elapsed during the stall have gone. Work completed for them is wasted, and their retries arrive as *new* load on an already-behind loop. Without deadline propagation (dropping requests whose deadline already passed) and retry budgets, this amplifies into collapse. ## Detecting it **Event-loop lag** is the primary signal. Schedule a repeating timer at interval P, and on each firing record `actual_interval - P`. Since a timer can only be late when the loop is occupied, this measures exactly the property you care about. Track p50/p99/max, not the mean; a single 400 ms stall per minute is invisible in a mean and obvious at max. Alert on the high percentiles. **Per-handler duration histograms.** Instrument dispatch: record start and end of each handler by type or route. This turns "the loop is stalling" into "the report-export handler takes 400 ms at p99". **Off-loop watchdog.** A separate thread updates a heartbeat counter from the loop and, when the loop fails to tick within a threshold, samples the loop thread's stack. Since the loop thread cannot report on itself while blocked, an outside observer is the only way to capture what it was doing. **Correlation.** Plot lag against request size, batch size, and collection counts. CPU stalls usually correlate with input size (a payload 100x larger, an unbounded result set), which points at the fix. **Distinguish causes.** Loop lag is agnostic about *why* the loop was unavailable. It could be application CPU work, a synchronous file or name-resolution call, a large serialisation, or a garbage-collection pause. Separate them by comparing lag with runtime pause metrics; if lag spikes align with collection pauses, the fix is allocation pressure, not handler structure. ## Fixing the shape Briefly, because interviews ask: bound the work per handler (chunk large iterations and yield between chunks so the loop can service I/O), cap input sizes so no request can generate an unbounded burst, move genuinely heavy CPU work off the loop, and shed load rather than queue it without limit. Add deadline awareness so work whose caller has already given up is dropped rather than executed. And treat the queue itself as a resource: unbounded queues convert an overload into a latency and memory problem instead of a clean rejection. ## Framing for an interview Say it as: *the event loop's operating characteristic is per-event service time; anything that makes one event slow is a global latency event, so the discipline is bounding service time and measuring lag directly rather than inferring it from response times.*
- Why is event-loop lag a better health signal than request latency alone?Request latency mixes together the time spent in downstream systems, the network, and the loop itself, so a spike is ambiguous. Loop lag isolates one thing: how long the loop was unavailable to dispatch, because a timer can only be late if the loop was busy. It also catches stalls that occur when no requests are in flight, such as a background job hogging the loop.
- Why can a stall in one instance make the whole cluster worse rather than better?Missed health probes cause the load balancer to remove the instance, redistributing its traffic to peers. If the stall was caused by a traffic characteristic — oversized payloads, an expensive query shape — the peers now receive that same traffic and stall too. Meanwhile clients whose deadlines expired retry, adding load on top of a system already behind.
- How do you keep a large in-memory computation from stalling the loop without moving it off the loop?Chunk it: process a bounded number of items, then re-schedule the remainder as a new task so the loop can poll for I/O and fire timers between chunks. Total wall-clock time increases slightly, but per-event service time stays bounded and the tail latency of everything else is preserved.
A single toll booth where the attendant answers a long phone call. Nobody in the line is doing anything wrong; the queue simply grows, and the drivers who leave and come back later make it worse.
saying these in an interview costs you the question
- Claiming only the slow request is affected
- Saying the runtime will preempt or timeslice a long handler
- Diagnosing purely from average latency, where an intermittent stall is invisible
- Assuming loop lag always means application CPU, ignoring garbage-collection pauses and synchronous file or name-resolution calls
- Responding to overload by enlarging queues, which increases latency and memory instead of shedding load