skip to content

A service built on a small pool of event-loop threads starts showing high p99 latency on many unrelated endpoints whenever one particular feature is used. How does a single blocking or CPU-heavy call inside a handler produce that pattern, and how would you locate it?

level: middleimportance: must knowfreq 55%

answer

  1. cooperative: nothing preempts a handler
  2. p99 blows up on unrelated endpoints
  3. loop lag = no-op task delay
  4. profile loop threads only
  5. offload + bounded queue + bulkhead

basics

~20 s

Loop threads are shared by all connections, so a handler that blocks or computes long stops every queued event behind it — latency appears everywhere, not on the guilty endpoint. Find it by measuring event-loop lag, instrumenting handler durations, and profiling loop threads.

solid answer

~60 s

Event loops are cooperative: a handler holds its thread until it returns or yields at an await point. Nothing preempts it back into the loop. So a handler that makes a synchronous call, reads a file, resolves a name, waits on a lock, or spends 50 ms compressing data holds one of very few loop threads, and every event already queued for that thread waits. The signature is **cross-cutting latency** — endpoints with no code in common get slower together, correlated with traffic to the offending feature — plus low CPU utilisation if the stall is I/O waiting, or one saturated core if it is computation. To locate it: 1. Instrument **event-loop lag**: enqueue a trivial task periodically and record the delay before it runs. Spikes prove the loop, not a downstream service. 2. Record **per-handler execution time** and alert on any handler exceeding a few milliseconds. 3. Profile or sample stacks **of the loop threads only** during a spike; the offending frame appears directly. 4. Enable a blocking-call detector if the runtime offers one. The fix is to move the work to a bounded offload pool and await its result.

code

text · 6 lines
text
every 100 ms, on each loop:
    t0 = now()
    schedule_on_loop(() -> record(now() - t0 - 100ms))

healthy: single-digit milliseconds
blocked loop: lag ~= duration of the offending handler

go deeper

for a junior

Explain that loop threads are shared and a handler that waits holds one, so other requests queue behind it.

for a middle

Give the symptom signature (cross-cutting p99, low CPU), list common hidden blockers, and name loop-lag measurement as the first diagnostic.

for a senior

Add the diagnostic ladder — loop lag, handler histograms, loop-thread profiling, blocking detectors — and the bounded, bulkheaded offload pool as the fix.

for a principal

Treat it as an isolation-boundary defect: define which work may run on shared threads, enforce it with guardrails and CI checks, and specify the degradation behaviour when the offload path saturates.

## Why the symptom looks unrelated to the cause An event loop is a cooperative scheduler. It pulls a ready event, calls a handler, and regains control only when that handler returns or suspends at a wait point. There is no preemption *into* the loop: the operating system can preempt the thread, but that does not hand the loop back — the thread resumes inside the same handler afterwards. Every connection assigned to that loop is therefore queued behind whatever the handler is doing. When a handler stalls for 100 ms, everything else on that loop is 100 ms late, regardless of which endpoint it belongs to. With, say, four loops, the observable pattern is: - p50 stays acceptable (most requests miss the stall) while **p99 explodes** — the classic tail-latency signature of a shared, non-preemptible resource. - **Unrelated endpoints degrade together**, in proportion to how often the offending code runs. - Latency correlates with the *feature's traffic*, not with any downstream dependency's own latency. - CPU is **low** if the stall is I/O waiting (threads parked, no work done) or **one-core-saturated** if it is computation. - Adding replicas helps only proportionally, and adding loop threads beyond core count does not help at all. That mismatch between the guilty endpoint and the suffering endpoints is what makes this hard: on-call engineers chase the slow endpoints and find nothing wrong with them. ## What is actually blocking The list is longer than people expect, and the entries most often missed are the boring ones: - Synchronous HTTP or database clients pulled in by a library you did not choose deliberately. - Regular-file access: configuration reload, certificate load, template read, log write to a full disk, or a page fault on memory-mapped data. - **Name resolution.** Many "async" clients still resolve host names synchronously. - Lock acquisition, especially a lock also held by non-loop threads. - Class or module loading on first use, which reads from disk. - CPU-heavy work: compression, cryptography, large serialisation, image processing, pathological regular-expression backtracking. - Constructing a client or connection pool lazily inside a handler. - Logging synchronously to a slow sink. ## Locating it, in order of cheapness **1. Event-loop lag.** Schedule a no-op task onto each loop every N milliseconds and measure how late it runs. This single metric separates "our loop is stuck" from "the downstream is slow", because a slow dependency does not delay the loop when the client is genuinely asynchronous. Export percentiles per loop and alert on the maximum. It is a few lines of code and is the highest-value observability you can add to an async service. **2. Handler duration histograms.** Time each handler invocation from dispatch to return-or-suspend. Any bucket above a few milliseconds is suspect. Tagging by handler name gives you the culprit directly. **3. Targeted profiling.** Sample stacks of the loop threads only, during a spike. Because loop threads should be nearly always inside the wait call, any other frame that shows up repeatedly *is* the bug. Continuous profilers make this trivial after the fact; a few thread dumps during a spike work in a pinch. **4. Runtime blocking detectors.** Many async runtimes ship an option that flags known blocking calls when they execute on a loop thread, either by instrumentation or by a watchdog thread that notices a loop has not returned for longer than a threshold. Enable it in staging permanently and under a flag in production. **5. Static and dependency review.** Grep for synchronous client types, file APIs and sleep calls in handler code paths, and audit transitive libraries. This catches the ones that only trigger on rare branches — cache miss, certificate refresh, first request after deploy. ## Fixing it - **Offload** unavoidable blocking work to a bounded, separately sized worker pool and wait for its result asynchronously. Keep one pool per dependency class so a single slow dependency cannot consume the capacity of others — a bulkhead. - **Chunk or offload CPU work**. If a computation cannot be moved, split it so the handler yields between slices; long unyielding computation is starvation even without I/O. - **Bound the offload queue** so saturation becomes fast rejection rather than unbounded memory growth and creeping latency. - **Do the slow setup eagerly** at startup — connection pools, certificates, compiled patterns, class loading — so no handler ever pays it. - **Add a guardrail**: a test or lint rule that fails the build if a known blocking type is referenced from handler packages, plus the loop-lag alert so a regression is caught in minutes rather than at the next incident. ## Saying it in an interview Name the mechanism (cooperative scheduling, no preemption back into the loop, shared threads), name the signature (cross-cutting p99 with low CPU), name the measurement (loop lag first, handler timings, loop-thread profiling), then the fix (bounded offload pool with bulkheads, eager setup, guardrails).

  • Why does increasing the number of event-loop threads rarely fix this?
    Loops are normally sized to core count because their job is CPU-bound dispatch; adding more does not create more cores, and each extra loop only widens the pool of victims before it too is blocked. If the blocking call is frequent enough to stall four loops, it will stall eight at twice the traffic. It also costs cache locality and, in sharded designs, breaks the assumption that a connection has a single owning thread.
  • How do you distinguish a stalled event loop from a genuinely slow downstream dependency?
    Event-loop lag is the discriminator: a slow dependency with a truly asynchronous client leaves the loop idle and lag near zero, while a blocked loop shows lag equal to the stall. The second clue is scope — a slow dependency degrades the endpoints that call it, whereas a blocked loop degrades everything that shares the thread. Correlating the spike with the offending feature's request rate rather than the dependency's own latency confirms it.

One cashier serves the whole shop by taking one item from each shopper in turn. If a shopper starts a phone call at the till, everyone's shopping stops — and the complaints come from people who never spoke to that shopper.

saying these in an interview costs you the question

  • Blaming the slow-looking endpoints instead of the code holding the loop.
  • Adding more event-loop threads as the fix.
  • Assuming an async-looking client library never blocks — name resolution and connection setup often do.
  • Treating CPU-heavy handlers as safe because they are not waiting on I/O.
  • Offloading to an unbounded queue, which converts a latency spike into memory growth and later timeouts.

context