skip to content

An asyncio payroll CSV importer times out intermittently — how do you tell a blocked loop from a hung await?

level: seniorimportance: should knowfreq 40%

answer

  1. Two diseases, one symptom
  2. Ask whether everything stalled together
  3. Timers cannot fire while a step runs
  4. Measure how late a sleep woke up
  5. Diff two task dumps, do not read one

basics

~20 s

Check whether everything stalls together. A blocked event loop freezes every task at once and makes timers fire late in a burst; a hung await leaves the loop responsive with one task stuck on the same frame.

solid answer

~40 s

The two failures look identical from the outside — requests time out — and are diagnosed by different evidence. Run a watchdog coroutine that sleeps a fixed interval and measures the drift with `time.monotonic()`: drift is how long the loop's thread was held, and it is near zero for a hung await. Turn on asyncio debug mode so a blocking step is named in the slow-callback warning. Then take two `asyncio.all_tasks()` dumps a few seconds apart: a blocked loop shows nothing moved anywhere, including the heartbeat, while a hung await shows one task on the same frame and the others advancing. The intermittency and the fan-out are the tell: if all 17 upstream calls time out in the same instant, the loop stopped, because expired deadlines all fire at once when it resumes.

code

python · 23 lines
python
import asyncio
import time


async def watchdog(interval=0.25):
    while True:
        start = time.monotonic()
        await asyncio.sleep(interval)
        drift = time.monotonic() - start - interval
        if drift > 0.1:
            print(f"loop was blocked for about {drift:.2f}s")


async def main():
    monitor = asyncio.create_task(watchdog(), name="loop-watchdog")
    await asyncio.sleep(0)  # let the watchdog reach its first await
    time.sleep(1.0)  # the blocking call we are hunting
    await asyncio.sleep(0.5)
    monitor.cancel()


asyncio.run(main())
# loop was blocked for about 0.75s

go deeper

for a junior

Know that a coroutine only gives up the loop at an await, so any synchronous call inside one stops every other task — a timeout does not prove the other service was slow.

for a middle

Explain the mechanics that separate the two cases: timers are loop callbacks and cannot fire during a blocking step, so a blocked loop produces a burst of simultaneous timeouts a hung await never does.

for a senior

Show the method under pressure: a drift watchdog for the timestamp, debug mode for the name, two task dumps diffed for the parked task, and the fan-out shape read as evidence rather than noise.

for a principal

Own what every service ships by default — named tasks, a permanent loop-drift metric with an alert, and an on-demand task dump — so this class of incident is a five-minute read rather than an archaeology project.

## Two failures, one symptom A payroll CSV import fans out to 17 dependent services per batch. Some runs finish; some report timeouts on several upstreams at once. "Timeout" is the symptom of at least two unrelated diseases, and treating them as one is how teams end up raising every deadline and shipping the bug to production intact. **A blocked loop.** Synchronous code ran inside a coroutine and held the loop's only thread. Nothing else ran: not the other 16 calls, not the heartbeat, not the timers. **A hung await.** The loop is fine and running other work; one task is parked on something that never completes — a socket read with no deadline, a lock nobody releases, a queue nobody feeds, a future nobody resolves. ## The signature of a blocked loop The loop cannot preempt. A coroutine yields control only at an `await`; between two awaits the runtime has no mechanism to interrupt it. So a large in-memory parse, a synchronous database driver, a blocking file read, a name resolution, or a first-use import that pulls in a large dependency all stop the world. The distinctive consequence is what happens to *time*. asyncio timers are loop callbacks. While a step runs, no callback fires, so deadlines do not expire on schedule — they expire the instant the loop resumes, all of them, together. That is why the report reads "several upstreams timed out simultaneously" even though those services were healthy and their calls were nowhere near their limits. A burst of unrelated, simultaneous timeouts after a quiet period is close to a fingerprint. Intermittency fits too: the import blocks only on the batches large enough to push a synchronous step past the deadline, which is why it correlates with file size rather than with any particular service. ## The instrument: a drift watchdog The cheapest reliable detector is a coroutine that sleeps a fixed interval and measures how late it woke: ```python start = time.monotonic() await asyncio.sleep(0.25) drift = time.monotonic() - start - 0.25 ``` A healthy loop drifts by a millisecond or two. If drift is 0.75 s, the loop's thread was held for roughly that long, and you now have a timestamp to correlate against your logs. Use `time.monotonic()`, not wall-clock time, so a clock adjustment cannot fake a stall. This costs almost nothing and, unlike debug mode, is cheap enough to leave running permanently and export as a metric. asyncio debug mode is the complement: it names the offender. Its slow-callback warning reports the handle and how long it ran, which turns "the loop stalled for 750 ms" into "this coroutine stalled it". ## The instrument: two dumps, diffed Take `asyncio.all_tasks()` and print each task's name and frame, wait a few seconds, do it again. - Blocked loop: taken *after* the fact, everything looks normally suspended — which is the point. The drift watchdog, not the dump, is what identifies this case, because during the block no dump can run at all. - Hung await: the loop is alive, so the dumps differ. Most tasks change; one keeps the same name and the same frame across every dump. That task is your target, and its frame names the resource. On 3.14, `asyncio.print_call_graph()` on that task shows the whole await chain instead of the single outermost frame `get_stack()` gives you. A count is informative on its own: unfinished task counts that climb monotonically mean work is arriving faster than it is completing, which is a third disease again — saturation rather than a stall. ## Reading the fan-out With 17 dependencies, the shape of the failure discriminates for free. One stuck task and 16 completed results is a hung await against one dependency; look at that peer, then bound the wait. Seventeen simultaneous timeouts with a matching drift spike is a blocked loop; look at your own synchronous code in the batch path. Seventeen timeouts spread over the window with no drift is genuine upstream slowness, and the async runtime is not the story at all. ## Making it diagnosable next time Most of the cost in an incident like this is instrumentation you did not have. Name every task at creation. Keep the drift watchdog permanently, exported as a metric with an alert threshold. Provide a way to trigger a task dump on a live process. Log unfinished task counts periodically. None of it is expensive, and all of it is much harder to add while the service is stalling.

  • Why do timeouts on unrelated calls all expire in the same instant when the loop is blocked?
    Deadlines are implemented as loop callbacks. While a step holds the thread, no callback runs, so nothing expires on schedule; the moment the step returns, the loop processes every deadline that passed during the block. Sixteen healthy services therefore appear to fail together, which is the strongest evidence you will get that the fault is inside your process, not upstream.
  • Successive dumps show one task on the same frame for minutes while everything else advances. Where do you look?
    At the resource named in that frame, not at the loop. It is typically a read with no deadline, a synchronization primitive nobody releases, a queue nobody feeds, or a future nobody resolves. Confirm from the peer's side — is the connection still open, did the reply ever leave — then bound the wait so the next occurrence surfaces as a fast, attributable failure instead of a permanent park.
  • Why is a drift watchdog worth keeping in production when asyncio debug mode already reports slow callbacks?
    They answer different questions at different costs. Debug mode names the culprit but adds per-task and per-iteration overhead, so it is rarely on everywhere. A watchdog coroutine costs one wake-up per interval, runs permanently, and turns loop stalls into a numeric metric you can alert on and correlate with request latency — it tells you *that* the loop stalled, and when, so you know to turn the expensive tooling on.

A blocked loop is a single cashier who stops to do long division: every queue freezes and everyone's patience expires the moment she looks up. A hung await is one customer waiting for a price check while the rest of the line moves.

saying these in an interview costs you the question

  • Blames the upstream service because a timeout fired
  • Assumes async code cannot block the loop
  • Adds more concurrent tasks to fix a blocked loop
  • Thinks the loop preempts a long CPU-bound step
  • Raises every timeout instead of finding the stall
  • Draws a conclusion from a single task dump

context