skip to content

Why does a blocking call inside one asyncio coroutine freeze the whole event loop?

level: middleimportance: must knowfreq 72%

answer

  1. Cooperative, not preemptive
  2. One thread turns the crank
  3. Control returns only at a suspending await
  4. Delay is additive across everything in flight
  5. async def alone grants no concurrency

basics

~20 s

An asyncio event loop drives every coroutine on one thread and only regains control at an await that suspends. A synchronous call — CPU work, time.sleep, a blocking read — holds that thread, so nothing else runs.

solid answer

~50 s

Scheduling in asyncio is cooperative, not preemptive. One loop iteration drains ready callbacks, polls the selector for ready file descriptors, and fires due timers; a Task only yields the thread at an `await` that actually suspends. A synchronous call inside a coroutine sits in the middle of that iteration, so for its whole duration the selector is not polled, timers are not checked, and every other task is stalled — not slowed, stalled. The damage is proportional and global: a 300 ms resize inside one request adds 300 ms to *every* request in flight, which is why tail latency collapses long before mean latency looks wrong, and why timeouts fire late rather than on time. The fixes are to keep synchronous stretches short, to chunk long CPU work with explicit `await asyncio.sleep(0)` yields, or to move the blocking call off the loop thread entirely.

code

python · 17 lines
python
import asyncio
import time

async def make_thumbnail():
    time.sleep(0.3)                 # blocking: holds the loop thread
    return "thumbnail"

async def heartbeat(start):
    for _ in range(3):
        await asyncio.sleep(0.05)   # asks for 50 ms, waits for the stall
        print(f"heartbeat at {time.perf_counter() - start:.2f}s")

async def main():
    start = time.perf_counter()
    await asyncio.gather(make_thumbnail(), heartbeat(start))

asyncio.run(main())

go deeper

for a junior

Be ready to name the rule: one loop, one thread, and code only yields at an await. Know that time.sleep() inside a coroutine stops everything and that asyncio.sleep() is its non-blocking counterpart.

for a middle

Explain the loop iteration — ready callbacks, selector poll, due timers — and show why a synchronous stretch inside a coroutine delays all three. Be able to describe await asyncio.sleep(0) as an explicit yield and its limits.

for a senior

Demonstrate the production reading: additive delay wrecks tail latency before the mean moves, timeouts fire late, health checks go quiet in bursts. Talk about budgeting the code between two awaits the way you budget time inside a lock.

for a principal

Own the boundary rule for a codebase: which clients are permitted on the loop thread, what maximum synchronous stretch is acceptable, and how CPU-heavy paths are kept off loop-bound services by design rather than by review.

**One thread, one loop, no preemption.** The reason is structural rather than subtle. An asyncio event loop runs on exactly one thread and repeats a fixed iteration: run the callbacks that are already ready, ask the selector which registered file descriptors are readable or writable (blocking at most until the next timer deadline), queue the callbacks those wake, and run the timer callbacks whose deadline has passed against the loop's monotonic clock. Coroutines participate through Tasks: a Task calls into the coroutine, the coroutine runs ordinary Python until it hits an `await` that suspends, and control returns to the loop with a note about what should resume it. Nothing in that picture can interrupt running Python code. There is no scheduler thread, no timer signal, no quantum. The loop regains control only when the currently running coroutine hands it back. So a synchronous call — `time.sleep(2)`, a blocking socket `recv`, a large image resize, a synchronous database driver, a `json.loads` on a huge payload — does not "block that coroutine". It blocks the crank. The selector is not polled; sockets that became readable stay unread; timers whose deadline passed are not noticed; `asyncio.sleep(0.05)` in some other task returns after 2.05 seconds. A worked case makes the shape obvious. An image-thumbnail worker consumes jobs and, inside `async def handle(job)`, calls a synchronous resize that takes about 300 ms. Under a light load nothing looks wrong. Under a hundred concurrent jobs, the service's 92nd-percentile budget blows out first, because delay is additive across everything in flight: while job A resizes, jobs B through Z are frozen, and each of them will in turn freeze the rest. Mean latency lies about this, because it averages the waiting into the many jobs that happened to sit near the front of a queue; the tail tells the truth. Two secondary symptoms confirm the diagnosis: timeouts that fire well past their deadline (timers are checked once per iteration, so a 300 ms stall shifts every deadline by up to 300 ms), and heartbeats or health checks that go quiet in bursts rather than degrading smoothly. ```python import asyncio, time async def resize(): time.sleep(0.3) # synchronous: the loop is stuck here async def heartbeat(): for _ in range(3): await asyncio.sleep(0.05) print("beat") # all three arrive after the 300 ms stall async def main(): await asyncio.gather(resize(), heartbeat()) asyncio.run(main()) ``` **`async def` grants nothing by itself.** Marking a function `async def` only makes it a coroutine function; it does not make its body concurrent, and `await` does not make a synchronous callee non-blocking. `await some_sync_wrapper()` still runs that code to completion on the loop thread. The only thing that helps is the code actually suspending. **Where suspension really happens.** Only awaits that are not already resolved give the loop a turn. Awaiting a future that is already done, or a coroutine that returns without suspending, keeps the thread the whole way. For a long CPU-bound stretch you can insert `await asyncio.sleep(0)` between chunks — a documented yield that reschedules the current task at the back of the ready queue and lets one loop iteration happen. That is a mitigation, not a cure: it converts one long stall into many short ones, and the total CPU still competes with everything else on that thread. Real CPU-bound work belongs off the loop thread — handed to a worker thread or a separate process — and pure waiting belongs in an async-aware client rather than a blocking one. **Free threading does not change this.** The free-threaded build, officially supported in 3.14 (PEP 779), removes the interpreter-wide lock so *threads* can run Python in parallel. A single event loop still drives its own tasks sequentially on its own thread; a coroutine has no thread of its own to be parallel with. Getting parallelism means several loops on several threads, or work moved off the loop, not a different build flag. The practical rule to state in an interview: **inside a coroutine, every line between two awaits is latency you are adding to every other task on that loop.** Budget those stretches the way you would budget time inside a lock.

  • Where exactly can an asyncio event loop switch from one task to another?
    Only at an `await` that actually suspends — one whose awaitable is not already resolved. Awaiting a completed future, or a coroutine that returns without suspending, never returns control to the loop. `await asyncio.sleep(0)` is the explicit way to yield a turn without waiting for anything.
  • Does the free-threaded build of 3.14 let coroutines on one loop run in parallel?
    No. Free threading (PEP 779, officially supported in 3.14) removes the interpreter-wide lock so threads can execute Python simultaneously, but a single event loop still steps its tasks sequentially on one thread. Parallelism needs several loops on several threads, or the work moved off the loop.
  • Why do asyncio timeouts fire late rather than not at all when the loop is blocked?
    Timers live on a heap keyed by the loop's monotonic clock and are checked once per iteration. A blocking stretch delays the next iteration, so every deadline that passed during it is noticed only afterwards — the callbacks still run, just shifted by roughly the length of the stall.

It is a single checkout lane where customers agree to step aside while they think. One customer who refuses to step aside stops the whole line, however short their errand looked.

saying these in an interview costs you the question

  • Claiming asyncio runs coroutines in parallel threads
  • Thinking await makes a synchronous function non-blocking
  • Believing time.sleep in a coroutine pauses only that coroutine
  • Expecting more tasks on the loop to restore throughput
  • Assuming CPU-bound work is fine as long as it is in an async def
  • Reading mean latency instead of the tail to spot a stalled loop

context