A chat-transcript archiver on async workers spikes latency on every open connection whenever one upload hits a slow synchronous decode. Why?
answer
- The slowdown is correlated, not isolated
- Nothing preempts a running coroutine
- One thread per worker holds the loop
- Debug mode reports slow callbacks
- Waiting goes to a thread, computing to a process
basics
~20 sEach async worker runs its coroutines on one thread with cooperative scheduling. A synchronous decode never yields, so the event loop cannot resume anything else until it finishes, and every connection that worker holds waits behind it.
solid answer
~50 sThe event loop in an async worker is a single thread that resumes coroutines only when the running one suspends at an `await`. A synchronous decode has no suspension point, so for its whole duration the loop runs nothing else — every request that worker holds shows the stall, which is why the latency spike is correlated across unrelated connections rather than isolated to the slow upload. To confirm it, enable asyncio debug mode (`asyncio.run(main(), debug=True)` or `PYTHONASYNCIODEBUG=1`), which logs callbacks that run too long, and check whether spikes align with those log lines. The fix is to stop running that work on the loop thread: push the decode to `asyncio.to_thread` if it releases the interpreter lock or is genuinely I/O, or to a `concurrent.futures.ProcessPoolExecutor` if it is pure-Python CPU work. The deeper lesson is that an async worker is only as good as the blocking-free discipline of everything it calls.
code
python · 17 linesimport asyncio
import time
async def other_connection():
start = time.perf_counter()
await asyncio.sleep(0.1)
return time.perf_counter() - start
async def slow_handler():
time.sleep(0.5)
return "archived"
async def main():
waited, _ = await asyncio.gather(other_connection(), slow_handler())
print(f"a 0.1s await actually took {waited:.2f}s")
asyncio.run(main())go deeper
Remember that an event loop runs on one thread and only switches tasks at an await. Declaring a function with async def does not make a blocking call inside it non-blocking.
Explain cooperative scheduling and name the offload tools: asyncio.to_thread for calls that wait, a process pool for calls that compute. Be able to say why the stall hits every connection on that worker rather than only the slow one.
Show the diagnostic path: notice that latency is correlated across unrelated requests, confirm with asyncio debug mode or a loop-lag metric, identify the offending call, then choose the offload that matches whether it waits or computes — and consider moving it off the request path entirely.
Treat it as a model-fit question, not a bug. Decide whether the dependency stack is non-blocking enough to justify async workers at all, mandate loop lag as a first-class service metric, and set the policy that keeps blocking calls out of handlers as the codebase grows.
## The symptom is the diagnosis The distinguishing feature here is *correlation*: when a request is slow, unrelated requests on the same worker are slow by the same amount, and requests on other workers are fine. That pattern points at a **shared serialising resource inside one worker** — and in an async worker the shared resource is the loop thread itself. ## Mechanically An event loop is a single thread running a queue of ready callbacks. A coroutine keeps that thread from the moment it is resumed until it suspends at an `await` on something not yet ready. There is **no preemption**: the loop cannot take the thread back. So any code path with no suspension point occupies the whole worker: - a synchronous driver call, - a large `bytes.decode` over a big payload, - a compression pass, - a regex over megabytes, - a `time.sleep`. Concretely for an archiver serving an eleven-person team's history: a transcript whose stored bytes were labelled one codec but written in another falls out of the fast path and into a re-decode with error handling over the entire file. That is several hundred milliseconds of pure Python and C work with no await in it, and every one of the worker's open connections — idle websockets, in-flight uploads, the health check — is frozen for that long. The health check timing out is often what pages someone, which sends people hunting for a network problem that does not exist. ## Confirming it rather than guessing - Run the service with **asyncio debug mode** enabled — `asyncio.run(main(), debug=True)`, or `PYTHONASYNCIODEBUG=1` in the environment. Debug mode logs a warning whenever a callback occupies the loop longer than a threshold, naming the callback; that log line is the direct evidence, and correlating its timestamps with the latency spikes closes the case. - A second, cruder confirmation that needs no code change: a monitoring coroutine that awaits a fixed short sleep in a loop and records how much longer than the requested interval it actually took. That number *is* **loop lag**, and it spikes exactly when the loop is held. If the delay were in the datastore instead, other coroutines would keep making progress and the lag would stay flat. ## Fixing it First separate the two cases, because they need different answers. - If the slow call *waits* — a synchronous network or file call — wrapping it in `asyncio.to_thread` is correct: CPython releases the interpreter lock around blocking I/O, so the worker's loop thread runs while the helper thread waits. - If the slow call *computes* in pure Python, a thread does not save you: the interpreter lock means the helper thread and the loop thread take turns executing bytecode, so the loop still stutters, just less obviously. That work belongs in a `concurrent.futures.ProcessPoolExecutor`, or off the request path entirely into a queue consumed by separate workers — which for an archiver's re-decode is usually the right shape, since nothing about it needs to happen while a client waits. ## Second-order effects worth naming Offloading to threads is not free: each offloaded call costs a thread from a bounded pool, and if the pool saturates, `asyncio.to_thread` calls queue — you have re-created the threaded model inside the async one, with an event loop's constraints and a thread pool's ceiling. And any *synchronous* work you keep on the loop by accident still hurts, so the discipline has to cover **third-party code** too: an ordinary-looking helper deep in a library can contain a blocking call, and the loop cannot tell you which. ## The structural conclusion This class of incident is what "a blocking driver rules out the async model" means. If the datastore, the queue client or the file-format library your service depends on offers only synchronous calls, an async worker gives you the constraints of cooperative scheduling with none of its benefit — you will wrap everything in threads and end up paying for both. In that situation **threaded workers are the honest choice**: they accept blocking code by design, and one slow call delays one request instead of all of them. Choose async when the whole dependency stack is non-blocking, and keep a loop-lag metric on the dashboard so that the first blocking call anyone introduces is visible within a deploy rather than after an incident.
- How would you measure loop lag without changing the handlers?Run a small background coroutine that awaits a fixed short sleep in a loop and records the difference between the requested interval and the elapsed time, exporting it as a metric. That gap is time the loop was unavailable. It needs no handler changes, works in production, and distinguishes a held loop from a slow dependency — a slow upstream leaves the lag flat while request latency rises.
- Why is asyncio.to_thread the wrong fix for a pure-Python CPU-bound decode?Because the interpreter lock still serialises bytecode in the default build. The helper thread and the loop thread alternate rather than run together, so the loop keeps stuttering — the stall is smeared out instead of removed. CPU work needs a separate interpreter: a process pool, or a background service consuming a queue. Threads are the right offload only when the call actually waits and releases the lock.
- If the archiver's only datastore driver were synchronous, what would you recommend?Not async workers. Wrapping every driver call in a thread gives you cooperative scheduling's constraints plus a thread pool's ceiling, and one saturated pool re-serialises everything. Threaded workers accept blocking code natively and confine a slow call to the request that made it. Reserve async for a stack that is non-blocking end to end, and revisit only if a non-blocking driver becomes available.
saying these in an interview costs you the question
- Blames the network or the datastore before measuring loop lag
- Thinks the event loop preempts a long-running coroutine
- Assumes marking a function async makes its body non-blocking
- Offloads pure-Python CPU work to threads and expects the stall to vanish
- Adds more async workers instead of removing the blocking call
- Believes only the slow request is affected by the stall