How does asyncio debug mode expose a callback that blocks the event loop?
answer
- The loop times what it runs
- Threshold attribute defaults to a tenth-second
- The log line names the handle
- Four ways to turn the mode on
basics
~20 sWith asyncio debug mode on, the event loop times every callback it runs and logs a warning naming any that exceeded slow_callback_duration, which defaults to 0.1 seconds. That message identifies the code hogging the loop thread.
solid answer
~40 sTurn debug mode on with `asyncio.run(main(), debug=True)`, `set_debug(True)` on the running loop, the `PYTHONASYNCIODEBUG=1` environment variable, or `-X dev`. In debug mode the loop measures each callback and, for any that runs longer than the loop's `slow_callback_duration` attribute (default `0.1` seconds), logs a warning through the `asyncio` logger of the form `Executing <Handle ...> took 0.256 seconds`, with the callback's function and source location in the handle's repr. That is the signature of synchronous work parked on the loop thread — blocking I/O, a slow driver call, or CPU-heavy parsing — because a cooperatively scheduled loop cannot run anything else meanwhile. Lower `slow_callback_duration` in staging to catch shorter stalls; debug mode also adds coroutine origin tracking and thread-affinity checks, and costs enough that it is not an always-on production setting.
code
python · 12 linesimport asyncio
import time
async def main() -> None:
loop = asyncio.get_running_loop()
loop.slow_callback_duration = 0.05
time.sleep(0.2) # synchronous call parked on the loop thread
await asyncio.sleep(0)
asyncio.run(main(), debug=True)go deeper
Know that an event loop runs one callback at a time, so any blocking call freezes everything, and that asyncio has a debug mode you switch on with asyncio.run(main(), debug=True) to be told which callback ran long.
Explain the mechanics: the loop times each handle, compares against slow_callback_duration (default 0.1 s), and logs through the asyncio logger. Name the four ways to enable debug mode and the other checks it brings.
Demonstrate the operational judgement — where you run debug mode, how you tune the threshold to keep the signal readable, why the warning may never appear if logging is misconfigured, and how you move the offending work off the loop thread.
Own the tradeoff between observability cost and coverage: whether stall detection belongs in the runtime, in metrics on loop lag, or in load testing, and what latency budget a single callback is allowed to consume across a fleet.
## The failure this catches An asyncio event loop is cooperative: it runs one callback at a time on one thread, and every other ready callback waits until that one returns. So a single synchronous call that takes 300 ms — a blocking database driver, a `time.sleep`, `json` or regex work over a large payload, a password hash, a DNS lookup through a synchronous resolver, even logging to a slow sink — freezes *everything*: timers fire late, sockets are not read, and throughput collapses even though CPU looks idle. The symptom at the edge is latency that grows with load for no visible reason. Debug mode's job is to name the offending callback instead of leaving you to guess. ## Turning it on Four equivalent entry points: - `asyncio.run(main(), debug=True)` — the usual one in application code. - `set_debug(True)` on the loop object returned by `asyncio.get_running_loop()`, which you can flip at runtime. - `PYTHONASYNCIODEBUG=1` in the environment — useful when you cannot edit the entry point, for example in a container. - `-X dev` (or `PYTHONDEVMODE=1`), which enables development mode; asyncio debug mode is part of it. `get_debug()` on the loop tells you whether it is currently on. ## The slow-callback warning With debug on, `run_once` timestamps each handle it executes and compares the elapsed time against `slow_callback_duration`, a plain float attribute on the loop object whose default is `0.1` seconds. Exceed it and the loop logs at WARNING through the `asyncio` logger: ``` WARNING:asyncio:Executing <Handle enrich_ticket() at triage/worker.py:88 created at triage/router.py:41> took 0.412 seconds ``` The handle's repr carries the callback and its source location, and for a coroutine step the repr names the Task and the coroutine, so you usually land on the right function immediately. Because it is a float attribute you can retune it live: set it to `0.05` in staging to surface shorter stalls, or raise it temporarily during a startup phase that legitimately does synchronous work and would otherwise flood the log. One operational trap: the message goes through the `asyncio` logger, so it obeys your logging configuration. If your app configures the root logger above WARNING, or attaches handlers that route `asyncio` elsewhere, you will never see it. Check that first when debug mode "produces nothing". ## What else debug mode does Slow-callback detection is the headline, but debug mode is a bundle: - **Coroutine origin tracking** — the never-awaited warning gains a traceback showing where the coroutine object was created. - **Thread-affinity checks** — calling non-thread-safe loop APIs from a thread other than the loop's raises instead of corrupting state silently. - **Retained tracebacks** — exceptions that are never retrieved from a future are logged rather than lost, and objects are wrapped with extra source information. - **Slow selector reporting** — a long poll in the I/O selector is logged the same way, which distinguishes "my code blocked" from "the loop waited on the OS". ## Reading the result A warning naming your own function is an immediate fix: move the work off the loop thread with `asyncio.to_thread()` or `run_in_executor()` on a bounded thread pool for blocking I/O, or a process pool for CPU-bound work, and give it a timeout. A warning naming a library callback usually means the library is synchronous under an async-looking API. A warning naming a garbage-collection-heavy step or an enormous serialization tells you the payload, not the call, is the problem. A subtlety worth stating in an interview: debug mode reports *duration*, not *cause*. It tells you which handle ran long, and the next step is a targeted measurement inside that function. It also cannot see a stall that happens between loop iterations for reasons outside the loop — a stop-the-world pause, CPU starvation from a noisy neighbour, or the process being swapped. ## The cost, and where to run it Debug mode adds per-callback timing, keeps origin stacks for every coroutine object created, and performs extra validity checks. That is real overhead on a hot path, plus log volume, which is why the documentation frames it as a development aid. The pragmatic pattern is: on by default in local development and in the test suite, on in staging, and in production either off or enabled deliberately for a bounded window on one instance while you chase a specific stall — with `slow_callback_duration` raised high enough that only genuinely pathological callbacks report.
- What is the default slow-callback threshold, and when would you change it?`slow_callback_duration` defaults to `0.1` seconds. Lower it — `0.05` or less — in staging or a test run to surface stalls that are individually small but add up under concurrency. Raise it when a phase such as startup or a scheduled batch legitimately performs long synchronous work and the warnings would drown out the real signal. It is a mutable float attribute on the loop, so you can retune it while the process runs.
- Why not simply leave asyncio debug mode enabled in production?It times every callback, records an origin stack for every coroutine object created, and performs extra thread and lifetime checks, so it costs both CPU and memory on the hot path and produces log volume. The usual compromise is debug on in development, tests and staging, and in production only for a bounded window on a single instance, with the slow-callback threshold raised so only pathological callbacks report.
- Debug mode names the handle, not the specific line that blocked. How do you narrow it down?Treat the warning as a pointer, then measure inside the named callback: bracket suspect sections with timers, or run that function alone with a deterministic profiler. In practice the shortlist is small — synchronous network or file I/O, a driver that is blocking under an async-looking API, serialization of a large payload, or CPU-heavy work — and confirming which one it is takes one targeted measurement.
saying these in an interview costs you the question
- Thinks debug mode makes blocking calls run concurrently
- Cannot name a way to enable it
- Says the threshold defaults to one second
- Recommends leaving debug mode on in production
- Blames network latency for a slow-callback warning
- Does not know the warning goes through a logger