skip to content

An asyncio service stops responding at peak but stays alive — how do you find the stuck coroutine?

level: seniorimportance: should knowfreq 46%

answer

  1. Classify by CPU before instrumenting
  2. Ask the process for an inventory
  3. Never queue the dump on the stalled loop
  4. One set of pending tasks, one stack each
  5. The task count is itself a finding

basics

~10 s

Get a task inventory out of the live process: a pre-installed signal handler that walks asyncio.all_tasks() and calls print_stack() on each. The dump shows every pending task and the await where it is suspended.

solid answer

~50 s

First classify the stall by CPU: one core pinned means something is running and never yielding; near-zero CPU means everything is waiting. Then dump state from inside the process. Pre-install a handler for a spare signal with `signal.signal` — not the loop's `add_signal_handler`, which schedules the callback on the very loop that is stuck — and have it iterate `asyncio.all_tasks(loop)` and call `print_stack(file=...)` on each pending task. Reading that dump answers the question: many tasks suspended at the same `await` on a lock, semaphore or queue means contention or deadlock; a task count that grows without bound means missing backpressure; no task making progress plus a pinned core means a synchronous call parked on the loop thread, which asyncio debug mode's slow-callback warning names directly. If the handler never runs, the interpreter is stuck inside a C call and you need an out-of-band mechanism.

code

python · 22 lines
python
import asyncio
import signal
import sys

_loop = None


def dump_tasks(signum, frame) -> None:
    tasks = asyncio.all_tasks(_loop)
    print(f"pending tasks: {len(tasks)}", file=sys.stderr)
    for task in tasks:
        task.print_stack(limit=8, file=sys.stderr)


async def main() -> None:
    global _loop
    _loop = asyncio.get_running_loop()
    signal.signal(signal.SIGUSR1, dump_tasks)
    await asyncio.sleep(3600)


asyncio.run(main())

go deeper

for a junior

Know that an event loop runs one callback at a time on one thread, and that asyncio.all_tasks() lists the tasks that have not finished. Being able to print those tasks is already useful.

for a middle

Explain how to get a stack out of a suspended task with print_stack(), what its top frame means, and why blocking work on the loop thread stalls every other task even when CPU looks fine.

for a senior

Show a working diagnosis: classify by CPU first, dump from a pre-installed signal handler rather than a loop callback, read the pattern in the dump — same await, growing count, or nothing suspended — and turn it into bounded concurrency and timeouts.

for a principal

Own the readiness question: which introspection hooks ship enabled in every service by default, what a stall costs before it is detected, and how loop-lag and pending-task metrics turn this from a live-debugging exercise into an alert.

## Step 1 — classify before you instrument A process that is alive but not serving has three distinct shapes, and cheap external signals separate them before you touch the code: - **One core pinned at 100%** — something is executing and never yielding. Either a synchronous blocking call is parked on the loop thread, or a genuine CPU-bound computation is running inside a coroutine. - **CPU near zero, connections accepted but never answered** — everything is suspended. A deadlock on an `asyncio.Lock`, an exhausted `asyncio.Semaphore` whose holders never release, a consumer that died leaving producers blocked on a full `asyncio.Queue`, or an awaited external call with no timeout. - **CPU near zero and connections refused** — the loop is not even accepting; look at the accept path and file-descriptor limits rather than at tasks. Consider a ticket-triage bot that serves fine all day and wedges at its 1,200-request-per-minute peak, still holding connections open while replies stop and the last few summaries it posted came back truncated. "At peak" is a strong hint: it points at a bounded resource — a connection pool, a semaphore, a queue — that only saturates under load. ## Step 2 — get a task inventory out of the live process `asyncio.all_tasks(loop)` returns the set of tasks for that loop that are **not yet done**, and `asyncio.Task.print_stack(limit=..., file=...)` prints that task's coroutine frames. Together they are the async equivalent of a thread dump. The trick is delivering the request into a process that is already wedged. The correct mechanism is a plain `signal.signal` handler installed at startup. Python runs it on the main thread between bytecodes, independently of the event loop's scheduling. The tempting alternative — the loop's `add_signal_handler` — is wrong here for a precise reason: it wakes the loop and *schedules* your callback as a loop callback, so if the loop is stalled on a blocking call the dump never runs. Diagnose the stuck component with a tool that does not depend on it. The handler should be cheap and allocation-light: iterate the task set, print each stack to `sys.stderr` with a small `limit`, and print the task count first, since the count alone is often the answer. ## Step 3 — read the dump `print_stack()` on a **suspended** task prints the coroutine's frames up to the `await` that suspended it — the top frame is where it is parked. On a task that has already failed, it prints that exception's traceback instead. Four patterns cover most real incidents: 1. **Many tasks parked at the same line** — an `await` on a lock, semaphore, queue or connection-pool acquire. That is contention or deadlock, and the top frames name the resource. 2. **A task count that keeps climbing** — a request handler creates tasks faster than they complete. There is no backpressure; the fix is bounded concurrency, not a faster loop. 3. **Nothing suspended anywhere interesting, plus a pinned core** — no task is at a meaningful `await` because one callback never returned. Correlate with asyncio debug mode's slow-callback warning, which names the handle directly. 4. **A handful of tasks parked on an outbound call with no visible timeout** — an upstream that stopped answering, propagating back through your loop as a stall. The count deserves emphasis. A dump showing 40,000 pending tasks tells you the design problem in one number, and no stack reading is required. ## Step 4 — when the handler never fires If you signal the process and nothing appears, the interpreter cannot reach a bytecode boundary — it is inside a C call that holds the GIL, or blocked in a system call the signal does not interrupt. Python-level introspection is out at that point. On 3.14 the remote-debugging entry point (`sys.remote_exec`, PEP 768) lets a privileged external process inject a script into the target to run at its next safe point, which recovers the same dump without a pre-installed handler; failing that, an out-of-band all-threads stack dump or a native debugger is the next step. ## Step 5 — turn the finding into a fix Blocking work moves off the loop thread with `asyncio.to_thread()` or `run_in_executor()` on a bounded pool. Every await of an external resource gets a timeout, so a stalled upstream becomes an error rather than an accumulating set of parked tasks. Unbounded task creation gets a semaphore or a fixed worker pool draining a bounded queue. And the dump handler stays installed in production permanently: it costs one signal registration, and it is the difference between a five-minute diagnosis and a restart that destroys the evidence. ## What an interviewer is listening for That you reach for evidence from inside the live process rather than restarting it; that you know `all_tasks` returns pending tasks for a specific loop; that you can explain why a signal handler beats a loop callback when the loop is the suspect; and that you read the *number* of tasks as carefully as their stacks.

  • Why install the dump handler with signal.signal instead of the event loop's add_signal_handler?
    `add_signal_handler` wakes the loop and schedules your callback as an ordinary loop callback. If the loop is the thing that is stuck — a blocking call on its thread, or a callback that never returned — the dump is queued behind the very problem you are diagnosing and never runs. A `signal.signal` handler runs on the main thread between bytecodes, so it does not depend on the loop making progress.
  • What does asyncio.all_tasks return, exactly, and what does it leave out?
    It returns a set of the tasks for a given loop that are not yet done, defaulting to the running loop. It excludes completed and cancelled tasks, coroutines that were never wrapped in a task, work running in a thread or process pool, and tasks belonging to a loop in another thread. So an empty or tiny set does not prove the process is idle — it proves nothing is pending on *that* loop.
  • You signal the process and no dump appears. What does that tell you?
    That the interpreter cannot reach a bytecode boundary, so the Python-level handler never runs — typically a C call holding the GIL or a system call the signal does not interrupt. Python-level introspection is unavailable; on 3.14 the remote-execution debugging entry point can inject a script at the next safe point, and otherwise you fall back to an out-of-band all-threads stack dump or a native debugger.
  • The dump shows tens of thousands of pending tasks. What is your reading?
    That the service creates work faster than it completes it and has no backpressure — the count is the finding, ahead of any individual stack. Each task carries a frame and its captured state, so memory climbs alongside latency. The fix is bounded concurrency: a semaphore around the spawn site or a fixed pool of workers draining a bounded queue, with timeouts so nothing parks forever.

It is a roll call in a stalled building: you are not looking for who is missing, you are asking every occupant which door they are standing at, and the answer is usually that they are all queued at the same one.

saying these in an interview costs you the question

  • Restarts the process before collecting any evidence
  • Registers the dump as a callback on the stalled loop
  • Thinks asyncio.all_tasks includes finished tasks
  • Believes a stalled loop always means slow network
  • Cannot distinguish suspended tasks from a hogged loop thread
  • Ignores the pending-task count and reads only stacks

context