skip to content

Crashes and Hangs

What to do when the process gives you nothing: a segfault with no traceback, a worker that stopped responding, or a death inside compiled code. Each symptom has its own capture technique.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

What does Python's "coroutine was never awaited" RuntimeWarning mean?

level: juniorimportance: must knowfreq 62%

answer

  1. The call returned something inert
  2. Nothing in the body ever ran
  3. The finalizer notices an unstarted frame
  4. Debug mode adds the creation traceback

basics

~20 s

Calling an async def function only builds a coroutine object; none of its body runs until something awaits it or wraps it in a task. When that unstarted object is garbage collected, CPython emits the warning.

solid answer

~40 s

`async def` defines a coroutine *function*. Calling it runs no code at all — it allocates a coroutine object that holds a not-yet-started frame. Only `await`, `asyncio.run()`, `asyncio.create_task()` or an `asyncio.TaskGroup` actually drives that object. If the last reference to an unstarted coroutine object goes away, its finalizer emits `RuntimeWarning: coroutine 'name' was never awaited`. The practical consequence is silence rather than a crash: the work simply never happens, so a service quietly returns incomplete results. Because the warning fires at collection time, the line it reports is where the object died, not where it was created — asyncio debug mode adds coroutine origin tracking so the warning also prints the creation traceback.

code

python · 14 lines
python
import asyncio


async def enrich(ticket_id: int) -> int:
    await asyncio.sleep(0)
    return ticket_id


async def main() -> None:
    enrich(17)  # no await: builds a coroutine object, runs nothing
    await asyncio.sleep(0)


asyncio.run(main())

go deeper

for a junior

Be able to say that calling an async def function returns a coroutine object and runs nothing, and that the warning means you forgot to await or schedule it. Know the two fixes: await it, or wrap it in a task.

for a middle

Explain the mechanics: the finalizer checks whether the frame ever started, which is why the warning arrives at garbage-collection time and reports a misleading line. Know that asyncio debug mode adds the creation traceback.

for a senior

Show how you keep this from reaching production: escalate the warning in CI, enable development mode in staging, and recognise the silent-partial-result failure shape it produces in a live service rather than a crash you would notice.

for a principal

Own the policy question — which warnings a service escalates to errors, whether debug mode runs in staging by default, and how much cost you accept in static checking and review to eliminate a class of silent-omission bugs.

## What a coroutine object actually is `async def` does not define something you call for effect. It defines a **coroutine function**. Calling that function evaluates none of the body: it allocates a *coroutine object*, a suspended computation wrapping a frame that has never started, and returns it. Something else must drive that object to completion — an `await` on it from inside another coroutine, a top-level `asyncio.run()`, or wrapping it in a Task via `asyncio.create_task()` or `asyncio.TaskGroup.create_task()`. Drop the object on the floor and the work never happens. Nothing raises, nothing blocks, nothing appears in a log unless the warning reaches you. This trips people who read `async def send(...)` as "a function that sends". It is closer to a factory: `send(...)` manufactures a *plan* to send, and the plan is inert until the event loop executes it. ## Where the warning comes from The coroutine type has a finalizer. When the last reference to a coroutine object disappears and CPython deallocates it, the finalizer asks whether the frame ever started executing. If it did not, it calls `warnings.warn()` with a `RuntimeWarning` whose message names the qualified coroutine: `coroutine 'triage.enrich' was never awaited`. `RuntimeWarning` is shown by default, so it usually lands on stderr — but two ordinary situations hide it. First, a process whose stderr is discarded or shipped to a channel nobody reads. Second, warning filters: `-W ignore`, a `warnings.simplefilter()` call at startup, or a logging bridge that captures warnings into a logger configured above WARNING. ## Why the reported location is useless by default The warning is emitted at *collection* time, not at the call site. The traceback line you get is wherever the object happened to die — often the tail of an unrelated function, or a garbage-collection pass in a completely different part of the program. That is exactly the gap asyncio debug mode closes: with debug enabled, coroutine origin tracking remembers the stack where each coroutine object was **created**, and the warning prints a "Coroutine created at" traceback pointing at the missing `await`. Enable it with `asyncio.run(main(), debug=True)`, the environment variable `PYTHONASYNCIODEBUG=1`, or `-X dev`, which turns on development mode and asyncio debug mode with it. ## The shapes the bug actually takes - A plain forgotten `await` — the majority case, and one a type checker catches only if the result is otherwise unused in a checkable way. - Handing a coroutine object to a synchronous API that expects a callable: a retry helper, a cache wrapper, or a mapping call that invokes `fn()` and stores whatever comes back. - Building coroutines in a comprehension and never passing them to `asyncio.gather()`. - A branch that awaits on one path and forgets on the other. - `with` where `async with` was required, so `__aenter__` is never driven. - Calling an async function from a genuinely synchronous context — a signal handler, a `__del__`, a synchronous framework hook — where there is no loop to await on. ## The failure mode is silence Picture a ticket-triage bot that builds three enrichment coroutines per ticket — classify, deduplicate, summarize — and awaits only two of them. Nothing fails. The bot keeps posting triage comments, but one field is quietly missing from every one of them: a silent truncation that passes review because the happy-path tests await everything. At a 1,200-request-per-minute peak the single warning line is buried in the log stream, if it is even enabled. This is why the warning is worth wiring into CI rather than trusting yourself to spot it in production. ## Fixes Await the object (`await enrich(t)`), gather a batch (`await asyncio.gather(*coros)`), or run a set concurrently under an `asyncio.TaskGroup` (3.11+), which also propagates failures instead of losing them. If you genuinely want fire-and-forget, `asyncio.create_task()` starts the coroutine — so it does *not* trigger this warning — but you must keep a strong reference to the returned Task until it completes, because the loop holds only a weak reference and a task can otherwise be collected mid-flight. ## Make it loud Run the test suite with `-W error::RuntimeWarning`. Because the warning is raised inside a finalizer, turning it into an exception does not propagate into your code; CPython prints `Exception ignored in:` with a traceback and continues. That is still a far better artifact than one line of text, and it is greppable in CI output. Running tests under `-X dev` additionally gives you the creation traceback. Finally, do not confuse this warning with `Task was destroyed but it is pending!`. That message means a Task *was* created and *did* start, and the loop was closed while it was still suspended — a different bug with a different fix.

  • Why does the warning point at a line that has nothing to do with the missing await?
    Because it is emitted from the coroutine object's finalizer, so the reported location is wherever the last reference was dropped and the object was deallocated — frequently a garbage-collection pass far from the call site. asyncio debug mode solves this by recording the stack at coroutine creation and printing it under a "Coroutine created at" header alongside the warning.
  • Does calling asyncio.create_task and never awaiting the task trigger the same warning?
    No. `create_task()` schedules the coroutine, so its frame starts and the never-awaited check does not fire. The related hazard is different: the event loop keeps only a weak reference to the task, so a fire-and-forget task with no strong reference can be garbage collected mid-flight. Store it in a module-level set and discard it from a done callback.
  • How would you make this warning fail a build rather than scroll past in a log?
    Run the suite with `-W error::RuntimeWarning` so the warning is escalated. It is raised inside a finalizer, so it prints as `Exception ignored in:` rather than propagating, but you get a full traceback that a CI step can grep for and fail on. Running under `-X dev` adds the coroutine's creation traceback to that output.

Calling an async def function is like filling out an order form: the form exists the moment you write it, but nothing is shipped until someone actually submits it.

saying these in an interview costs you the question

  • Says calling an async def function starts running it
  • Thinks the warning means the coroutine crashed
  • Believes the reported line is where the await was forgotten
  • Claims a missing await raises an error at the call site
  • Confuses it with "Task was destroyed but it is pending"
  • Assumes a fire-and-forget task never needs a strong reference

context

open as a page

What causes a Python RecursionError, and what do sys.getrecursionlimit and sys.setrecursionlimit control?

level: juniorimportance: must knowfreq 58%

basics

~20 s

CPython counts the Python frames stacked on the current thread and raises RecursionError once that count passes the ceiling sys.getrecursionlimit() reports, which is 1000 on a fresh interpreter. sys.setrecursionlimit() moves the ceiling; it does not enlarge the thread's real stack.

open as a page

Why does a segfault inside a compiled Python extension module leave no traceback?

level: middleimportance: must knowfreq 30%

basics

~20 s

A SIGSEGV kills the process at the operating-system level before CPython can run any Python code, so there is no unwinding, no except clause and no traceback - only a shell status of 139 and possibly a core file.

open as a page

A Python worker pool renders no more invoices — how do you read an all-threads stack dump to find the stuck frame?

level: seniorimportance: must knowfreq 52%

basics

~10 s

Take two dumps a few seconds apart. Identical stacks at near-zero CPU means blocked; a pinned core means spinning. Then read each worker's innermost frames and find the one thread that is not waiting.

open as a page

What does enabling Python's faulthandler give you when a process dies from a segfault?

level: juniorimportance: should knowfreq 26%

basics

~20 s

It installs handlers for fatal signals such as SIGSEGV and SIGABRT. When one fires, the interpreter writes a Python stack for each thread to stderr and then lets the process die exactly as it would have.

open as a page

What does threading.enumerate() return, and what does it tell you about a process that has gone quiet?

level: juniorimportance: should knowfreq 30%

basics

~20 s

threading.enumerate() returns a list of every Thread object alive right now, including the main thread, daemon threads and stand-ins for threads started by native code. It gives you names and identifiers, not what each thread is executing.

open as a page

How does asyncio debug mode expose a callback that blocks the event loop?

level: middleimportance: should knowfreq 44%

basics

~20 s

With asyncio debug mode on, the event loop times every callback it runs and logs a warning naming any that exceeded slow_callback_duration, which defaults to 0.1 seconds. That message identifies the code hogging the loop thread.

open as a page

How does faulthandler.register arm a Python process to dump every thread's stack on demand?

level: middleimportance: should knowfreq 42%

basics

~20 s

faulthandler.register(signal.SIGUSR1) installs a C-level handler for a spare signal, so sending that signal makes the process print every thread's Python stack to a file and keep running. It is Unix-only, and it works even while Python code is blocked.

open as a page

An asyncio service stops responding at peak but stays alive — how do you find the stuck coroutine?

level: seniorimportance: should knowfreq 46%

basics

~10 s

Get a task inventory out of the live process: a pre-installed signal handler that walks asyncio.all_tasks() and calls print_stack() on each. The dump shows every pending task and the await where it is suspended.

open as a page

How would you wire faulthandler into a service whose workers die and hang silently?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Set PYTHONFAULTHANDLER in the process environment so crashes during startup are covered too, make sure the destination stream is captured and retained, assert faulthandler.is_enabled() at boot, and arm a dump_traceback_later watchdog around the work that hangs.

open as a page

Exit status 137 and no traceback: how do you tell a kernel OOM kill of a Python process from a MemoryError?

level: seniorimportance: should knowfreq 47%

basics

~10 s

A MemoryError is an ordinary Python exception, so the interpreter is alive to print a traceback and run handlers. A kernel OOM kill is SIGKILL from outside: no traceback, no cleanup, exit status 137.

open as a page

A clinical-lab result loader dies with SIGSEGV inside a compiled extension - how do you capture a core file and read a Python stack from it?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Raise the core-size limit, find where the kernel writes cores, reproduce the crash, then open the core in gdb against the exact interpreter binary that produced it and use CPython's helper to print Python frames.

open as a page

What does Python's os.abort() do, and how does it differ from sys.exit()?

level: juniorimportance: nice to knowfreq 14%

basics

~20 s

os.abort() raises SIGABRT and ends the process in the hardest way available: no exception, no finally blocks, no atexit callbacks, exit status 134, and a core file if the limit allows. sys.exit() merely raises SystemExit, which unwinds normally.

open as a page

How does faulthandler.dump_traceback_later work as a watchdog for a wedged call?

level: middleimportance: nice to knowfreq 14%

basics

~10 s

It arms a watchdog thread that, after the given number of seconds, dumps every thread's Python stack and optionally kills the process. Call faulthandler.cancel_dump_traceback_later() when the guarded work finishes in time.

open as a page

What does threading.stack_size do, and how does it let deeper recursion run in a Python worker thread?

level: middleimportance: nice to knowfreq 16%

basics

~20 s

threading.stack_size(size) sets the stack size requested for threads started after the call, and returns the previous value. Pair it with a raised sys.setrecursionlimit() to run deep recursion in that worker; the main thread's stack is fixed by the OS.

open as a page

Why can a stack dumper built on sys._current_frames() miss the very hang it was written to diagnose?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

A dumper built on sys._current_frames() is ordinary Python: its thread must be scheduled and must run bytecode. If another thread is wedged inside a long native call holding the interpreter lock, the dumper never runs at all.

open as a page