skip to content

Fire-and-Forget Task Pitfalls

A task nothing holds a reference to can be collected mid-flight, and a failure inside it surfaces only when the object dies. Interviewers ask how you keep background work both alive and visible.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

3

Why does asyncio.create_task() require you to keep a reference to the Task it returns?

level: juniorimportance: must knowfreq 55%

answer

  1. Ask who else is holding it
  2. The loop's bookkeeping is not ownership
  3. Registry is a weakref.WeakSet
  4. Suspended task collected, work never resumes
  5. Own it in a set, discard when done

basics

~20 s

The event loop registers running tasks only in a weak set, so the Task object create_task() hands back may be the only strong reference. Drop it and the task can be garbage-collected while still suspended, and the work never finishes.

solid answer

~40 s

`asyncio.create_task()` schedules the coroutine and returns a `Task`. The loop's own bookkeeping — what `asyncio.all_tasks()` reads — is a `weakref.WeakSet`, so it does not own the task. While a step is queued to run, or the task is parked on a timer or a socket the loop watches, something else transiently references it, which is why the bug is intermittent rather than constant. But a task whose only remaining references are inside its own cycle can be reclaimed at any moment: the coroutine simply stops mid-flight, and you get at most a `Task was destroyed but it is pending!` line on stderr. The fix is to own the lifetime: put the task in a long-lived set and register `set.discard` with `Task.add_done_callback()` so it is dropped when it completes.

code

python · 17 lines
python
import asyncio

_background = set()

async def flush(name):
    await asyncio.sleep(0.05)
    print("flushed", name)

async def main():
    for i in range(3):
        task = asyncio.create_task(flush(f"batch-{i}"), name=f"flush-{i}")
        _background.add(task)
        task.add_done_callback(_background.discard)
    await asyncio.sleep(0.2)
    print("still held:", len(_background))

asyncio.run(main())

go deeper

for a junior

Recall the one-line rule: keep the object asyncio.create_task() returns, because the loop only holds weak references to tasks. Be able to say what goes wrong if you do not — the task can be collected mid-run and the work is silently lost.

for a middle

Explain the mechanics: a weakref.WeakSet registry, the incidental references (queued step, timer, selector) that make failure intermittent, and the set-plus-add_done_callback(set.discard) idiom that gives the task an owner without growing forever.

for a senior

Demonstrate that you have debugged this in production: work that thins out under load rather than failing, Task was destroyed but it is pending! in the logs, and async cleanup in finally that never completes because the coroutine was closed during garbage collection.

for a principal

Own the convention rather than the incident. Decide where fire-and-forget spawning is even allowed in a codebase, provide a single supervised spawn helper that names and tracks tasks, and treat bare asyncio.create_task() calls at review time as a defect class rather than a style preference.

### What `create_task` actually gives you `asyncio.create_task(coro)` wraps a coroutine object in an `asyncio.Task`, schedules its first step on the running loop, and returns the `Task`. That return value is not a receipt you may throw away — under CPython it is frequently the only strong reference to the running work. ### The registry is deliberately weak asyncio keeps a module-level registry of live tasks so that `asyncio.all_tasks()` can enumerate them. In CPython 3.14 that registry is a `weakref.WeakSet` (plus a small strong set used only for eagerly-started tasks while they run synchronously). A weak reference does not keep its target alive. The design is intentional: if the loop owned every task strongly, an abandoned or leaked task would pin its coroutine frame, its locals and everything they reference for the lifetime of the process, and the loop would become a memory sink. Bookkeeping is not ownership — the library expects the caller to own the task, and the `create_task` documentation says so explicitly. ### Why the bug is intermittent A task usually has *some* other reference, which is why sloppy code survives testing. When a step is queued, the loop's ready queue holds a callback bound to the task. When the coroutine is parked in `asyncio.sleep()`, the loop's timer heap holds a handle that reaches the task. When it is blocked on a socket the loop is watching, the selector registration reaches it. Each of these disappears the moment that particular wait ends. The collectible shape is a task suspended on a future that nothing outside the task itself reaches — a reference cycle, exactly what the generational cycle collector exists to reclaim. Then collection happens at an arbitrary time: after some unrelated allocation churn triggers a gen-2 pass, minutes into a run, on one machine and not another. That is what makes this class of defect so unpleasant. It does not fail; it *thins out*, dropping a small fraction of background work under load while every unit test passes. ### What you see when it happens Usually nothing useful. If the task was still pending, its finalizer reports `Task was destroyed but it is pending!` through the loop's exception handler, which by default logs at ERROR on the `asyncio` logger — a line that is easy to filter away and carries no application context. Worse, finalizing the coroutine object throws `GeneratorExit` in at the suspended `await`, so `finally` blocks run *during garbage collection*, off the loop, at an unpredictable moment. Any `await` inside such a `finally` cannot complete: the interpreter reports that the coroutine ignored `GeneratorExit`, and asynchronous cleanup — closing a connection, flushing and closing a file handle — silently does not happen. So a collected task can leak the very resource its cleanup code was written to release. ### The idiom that fixes it Hold a strong reference for exactly as long as the task runs: ```python _background: set[asyncio.Task] = set() task = asyncio.create_task(work(), name="flush-batch") _background.add(task) task.add_done_callback(_background.discard) ``` The set is the ownership; the done callback is what stops that set from growing without bound, since `Task.add_done_callback()` fires once the task finishes, whether it returned, raised or was cancelled. Note the direction of ownership: registering a callback does **not** keep the task alive — the task holds its callbacks, not the reverse. Passing `name=` costs nothing and makes `Task.get_name()` useful when you later log about the task. Also worth internalising: a task you `await`, or one owned by a scope that awaits its children, is already referenced by the awaiting frame for its whole life, so the hazard is specific to *fire-and-forget* spawning — work you start and walk away from. Where the work's lifetime is genuinely the lifetime of a block, a scoped group is the structurally simpler answer; where it truly outlives the caller, the set-plus-discard idiom is the minimum bar. ### Interview framing The strong answer names three things: the registry is weak, the intermittency comes from incidental references that come and go, and the fix is an owning container with a done callback to drain it. A candidate who says the loop keeps every task alive has an incorrect mental model of asyncio ownership, and that model produces exactly the leaks and dropped work that make async services hard to trust.

  • Is an unreferenced asyncio task guaranteed to be collected?
    No, and that is what makes it dangerous. While a step is queued, or the task waits on a timer or a watched socket, the loop transiently reaches it and it survives. Collection happens only when nothing outside the task's own cycle refers to it and the cycle collector runs. So the same code drops work on one deployment and not another, and never in a short unit test.
  • Does calling Task.add_done_callback() by itself keep the task alive?
    No. The task holds a list of its callbacks, not the other way round, so a callback creates no reference back to the task from anywhere durable. The callback is useful for draining an owning container or retrieving the result, but the strong reference has to be something you keep — typically a module-level or owner-held set.
  • Is a local variable inside the spawning coroutine a sufficient reference?
    Only while that frame is alive. If the spawner assigns the task to a local and then returns, the frame dies and the reference goes with it, which is the usual way this bug is written. The reference has to outlive the spawner: a set on the owning object, or a scope that awaits the task before returning.

The loop's task list is a guest register, not a coat check: it records who is in the building but holds nothing of yours, so if you let go of the ticket, the coat is thrown out.

saying these in an interview costs you the question

  • Claims the event loop keeps a strong reference to every task
  • Says a task, once started, always runs to completion
  • Thinks add_done_callback keeps the task from being collected
  • Assigns the task to a local in the spawner and calls it owned
  • Believes create_task returns nothing worth keeping
  • Assumes finally blocks still do async cleanup after collection

context

open as a page

When does asyncio log 'Task exception was never retrieved', and why so late?

level: middleimportance: should knowfreq 50%

basics

~20 s

Nothing is reported when the task raises: the exception is stored on the Task and flagged unconsumed. Only when that Task object is finalized without anyone calling result() or exception() does its finalizer log the message.

open as a page

How do you keep fire-and-forget asyncio tasks alive and their failures visible in a 6-hour log-ingest run?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Give background work an owner: a spawn helper that names each task, keeps it in a set, discards it when done, and retrieves its outcome in a done callback so failures are logged with context. Bound the in-flight count.

open as a page