skip to content

How do you keep fire-and-forget asyncio tasks alive and their failures visible in a 6-hour log-ingest run?

level: seniorimportance: should knowfreq 40%

answer

  1. Two defects wearing one coat
  2. Lifetime and visibility fixed together
  3. Name it, hold it, discard it, inspect it
  4. Held forever means never reported
  5. Async cleanup in a collected task cannot finish

basics

~20 s

Give background work an owner: a spawn helper that names each task, keeps it in a set, discards it when done, and retrieves its outcome in a done callback so failures are logged with context. Bound the in-flight count.

solid answer

~50 s

Fire-and-forget spawning has two independent failure modes that must be fixed together. A task nothing references can be collected mid-flight, so it needs an owning container; a task nobody inspects reports its failure only when finalized, so holding it forever hides the error instead. The pattern that satisfies both is one supervised spawn helper: `asyncio.create_task(coro, name=...)`, add to a set, `Task.add_done_callback()` twice — once to discard from the set, once to check `Task.cancelled()` and call `Task.exception()` and log with `Task.get_name()`. Over a long run add a bound — a semaphore or a fixed worker set reading a queue — because an unbounded spawn rate turns the owning set into a slow leak. And do not rely on `finally` inside a detached task for closing files or connections: if the task is collected, that cleanup runs during garbage collection and any `await` in it cannot complete.

code

python · 35 lines
python
import asyncio
import logging

log = logging.getLogger("ingest")

class Background:
    def __init__(self) -> None:
        self._tasks: set[asyncio.Task] = set()

    def spawn(self, coro, name: str) -> asyncio.Task:
        task = asyncio.create_task(coro, name=name)
        self._tasks.add(task)
        task.add_done_callback(self._tasks.discard)
        task.add_done_callback(self._report)
        return task

    @staticmethod
    def _report(task: asyncio.Task) -> None:
        if task.cancelled():
            return
        exc = task.exception()
        if exc is not None:
            log.error("background task %s failed", task.get_name(), exc_info=exc)

async def rotate_segment(n: int) -> None:
    await asyncio.sleep(0.01)
    raise OSError(f"segment {n} left open")

async def main() -> None:
    bg = Background()
    for n in range(2):
        bg.spawn(rotate_segment(n), name=f"rotate-{n}")
    await asyncio.sleep(0.1)

asyncio.run(main())

go deeper

for a junior

At minimum, know that background work started and forgotten needs someone to keep the task object and someone to look at how it ended. Recall the two symptoms: work that quietly does not happen, and failures that never appear in the log.

for a middle

Be able to build the helper: name the task at creation, add it to a set, discard it in a done callback, and inspect the outcome with a cancellation guard before calling Task.exception(). Explain why the set alone hides errors and the callback alone loses the task.

for a senior

Diagnose it from symptoms — descriptors climbing, output short, logs clean — and explain the collected-task cleanup trap that leaves resources open. Add a concurrency bound and move resource ownership out of the detached task, and say how you verified the fix.

for a principal

Own the policy: one supervised spawn API for the codebase, background task count and failure count as metrics rather than log lines, and an explicit rule about which work is allowed to outlive its caller at all. Decide what a failed background unit does to the status of the job around it.

### The failure being described A nightly log-ingest run spawns a segment flush with `asyncio.create_task()` for each rotated file and moves on. Six hours later the run reports success, output is short by a few segments, and one file descriptor per missing segment is still open. Nothing in the log explains it. This is the canonical fire-and-forget incident, and it is two defects wearing one coat. ### Defect one: nothing owns the task asyncio's live-task registry is a `weakref.WeakSet`, so the object returned by `create_task()` may be the only strong reference. A task suspended on a future that nothing else reaches is a collectible cycle; when it is reclaimed the coroutine stops where it stands. On a short run this almost never happens, which is why it survives review — the timer and selector registrations of a busy task keep it incidentally alive. Over six hours, with collector passes triggered by allocation churn, it happens enough to matter. The collection also explains the leaked descriptors. Finalizing a suspended coroutine throws `GeneratorExit` at the `await`, so a `finally:` block runs — but during garbage collection, off the loop. An `await writer.wait_closed()` or an async context manager exit inside that `finally` cannot complete; the interpreter reports that the coroutine ignored `GeneratorExit`, and the resource stays open. **Cleanup written as async code inside a detached task is not a guarantee.** ### Defect two: nobody consumes the outcome A task that raises stores the exception and flags it unconsumed. The `Task exception was never retrieved` line is emitted by the task's finalizer through the loop's exception handler, at ERROR on the `asyncio` logger. That means the report is tied to garbage collection, not to failure, so it lands at an unrelated moment with no batch or file context — and if you fixed defect one by holding tasks in a set you never drain, the finalizer never runs and the failure is reported *never*. Fixing lifetime naively hides the errors. ### One helper that fixes both Centralise spawning so the policy exists in exactly one place: * **Name every task.** `asyncio.create_task(coro, name=f"flush-{segment}")` makes `Task.get_name()` meaningful in every later log line and in any live task listing. * **Own it.** Add the task to a set held by the component that owns the work, not a global one nobody remembers. * **Drain it.** `task.add_done_callback(self._tasks.discard)` keeps the set at the size of the in-flight population instead of the run's total. * **Retrieve the outcome.** A second callback returns early when `Task.cancelled()` is true, then calls `Task.exception()` and logs with `exc_info` and the task name. This consumes the exception, so the anonymous asyncio line disappears and your contextual line replaces it. * **Decide, don't just log.** For an ingest run, a failed segment usually should mark the run degraded or increment a failure counter that a monitor reads. A silent ERROR that nobody alerts on is only marginally better than no message. ### Bounding the fan-out Owning the tasks means holding them, so spawn rate becomes a memory question. If the run rotates a segment every few seconds for six hours, a helper that never limits concurrency accumulates whatever fraction is slow. Two straightforward bounds: acquire an `asyncio.Semaphore` inside the spawned coroutine so only N flushes run at once, or stop spawning per item entirely and run a fixed set of worker tasks pulling from an `asyncio.Queue`. The second is usually the better shape for a steady pipeline, because it makes the concurrency a constant of the design rather than a consequence of input rate, and it turns backpressure into queue depth you can measure. ### Resource ownership belongs outside the task Given that async cleanup in a collected task cannot be trusted, keep the resource's lifetime with a component that outlives any single background task: open the destination once and hand the task a handle, or perform the close in code that is definitely awaited rather than in a detached task's `finally`. This is a design rule, not a workaround — it also survives cancellation and process signals, which are the other two ways a detached task fails to reach its cleanup. ### Prefer scope where the work is scoped Most of what gets spawned fire-and-forget does not actually need to outlive its caller; it was written that way to avoid waiting. If the work is scoped to a block, a scoped group that owns its children removes both defects by construction — references and failure propagation come with the scope. Reserve genuine fire-and-forget for work whose lifetime really is longer than the caller's, and accept that such work needs a supervisor. ### What a strong answer sounds like Name both defects, show that they pull against each other, and describe the single helper that resolves them: named tasks, an owning set, a discard callback, an outcome callback that retrieves the exception, a concurrency bound, and resource ownership outside the task. Then say how you would have found it: descriptor count climbing, output short, no errors — the signature of unsupervised background work.

  • Why bound the number of in-flight background tasks if each one is short?
    Because owning them means holding them, and the owning set grows to whatever fraction is slow at any instant. A steady spawn rate over a long run plus occasional slow flushes is a memory profile, not an anomaly. An `asyncio.Semaphore` inside the coroutine, or a fixed group of workers reading an `asyncio.Queue`, makes concurrency a design constant and turns backpressure into a queue depth you can measure.
  • How would you catch this class of bug before it reaches a long-running job?
    Make the invariant checkable rather than hoping for it. Forbid bare `asyncio.create_task()` outside the supervised helper and enforce it in review or a lint rule, export the size of the in-flight set as a gauge, and assert in tests that the set returns to zero after a workload. In staging, a run where output count is short while no error is logged is the exact signature to alert on.
  • Where should a background task's file handle or connection be closed, if not in its own finally block?
    In code whose completion you actually await, owned by a component that outlives any single task — open once and hand the task a handle, or close in the supervising scope. A detached task's `finally` can run during garbage collection, where any `await` fails, and it is equally fragile under abrupt cancellation, so it is the wrong place for the only copy of your cleanup.

saying these in an interview costs you the question

  • Fixes lifetime with a set but never retrieves outcomes
  • Assumes a finally block always closes the resource
  • Says the loop reports background failures on its own
  • Spawns one task per input with no concurrency bound
  • Treats an ERROR log line as sufficient failure handling
  • Uses fire-and-forget for work scoped to the caller

context