skip to content

Why choose asyncio.TaskGroup over asyncio.gather(return_exceptions=True) when fanning out a conversion batch?

level: seniorimportance: should knowfreq 46%

answer

  1. Ask what partial output is worth
  2. Failure as a value versus failure as a raise
  3. One aborts the siblings, one does not
  4. Errors returned as data get dropped
  5. Per-child try keeps the boundary

basics

~20 s

They encode opposite failure contracts. A task group is all-or-nothing: the first error cancels the siblings and raises. gather with return_exceptions=True is best-effort and hands failures back as values in a results list, where they are easy to drop.

solid answer

~40 s

Pick by what a partial result means. `asyncio.TaskGroup` is all-or-nothing — the first failure cancels the other children, nothing outlives the block, and the error propagates as an `ExceptionGroup` you cannot ignore by accident. `asyncio.gather(..., return_exceptions=True)` is best-effort: every awaitable runs to completion and failures come back as exception *objects* in a positional list, so the caller must remember to check them; the classic incident is a batch whose result list was scanned for successes and never for `isinstance(r, Exception)`. For a pipeline whose stages are meaningless apart, use the group. For a large batch where a handful of bad rows must not sink the rest, I still prefer the group but put the `try`/`except` inside each child, keeping the structural guarantee while making the per-item error policy explicit.

code

python · 21 lines
python
import asyncio

async def convert(page: int) -> str:
    if page == 2:
        raise ValueError("page 2 is corrupt")
    await asyncio.sleep(0.1)
    return f"page-{page}"

async def main() -> None:
    out = await asyncio.gather(*(convert(p) for p in range(4)),
                               return_exceptions=True)
    print("gather kept going:", out)

    try:
        async with asyncio.TaskGroup() as tg:
            for p in range(4):
                tg.create_task(convert(p))
    except* ValueError as eg:
        print("taskgroup aborted:", eg.exceptions)

asyncio.run(main())

go deeper

for a junior

Know the headline difference: a task group stops the other tasks when one fails and raises, while a best-effort fan-out lets everything finish and returns the failures as items in a list.

for a middle

Explain the mechanics of each contract — cancellation and a grouped raise versus a positional list containing exception objects — and why bare gather without the flag can leave tasks running after an error.

for a senior

Demonstrate the judgement call on real work: decide from whether partial output is shippable, and show the hybrid where per-child error handling inside a group buys tolerance without giving up the boundary.

for a principal

Own the standard: which fan-out shape is the default in this codebase, how failures returned as data are guaranteed to reach alerting, and how existing bare-gather call sites get migrated.

### Two different contracts `asyncio.gather(*aws, return_exceptions=True)` and `async with asyncio.TaskGroup()` both fan work out, but they promise opposite things about failure: | | `gather(..., return_exceptions=True)` | `TaskGroup` | |---|---|---| | First failure | ignored; everything keeps running | siblings cancelled | | Result | a list, positionally aligned, with exception *objects* in the failed slots | no results; raises `ExceptionGroup` | | After the call | all awaitables have settled | all children have settled | | Shape | best-effort batch | all-or-nothing operation | So the question is never "which is better" — it is *what does this piece of work mean when part of it fails?* ### Choose the group when partial success is meaningless Take a document-conversion queue that fans out one task per stage — fetch the source, render the pages, write the manifest. If the render fails, the manifest is garbage, so continuing to fetch is waste. `TaskGroup` gives you exactly the semantics you want: the first error cancels the rest, the block cannot be left behind, and the failure arrives as an `ExceptionGroup` you cannot silently ignore, because an unhandled group propagates like any other exception. That last property matters more than it sounds: `return_exceptions=True` turns failures into *values*, and values get dropped. The classic production incident is a batch job whose results list was scanned for successes and never checked for `isinstance(r, Exception)`, so a month of failures was invisible. ### Choose best-effort when partial success is the product Now take the other shape: a nightly conversion of a **6,800-row batch**, where twelve corrupt rows must not cost you the other 6,788. Here fail-fast is actively wrong. You have two decent options: 1. `gather(..., return_exceptions=True)` and then *deliberately* partition the returned list into successes and exceptions, reporting the failures. 2. Keep the `TaskGroup` — for its no-orphans guarantee — and push the `try`/`except` **inside each child**, so each task returns a success-or-failure record and no exception ever reaches the group. Option 2 is usually the better production answer, because you keep the structural guarantee and make the error policy explicit per item, at the cost of one wrapper coroutine. The rule of thumb: a group is for work that fails *together*; a per-item `try` inside the group is how you buy per-item tolerance without giving up the boundary. ### The option nobody should pick `gather()` *without* `return_exceptions` is the genuinely dangerous one. The first exception propagates to the caller immediately, and the remaining awaitables **are not cancelled** — they keep running as orphans, still holding connections, still writing rows for an operation the caller has already given up on. Every "why is this job still hitting the database after the request 500'd" mystery starts there. If you inherit that call, replacing it with a group is nearly always the right change. ### Two accumulation traps that ride along Best-effort code tends to grow an accumulator, and accumulators attract two bugs. The first is a mutable default: `def collect(row, results=[])` evaluates that list **once**, at `def` time, so the same list is shared by every call and the second nightly run reports the first run's rows as well. Take `None` as the default and build the list inside. The second is ordering: `gather` results are positional, matching the order of the awaitables you passed, *not* completion order — code that zips results back to inputs by arrival will mis-attribute failures. ### How to say it in an interview Lead with the semantics, not the API: "TaskGroup is all-or-nothing with a hard structural guarantee; `return_exceptions=True` is best-effort and hands failure back as data. I pick by whether partial output is shippable, and if it is, I still prefer a group with per-child error handling so nothing outlives the block." Then note that `gather` remains the right tool when you specifically need the positional result list, and that neither tool leaves work running behind it *except* bare `gather` on the error path. `TaskGroup` and `ExceptionGroup` both landed in **Python 3.11**; on 3.10 and earlier `gather` is all you have, which is why so much existing code is written that way. The behaviour above is current on 3.14.

  • What is wrong with plain `asyncio.gather()` without return_exceptions on the error path?
    The first exception propagates to the caller immediately, but the remaining awaitables are not cancelled — they keep running as orphans, holding connections and applying side effects for an operation the caller has already abandoned. That is the source of most "why is this job still writing rows after the request failed" mysteries. A task group has no such gap.
  • How do you keep partial results and still get the group's no-orphan guarantee?
    Wrap each unit of work in a coroutine that catches its own expected errors and returns a success-or-failure record, then spawn those inside the group. No exception ever reaches the group, so nothing is cancelled, you still get a per-item outcome for every row, and the block still cannot be left behind. The error policy becomes explicit and per item instead of implicit in a flag.
  • If you do use return_exceptions=True, what is easy to get wrong about the results list?
    Two things. It is positional — aligned with the order of the awaitables you passed, not completion order — so attributing failures by arrival mis-labels them. And the exceptions are ordinary values, so any code path that only looks for successes silently discards every failure; partition the list explicitly with an `isinstance` check and report the failures.
  • How would you accumulate per-row outcomes across a nightly batch without a stale-state bug?
    Never as a mutable default parameter. `def collect(row, results=[])` evaluates that list once, at definition time, so every call shares it and the second night's report includes the first night's rows. Default to `None` and build the list inside the call, or pass the accumulator in explicitly.

saying these in an interview costs you the question

  • Calls one strictly better rather than asking what failure means
  • Thinks gather cancels the other awaitables on first error
  • Treats return_exceptions=True as error handling by itself
  • Expects TaskGroup to hand back a list of results
  • Matches gather results to inputs by completion order
  • Accumulates batch results in a mutable default argument

context