Two of five concurrent children fail within milliseconds of each other, and cancelling the remaining three produces cancellation errors of its own. How do you report this to the caller without losing information or blaming the wrong cause?
answer
- causes vs consequences: drop scope-caused cancellations
- first genuine failure = primary
- extras as suppressed, never dropped
- keep child identity + stack + correlation id
- dedupe identical causes, log once at scope
basics
~20 sPick the first genuine failure as the primary cause, attach the second as a suppressed or secondary cause, and drop the cancellation errors entirely — they are consequences of the first failure, not causes. Keep each error's identity (which child) and never let a wrapper hide the type callers catch on.
solid answer
~60 sThe caller expects one outcome; you have several errors of two different kinds, and they must not be treated alike. **Classify first.** Independent failures are causes. Cancellation outcomes from siblings you yourself cancelled are *consequences* — filter them out rather than reporting them, otherwise every incident reads "cancelled" and the real bug is invisible. **Then choose a reporting shape:** - *First-failure-wins with suppressed causes* — the earliest genuine failure is primary; later independent ones are attached as secondary. Best default: it preserves the type callers catch on while keeping the rest. - *Composite/aggregate error* carrying all causes — honest when failures are truly peers, but breaks catch-by-type upstream, so unwrap when there is only one cause and expose the list for inspection. - *Per-child results* — return an outcome per child and let the caller decide; this is really a partial-results policy. **Then preserve context:** each error keeps which child produced it, its own stack, and the correlation id. Log once at the scope, not in every child, or five failures become fifty log lines that look like five incidents.
code
text · 11 linesoutcomes = join_all(children) # [ok, fail(B, TimeoutErr), fail(D, ParseErr), cancelled(A), cancelled(E)]
failures = [o for o in outcomes if o.failed]
if scope.cancelled_by_me:
failures = [f for f in failures if not f.is_cancellation] # drop consequences
failures.sort(by=observed_order)
primary = failures[0] # TimeoutErr from child B
for f in failures[1:]:
primary.add_suppressed(f) # ParseErr from D survives
raise primary # caller can still catch TimeoutErrgo deeper
Say the first real failure is reported and the others are kept attached to it, and that cancellation errors from siblings are not the cause.
Explain causes vs consequences, first-failure-wins with suppressed extras, and keeping which child failed plus its own stack.
Weigh aggregate versus primary-with-suppressed against caller ergonomics (catch by type, retry decisions), deduplicate identical causes from a shared dependency, log once at the scope, and make ordering deterministic when failures are truly concurrent.
Standardise the error contract across the platform — attribution of cancellation, retained-cause structure, correlation ids, and one logging point — so incident triage is uniform and error-rate metrics are not inflated by consequences.
## The core problem A concurrent scope has N children but exactly one way to return. When several fail, something must decide what the caller sees. Doing this badly is one of the most expensive debugging taxes in a fan-out system: the information that would have identified the root cause is discarded at exactly the moment it existed. ## Step 1: separate causes from consequences After the first failure, a fail-fast scope cancels the remaining children, and those children terminate with cancellation outcomes. Those are **not failures of the system** — they are the scope doing its job. Rules: - A cancellation outcome caused by the scope's own cancellation is **discarded**, not reported and not counted as an error. - It must never become the primary cause. A scope that reports "cancelled" as the reason a request failed has hidden the actual bug, and every dashboard fills with a symptom. - The exception is a cancellation that arrives from *outside* (client disconnect, shutdown, deadline). If the scope was cancelled externally, cancellation genuinely is the outcome — but even then, a real child failure that happened before it usually deserves to be reported alongside, because it may be why the deadline was missed. Distinguishing the two requires knowing *who* cancelled: the scope tracks whether it initiated cancellation itself, so cancellation outcomes arriving after that point are attributable and filterable. ## Step 2: choose a reporting shape **First-failure-wins with suppressed causes.** The earliest genuine failure becomes the primary error; subsequent independent failures are attached to it as secondary/suppressed causes. Advantages: the caller's `catch SpecificError` still works, the stack of the primary is intact, and nothing is lost because the extras travel along. This is the best default and is what most structured-concurrency implementations do. **Composite (aggregate) error.** A single error type holding a list of causes. Honest when the failures are genuine peers — five independent shard writes, three of which failed for different reasons. Two costs: (a) it breaks type-based handling, because upstream code catching a specific error type no longer matches — mitigate by unwrapping when there is exactly one cause and by exposing the causes for inspection; (b) message formatting explodes if you concatenate five stacks into one line. **Per-child results.** Instead of exceptions escaping, each child returns a success-or-failure value, and the scope returns the list. This does not aggregate errors at all — it moves the decision to the caller, and is really a partial-results policy rather than an error-reporting one. ## Step 3: keep identity and context Whatever the shape, each retained error needs: - **Which child** produced it — a stable task name or index. "Timeout" with no indication of which of five downstream calls timed out is nearly useless. - **Its own stack/trace**, not a re-thrown wrapper that starts at the scope. - **The correlation id** of the originating request, so the scope's report and the children's own traces can be joined. ## Step 4: deduplicate and rate-limit When a shared dependency dies, all five children fail with the same error. Reporting five identical causes is noise; collapse duplicates by type-and-message with a count ("connection refused ×5"). Similarly, log **once at the scope**, not inside every child: per-child logging plus scope logging turns one incident into a burst that looks like many, distorts error-rate metrics, and can itself become a load problem during an outage. ## Step 5: get the ordering right "First" should mean first *observed failure*, which is well-defined even without a global clock because the scope records arrivals in order. Do not order by which cancellation finished first, and do not let a slow child's failure that arrived during wind-down displace the original trigger. When two failures are genuinely concurrent and the ordering is arbitrary, say so — pick one deterministically (lowest child index) and keep the other attached, rather than letting scheduling noise change which error users see between runs. ## Step 6: think about what the caller does next The report shape should support the caller's decision: - **Retry?** Only if all retained causes are retryable. One permanent failure among four transient ones means do not retry — which the caller can only determine if the permanent one was not discarded. - **Which downstream to blame?** Needs per-child identity. - **Degrade?** Needs to know which parts succeeded, which pushes toward per-child results. ## Anti-patterns - Reporting the cancellation of siblings as the failure cause. - Keeping only the *last* error to arrive — the last one is usually the least informative, since it is often a consequence. - Wrapping everything in a generic error type that upstream cannot inspect, forcing string matching on messages. - Losing the child identity so all five failures look identical. - Logging in the child and rethrowing, then logging again at the scope, and again at the top-level handler.
- Why not always wrap every failure in a single aggregate error type?Because it breaks type-based handling: upstream code that catches a specific error to retry or map it to a status code no longer matches, and people fall back to matching on message strings. Aggregates are appropriate when the failures are genuine peers, but they should unwrap to the single cause when there is only one, and always expose the causes for inspection so callers can decide by type rather than by text.
- How do you tell a cancellation that is a consequence from one that is the real outcome?By attribution: the scope knows whether it initiated the cancellation itself after a child failed, or whether the signal came from outside — a client disconnect, a deadline, or shutdown. Cancellation outcomes observed after the scope's own cancellation are consequences and are filtered out; an externally originated cancellation is a genuine outcome and is reported, though any real child failure observed before it is usually worth reporting alongside it.
saying these in an interview costs you the question
- Reports a sibling's cancellation as the root cause of the request failure
- Keeps only the last error to arrive rather than the first genuine failure
- Discards the second independent failure entirely instead of attaching it
- Wraps everything in a generic error type upstream cannot inspect by type
- Logs the failure in every child and again at the scope, multiplying one incident into many