You start three concurrent child tasks and wait for all of them. One fails after 50 ms while the others would run for another 10 seconds. What should happen to the siblings, and why is that the default in structured concurrency?
answer
- outcome already decided → stop the rest
- cancel siblings, then join, then report
- 50 ms not 10 s
- fail-fast = policy, cancellation = mechanism
- only as fast as siblings are cancellable
basics
~20 sBy default the scope fails fast: the first failure determines the overall result, so the siblings are cancelled immediately, the scope waits for them to finish unwinding, and the failure is reported to the caller after about 50 ms instead of 10 seconds.
solid answer
~60 sWith an all-must-succeed policy, the moment one child fails the combined result is already decided — nothing the other two produce can change it. So the default is **fail-fast sibling cancellation**: cancel the remaining children, wait for them to actually terminate and clean up, then propagate the failure. Why it is the default: - **Latency.** The caller learns in ~50 ms rather than 10 s. Under a request timeout, that is the difference between a fast error and a client timeout. - **Resources.** Ten more seconds of connections, CPU and outbound calls buys nothing. - **No orphans.** Because the scope joins the cancelled siblings before returning, the caller's error arrives after the concurrent work is genuinely finished — no background tasks still writing. - **Locality.** The failure surfaces at the call site, where the caller's normal error handling already applies. Two caveats: cancellation is cooperative, so "cancel then join" takes as long as the slowest sibling's checkpoint plus cleanup; and fail-fast is a *policy*, not a law — when partial results are valuable you deliberately choose a different one.
code
text · 6 linest=0 spawn A (10s), B (fails @50ms), C (10s)
t=50ms B fails -> scope outcome decided: FAILURE
t=50ms scope cancels A and C
t=70ms A and C hit checkpoints, clean up, terminate
t=70ms scope exits, raising B's failure to the caller
# (naive 'wait for all' would raise the same error at t=10s)go deeper
Say the overall result is already known to be a failure, so the other tasks are cancelled and the error is returned quickly instead of after ten seconds.
Add that the scope joins the cancelled siblings before returning (no orphans), that the first failure is the reported cause, and that fail-fast is a default policy you can change.
Quantify it — total time is first failure plus worst sibling wind-down — and discuss uncancellable siblings defeating it, grace deadlines with escalation, side effects that cancellation cannot undo, and interaction with in-child retries.
Frame the completion policy as an explicit design parameter per call site (all-succeed, any-succeed, quorum, best-effort), with fail-fast as the safe default, and make cancellability and grace budgets a requirement rather than an aspiration.
## The situation A scope starts children A, B, C and its contract is "give me all three results". At t=50 ms, B throws. The scope's result is now determined: it cannot produce all three, so it will fail. The only remaining question is what to do with A and C. ## Three possible behaviours 1. **Wait for everyone, then report.** The scope lets A and C run to completion (10 s), then raises B's failure. Correct, but the caller waited 10 seconds for an answer known at 50 ms, holding a request slot, connections and memory the whole time. Under a 2-second client timeout, the caller has already given up and the work is pure waste. 2. **Report immediately, leave siblings running.** The caller gets the error at 50 ms but A and C are now orphans: still holding pooled connections, still issuing outbound calls, still able to write. The caller sees a failed request while side effects continue behind it, and shutdown cannot account for them. 3. **Fail fast: cancel siblings, join them, then report.** Signal A and C, wait for them to reach a checkpoint and clean up, then propagate B's failure. Latency is 50 ms plus the wind-down; no orphans remain. Structured concurrency chooses (3), because (1) wastes time and (2) breaks the ownership invariant that no task outlives its scope. ## Why fail-fast is the right default - **The result is already decided.** With an all-must-succeed contract, continuing sibling work is provably useless. Spending resources on a computation whose output will be discarded is the definition of waste. - **Latency compounds.** In a fan-out tree, waiting-for-everyone at each level multiplies: the request's latency becomes the maximum across every branch even when a failure at one leaf doomed the whole thing early. Fail-fast collapses it to the *first* failure. - **Failure locality.** Because the scope propagates the error at its exit, the caller sees it in its own stack with its own context, and existing try/catch, retry and fallback logic applies unchanged. There is no separate asynchronous error channel to remember to wire up. - **Bounded resource use.** Connections, file handles and memory held by doomed siblings are returned at the moment the outcome becomes known. - **Shutdown correctness.** Joining the cancelled siblings means that when the caller receives the error, the concurrent work is genuinely over — nothing continues to mutate state behind an error response. ## The mechanism, and what it costs Fail-fast is a *policy* implemented with cancellation as its *mechanism*: the failure propagates upward to the scope, and the scope propagates cancellation downward to the remaining children. Because cancellation is cooperative, the scope's total time is: **time-to-first-failure + max over surviving siblings of (checkpoint gap + cleanup)** So a sibling with no checkpoints, or with an unbounded shielded cleanup, erases most of the benefit: you fail fast in intent but still wait ten seconds in practice. Fail-fast is only as fast as the siblings are cancellable, and a grace deadline with escalation is needed for the pathological cases. ## Details that trip people up - **The scope must still join.** Cancelling and returning immediately is behaviour (2) in disguise. "Fail fast" means fast to *decide*, not skip the wait. - **Which error do you report?** The first observed failure is the natural primary cause. The cancellations that follow are consequences and must not be reported as the reason — otherwise every incident log says "cancelled" and the real cause is hidden. When two children fail genuinely independently, both should be preserved rather than one silently dropped. - **Side effects already committed are not undone.** Cancelling A stops future work; it does not roll back the row A already wrote. If the operation needs atomicity across children, you need compensation or an idempotent retry path — fail-fast is not a transaction. - **Retries interact.** If a child is internally retrying a transient error, the scope may cancel it mid-retry. That is usually correct, but it means the sibling's own retry budget is subordinate to the scope's outcome, which should be a conscious decision. ## When fail-fast is the wrong default Whenever the contract is not all-must-succeed: querying several replicas where the first success suffices; aggregating optional enrichments where a missing one just degrades the response; scatter-gather with a quorum; best-effort collection under a deadline. Those need a deliberately different completion policy in which a child's failure becomes a *value* rather than an escaping error. The structure — scope owns children, scope joins them, no orphans — stays identical; only the policy changes.
- If the scope cancels the siblings, why must it still wait for them before returning the error?Because returning early leaves them running: they still hold connections, still perform side effects, and the caller now sees a failed request while work continues behind it. Waiting is what preserves the invariant that no task outlives its scope, so when the error surfaces, all the concurrent work is genuinely finished. In practice the wait is bounded by a grace deadline, after which the scope escalates rather than blocking forever.
- The siblings do not stop for ten seconds even though they were cancelled. What have you actually built?A scope that decides quickly but still returns slowly — fail-fast in intent only. The cause is siblings that are not cancellable: compute loops without checkpoints, blocking calls that ignore the signal, or unbounded shielded cleanup. The fixes are on the child side (add checkpoints, use interruptible or timed calls, bound cleanup) plus a grace deadline at the scope that escalates by closing the resources the stragglers are blocked on.
Three cooks are preparing one dish; when the sauce is ruined and cannot be remade, you stop the other two rather than let them finish a plate that will be thrown away.
saying these in an interview costs you the question
- Reports the failure to the caller while siblings keep running in the background
- Lets all siblings run to completion first and calls that fail-fast
- Reports a sibling's cancellation as the root cause instead of the original failure
- Assumes cancelling siblings undoes the side effects they already committed
- Applies fail-fast to a scatter-gather where partial results were the point