skip to content

Sometimes you want the opposite of fail-fast: query five replicas and return the first good answer, or gather whatever four of six services returned before a deadline. How do you design a concurrent scope that deliberately collects partial results, and what guarantees must still hold?

level: seniorimportance: should knowfreq 42%

answer

  1. policy vs structure: only the policy changes
  2. failure as a value, not an escaping error
  3. all / any / k-of-n / deadline / collect-all
  4. losers still cancelled AND joined
  5. label partial as partial; don't cache it

basics

~20 s

Separate the completion policy from the structure: make each child's failure a value rather than an escaping error, so the scope decides. Common policies are all-succeed, any-succeed, k-of-n quorum and best-effort-by-deadline. The scope still owns and joins every child, and partial results must be labelled as partial.

solid answer

~60 s

Keep the structure, change the **completion policy**. The scope still owns every child and still joins them; what changes is when it decides it is done and how a child's failure is treated. The enabling move: a failing child returns a **result value** (success or failure) instead of letting an exception escape. Failure then no longer automatically dooms the scope — the policy decides. Useful policies: - **all-succeed** (fail-fast) — the default. - **any-succeed** — first success cancels the rest; fail only if all fail, reporting the aggregated causes. - **k-of-n quorum** — finish at k successes or when success becomes impossible. - **best-effort by deadline** — at the deadline, cancel stragglers and return whatever arrived. Invariants that survive any policy: no orphans (the scope joins even the losers it cancels), resources released, and — the one people forget — **the result is explicitly labelled partial**. Returning 4-of-6 as though it were complete is how partial data ends up cached, averaged, or written back as if authoritative. With hedged duplicate requests, also require idempotency: several replicas may have done the work.

code

text · 14 lines
text
# any-succeed
with scope() as s:
    hs = [s.spawn(fetch, r) for r in replicas]
    first = s.await_first_success(hs)
    s.cancel_rest(hs)            # scope exit still JOINS the cancelled ones
    if first is NONE: raise AggregateError(all_causes(hs))

# best-effort by deadline
with scope() as s:
    hs = [s.spawn(call, svc) for svc in services]
    s.wait_until(deadline)
    s.cancel_outstanding(hs)     # stop them; do not merely stop waiting
    done = [h.value for h in hs if h.succeeded]
    return Partial(results=done, attempted=len(hs), complete=(len(done)==len(hs)))

go deeper

for a junior

Say the policy changes, not the structure: capture each child's success or failure as a value, decide from those, and still wait for every child.

for a middle

Name the policies (all-succeed, any-succeed, k-of-n, deadline-bounded, collect-all), explain cancelling the losers, and that the result must say it is partial.

for a senior

Discuss deadline derivation from the caller's remaining budget, counting tolerated failures, aggregating causes when everything fails, idempotency and load cost of hedging, and the concrete downstream harms of unlabelled partial data.

for a principal

Make the policy an explicit, reviewed choice per call site with a degradation contract: what partial means for callers, what may not be cached or reconciled, and what signals the system emits when it degrades so silent degradation cannot become permanent.

## The idea: structure is fixed, policy is a parameter Structured concurrency gives one non-negotiable structure — a scope owns its children, and it does not exit until every child has terminated. Layered on top is a **completion policy**: when does the scope stop waiting, and what does a child's failure mean? Fail-fast/all-succeed is only the default. Real systems need several. ## Turning failure into a value The mechanical enabler is to stop letting child exceptions escape. Each child completes with an outcome — success(value) or failure(error) — captured by the scope. Because nothing escapes, nothing propagates automatically, and the scope's policy is free to interpret the mix. This is the difference between "a child failed, therefore we failed" and "a child failed, which is one data point in a decision". ## The policy catalogue **all-succeed (fail-fast).** First failure decides; cancel the rest. Correct whenever every part is required. **any-succeed (invoke-any / first-success).** Start n attempts, take the first success, cancel the rest. Fail only when all n fail — and then report the aggregated causes, because "all replicas failed" with only one error shown is a bad incident. Used for replicated reads, multi-region lookups, and hedged requests. Two subtleties: a *fast failure* must not be mistaken for a decision (one replica refusing connection in 1 ms says nothing about the others), and you must decide whether a non-retryable error (authorisation denied, malformed request) should short-circuit the whole set rather than waiting for four more identical rejections. **k-of-n quorum.** Finish when k successes arrive; fail early when n − failures < k, since success has become impossible. Useful for consensus-style reads and for confidence-weighted aggregation. **best-effort by deadline.** Wait until a deadline, then cancel whatever is outstanding and return what arrived. This is the classic "render the page with the modules that answered" pattern. The deadline is the *policy*, not an error. **collect-all-outcomes.** Never fail the scope; return every child's outcome and let the caller decide. Right for batch jobs where per-item failures are normal and expected. ## What must still hold, whatever the policy 1. **No orphans.** Winners cancel losers, and the scope still joins the losers before returning. Best-effort under a deadline is the most commonly botched case: returning at the deadline while stragglers keep running gives you exactly the leak fail-fast was designed to avoid. Returning "partial" means the outstanding work was *stopped*, not merely ignored. 2. **Resources released.** Each cancelled loser runs its cleanup; the scope's grace budget covers it. 3. **Failures are not silently swallowed.** Even when the policy tolerates them, they must be counted and attributed — per-child failure metrics and, on total failure, the aggregate of causes. A quiet degradation nobody measures becomes a permanent one. 4. **Partiality is explicit in the result.** The single most damaging bug in this area is returning a partial result with the same shape as a complete one. Downstream then averages over a subset, treats missing items as deleted, caches an incomplete list, or writes the subset back as authoritative. Return a completeness marker — counts of attempted vs succeeded, or a flag — and enforce the consequences: do not cache partial results (or cache them with a short TTL and a distinguishing key), do not use them for delete-by-diff or reconciliation, and surface degradation to the caller. 5. **Idempotency where attempts duplicate.** Hedging and any-succeed mean several replicas may perform the operation. That is fine for reads and dangerous for writes unless the operation is idempotent or deduplicated by a request key. The cost is also real: n× load on downstreams, so hedge on a delay (only send the second attempt after p95) rather than always. ## Choosing between them Ask what the caller can do with a partial answer. If a missing part makes the response wrong (a payment total, an authorisation decision), the policy is all-succeed. If it makes the response merely worse (a recommendations panel, an enrichment field), best-effort with an explicit degradation signal is better than failing the whole request. If several sources are equivalent, any-succeed converts variance in one dependency into a much tighter latency distribution. If correctness needs agreement, quorum. ## Interaction with deadlines and supervision A best-effort policy needs a deadline to be meaningful, and that deadline should be derived from the caller's remaining budget rather than a hard-coded constant, so that nested scopes do not each consume the full allowance. Supervision is a different axis: it decides what happens to a *long-lived* child that dies (restart, ignore, escalate), while completion policy decides how a *bounded* set of children combine into one result.

  • What is the most common bug in a best-effort partial-results implementation?
    Returning at the deadline without cancelling and joining the outstanding children. The caller gets its partial answer while the stragglers keep holding connections and producing side effects, which is precisely the orphan problem structured concurrency exists to prevent. The second most common is returning the partial result in the same shape as a complete one, so downstream code caches it, averages over it, or reconciles against it as if it were authoritative.
  • When is an any-succeed or hedged policy a bad idea?
    When the operation is not idempotent, because several attempts may each apply a write, or when the downstream cannot absorb the multiplied load — hedging every request can add close to n× traffic exactly when the system is already slow. Mitigations are hedging only after a delay near the p95 latency, capping the hedge rate, and deduplicating by a request key so repeated attempts collapse to one effect.

Sending five couriers with the same letter: you keep the first receipt and recall the others — but you must actually recall them, and you must record that only some routes were confirmed.

saying these in an interview costs you the question

  • Returns partial results at the deadline while the remaining tasks keep running
  • Returns a partial result with the same shape as a complete one, so callers cannot tell
  • Caches or reconciles against partial data as if it were authoritative
  • Swallows tolerated child failures without counting or attributing them
  • Hedges duplicate attempts against a non-idempotent write
  • Treats one fast connection refusal as proof that the whole replica set is unavailable

context