Two of eight parallel worker agents fail — should the orchestrator synthesize, retry, or abort?
answer
- partial failure is the normal case at width
- classify transient versus deterministic
- bounded retry, then degrade or abort
- name the gap, do not hedge vaguely
- completeness tasks abort; survey tasks degrade
basics
~20 sIt depends on whether the missing slices are load-bearing. Retry the two once with a bounded budget; if they still fail, synthesize from the six and state the gap explicitly, and abort only when the task is invalid without complete coverage. Never silently answer from partial data.
solid answer
~50 sFan-in needs an explicit **partial-failure policy**, decided before the run rather than improvised by the lead. The usual ladder is: **retry** the failed workers once, bounded, ideally with a narrowed scope or a cheaper tool path, since many failures are transient — a timeout, a rate limit, a tool error. If the retry fails, ask whether the missing pieces are load-bearing. For a survey-shaped task, **synthesize from six and flag the gap by name** — say which two councils are missing and why — because a partial answer with a declared hole is useful. For a task where completeness is the point, such as a compliance sweep or a security review, **abort or escalate**, because a confident-sounding answer built on 75% coverage is worse than no answer. The unacceptable outcome is the lead quietly writing as though eight returned.
go deeper
Know that some workers will fail in any fan-out and that the lead must not pretend otherwise. Being able to say "retry, then report the gap" is enough at this level.
Explain why partial failure is the common case at width, distinguish transient from deterministic failures, and describe the retry-then-degrade-or-abort ladder with a deadline on the gather step.
Demonstrate return validation for silent failures, a coverage statement in the output, and a clear rule for which task classes must abort rather than degrade. Mention retry safety and side effects unprompted.
Own the reliability framing: define the degradation contract your product offers, decide what coverage level is publishable, and argue why writes stay single-threaded so the retry path stays safe.
## Why this needs a policy, not a judgment call In a fan-out of eight, the probability that all eight succeed is the product of eight independent success rates. At 97% per worker that is about 78% of runs completing cleanly; at 90% it is under half. Partial failure is therefore the *normal* case at any real width, not an exception, and leaving the lead to improvise a response means the behaviour changes run to run. ## Classify the failure first Different failures deserve different handling: - **Transient** — timeout, rate limit, tool error, truncated response. Retry is likely to work; retry with backoff. - **Deterministic** — the source is missing, malformed, or access is denied. Retrying the same dispatch reproduces the failure exactly; either change the approach or record the gap. - **Budget exhaustion** — the worker hit its iteration or token cap mid-task. Often the best return is what it had at exhaustion, which is why workers should be instructed to report partial findings rather than fail silently at the cap. - **Silent failure** — the worst kind: the worker returned something well-formed and confidently empty, or answered a different question. This does not look like a failure at fan-in unless the lead validates returns against the contract. That last case is the argument for **validating returns**: check required fields, non-empty evidence, and that the return references the scope the worker was given. A worker that returns 200 fluent words about the wrong city is a failure the lead will otherwise synthesize straight into the answer. ## The policy ladder 1. **Bounded retry.** Once, maybe twice, with backoff and a cap. Uncapped retries turn a partial failure into a cost incident. Narrowing the scope or downgrading the tool path on the retry often converts a failure into a partial success. 2. **Degrade with a declared gap.** Synthesize from what returned and name what is missing, specifically: which slices, why, and what the reader should not conclude. Generic hedging ("some sources were unavailable") is nearly useless; a named gap lets a human decide whether to care. 3. **Escalate.** Hand the failed slice to a human, to a different tool path, or to a more capable model rather than dropping it. 4. **Abort.** Reserve for tasks where partial coverage is not merely weaker but wrong — anything that answers "is there any instance of X", any compliance or safety sweep, any decision gate. A negative result from an incomplete sweep is a false negative dressed as an answer. ## Timeouts and stragglers A fan-out's latency is the slowest worker's, so the policy also needs a **deadline**: at time T the lead proceeds with what it has and marks the stragglers as gaps. Without one, a single stuck worker can hold an otherwise complete run indefinitely. Pairing a deadline with an instruction to workers to return partial findings on approaching their budget converts most stragglers into usable, if thinner, returns. ## Idempotence and side effects Retries are only safe if the worker's action is repeatable. This is one of the practical reasons the surviving version of this pattern keeps **writes and side effects single-threaded** with the lead or one designated writer: read-only workers can be retried freely, whereas a worker that half-applied a change and then failed leaves state you now have to reason about before retrying. If workers must act, they need idempotency keys or a compensating path. ## What to surface to the user Whatever the policy, the final answer should carry the coverage story: eight slices attempted, six returned, these two missing for this reason. That single line converts an unquantified answer into one a reader can calibrate, and it is the difference between a degraded system and an unreliable one.
- What makes retrying a failed worker unsafe?Side effects. A read-only worker can be retried freely, but one that half-applied a change before failing leaves state you must reconcile first — re-running may duplicate the action. This is a practical reason to keep writes with the lead or a single designated writer and let parallel workers only read, analyse and report; if workers must act, they need idempotency keys or a compensating path.
- How does the lead detect a worker that failed without erroring?By validating returns against the contract rather than trusting them. Check the required fields are present, the evidence is non-empty, and the return actually references the scope that worker was assigned. A confidently written brief about the wrong source is the most damaging failure mode precisely because nothing in the transport layer flags it.
- Where does a deadline fit into the gather step?A fan-out's latency equals its slowest worker, so the lead should proceed at a fixed deadline with whatever has returned and mark stragglers as gaps. Combined with instructing workers to report partial findings as they approach their own budget, this converts most stragglers into thin but usable returns instead of an indefinitely stalled run.
saying these in an interview costs you the question
- Writing the final answer as though all workers returned
- Retrying failed workers without a cap
- Hedging with "some sources unavailable" instead of naming the gap
- Aborting the whole run on any single worker failure
- Retrying a worker that already performed a side effect