skip to content

How should an orchestrator handle a subagent that stalls and never returns?

level: seniorimportance: should knowfreq 42%

answer

  1. stalled and slow look identical from outside
  2. the supervisor decides, not the worker
  3. time since last progress, not total time
  4. return four of five, gap marked
  5. cancelled does not mean dead

basics

~20 s

Termination belongs to the orchestrator, not the worker. Dispatch every subtask with a deadline, cancel on expiry, and degrade gracefully: return the completed results with the stalled scope explicitly marked unresolved rather than blocking or failing the whole run.

solid answer

~50 s

A stalled worker is externally indistinguishable from a slow one, so the orchestrator must decide, not the worker — a worker that is looping or wedged is exactly the component least able to judge that it should stop. Attach a deadline when you dispatch, and prefer a progress signal over pure wall-clock: a worker that has emitted no new tool call or artifact for N seconds is stalled even if its total time is within budget. On expiry, cancel it, optionally retry once with backoff if the work is idempotent, then **degrade gracefully** — return the four subagents that finished plus an explicit gap for the fifth, rather than blocking indefinitely or discarding good work. Two details bite: the cancelled worker may still be alive, so guard writes with a generation token, and the partial result must be labelled as partial, since a silently truncated answer is worse than a slow one.

go deeper

for a junior

Know that subtasks are dispatched with a deadline and that something above the worker has to notice when it never comes back. Say plainly why the stuck worker cannot be trusted to report itself stuck.

for a middle

Explain why wall-clock deadlines alone are insufficient and how a no-progress threshold based on streamed output or tool-call boundaries separates a stalled worker from a slow one.

for a senior

Walk the full policy: cancel, bounded retry only if idempotent, reassign or downscope, then degrade with the missing scope explicitly marked. Raise the zombie-write problem and fence it with a generation token.

for a principal

Own the tradeoff between latency, completeness and spend — how tight deadlines trade truncated work against a wedged worker burning an order-of-magnitude token budget — and set the policy for when a partial answer may be shipped at all.

## Why the orchestrator owns termination In a fleet, the component that has stalled is the component least equipped to notice. An agent looping between two tools, waiting on a hung connection, or wedged mid-generation has no reliable outside view of itself. Time-based self-limits inside the worker help, but they fail exactly in the cases you care about — the process that is dead, partitioned, or blocked on a call that never returns. So termination authority sits with the orchestrator, which dispatched the work, holds the deadline, and owns the user-visible result. This is a straightforward distributed-systems position: the supervisor decides when a subordinate is done, and "done" includes "given up on". ## Detecting a stall A wall-clock deadline per subtask is the baseline and is easy to get wrong, because legitimate agent work is highly variable — one search returns in two seconds, another crawls a large repository for three minutes. A deadline tight enough to catch stalls quickly will kill honest work; one loose enough to be safe leaves the run hanging. Progress signals are better. Streamed output, tool-call boundaries, or explicit heartbeats give a second clock: time since last observable progress. A worker silent for longer than the progress threshold is stalled regardless of how much of its total budget remains. Combine both — a generous overall deadline plus a tight no-progress threshold — and you catch wedged workers fast without truncating slow-but-healthy ones. A third signal is semantic: repeated identical states. A worker issuing the same tool call with the same arguments in a cycle is not progressing even though it is emitting activity. That looks like liveness to a heartbeat and like a loop to anyone reading the trace. ## What to do when the deadline fires **Cancel explicitly.** Abandon the request rather than merely stopping waiting, so you stop paying for generation and free the slot. **Decide whether to retry.** If the subtask is idempotent and the failure looks transient, one retry with backoff is reasonable — with a lower attempt cap than for ordinary errors, since a stall usually means something structural. If the work has already applied external side effects and is not idempotent, do not retry; escalate. **Reassign or downscope.** Sometimes the right move is to re-dispatch the subtask with a narrower scope or to a different worker configuration, on the theory that the original prompt or tool set is what wedged it. **Degrade gracefully.** This is the part teams skip. If four of five parallel subagents returned, the run should produce the four results plus an explicit statement that the fifth's scope is unresolved. Blocking forever is the worst outcome; failing the whole run discards good work; silently omitting the fifth's scope is the dangerous outcome, because the consumer reads a complete-looking answer that is missing a fifth of its evidence. Partial results must be labelled partial, with the missing scope named, so both a human and a downstream agent can decide whether the partial answer is usable. **Decide whether the run can continue.** If the stalled subtask was on the critical path — later steps depend on its artifact — degradation is not available and the orchestrator should fail that branch explicitly and stop dependents. If it was one of several independent contributions, continue. ## The zombie problem Cancelling a worker does not guarantee it is dead. A partitioned or slow worker can return long after its deadline and write an artifact, clobbering work the orchestrator has since done or reintroducing state the run has moved past. The standard guard is a fencing mechanism: each dispatch carries a generation or epoch number, the artifact store accepts writes only from the current generation, and a late write from an older generation is rejected. Without it, timeouts create a second writer for an artifact you had carefully given one owner. ## Blast radius and budget Orchestrator-owned termination is also a cost control. Multi-agent orchestration burns dramatically more tokens than single-turn chat — research-style fan-out runs at roughly an order of magnitude above ordinary usage — so a wedged worker consuming its full budget in a loop is a real expense, not just latency. The orchestrator should hold the overall run deadline, kill the run when it expires, and return whatever has been assembled. ## What good looks like in the trace After the fact you want to be able to read: dispatched at T with deadline D, last progress at T+x, cancelled at T+D, retried once, failed, marked scope unresolved, run completed with four of five contributions. Every one of those transitions is orchestrator-side and recorded. A run where a worker simply never appears again, with no record of the decision to give up on it, is a system without termination policy — it just has a timeout somewhere in a client library.

  • How do you distinguish a stalled subagent from a legitimately slow one?
    Use time since last observable progress rather than total elapsed time. Streamed tokens, tool-call boundaries or heartbeats give that second clock, so a worker silent for the no-progress threshold is stalled even with budget remaining. Repeated identical tool calls are a third signal: activity without progress. A generous overall deadline plus a tight no-progress threshold catches wedged workers without killing honest long work.
  • What should the orchestrator actually return when one of five parallel subagents never completes?
    The four completed contributions, plus an explicit, machine-readable marker that the fifth's scope is unresolved and why. Blocking is the worst outcome and failing the whole run discards good work, but silently dropping the scope is the most dangerous, because the answer looks complete while missing a fifth of its evidence. Downstream consumers need the gap named to judge usability.
  • Why can a cancelled subagent still corrupt the run, and how do you prevent it?
    Cancellation stops you waiting; it does not prove the worker is dead. A partitioned worker can return later and write an artifact the run has already moved past. Fence it: each dispatch carries a generation number, and the store accepts writes only from the current generation, rejecting late writes from older ones. Otherwise your timeout policy quietly creates a second writer.

saying these in an interview costs you the question

  • Letting the worker decide when it has stalled
  • Using only a wall-clock deadline, with no progress signal
  • Blocking the whole run until the stalled subagent returns
  • Dropping the failed subagent's scope without telling the consumer
  • Assuming a cancelled worker can no longer write anything

context