A subagent dies mid-task; how does the orchestrator resume without redoing work?
answer
- record completed steps outside the process
- resume at step 7, not step 1
- failures are ambiguous, not clean
- at-least-once means idempotent steps
- backoff with jitter, cap the attempts
basics
~20 sPersist completed steps and their outputs to a durable store outside the agent's process, keyed by step. On restart, skip steps already recorded and re-run only the rest — which requires every step's side effects to be idempotent, since a step may have applied before it died.
solid answer
~50 sTreat the run as a sequence of recorded steps rather than one long conversation. After each completed step the orchestrator writes a durable checkpoint — the step identifier, its output or artifact reference, and enough of the plan to continue — to storage that survives the process. When a subagent is killed at step 7 of 12, the orchestrator relaunches it against that checkpoint and it resumes at step 7 instead of step 1, which matters because replaying a long agent trajectory is expensive and nondeterministic. The hard part is not persistence but **idempotency**: a step that failed ambiguously may already have applied its side effect, so retries need idempotency keys, conditional writes, or a natural upsert. Pair this with bounded retries and exponential backoff with jitter, and a per-step attempt cap so a genuinely poisoned step escalates instead of looping forever.
code
python · 22 linesimport json, os, time
def run_step(step_id, work, checkpoint_path, max_attempts=3):
done = json.load(open(checkpoint_path)) if os.path.exists(checkpoint_path) else {}
if step_id in done:
return done[step_id]
delay = 1.0
for attempt in range(max_attempts):
try:
result = work(idempotency_key=step_id)
break
except Exception:
if attempt == max_attempts - 1:
raise
time.sleep(delay)
delay *= 2
done[step_id] = result
with open(checkpoint_path, "w") as f:
json.dump(done, f)
return resultgo deeper
Know that long agent runs save progress so a crash does not lose everything, and that the saved state has to live somewhere outside the agent itself.
Explain what a checkpoint contains — step identity, output or artifact reference, plan position — and why resuming beats replaying: side effects were already applied and a resampled run will not retrace the original.
Show that ambiguous failures force at-least-once semantics, and give the concrete idempotency mechanisms: keys, conditional writes on a known version, write-ahead intent records. Add bounded retries with jittered backoff and a per-step attempt cap.
Decide where checkpointing earns its complexity — long, expensive, externally-visible runs — and where re-running is simply cheaper. Own the interaction with write ownership, including generation tokens that stop a resurrected worker from clobbering the resumed run.
## The problem Agent runs are long, expensive and nondeterministic. A subagent that has done eleven tool calls and produced a partial artifact represents real spend. When its process is killed — an out-of-memory kill, a host restart, a provider outage, a network partition — restarting from zero throws that away, and because the model is sampled, the second run will not retrace the first. Long agentic work therefore needs the same property that long-running workflows in any distributed system need: the ability to resume from a recorded position. ## What a durable checkpoint is A checkpoint is a record, written to storage outside the agent's process, that captures enough state to continue. Durable means it survives the death of the thing that wrote it: a database row, an object in a store, a file in a shared volume — not in-memory conversation state, and not the agent's own context window, which dies with it. What to record: - **Step identity and status.** Which unit of work, and whether it completed. - **The step's output**, or a reference to the artifact it produced. Small results inline; large ones by reference. - **The plan position.** Which steps remain and what they depend on. - **Provenance.** Who produced it, when, under what attempt number. What generally not to record: the entire raw trajectory. Replaying thousands of tokens of prior reasoning to "restore" an agent is expensive and often counterproductive — the resumed agent works better from a compact summary of what was decided plus the artifacts produced than from a verbatim transcript. Checkpoint at the granularity of decisions and side effects, not at every token. A useful default is: after each subtask completes, and immediately before and after any irreversible external action. ## Resume, replay, and the cost of getting it wrong On relaunch the orchestrator reads the checkpoint, skips steps already marked complete, and dispatches from the first incomplete step. The subagent receives the accumulated outputs as inputs. This is resume. Replay — re-running earlier steps for their side effects — is almost never what you want in agent systems, precisely because side effects were already applied. Which leads to the real difficulty. ## Idempotency is the hard requirement A failure is usually *ambiguous*: the orchestrator knows the call did not return, not whether it took effect. The request may have been applied and the acknowledgement lost. So retry semantics are at-least-once, and correctness comes from making steps idempotent rather than from pretending delivery is exactly-once. Practical mechanisms: - **Idempotency keys.** The step generates a stable key from the task and step identity, and the downstream system deduplicates on it. Many external services support this directly. - **Conditional writes.** Write only if the artifact is still at the version you read, so a duplicate attempt fails harmlessly instead of double-applying. - **Naturally idempotent operations.** Setting a value rather than incrementing it; upserting rather than inserting; declaring desired state rather than issuing a delta. - **Write-ahead intent.** Record "about to do X with key K" before doing it. On resume, look up K to determine whether X happened, rather than guessing. Where a step cannot be made idempotent — sending a notification, charging money, destroying infrastructure — do not retry it automatically. Record it as needing a decision and escalate. ## Retry policy at the coordination layer Retries belong in a policy, not in the model's judgement. Bound the attempts per step, back off exponentially with jitter so a fleet of subagents does not synchronize into a thundering herd against a rate-limited provider, and distinguish retryable failures — timeouts, transport errors, throttling — from failures that will recur no matter how often you retry, such as a malformed request or a permission denial. A per-step attempt cap keeps a poisoned step from consuming the whole run's budget; when it trips, the orchestrator marks that branch failed and continues with the rest rather than dying. ## Where this sits architecturally Checkpoint state belongs to the orchestrator, not to the worker. A worker that owns its own resume state cannot be resumed when it is the thing that died. The durable-graph frameworks that took hold by 2026 make exactly this choice: the graph and its checkpointer live outside the agent, so a crashed or paused run can be picked up later, and the same machinery serves both crash recovery and deliberate pauses. One subtlety worth naming: a resumed worker and a presumed-dead original can both be alive at once. If the original wakes up and writes, it clobbers the resumed run's work. Guard against it with a generation or epoch token on writes — the store accepts writes only from the current generation — which is also why single-writer discipline and resume policy have to be designed together. ## What it costs Checkpointing adds storage, write latency at each boundary, and design work to make steps idempotent. For short, cheap, reversible runs it is not worth it — just re-run the task. It becomes worth it when runs are long, expensive, touch external systems, or must survive deliberate interruption.
- What makes a checkpoint durable enough to resume a dead subagent from?It must live outside the process that dies: a database row, an object store entry, a shared file — written and flushed before the step is considered complete. In-memory conversation state and the agent's own context window do not qualify, because they vanish with the worker. The checkpoint should also be owned by the orchestrator, since a worker cannot restore itself.
- Why does resumption depend on side effects being idempotent?Because a failure is usually ambiguous: the call did not return, but it may still have been applied and the acknowledgement lost. Retrying is therefore at-least-once. Idempotency keys, conditional writes on a known version, or naturally idempotent upserts make a duplicate attempt harmless. Steps that cannot be made idempotent — payments, destructive operations — should escalate instead of retrying.
- What stops a presumed-dead subagent from clobbering the resumed run if it wakes up?A generation or epoch token. Each relaunch increments it, and the artifact store accepts writes only from the current generation, so the zombie's late write is rejected rather than applied. Without it you have two writers for the same artifact — which is exactly the situation single-writer discipline exists to prevent.
saying these in an interview costs you the question
- Restarting a crashed agent from step one and calling that recovery
- Storing resume state inside the agent that might die
- Retrying every failure, including permission denials and malformed requests
- Assuming a failed call definitely did not take effect
- Retrying without backoff or jitter across a fleet of subagents