skip to content

How do you pause an AI agent for human approval that may take days?

level: seniorimportance: must knowfreq 58%

answer

  1. the wait may outlive the process
  2. state lives entirely in the conversation
  3. persist, exit, resume later
  4. checkpoint keyed by a run id
  5. replay can repeat side effects

basics

~20 s

Persist the loop's state to a durable checkpoint, emit the approval request, and let the process exit. When the human answers — hours or days later, on a different machine — load the checkpoint, inject the decision, and resume from the interrupt point. Never block a thread.

solid answer

~50 s

The naive implementation blocks the worker until a callback arrives, which fails the moment you deploy, restart, or scale down: the whole run lives in process memory and disappears. The durable pattern is interrupt-and-resume. At the gate the harness raises an interrupt, a checkpointer writes the complete loop state — message history, pending tool call and its exact arguments, plan, scratchpad, cursor — to durable storage under a run id, and the process exits holding nothing. The approval request carries that run id. When the human answers, a resume path loads the checkpoint, injects the decision as the interrupted step's result, and the loop continues. Two things bite. Resume replays from a checkpoint boundary, so any side effect executed before the interrupt inside that same step can run twice — put the gate before side effects or make them idempotent. And the world moved while you were paused, so revalidate preconditions on resume instead of trusting a three-day-old observation.

code

python · 18 lines
python
import json, uuid

CHECKPOINTS = {}  # stands in for durable storage

def suspend_for_approval(state, pending_call):
    run_id = str(uuid.uuid4())
    CHECKPOINTS[run_id] = json.dumps({"state": state, "pending": pending_call})
    return run_id  # process may now exit; nothing is held in memory

def resume(run_id, decision):
    saved = json.loads(CHECKPOINTS[run_id])
    if decision != "approve":
        return {"executed": False, "reason": "rejected"}
    return {"executed": True, "call": saved["pending"]}  # stored args, not a new model turn

rid = suspend_for_approval({"messages": []},
                           {"tool": "wire_transfer", "args": {"usd": 75000}})
print(resume(rid, "approve"))

go deeper

for a junior

Know that an agent waiting for approval must save its state somewhere durable and stop, rather than sitting in memory waiting, because the wait can outlast the program.

for a middle

Explain the mechanics: serialize message history and the pending call under a run id, exit, then load and resume when the decision arrives, possibly on a different process.

for a senior

Demonstrate the production concerns — idempotent side effects across step replay, revalidating preconditions after a multi-day pause, reassigning or expiring stale runs, and the retention rules checkpoints inherit.

for a principal

Own the platform decision: whether durable execution is a shared capability every agent inherits or something teams reimplement, and what that means for tenant isolation, retention and operability across a fleet of long-running runs.

## Why the pause is the hard part Deciding *what* to gate is a policy problem. Actually stopping is an engineering problem, and it is the one interviewers probe, because an agent's entire state is a conversation: an ordered list of messages, a pending tool call, a plan, whatever notes it wrote to a scratchpad. None of that is in a database row you can pick up later unless you deliberately put it there. The span of a human approval is the forcing constraint. An approver may be in a meeting, asleep in another timezone, or on leave until Thursday. Any design that assumes the wait is measured in seconds is wrong at the first real incident. ## What blocking costs you The first implementation everyone writes is a blocking call inside the loop: post the request, wait on a future or poll a table, continue when the answer lands. It works in a demo and fails in production for four reasons. - **Deploys and restarts destroy the run.** Ship on Tuesday afternoon and every pending approval evaporates, mid-plan, with the user's work lost. - **Resources are pinned.** A worker, a connection, sometimes a sandbox, held for three days per pending decision. Concurrency collapses to the number of humans who have not answered yet. - **It does not survive scale-down or preemption.** Autoscalers and spot instances assume workers are disposable. - **There is no external handle.** Nobody can list what is pending, reassign it, or cancel it, because the state exists only inside one process's stack. ## The interrupt-and-checkpoint pattern The durable version inverts control. When the tool router hits a gated call it does not wait; it *suspends*: 1. Serialize the loop state to durable storage keyed by a run id: full message history including the assistant turn containing the pending call, the tool name and literal arguments, the plan or todo list, scratchpad references, and where in the graph or loop execution stopped. 2. Write an approval request row — run id, rendered payload, requester, policy that fired, deadline — and notify the approver. 3. Return. The process is now free; nothing is held. Resumption is a separate entry point, usually an HTTP handler behind the approval UI: load the checkpoint by run id, validate that the decision is still admissible, splice the decision in as the outcome of the interrupted step, and run the loop forward. It may execute on a different pod, a different deploy, a different week. Production agent frameworks expose exactly this shape. In LangGraph, `interrupt()` raises inside a node while a checkpointer — commonly Postgres-backed — persists the graph state under a thread id, and a resume command carries the human's value back into the same point. The names differ elsewhere; the pattern does not. The checkpoint store is the mechanism that makes a pause survivable, and everything else is ergonomics. ## Replay and idempotency The subtle failure is double execution. Checkpoints are taken at step boundaries, so resuming typically re-executes the interrupted step from its start, not from the exact instruction after the interrupt. If that step posted an invoice, sent a webhook, or incremented a counter *before* reaching the gate, resume repeats it. Two defences, and you want both: structure steps so the approval gate comes before any side effect in that step, and make external calls idempotent with a key derived from the run id and call index so a repeat is a no-op. This is also why the approval decision itself must be recorded transactionally — a resume that half-applies is worse than one that fails cleanly. ## The world moves while you are paused A three-day pause invalidates the observations the plan was built on. The invoice was already paid manually. The index the agent wanted to drop is now serving a hot query path. The customer closed their account. Time-of-check to time-of-use is not a theoretical concern at human latency; it is the normal case. So resume should revalidate rather than trust: re-read the preconditions the gated action depends on, compare them to what the approver saw, and if they diverge materially, fail the resume back to a fresh approval rather than executing quietly. Pair that with an expiry on the request itself, so a decision cannot be applied against a world that no longer resembles the one it was made in. ## Operational surface Once runs are durable they become inspectable, and that is most of the payoff. You can list every pending approval with its age, reassign one when the owner goes on leave, cancel a stale run, resume in bulk after an outage, and replay a checkpoint into a debugger to see exactly what the agent was about to do. Checkpoints also need lifecycle care: they contain conversation history and tool arguments, so they inherit your retention, encryption and tenant-isolation rules, and abandoned ones need a reaper.

  • What exactly has to be in the checkpoint for a resume to be correct?
    Everything the next model call needs plus everything the executor needs: the full message history including the assistant turn that proposed the gated call, the tool name and its literal arguments, the plan or todo state, references to any scratchpad or files the agent wrote, and the execution cursor. Anything you reconstruct instead of storing — a cached retrieval, a computed budget — has to be either deterministic or explicitly refreshed on resume.
  • How do you make sure a resume executes the action the human actually approved?
    Resume from the stored arguments, not from a fresh model turn. Re-prompting the model after approval invites it to generate a slightly different call, which means the approval binds to nothing. Load the pending call from the checkpoint, verify it still satisfies policy and that its precondition checks pass, then execute that exact payload and record the linkage between the approval record and the executed call.
  • An approval sits unanswered while the agent's underlying model version is deprecated. What do you do on resume?
    Treat it like any other precondition drift. A checkpoint pins conversation state, not the runtime, so resuming on a different model can change behaviour after the approved step. For the approved action itself that is fine — you execute stored arguments, not a new inference. For everything after, either resume on a pinned version if you still have it, or expire the request and force a fresh plan rather than continuing a trajectory the new model did not author.

saying these in an interview costs you the question

  • Blocks a worker thread or process while waiting for the human
  • Keeps pending-approval state only in memory or in a local queue
  • Re-prompts the model after approval instead of executing stored arguments
  • Ignores that a replayed step can repeat side effects
  • Assumes the world is unchanged when a run resumes days later

context