skip to content

How would you design a LangGraph approval gate for irreversible agent actions?

level: principalimportance: should knowfreq 40%

answer

  1. Gate by blast radius, not by node
  2. Pause before the effect, always
  3. Payload is a UI contract
  4. Nobody answers is a design case
  5. Approver identity lives outside state

basics

~20 s

Gate by risk, not by node: interrupt only for actions that are irreversible or above a threshold, pause before the effect happens, ship the reviewer a payload they can actually judge, and accept approve, edit and reject-with-reason. Back it with a durable checkpointer and a timeout policy.

solid answer

~60 s

Start from the action, not the graph. Classify tools by blast radius and interrupt only for the irreversible or high-value ones — gating everything trains reviewers to click approve without reading, which is worse than no gate. Put the pause **before** the side effect, either with `interrupt_before` on the tool node or a conditional `interrupt()` inside a wrapper, and keep the gating node free of effects so node replay on resume cannot duplicate anything. Design the payload as a product surface: action name, arguments, the agent's rationale, and the identifiers a reviewer needs, with secrets excluded. Support three answers — approve, edit (rewrite arguments via `update_state`), and reject with a reason fed back as a `ToolMessage` so the model can re-plan. Because a pause is durable state on a thread, use a persistent checkpointer, treat thread ids as queue items, and define what happens when nobody answers: expire the thread and take the safe default rather than leaving it open forever. Record who approved outside the graph, since checkpoints store values, not actors.

go deeper

for a junior

Know that agents which can do irreversible things should ask a person first, and that LangGraph implements that as a pause before the acting node rather than a blocking wait.

for a middle

Be able to build the gate: pause before the tool node, surface the proposed call, and support approve, edit and reject, answering a rejected call with a ToolMessage so the model is not left with a dangling request.

for a senior

Show operational thinking — persistent checkpointer, thread ids as queue items, timeout and expiry policy, no side effects in the gating node because of replay, and approval latency accounted for separately in your SLAs.

for a principal

Own the policy: which action classes justify a human at all, thresholds versus blanket gates, separation of duties and audit outside graph state, and when a scoped credential, spend cap or reversible operation is a better control than spending a reviewer's attention.

## Frame it as a control, not a feature The interview question is not "can you call `interrupt()`" — it is whether you can design a control that a real operations team will run for a year. Five decisions matter. ## 1. What gets gated Enumerate the agent's tools and sort them by reversibility and blast radius. Read-only lookups: never gate. Reversible writes (draft created, ticket opened): usually not, or gate above a threshold. Irreversible or externally visible actions (payment, email to a customer, deletion, deploy): gate. Money and record counts deserve thresholds rather than blanket rules — approve under $50 automatically, interrupt above. The failure mode of over-gating is well documented in any alerting system: reviewers habituate, approve reflexively, and the gate becomes a latency tax that provides no safety. Under-gating is the obvious failure. Getting the boundary right is the actual engineering. ## 2. Where the pause sits It must sit **before** the effect. Two placements work. `interrupt_before=["tools"]` stops the run with the model's proposed call sitting in state, unexecuted — simple, but unconditional. A wrapper node that inspects the proposed call and calls `interrupt()` only when policy says so gives you risk-proportional gating plus a structured payload, at the cost of a little code. Whatever you choose, keep the gating node's body free of side effects. Resuming replays the node from its first line, so anything above the `interrupt()` call runs twice. The disciplined shape is: decide in the gate node, act in the next node. ## 3. The payload and the decision vocabulary The payload is an API between your graph and your approval UI, so version it and keep it stable. Include the tool name, the arguments in full, a human-readable summary of what will happen, the agent's stated reason, and the identifiers (thread id, request id, customer) a reviewer needs to check the claim. Exclude credentials and anything the reviewer is not cleared to see — the payload is persisted in the checkpointer and rendered in a UI. Accept at least three decisions. **Approve** proceeds. **Edit** rewrites the arguments — `update_state` with a corrected message reusing the original message id, then continue. **Reject with reason** must feed a `ToolMessage` back against the pending `tool_call_id` so the model sees a real tool result and can re-plan; leaving a requested call unanswered breaks the next model turn. A fourth, **escalate**, is worth having when a reviewer is not authorized for this class of action. ## 4. Durability, latency and expiry A LangGraph pause holds no thread and no connection — it is rows in the checkpointer. That is the property the whole design leans on: use a persistent checkpointer so a deploy or crash does not lose pending approvals, and treat `thread_id` as the item id in the approval queue. Approvals take hours; the graph does not care. But *nobody answers* is a state you must design, not discover. Pick an expiry per action class, run a sweeper over pending threads, and define the timeout outcome — for irreversible actions the safe default is auto-reject with a reason, never auto-approve. Notification matters too: a queue nobody is paged for is a queue with a growing tail. And the human latency is now inside your end-to-end SLA, so report those runs separately from autonomous ones or your p99 becomes meaningless. ## 5. Authorization and audit Graph state records values, not actors. Authorization belongs in the service that calls `resume` — verify the approver is entitled to approve *this* action class and this tenant before the resume ever reaches LangGraph, and write an audit record (who, when, what payload, what decision, what edits) in your own store, keyed by thread id and checkpoint id. Checkpoint history gives you the before-and-after state for free, which is genuinely useful evidence, but it does not tell you who did it. Separation of duties is worth stating: the person who launched the run should generally not be the person who approves its irreversible step. ## 6. Know when a gate is the wrong tool Approval is expensive — it costs a human's attention on every invocation. Cheaper controls often dominate: scoping credentials so the destructive action is impossible, spend caps enforced by the downstream system, dry-run-then-diff, making the operation reversible with an undo window, or sampling for post-hoc review instead of blocking every call. A principal-level answer names the gate as one instrument in that set and explains the conditions under which it is the right one: high-consequence, low-frequency, and genuinely un-undoable actions where a human can add judgment the model cannot.

  • What should happen when an approval request expires with no human response?
    Define it per action class and enforce it with a sweeper over pending threads. For irreversible actions the safe default is auto-reject: resume the thread with a rejection, feed a ToolMessage explaining the timeout so the model can report or re-plan, and alert the owner. Auto-approve on timeout inverts the control entirely — it means the gate stops protecting you exactly when the team is overloaded, which is when you need it most.
  • Why is it risky to gate every tool call rather than a risky subset?
    Habituation. When most requests are trivially fine, reviewers stop reading and approve reflexively, so the dangerous one passes with the same click as the harmless ones — you have paid full latency cost for near-zero safety. It also makes the human the throughput ceiling for the whole agent. Gate on reversibility and value thresholds so that seeing a request is itself a signal that something warrants attention.
  • Where do you enforce that the approver is allowed to approve this action?
    In the service that calls resume, before LangGraph is involved. Checkpoints store values, not actors, so the graph has no notion of identity and cannot authorize anything. Your endpoint authenticates the reviewer, checks entitlement for that action class and tenant, writes an audit record keyed by thread and checkpoint id, and only then invokes the graph with the resume value. Treat the resume endpoint as privileged and rate-limited.

saying these in an interview costs you the question

  • Gates every tool call and calls it defence in depth
  • Auto-approves when an approval request times out
  • Relies on graph state to record who approved
  • Places the pause after the side effect has run
  • Puts the API call above interrupt() in the same node

context