skip to content

Why does a LangGraph node re-run from the top after interrupt(), and what breaks?

level: seniorimportance: must knowfreq 50%

answer

  1. State is saved, stack is not
  2. Above the line runs twice
  3. Effects belong in a later node
  4. Interrupt first, compute after
  5. Idempotency key when unavoidable

basics

~20 s

LangGraph checkpoints state, not the Python stack, so resuming replays the interrupted node function from its first line. Any side effect executed before interrupt() — an API call, a write, an email — therefore happens twice. Put interrupt() first, or isolate side effects in a separate node.

solid answer

~60 s

When `interrupt()` pauses a node, LangGraph saves the graph's **state channels**, not the interpreter frame. There is no way to re-enter a half-finished Python function, so on `Command(resume=...)` the runtime re-invokes the whole node function against the same checkpointed input; the `interrupt()` call then returns the stored resume value instead of raising. Everything above that line runs a second time. If those lines charged a card, sent a message, inserted a row, or appended to an external log, you get a duplicate — and if they computed something non-deterministic (a timestamp, a UUID, a fresh LLM call), the second value is the one that survives, so state may not match what the human approved. The fixes are structural: call `interrupt()` as early as possible in the node, move side effects into a node *after* the approval node so the pause never straddles them, and make unavoidable effects idempotent with a stable key. The same replay applies to a node that contains a subgraph, and to the ordering of multiple `interrupt()` calls, whose resume values are matched positionally.

code

python · 16 lines
python
# Hazard: charge_card() runs on the first pass AND again on resume.
def pay(state: State) -> dict:
    receipt = charge_card(state["amount"])          # side effect too early
    ok = interrupt({"amount": state["amount"]})
    return {"receipt": receipt, "paid": ok == "yes"}


# Fix: the gate holds no side effect; the effect lives in its own node.
def approve(state: State) -> dict:
    return {"approved": interrupt({"amount": state["amount"]}) == "yes"}


def pay_now(state: State) -> dict:
    if not state["approved"]:
        return {"receipt": None}
    return {"receipt": charge_card(state["amount"])}

go deeper

for a junior

Remember that resuming replays the node from its beginning, not from the interrupt line. That means anything written above the interrupt runs a second time — so keep such nodes free of real-world effects.

for a middle

Explain the mechanism: LangGraph persists state channels, not interpreter frames, so the function is re-invoked and interrupt() returns the stored value on the second pass. Be able to point at which lines run twice and which run once.

for a senior

Demonstrate that you design around it — approval node separate from effect node, interrupt() first, non-deterministic values computed after the pause, idempotency keys for anything you cannot move. Mention model calls above the line as a real cost and correctness issue, not just a style point.

for a principal

Frame replay as the standard cost of durable execution and set the engineering rules that follow: which effects may live in graph nodes at all, how idempotency keys are derived, and how in-flight threads are drained before the interrupt shape of a node changes.

## The mechanic LangGraph's durability model is state-based: at each superstep it persists the values of the state channels to the checkpointer. It does not, and cannot portably, persist a Python call stack — a paused run may be resumed by a different process, on a different machine, days later, possibly after a redeploy. So "resume" means: load the checkpointed state, and call the pending node's function again from the beginning. Inside that second call, `interrupt()` behaves differently. The runtime carries a list of resume values for the task; the first `interrupt()` call consumes the first value and returns it rather than raising. Execution then proceeds past the line that previously paused, reaches the node's `return`, and its update is committed for the first time. That asymmetry is the whole hazard: **the code above `interrupt()` executes twice, the code below it once.** ## What actually breaks **Duplicated external effects.** A node that posts to an API, writes a row, sends a Slack message or increments a counter *before* asking for approval does that thing on the first pass — while the human is still deciding — and again on resume. In an approval flow this is exactly backwards: the point of the gate was to not do the thing yet. **Non-determinism drift.** Values computed before the interrupt are recomputed. `datetime.now()`, `uuid4()`, a random sample, or another LLM call will produce different results on the replay, and the replayed value is what lands in state. The human approved a payload derived from the first computation; the graph proceeds with the second. Silent, and very hard to debug. **Wasted cost and latency.** An expensive retrieval or a large model call sitting above the interrupt is paid for twice per approval. **Mismatched interrupt ordering.** Because multiple `interrupt()` calls in one node are matched to resume values positionally, adding, removing or reordering interrupt calls in that node while threads are paused will hand stored answers to the wrong questions on resume. **Subgraph replay.** If a node wraps a subgraph and something inside the subgraph interrupts, resuming replays from the parent node's entry as well as re-entering the subgraph at the interrupted point — so parent-level work above that call is subject to the same double execution. ## The patterns that fix it **Interrupt first.** Make `interrupt()` the first meaningful statement in the node. If nothing precedes it, nothing is duplicated. Pass into the payload only data that already lives in state, so you do not need to compute anything before asking. **Separate the gate from the effect.** The cleanest structure is two nodes: an `approval` node whose entire body is an `interrupt()` and a decision, and an `execute` node downstream that performs the side effect. The pause now lives on a node boundary; the executing node runs exactly once because it is only scheduled after the approval node has committed. **Compute after, not before.** Anything non-deterministic that must be part of the outcome should be computed below the interrupt line, so it happens once and is what the run actually uses. If the human must *see* a generated value, generate it in the previous node and read it from state. **Idempotency keys.** When an effect genuinely cannot be moved — a legacy client, a batched write — give it a deterministic key derived from state (thread id plus a step identifier) and let the downstream system dedupe. That converts an at-least-once execution into an effectively-once one. **Pin the interrupt shape while threads are open.** Treat the number and order of `interrupt()` calls in a node as part of a wire contract. Draining or migrating in-flight threads before changing them avoids delivering stale answers to renumbered questions. ## Interview framing Interviewers ask this because it separates people who have run LangGraph in production from people who have run the quickstart. The tell of a strong answer is that you do not describe it as a bug: replay is the price of a pause that survives a process restart, and every durable-execution system pays some version of it. What you are expected to own is the discipline that follows — treat the region above `interrupt()` as code that must be safe to run more than once, and put anything that is not into its own node.

  • Does the same double-execution risk apply to a node that only calls an LLM before interrupting?
    Yes, and it costs you. The model call above the `interrupt()` line is re-issued on resume, so you pay tokens twice and may get a different completion than the one the human reviewed — the replayed answer is the one that lands in state. Either move the call into the preceding node and read its result from state, or cache it deterministically so the replay is free and stable.
  • How would you make an unavoidable pre-interrupt write safe?
    Give it a deterministic idempotency key built from data already in the checkpoint — typically the thread id plus a stable step name — and let the target system reject or collapse the duplicate. Do not derive the key from a timestamp, a UUID generated in the node, or anything else that changes on replay, because then the second attempt looks like a new operation and the dedupe never fires.
  • What is the risk of editing a node's interrupt() calls while threads are paused?
    Resume values are matched positionally within the node's task, so if you insert, remove or reorder `interrupt()` calls, a paused thread's stored answers get delivered to different questions on resume. The run continues confidently with wrong values. Treat the interrupt sequence in a node as a versioned contract: drain or migrate open threads before changing it, and log the interrupt payload so mismatches are visible in traces.

saying these in an interview costs you the question

  • Assumes execution continues on the line after interrupt()
  • Puts API calls or writes above the interrupt in the same node
  • Treats non-deterministic values computed pre-interrupt as stable
  • Calls the replay a LangGraph bug rather than a durability tradeoff
  • Thinks the checkpointer stores the Python call stack

context