skip to content

How does LangGraph's checkpointer make a crashed run resumable, and what can still be lost?

level: seniorimportance: should knowfreq 44%

answer

  1. state and pending tasks both persisted
  2. successful siblings are not redone
  3. a knob trades latency for safety
  4. default writes in the background
  5. the world is not rewound with the state

basics

~20 s

Checkpoints written at each super-step, plus task-level pending writes, let a re-invoke on the same thread_id with None input continue from the last durable point. What is lost depends on the durability mode and on side effects: re-executed nodes repeat theirs.

solid answer

~50 s

After each super-step LangGraph persists the full state and the tasks scheduled next; within a step it also records the writes of tasks that finished, so if one parallel branch throws, the siblings that succeeded do not have to run again. Recovery is then simply `graph.invoke(None, {"configurable": {"thread_id": same_id}})` — the graph loads the last checkpoint and continues from its pending tasks. The gap is controlled by the `durability` argument on `invoke`/`stream` in LangGraph 1.x: `"exit"` writes only when the run ends (fastest, nothing recoverable after a crash), `"async"` writes in the background while the next step starts (the default — a small window where the very last step may be missing), and `"sync"` blocks until each checkpoint is committed (slowest, tightest guarantee). What is never recovered is external side effects: the node that resumes runs from its start, so anything it already did to the outside world happens twice unless you made it idempotent.

code

python · 9 lines
python
cfg = {"configurable": {"thread_id": "job-123"}}

# strongest guarantee: commit each checkpoint before the next step
graph.invoke({"steps": []}, cfg, durability="sync")

# after a crash, in a new process: continue where it stopped
snapshot = graph.get_state(cfg)
if snapshot.next:              # tasks still pending on this thread
    graph.invoke(None, cfg)    # resume; completed work is not redone

go deeper

for a junior

Know that re-invoking the same thread_id with None as input continues an unfinished run rather than starting over, because the checkpointer stored where it stopped.

for a middle

Explain the two records that make it work — the per-super-step checkpoint and the task-level pending writes — and name the durability modes and what each trades away.

for a senior

Show that resumability ends at the process boundary: re-executed nodes repeat external effects, so idempotency is required. Discuss node granularity as recovery granularity and how you cap retries on a poisoned thread.

for a principal

Own the recovery policy end to end: which workloads justify sync durability, how unfinished threads are discovered and re-driven, attempt limits and dead-lettering, and how the resulting at-least-once execution is made safe for downstream systems.

## The recovery model LangGraph's fault tolerance is not magic rollback; it is bookkeeping. Two records make a crashed run resumable. **Checkpoints.** At the end of every super-step the checkpointer stores the full channel state and the `next` tasks. That pair is sufficient to reconstruct where the run was. **Pending writes.** Within a super-step, several tasks can run — parallel branches, a fan-out, multiple tool calls. If task A completes and task B raises, LangGraph records A's output as a pending write attached to the current checkpoint (`put_writes` on the checkpointer interface). On resume, A is not re-executed; only B is retried. Without this, one flaky branch would force every sibling to repeat, doubling cost and duplicating side effects. ## Resuming Recovery is deliberately unremarkable: `graph.invoke(None, {"configurable": {"thread_id": "job-123"}})` The `None` input says "add nothing, continue." LangGraph loads the latest checkpoint for the thread, sees non-empty `next`, and schedules exactly those tasks. Practically this means your crash-recovery worker needs only the thread ids that are not finished — which you can find by reading `get_state(config).next` and checking whether it is empty. ## Durability modes Checkpointing costs a database round-trip per super-step, so LangGraph 1.x lets you choose where you sit on the safety/latency curve with the `durability` argument on `invoke`, `stream` and their async twins: - **`"exit"`** — checkpoints are persisted only when the run finishes. Fastest, and the right choice for short, cheap, retriable pipelines where restarting from the top is acceptable. A crash mid-run leaves nothing to resume. - **`"async"`** — the default. The checkpoint write is issued in the background while the next super-step begins. Almost all of the durability, almost none of the latency, at the price of a narrow window in which a hard crash loses the most recent step. On resume you redo that step. - **`"sync"`** — the run blocks until each checkpoint is committed before proceeding. The strongest guarantee, and what you want when a step is expensive or externally observable enough that redoing it is worse than the added latency. The mode is per-invocation, so a system can use `"sync"` for a long approval workflow and `"exit"` for a cheap classification graph against the same checkpointer. ## What is genuinely lost Three categories, and naming them is what a strong answer does. 1. **The last unwritten step.** Under `"async"` a crash can lose the step in flight; under `"exit"` it loses everything. Recovery redoes that work. 2. **External side effects, in the wrong direction.** LangGraph rewinds *its* state, never the world. A node that sent an email, wrote a row, or called a payment API and then crashed before its checkpoint landed will run again from its start and repeat the effect. This is the single most important production consideration in checkpointed agents: **nodes must be idempotent**, or must guard themselves with a key derived from the thread and step, or must confine their effect to the last thing they do so the window is minimal. 3. **Anything not in state.** Values held in a node's local variables, module globals, caches or open connections are gone. Only what a node returns into the state schema, and therefore what the serialiser can persist, survives. This is also why state must be serialisable — a live client object cannot be checkpointed. ## Retries and poison threads Resuming is not automatically safe from repetition: if a node fails deterministically — a bad prompt, a permanently 400ing tool call — a naive supervisor that keeps re-invoking the thread will loop forever, re-paying for the prefix each time. Real deployments cap resume attempts per thread, record the attempt count somewhere durable (often in state itself, so it is checkpointed), and move exhausted threads to a dead-letter path for human inspection. `get_state` on such a thread is your forensic surface: `values` shows what the agent believed and `next` shows the task that keeps dying. ## Interaction with graph shape Super-step granularity means recovery granularity. A graph decomposed into many small nodes recovers close to where it failed; a graph with one enormous node that makes six LLM calls internally re-does all six on resume, because LangGraph only knows about node boundaries. If resumability matters, node granularity is a design decision, not a style preference.

  • Why do pending writes matter when a single branch of a parallel fan-out fails?
    Because without them the whole super-step would be unfinished, and every task in it would re-run on resume. LangGraph records the outputs of tasks that completed as writes attached to the current checkpoint, so only the failed task is retried. That saves the cost of the successful branches and, more importantly, stops their side effects from happening twice.
  • When would you deliberately choose durability="exit" in production?
    For short, cheap, deterministic graphs where restarting from the top costs less than a database round-trip per step — a classification or extraction pipeline of two or three nodes with no external writes, run at high volume. You are trading resumability you would never use for real throughput. It is the wrong choice the moment a run is long, expensive, or externally observable.
  • How do you stop a deterministically failing node from being resumed forever?
    Cap it. Keep an attempt counter in the graph state so it is checkpointed with everything else, increment it on entry to the failing region, and route to a terminal or dead-letter node once it exceeds the limit. External supervisors that blindly re-invoke unfinished threads will otherwise re-pay for the replayed prefix on every attempt and never make progress.

saying these in an interview costs you the question

  • Assuming a crash-resume undoes what the failed node already did externally
  • Thinking every task in a step re-runs when one sibling fails
  • Believing checkpointing is free and always fully synchronous
  • Expecting local variables or open clients to survive a resume
  • Retrying a poisoned thread forever with no attempt cap

context