skip to content

How do you replay a LangGraph run from an earlier checkpoint, and what re-executes?

level: seniorimportance: should knowfreq 48%

answer

  1. history is a chain, not a value
  2. each snapshot knows its own address
  3. earlier steps come back from storage
  4. later steps genuinely run again
  5. the old branch is not erased

basics

~20 s

Read the thread's history with graph.get_state_history(config), take the StateSnapshot you want, and invoke with its config — which carries a checkpoint_id. Steps recorded before that checkpoint are replayed from storage, not re-run; everything after it executes again, appended as a new branch.

solid answer

~50 s

`graph.get_state_history(config)` yields the thread's `StateSnapshot` objects, newest first. Each snapshot's `.config` contains not just the `thread_id` but also a `checkpoint_id` identifying that exact point. Calling `graph.invoke(None, snapshot.config)` tells LangGraph to start there: because the checkpoint already records the state and the parent chain, the earlier super-steps are **replayed** — LangGraph loads their recorded results rather than calling the nodes again — and execution genuinely resumes only from the chosen checkpoint onward. The new checkpoints are written as a fork; the original history is not overwritten, so a thread can carry several divergent branches. In practice this is how you debug an agent that took a wrong turn eight steps in: find the snapshot where `values` was still correct, and re-run only from there instead of replaying the whole conversation and paying for all the tokens again.

code

python · 11 lines
python
cfg = {"configurable": {"thread_id": "thread-1"}}

history = list(graph.get_state_history(cfg))  # newest first
for snap in history:
    print(snap.metadata.get("step"), snap.next, snap.values)

target = history[2]                 # a known-good earlier point
print(target.config)                # includes checkpoint_id

# prefix is replayed from storage; execution resumes at the target
graph.invoke(None, target.config)

go deeper

for a junior

Know that a thread keeps a history of snapshots and that get_state_history lets you read them. Being able to say that you can restart from an earlier point is enough at this level.

for a middle

Explain that a snapshot's own config carries the checkpoint_id used as the resume target, and that earlier steps are replayed from storage while later ones actually re-execute.

for a senior

Demonstrate the operational consequence: replay does not undo external side effects, and re-executed nodes repeat theirs. Show how you pick a rewind point from values and next when debugging a bad trajectory.

for a principal

Own idempotency as a design rule for any node with external effects, and decide policy on branch accumulation — how many forks a thread may carry, who may rewind production threads, and how that interacts with audit and retention.

## What time travel means here Because a checkpointer writes a full state snapshot at every super-step and links each one to its parent, a thread is not a single current value — it is a chain. Time travel is the ability to point at any link in that chain and continue from it. The two operations are *inspect the history* and *resume from a chosen point*. ## Reading the history `graph.get_state_history(config)` takes a config identifying a thread and yields `StateSnapshot` objects in reverse chronological order. Each snapshot exposes: - `values` — the full channel state at that step, - `next` — the tasks that were about to run, - `config` — a config that includes this checkpoint's `checkpoint_id` (and `checkpoint_ns` when subgraphs are involved), - `parent_config` — the config of the preceding checkpoint, - `metadata` — step number and write source, - `created_at` and `tasks`. `graph.get_state(config)` is the same thing for just the latest checkpoint. Neither method executes anything; they are pure reads against the checkpointer. The important detail is that `snapshot.config` is not the same object you passed in. It is enriched with the checkpoint identity, which is precisely what makes it usable as a resume target. ## Resuming from a specific checkpoint Hand that enriched config back to the graph: `graph.invoke(None, snapshot.config)` The `None` input means "add nothing new, just continue." LangGraph resolves the checkpoint, restores its channel values and its `next` tasks, and proceeds. The part interviewers probe is what actually runs. Steps whose results are already recorded before the target checkpoint are **replayed**: LangGraph knows they completed and reuses their recorded effect on state rather than invoking the node functions again. No LLM calls, no tool calls, no cost. Execution — real node invocation — begins at the chosen checkpoint's pending tasks and continues forward. This distinction has a sharp consequence: **replay restores state, it does not restore the world.** Any side effect a replayed node performed — a row inserted, an email sent, a payment captured — happened once and is not undone by rewinding the graph. Conversely, a node that *does* re-execute after the fork will perform its side effects a second time. Designing nodes to be idempotent, or to write through an idempotency key derived from the thread and step, is what makes time travel safe on anything that touches the outside world. ## Forking, not rewriting Resuming from an old checkpoint does not delete the checkpoints that came after it. The new run's checkpoints are appended with the target as their parent, so the thread's history becomes a tree rather than a line. Both branches remain readable through `get_state_history`. That is a feature — you can compare what the agent did on each path — and a cost, since a thread you replay repeatedly accumulates several branches' worth of rows. ## What this is actually used for 1. **Debugging a bad trajectory.** An agent looped or hallucinated at step 9. Rather than re-running the whole conversation with a changed prompt, find the last good snapshot, change the code or configuration, and resume from there. Only the tail re-executes. 2. **Recovering from a transient failure.** A tool timed out three steps ago and poisoned the state; rewind past it and continue. 3. **Exploring alternatives.** Branch the same thread twice from one decision point and compare outcomes — cheap A/B on an agent's behaviour, using recorded prefixes for both branches. 4. **Post-mortem inspection.** Reading `values` and `next` at each step tells you exactly what the agent believed and what it intended, which is far more legible than reconstructing it from logs. ## Subgraphs When a graph nests subgraphs, their checkpoints live under a checkpoint namespace within the same thread. `graph.get_state(config, subgraphs=True)` includes the nested snapshots, so a rewind can target a point inside a nested run rather than only the outer step boundaries. ## Common misconceptions - *"Replaying re-runs everything from the start."* It does not; that is the whole point, and it is the difference between a cheap rewind and paying for the full token history again. - *"You need the raw checkpoint id string."* You can construct the config by hand, but the snapshot's own `.config` already carries it and is the intended handle. - *"Rewinding erases the branch you left."* It does not — history forks and both branches persist. - *"Time travel undoes side effects."* It only rewinds LangGraph's state. Everything your nodes did to external systems stands.

  • If a replayed prefix includes a node that charged a customer, does rewinding refund it?
    No. Replay only restores LangGraph's recorded state; the node is not called again, and nothing it did externally is reversed. Worse, a node that sits after the fork point will execute a second time and can charge again. Anything with external side effects needs idempotency — a key derived from thread and step, or a check-before-act — before time travel is safe in production.
  • Does resuming from an old checkpoint delete the checkpoints that came after it?
    No. The resumed run writes new checkpoints whose parent is the target, so the thread's history becomes a tree with both branches intact and readable through get_state_history. That is deliberate — you can compare trajectories — but it means a thread you rewind repeatedly accumulates several branches of rows, which matters for storage growth.
  • How do you find the right checkpoint to rewind to without guessing?
    Walk get_state_history and read each snapshot's values and next. values shows what the agent believed at that step and next shows what it was about to do, so the last snapshot whose values are still correct is your target. metadata carries the step number, which makes it easy to correlate with a trace or log line from the same run.

saying these in an interview costs you the question

  • Believing replay re-executes every node from the beginning
  • Assuming rewinding undoes side effects the agent already performed
  • Thinking a fork overwrites or deletes the later checkpoints
  • Passing only thread_id and expecting the graph to start at an older step
  • Confusing get_state_history with a log of LLM messages

context