skip to content

Checkpointing & Persistence

You will learn persistence: the checkpointer interface with MemorySaver, SqliteSaver and PostgresSaver, thread-level versus run-level state, replaying from a checkpoint, and state snapshots for fault tolerance. Interviewers ask because checkpointing is what makes a long agent run resumable after a crash and gives each conversation thread its own durable history.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

questions

6

In LangGraph, what does compiling with a checkpointer add, and what is thread_id for?

level: middleimportance: must knowfreq 78%

answer

  1. state survives between invocations
  2. one snapshot per execution step
  3. conversations are keyed, not shared
  4. the key lives under configurable
  5. thread_id selects which history to resume

basics

~20 s

A checkpointer makes LangGraph save a snapshot of the graph's state after every super-step. The thread_id you pass in config names the conversation those snapshots belong to, so the next invoke resumes that thread's accumulated state instead of starting empty.

solid answer

~40 s

Without a checkpointer, a compiled LangGraph graph is stateless across calls: state lives only for the duration of one `invoke`, and the next call starts from whatever input you hand it. Passing `builder.compile(checkpointer=...)` turns on persistence — after every super-step LangGraph writes a checkpoint containing the full channel state, the pending next tasks, and metadata. Those checkpoints are keyed by a **thread**, and you select the thread with `config = {"configurable": {"thread_id": "..."}}`. Invoke twice with the same thread_id and the second run starts from the first run's final state, with your new input merged in through the state schema's reducers. Invoke with a different thread_id and you get an independent, empty conversation. So the checkpointer supplies durability and the thread_id supplies isolation — one conversation or one long-running job per thread.

code

python · 22 lines
python
from typing import Annotated, TypedDict
from operator import add
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import InMemorySaver

class State(TypedDict):
    steps: Annotated[list[str], add]

def step(state: State):
    return {"steps": ["ran"]}

builder = StateGraph(State)
builder.add_node("step", step)
builder.add_edge(START, "step")
builder.add_edge("step", END)

graph = builder.compile(checkpointer=InMemorySaver())

cfg = {"configurable": {"thread_id": "thread-1"}}
print(graph.invoke({"steps": []}, cfg))  # {'steps': ['ran']}
print(graph.invoke({"steps": []}, cfg))  # {'steps': ['ran', 'ran']}
print(graph.get_state(cfg).next)          # () - nothing left to run

go deeper

for a junior

Know the two moving parts by name: a checkpointer passed to compile(), and a thread_id passed under config["configurable"]. Be able to say that together they make a second invoke continue the first one.

for a middle

Explain that a checkpoint is written after every super-step and holds the full channel state plus the next scheduled tasks, and that new input merges through reducers rather than replacing state. Know get_state and get_state_history.

for a senior

Show you use thread_id as a real isolation key mapped to a session or job id, and that you read StateSnapshot.next when diagnosing a thread that stalled. Be clear about which state is thread-scoped and which is not.

for a principal

Own the thread-identity scheme across the system: what a thread means in your domain, who may read it, how it maps to tenancy and auth, and whether persistence lives in your process or in a managed runtime that configures it for you.

## The problem persistence solves A compiled LangGraph graph is, by default, a pure function of its input. You call `graph.invoke(input, config)`, the graph runs its nodes, and the final state is returned and then forgotten. That is fine for a one-shot pipeline and useless for a chat agent, an approval workflow, or any job that must survive a process restart. Checkpointing is the feature that converts the graph from stateless to stateful. ## Turning it on You attach a checkpointer at compile time: `graph = builder.compile(checkpointer=InMemorySaver())` `InMemorySaver` (from `langgraph.checkpoint.memory`, also exported under the older name `MemorySaver`) is the process-local implementation; `SqliteSaver` and `PostgresSaver` are the durable ones. All of them implement the same `BaseCheckpointSaver` interface, so swapping backends is a one-line change and nothing about the graph definition changes. ## What a checkpoint is LangGraph executes as a sequence of **super-steps** (the Pregel model): in each super-step, every task scheduled for that step runs, their state updates are applied through the channel reducers, and then the next set of tasks is computed. A checkpoint is written at the end of each super-step. It records: - the full values of every channel in the state schema at that point, - the `next` tasks — which nodes are scheduled to run in the following step, - metadata such as the step number and the source of the write, - a pointer to the parent checkpoint, which is what makes the history a linked chain. Because the whole state is written, not a diff, a single checkpoint is enough to reconstruct the run at that instant. That is also why state size is a real cost — see the operational question on this topic. ## Threads Checkpoints do not float free; they belong to a thread. The thread is named by you, in the `configurable` sub-dictionary of the config object: `config = {"configurable": {"thread_id": "user-42-session-7"}}` This is the single most common beginner mistake — putting `thread_id` at the top level of the config instead of inside `configurable`. LangGraph will not find it there, and you get a fresh, empty run every time, or an error telling you a checkpointer requires a thread_id. A thread is your unit of isolation and your unit of resumption. One chat session, one document being processed, one ticket being triaged — each gets its own thread_id, typically derived from an ID you already own (session id, user id plus conversation id). Two threads never see each other's state. ## Resuming a thread On the second call with the same thread_id, LangGraph loads the latest checkpoint for that thread and uses its channel values as the starting state. Your new input is not a replacement — it is applied as an update through the same reducers the graph uses internally. For a message-list channel with an appending reducer, that means the new user message is appended to the existing history rather than clobbering it. This is why a LangGraph chat agent needs no separate conversation-memory object: the state channel plus the checkpointer *is* the memory. Calling `graph.invoke(None, config)` — a `None` input on an existing thread — means "continue from where this thread left off without adding anything," which is the shape used to resume after a crash or a pause. ## Inspecting the thread Two methods read persisted state without running anything: - `graph.get_state(config)` returns a `StateSnapshot` for the latest checkpoint of the thread, with fields including `values`, `next`, `config`, `metadata`, `created_at`, `parent_config` and `tasks`. - `graph.get_state_history(config)` yields the thread's snapshots, most recent first. These are the debugging surface: `values` tells you what the agent believes, and `next` tells you what it is about to do. If `next` is empty, the run finished; if it is non-empty on a thread that is not running, the thread is paused or crashed mid-flight. ## Namespacing When a graph contains subgraphs, their checkpoints are stored under a checkpoint namespace (`checkpoint_ns`) inside the same thread, so a parent thread carries the nested state too rather than losing it. `graph.get_state(config, subgraphs=True)` surfaces the nested snapshots as well. ## Deployment note When a graph is served by LangGraph's managed server rather than embedded in your own process, persistence is configured for you and you do not pass a checkpointer at compile time. In self-hosted or embedded use, attaching one is your job.

  • If I invoke the same thread a second time with a new user message, does that replace the state or add to it?
    It is applied as an update, not a replacement. LangGraph loads the latest checkpoint's channel values and then merges your input through the state schema's reducers. For a message channel with an appending reducer, the new message is appended to the existing history; for a plain overwrite channel, the new value does replace the old one. The reducer you declared decides.
  • What does an empty `next` field on a StateSnapshot tell you?
    That no tasks are scheduled after that checkpoint — the run reached its end. A non-empty `next` on a thread that is not currently executing means the graph stopped part-way: it was paused, or the process died before those tasks ran. That field is the first thing to read when diagnosing a thread that appears stuck.
  • Do I need a checkpointer just to use a graph with cycles?
    No. Cycles, conditional edges and loops all work on a stateless compiled graph within a single invocation, because state is held in memory for the duration of the run. The checkpointer is only needed when state must outlive the call — multi-turn conversations, resumption after a crash, inspecting history, or pausing mid-run.

saying these in an interview costs you the question

  • Thinking a graph remembers state without a checkpointer attached
  • Putting thread_id at the top of config instead of under configurable
  • Believing a checkpoint is written only when the run ends
  • Assuming new input overwrites persisted state rather than reducing into it
  • Treating thread_id as optional when a checkpointer is compiled in

context

open as a page

How do you choose between LangGraph's InMemorySaver, SqliteSaver and PostgresSaver?

level: middleimportance: should knowfreq 58%

basics

~20 s

InMemorySaver keeps checkpoints in process memory — fine for tests and notebooks, lost on restart. SqliteSaver writes to a local file, suitable for single-process apps. PostgresSaver is the production choice: shared, durable, and usable from many workers at once.

open as a page

How does LangGraph's checkpointer make a crashed run resumable, and what can still be lost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Checkpoints written at each super-step, plus task-level pending writes, let a re-invoke on the same thread_id with None input continue from the last durable point. What is lost depends on the durability mode and on side effects: re-executed nodes repeat theirs.

open as a page

In LangGraph, when do you need a BaseStore instead of a checkpointer?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A checkpointer persists one thread's state and history, and nothing crosses thread boundaries. A BaseStore is namespaced key-value storage shared across threads, so facts a user should keep between separate conversations belong there, not in checkpointed state.

open as a page

How do you replay a LangGraph run from an earlier checkpoint, and what re-executes?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Read the thread's history with graph.get_state_history(config), take the StateSnapshot you want, and invoke with its config — which carries a checkpoint_id. Steps recorded before that checkpoint are replayed from storage, not re-run; everything after it executes again, appended as a new branch.

open as a page

How do you keep LangGraph checkpoint storage from becoming a production liability?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat checkpoints as a real dataset. Every super-step writes the whole state, so rows grow with threads times steps times state size: keep large payloads out of state, choose durability per workload, and run explicit retention, encryption and capacity planning on the checkpoint database.

open as a page