How do you choose between LangGraph's InMemorySaver, SqliteSaver and PostgresSaver?
answer
- same interface, different durability story
- one is a dict in RAM
- one is a file on one box
- one is shared by many workers
- Postgres needs a one-time table setup
basics
~20 sInMemorySaver keeps checkpoints in process memory — fine for tests and notebooks, lost on restart. SqliteSaver writes to a local file, suitable for single-process apps. PostgresSaver is the production choice: shared, durable, and usable from many workers at once.
solid answer
~50 sAll three implement the same `BaseCheckpointSaver` interface, so the choice is purely operational. `InMemorySaver` (from `langgraph.checkpoint.memory`) stores checkpoints in a dict — zero setup, no durability, no sharing between processes, and it grows until the process dies; use it in tests and demos. `SqliteSaver` persists to a file, which survives restarts but ties you to one machine and serialises writes, so it fits single-process desktop or local tooling. `PostgresSaver` is what you deploy behind a horizontally scaled service: many workers can pick up any thread, and you get real backup and retention. Postgres needs `checkpointer.setup()` run once to create its tables. Each backend also ships an async twin — `AsyncSqliteSaver`, `AsyncPostgresSaver` — and you must match the flavour to how you drive the graph: an async saver under a synchronous `invoke` will fail rather than silently degrade.
code
python · 18 linesfrom langgraph.checkpoint.postgres import PostgresSaver
from langgraph.graph import StateGraph, START, END
from typing import TypedDict
class State(TypedDict):
answer: str
builder = StateGraph(State)
builder.add_node("reply", lambda s: {"answer": "ok"})
builder.add_edge(START, "reply")
builder.add_edge("reply", END)
DB_URI = "postgresql://postgres:postgres@localhost:5432/langgraph"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # one-time: creates the checkpoint tables
graph = builder.compile(checkpointer=checkpointer)
print(graph.invoke({"answer": ""}, {"configurable": {"thread_id": "t1"}}))go deeper
Be able to name the three savers and say plainly which one survives a restart. Knowing that InMemorySaver is for tests and Postgres is for production already answers the screening version of this question.
Explain that all three implement the same BaseCheckpointSaver interface so the choice is operational, and know the two Postgres gotchas: setup() creates the tables, and from_conn_string is a context manager.
Show the deployment reasoning — replica count, connection pooling, async saver under an async server, and state serialisability. Be ready to say what actually goes wrong with the wrong choice rather than reciting the table.
Own where agent state lives relative to the rest of the system: its own database or shared with OLTP, who backs it up, what retention and encryption apply to it, and whether a custom checkpointer over existing infrastructure is worth the maintenance.
## One interface, three operational profiles Every LangGraph checkpointer implements `BaseCheckpointSaver`. The interface is small: `put` writes a checkpoint, `put_writes` records the results of individual tasks within a step, `get_tuple` fetches a checkpoint for a config, and `list` enumerates a thread's history — each with an async counterpart (`aput`, `aput_writes`, `aget_tuple`, `alist`). Because the graph only ever talks to that interface, choosing a backend is a deployment decision, not an architecture decision, and swapping one for another is a one-line change at `compile()`. ## InMemorySaver Shipped in the base `langgraph-checkpoint` package as `langgraph.checkpoint.memory.InMemorySaver` (the older name `MemorySaver` is still exported). Checkpoints live in a Python dict inside the process. - **Good for**: unit tests, notebooks, tutorials, and any single-run script where you want `get_state_history` to work. - **Bad for**: anything real. State vanishes on restart, is invisible to a second worker, and accumulates without bound — every super-step of every thread stays resident, so a long-lived service using it is a memory leak with extra steps. A common trap is a demo that "works" with `InMemorySaver` behind two Uvicorn workers: requests land on different processes and users see their conversation randomly reset. That is not a bug in LangGraph; it is the wrong backend. ## SqliteSaver From the `langgraph-checkpoint-sqlite` package: `langgraph.checkpoint.sqlite.SqliteSaver`, with `AsyncSqliteSaver` for async graphs. Typically constructed with `SqliteSaver.from_conn_string("checkpoints.sqlite")`, which is a context manager. - **Good for**: local agents, CLI tools, desktop apps, and integration tests that need durability across process restarts without standing up a database. - **Bad for**: multi-process or multi-container deployments. SQLite's write model serialises writers, and a file on a container's ephemeral disk is not durable in any meaningful sense. Note that `from_conn_string(":memory:")` gives an in-memory SQLite database — durable-looking API, non-durable storage, another way to accidentally ship a demo. ## PostgresSaver From `langgraph-checkpoint-postgres`: `langgraph.checkpoint.postgres.PostgresSaver` and `AsyncPostgresSaver`, built on psycopg. Two setup facts matter in interviews: 1. **`setup()` must be called once.** The saver creates and migrates its own tables; forget it and the first write fails on a missing relation. Run it at deploy time or at startup behind a guard — it is idempotent, but it takes DDL locks, so having every worker race to run it on boot is untidy. 2. **`from_conn_string` is a context manager.** It opens a connection and closes it on exit, which is right for a script and wrong for a long-lived server that must keep the saver alive for the process's lifetime. For a service, construct the saver over a connection pool you own and manage its lifecycle with the application. Postgres is the answer when several replicas of your service must be able to continue any thread, when you need backups and point-in-time recovery of agent state, and when checkpoint data falls under retention or privacy rules that a file on a pod cannot satisfy. ## Sync versus async The async savers are not a nicety. If your graph is driven with `ainvoke`/`astream`, use the async saver so checkpoint writes do not block the event loop; if you drive it synchronously, use the sync one. Mixing them raises rather than degrading — an async-only saver has no synchronous `get_tuple` to call. In an ASGI service (FastAPI and friends), that means `AsyncPostgresSaver` plus an async connection pool sized against your worker concurrency, because every super-step of every in-flight run performs a write. ## Serialization All backends serialise state through a `serde` component before storing it, and it is settable on the saver. Practically this means your state must be serialisable: TypedDict and Pydantic values, messages and plain Python data are fine; open sockets, database handles, file objects and arbitrary class instances are not. State that must reference such a thing should store a key or URI and re-resolve it inside the node. ## How to answer the question The expected shape is a decision rule, not a feature list: memory for tests, SQLite for a single process that must survive restarts, Postgres for anything horizontally scaled or subject to retention rules — plus the two setup gotchas (`setup()` and the context-manager lifecycle) and the sync/async pairing. Mentioning that the interface is identical, so the migration path is trivial, is what separates a confident answer from a recited table.
- What breaks first if you deploy InMemorySaver behind two service replicas?Thread continuity. Each replica holds its own dict, so a user whose next request is load-balanced to the other pod resumes an empty thread and the conversation appears to reset at random. Nothing errors — it just silently loses state. Memory growth is the second failure: nothing evicts old checkpoints, so the process bloats until it is restarted, which discards everything again.
- Why does PostgresSaver need setup() and where should you call it?The saver owns its own schema and creates the checkpoint tables and any pending migrations. Without setup() the first write fails on a missing relation. Call it once during deployment or a startup migration step rather than in every worker's boot path — it is idempotent, but it takes DDL locks, so concurrent replicas racing to run it is avoidable noise.
- Can you write your own checkpointer for a store LangGraph doesn't ship?Yes — subclass BaseCheckpointSaver and implement get_tuple, list, put and put_writes, plus the async variants if you drive the graph asynchronously. The graph only depends on that interface. The hard parts are honouring the parent-checkpoint chain so history and replay stay coherent, and recording task-level writes correctly so partially completed steps resume properly.
saying these in an interview costs you the question
- Calling InMemorySaver production-ready because state persists between calls
- Deploying a memory-backed saver behind multiple replicas or workers
- Forgetting PostgresSaver.setup() and blaming a connection error
- Holding from_conn_string's context manager open incorrectly in a long-lived service
- Using a sync saver under ainvoke and expecting it to just work