How do you keep LangGraph checkpoint storage from becoming a production liability?
answer
- full state written every step
- biggest payloads are the real cost
- one knob trades writes for safety
- nothing expires by itself
- it is a transcript store, legally speaking
basics
~20 sTreat checkpoints as a real dataset. Every super-step writes the whole state, so rows grow with threads times steps times state size: keep large payloads out of state, choose durability per workload, and run explicit retention, encryption and capacity planning on the checkpoint database.
solid answer
~50 sThe failure mode is quiet: a checkpointer persists the **full** state at every super-step, so a chat agent carrying a growing message list writes an ever-larger row on every step of every thread. Volume is roughly threads x steps x state size, and it is dominated by state size, which is the lever you actually control — keep retrieved documents, file contents and blobs out of state and store references instead. Then decide durability per workload rather than globally: `"sync"` for expensive long-running runs, `"exit"` for short high-volume ones, since each mode changes both write volume and per-step latency. Operationally, checkpoints hold the raw conversation, so they inherit your PII, encryption and retention obligations; you need a deletion path for finished or expired threads, and you should size the connection pool against concurrent runs, not request rate, because every super-step is a write.
go deeper
Know that checkpoints accumulate in a real database and nothing deletes them for you, and that big values in state are written again on every step.
Explain the volume model — threads times steps times state size — and why keeping large payloads out of state, using references instead, is the main lever you control.
Show the operational picture: durability chosen per workload, pool sizing against concurrent runs, separating the checkpoint database, and a concrete deletion path for expired threads.
Own it as a data-governance problem. Checkpoints are transcripts: set classification, encryption and retention, decide who may read thread history, plan capacity, and make state-size and thread-lifetime limits an architectural rule rather than a code-review opinion.
## Why this bites late Checkpointing looks free in development. A demo has ten threads, five steps each, and a state of a few kilobytes. In production the same code writes on every super-step of every concurrent run, forever, and nothing in the framework prunes it. Teams typically discover this as a database that grew a hundred times in a quarter, or as p99 latency creeping up on a chat endpoint that got no slower in the model. ## The volume model Approximate the write volume as **threads x super-steps per thread x state size per checkpoint**, plus task-level write records within steps, plus a branch per time-travel rewind. Three of those four terms are set by design decisions: - **State size** is the dominant term and the one people ignore. Because a checkpoint stores the full state rather than a diff, a 2 MB blob parked in state is re-serialised and re-written at every step. Twenty steps means 40 MB for one run. The fix is architectural: state carries an object key, a document id, or a URI, and the node fetches the payload when it needs it. Retrieved chunks, uploaded files, images and large tool outputs should almost never live in state. - **Steps per thread** is graph granularity. Fine-grained nodes buy precise crash recovery and cost more checkpoints; coarse nodes are cheaper to persist and redo more work on resume. That is a genuine tradeoff to make deliberately, not accidentally. - **Thread lifetime** is product design. A conversation that never ends is a checkpoint chain that never stops growing, and each of its steps re-writes an ever-longer message history — quadratic in the conversation length. Bounding a thread (session windows, summarisation into a smaller state channel, or moving durable facts to a cross-thread store) is the standard mitigation. ## Durability as a per-workload knob In LangGraph 1.x, `invoke` and `stream` accept `durability` with values `"exit"`, `"async"` (default) and `"sync"`. This is a capacity lever, not only a safety one: - High-volume, short, cheap graphs — classification, extraction, routing — often justify `"exit"`: one write per run instead of one per step, at the cost of resumability you would never have used. - Long, expensive, externally observable runs justify `"sync"`: the added per-step latency is trivial next to redoing an LLM call or a human approval. - Most interactive work sits on the default `"async"`. Because the argument is per-invocation, one service can apply all three against the same checkpointer, and choosing per route is a cheap win. ## Retention and deletion Nothing expires on its own. You need a policy and something that enforces it: - Decide how long a finished thread stays queryable. Support and audit usually want some window; forever is not a policy. - Have a way to delete a thread's checkpoints outright — recent checkpointer versions expose a thread deletion operation on the interface, and where a version does not, deleting by thread key in the backing store is the fallback. This is not optional if you are subject to data-deletion requests, because the checkpoint database holds the conversation. - Remember that time travel forks rather than overwrites, so threads that are repeatedly rewound carry several branches. Rewind-heavy workflows need tighter retention than their thread count suggests. ## Privacy and security Checkpoints are the highest-fidelity record of what a user said and what the agent did — usually more complete than your logs. Consequences to state plainly: - The checkpoint database is in scope for the same classification, encryption-at-rest and access-control rules as your primary user data, and often stricter, since it captures raw prompts. - Access to `get_state_history` is effectively access to conversation transcripts; whatever admin surface exposes it needs authorisation, not just an internal URL. - Serialisation is the boundary at which secrets leak: a credential or token pulled into state to pass between nodes is written to disk at every subsequent step. Pass identifiers and resolve secrets inside the node instead. ## Database placement and connections Agent checkpoints have a write-heavy, high-churn profile that is unlike most OLTP tables, so putting them in their own database or at least their own schema keeps vacuum, bloat and backup pressure away from your product tables — and lets you set retention independently. Connection sizing is the other common miss: with an async saver, every super-step of every concurrent run issues a write, so the pool must be sized against **concurrent in-flight runs multiplied by their step rate**, not against HTTP requests per second. Under-provisioning shows up as runs stalling on connection acquisition, which looks like model latency in a naive trace. ## What a strong answer sounds like Not "add an index". It is: bound the state, bound the thread, choose durability per workload, isolate and size the database, and own retention and deletion of what is, in effect, a transcript store. The framework gives you durability; the liability is yours to manage.
- Why does an ever-growing message list make checkpoint cost grow faster than linearly?Because every super-step writes the full state, and the state includes the entire message list so far. Step one writes one message, step twenty writes twenty — total bytes grow with the square of the conversation length rather than with its length. Bounding it means windowing or summarising the message channel, or moving durable facts out to cross-thread storage so the live state stays small.
- How should the connection pool for a Postgres checkpointer be sized?Against concurrent in-flight graph runs and their step rate, not against HTTP request rate. Every super-step of every active run performs a write, so a service with modest request volume but long multi-step runs can saturate a pool that looks generously sized. Starvation manifests as runs stalling on connection acquisition, which is easy to misread as model latency unless the pool is instrumented.
- What is the risk of exposing get_state_history on an internal admin page?It is a transcript viewer. Checkpoints record the full conversation and the agent's intermediate reasoning state, usually in more detail than application logs, so the endpoint carries the same access-control and audit requirements as reading user messages directly. Being internal is not authorisation, and per-thread access should be checked against the same tenancy rules the product enforces.
- Where should the checkpoint database sit relative to your application database?Preferably separate, or at minimum in its own schema. The workload is write-heavy and high-churn with a short useful life, which is a poor neighbour for product tables in terms of bloat, vacuum pressure and backup size. Separation also lets you apply independent retention and encryption policy to what is effectively a conversation archive.
saying these in an interview costs you the question
- Assuming checkpoints are pruned automatically by the framework
- Storing retrieved documents or file blobs directly in graph state
- Treating the checkpoint database as low-sensitivity operational data
- Sizing the connection pool by request rate instead of concurrent runs
- Running one endless thread per user and calling it long-term memory