skip to content

How do you decide what belongs in LangGraph graph state versus outside it?

level: principalimportance: should knowfreq 34%

answer

  1. state is a run's ledger, not its storage
  2. later reader or resume, else local
  3. references travel, payloads do not
  4. handles and secrets never serialize well
  5. accumulating keys need a growth policy

basics

~20 s

Put in state only what a later node must read or what must survive a pause: identifiers, decisions, accumulated results. Keep large payloads, live clients and secrets out — state is snapshotted every superstep, so its size is a per-step cost and everything in it is exposed.

solid answer

~50 s

The test is simple to state and hard to apply: a value belongs in state if a *later* node needs it, or if the run must be able to resume with it. Everything else should live outside. Three categories almost never belong. Large payloads — full documents, dataframes, images — because state is captured after every superstep, so payload size becomes a per-step write and serialization cost; store an id or URI and fetch inside the node that needs it. Live handles — database connections, HTTP clients, model clients — because they do not serialize and are not meaningful across a resume; construct them in the node or inject them as runtime context. Secrets, because state is snapshotted, streamed and logged. The fourth trap is unbounded growth: any key with an accumulating reducer grows every loop iteration, so message history and tool results need an explicit trimming policy, decided when you design the schema rather than after the first context-window failure.

go deeper

for a junior

Know the basic rule: state carries what a later node needs to read. A value used only inside one node should be a local variable there.

for a middle

Explain why state size is a per-step cost rather than a one-off, and why live clients and connections cannot live in state once persistence is involved.

for a senior

Design for the hundredth iteration: bound accumulating keys with an explicit trimming policy, keep references rather than payloads, and keep secrets and raw third-party responses out of anything that gets snapshotted or traced.

for a principal

Own the state schema as the system's contract and blast radius: one writer per key, narrow views per node, dull data that is boring when leaked, and a growth policy agreed before the graph is built.

## Why the question has teeth In a small graph, state is free and you put everything in it. In a production graph — long-running, looping, resumable, multi-tenant — the state schema is the most consequential design artifact you own. It determines what a run costs per step, what survives a pause, what leaks into logs and traces, and how tightly your nodes are coupled to one another. LangGraph makes state so convenient that the default failure is putting too much in it. ## The inclusion test A value belongs in state if either is true: 1. **A later node reads it.** State is the only channel between nodes; edges carry control, not data. 2. **It must survive an interruption.** If a run can pause and resume, whatever the resumed run needs to continue must be in state, because nothing else about the in-process world comes back. If neither holds, the value is a local variable inside a node, and it should stay one. A key that exists only because "we computed it and it seemed a shame to throw it away" is pure cost. ## What to keep out **Large payloads.** After every superstep, the accumulated state is the thing captured and handed forward. A key holding the full text of twenty retrieved documents means that payload is written, and possibly serialised and stored, once per step — for the whole run, not once. The standard fix is indirection: keep document ids, a URI, or a cache key in state, and have the node that needs the bytes fetch them. State then describes the run; it does not store the corpus. **Live handles.** Connections, sessions, clients and file handles do not serialise and are meaningless after a resume — a reconstructed run cannot revive a socket. Build them inside nodes, hold them at module scope, or pass them through the graph's runtime configuration (in LangGraph 1.x, `context_schema` declares that non-state runtime context). Putting a client in state turns a graph that runs fine in memory into one that fails the moment you enable persistence. **Secrets and raw third-party responses.** State is snapshotted, streamed to observability tooling, and printed in debugging. Anything in it should be safe in all three places. A raw tool response with a token embedded, or a full user record when the graph only needs the user id, is a leak waiting for the first support engineer who prints the state. **Anything derivable cheaply.** A key that is a pure function of another key is a consistency bug in waiting: one node updates the source and forgets the derived copy. Compute it where you need it. ## The growth problem Accumulating channels are the second scaling failure. A message list, a tool-results list, a notes list — each grows on every loop iteration and never shrinks. In a graph that loops twenty times, the state carried into the twentieth step is far bigger than the first, and if that key is rendered into the prompt, cost and latency grow with it until the context window ends the run. So growth policy is part of schema design, not a later patch. Decide up front: does this key keep everything, keep the last N, keep a rolling summary, or drop entries once a downstream node has consumed them? Message history has purpose-built support for deletion, which is the hook a trimming policy hangs on. "We will deal with it when it gets long" means dealing with it during an incident. ## State as a coupling surface Every key in the shared schema is readable by every node that declares it. A schema that has become a flat bag of thirty keys means any node can read anything, so nobody can reason about who depends on what, and renaming a key is a graph-wide change. Two mitigations: narrow the schema a node sees to the keys it actually uses, and give each key a single obvious writer. A key with five writers and no reducer is a conflict waiting to happen; a key with five writers *and* a reducer is usually a sign that five nodes are doing one job. ## Multi-tenancy and blast radius In a shared service, ask what happens if a state snapshot is mixed up between runs. State that holds only ids and decisions is boring when leaked; state that holds a full customer record is an incident. Designing state to be *dull* — identifiers, flags, small results — is a cheap way to bound the blast radius of every other bug in the system. ## A working heuristic Sketch the state schema before the graph. For each key, answer three questions: which node writes it, which node reads it, and how large does it get at the hundredth iteration. Keys that fail any of the three do not belong.

  • What is the concrete cost of keeping full documents in state instead of ids?
    State is captured after every superstep, so the payload is carried, and where persistence is enabled written, once per step rather than once per run. A twenty-step graph pays twenty times. It also inflates traces and any returned payload. Keeping ids or URIs in state and fetching inside the consuming node makes the per-step cost proportional to the run's decisions, not to the corpus.
  • Where should a database client or model client live if not in state?
    Outside the state channels: constructed inside the node, held at module scope, or supplied as runtime context that the graph declares separately from state. A live handle cannot be serialised and has no meaning after a resume, so putting it in state makes an in-memory graph break the moment persistence is switched on.
  • How do you keep an accumulating message list from growing without bound?
    Decide the policy when you write the schema: keep the last N turns, keep a rolling summary plus recent turns, or delete entries once consumed. Message state supports explicit removal, so trimming is expressed as an update through the same reducer rather than by reaching around the state model. What you must not do is leave it unbounded and discover the limit in production.
  • How do you tell a state schema has become a coupling problem?
    Look for keys with no single obvious writer, keys read by nodes that have nothing to do with each other, and keys nobody can explain. When renaming a key means touching most nodes, the schema is a shared global rather than a contract. Narrow each node's view to what it uses and give each key one writer to pull it back.

saying these in an interview costs you the question

  • Puts full retrieved documents in state by default
  • Stores a live client or connection in state
  • Treats message history as naturally bounded
  • Adds keys because a value might be useful later
  • Assumes state is private and never logged

context