skip to content

How do you keep internal keys out of a LangGraph graph's input and output?

level: seniorimportance: should knowfreq 36%

answer

  1. one graph, more than one schema
  2. public surface versus working keys
  3. filter what invoke gives back
  4. node type hints can narrow the view
  5. internals stay channels, just not public

basics

~20 s

Declare narrower schemas alongside the overall state: pass input_schema and output_schema to StateGraph so callers supply only the public inputs and invoke returns only the public results. Scratch keys stay internal, declared on the overall schema or on a node's own annotated schema.

solid answer

~50 s

In LangGraph 1.x, `StateGraph` takes the overall state schema plus optional `input_schema` and `output_schema`. The overall schema still defines every channel the graph uses, but the input schema declares what a caller is expected to pass, and the output schema filters what `invoke` returns — so scratch keys such as retry counters, raw tool payloads or intermediate scores never appear in the public result. Beyond that, a node can declare an even narrower view by type-hinting its state parameter with a different schema; keys that appear only there become channels no other node reads, which is the idiomatic way to keep private working data out of the shared surface. The payoff is a real contract: callers cannot depend on internals, the returned payload stays small, and you can refactor the middle of the graph without breaking consumers — the same reason you would not return your whole ORM entity from an HTTP handler.

code

python · 32 lines
python
from typing import TypedDict
from langgraph.graph import StateGraph, START, END


class InputState(TypedDict):
    question: str


class OutputState(TypedDict):
    answer: str


class OverallState(InputState, OutputState):
    attempts: int
    raw_tool_payload: str


def solve(state: OverallState) -> dict:
    return {
        "answer": f"about {state['question']}",
        "attempts": 1,
        "raw_tool_payload": "...large blob...",
    }


builder = StateGraph(OverallState, input_schema=InputState, output_schema=OutputState)
builder.add_node("solve", solve)
builder.add_edge(START, "solve")
builder.add_edge("solve", END)

# {'answer': 'about schemas'} — internals stayed internal
print(builder.compile().invoke({"question": "schemas"}))

go deeper

for a junior

Know that by default a graph returns its whole state, and that StateGraph accepts separate input and output schemas to narrow what callers pass and receive.

for a middle

Explain that narrowing filters the boundary only — internal keys remain real channels with their reducers — and that a node's own type hint decides which keys it sees.

for a senior

Justify the split as contract design: stable public surface, bounded response payload, no accidental leaking of raw tool output, and freedom to refactor the middle of the graph.

for a principal

Own the graph as a callable component: define the input and output contracts every terminating path must satisfy, and treat tightening them later as a versioned change to consumers, not a refactor.

## The problem A state schema tends to grow. It starts as `{question, answer}` and ends up carrying a retry counter, the raw tool response, a confidence score, a scratch list the reranking node needs, and two flags. All of it is legitimate working data — and all of it comes back to the caller from `invoke`, because by default the graph's output is the whole state. Now every internal key is part of your public contract, and renaming one breaks a consumer. ## Three schemas, not one LangGraph 1.x lets you separate them at construction time. The first positional argument is the **overall** state schema — the union of every channel the graph uses. `input_schema` narrows what the caller is expected to provide. `output_schema` narrows what comes back: after the final superstep, LangGraph filters the accumulated state down to those keys before returning it. The important nuance: narrowing input or output does **not** delete channels. Internal keys still exist, still merge with their reducers, still flow between nodes. You have only changed what crosses the graph's boundary in each direction. ## Private state between nodes A second, finer mechanism: a node's state parameter can be type-hinted with a schema that is not the overall one. LangGraph reads that hint to decide which keys the node sees. Keys that appear only in one pair of nodes' schemas become effectively private channels — node A writes them, node B reads them, and nobody else in the graph has them in view. This is how you hand a large intermediate artifact from a producer node to a consumer node without adding it to the surface every other node reads. Use it sparingly. It is powerful and it is also invisible in the topology diagram: the coupling between A and B lives in a type hint, not in an edge, so a reader tracing the graph will not see it. Reach for it when the intermediate is genuinely private and awkward, not as a default style. ## Why bother Four concrete payoffs. **Contract stability.** Consumers depend on the output schema only. Adding, renaming or deleting an internal key becomes a non-breaking change. Without the split, every field is public by accident. **Payload size.** The returned state is often serialised — into an HTTP response, a log line, a trace. A graph that accumulates raw retrieved documents in state will return megabytes if you let it. Filtering at the output boundary is one line. **Leak prevention.** Internal state legitimately holds things a caller should never see: raw API responses, intermediate prompts, tool credentials threaded through a step, another user's cached result in a shared node. Returning the whole state exposes all of it. This is the same reasoning behind not serialising a database entity straight to an API response. **Callability.** A graph with clear input and output schemas composes: it can be embedded as a node inside a bigger graph, or wrapped as a tool, because there is a defined thing to pass in and a defined thing that comes back. ## Practical shape A common layout is three `TypedDict`s: `InputState` with the one or two keys a caller supplies, `OutputState` with the answer plus whatever metadata you have decided to support, and `OverallState` that inherits from both and adds the working keys. Inheritance keeps them consistent — you cannot drift the output key's type away from the internal one. ## Caveats Keys in the input schema are what a caller is *expected* to pass; supplying anything else is not a contract you should build on. Keys in the output schema that no node ever wrote come back unset, so it is easy to promise a field the graph does not always produce — treat the output schema as an interface you are obliged to satisfy on every path, not as an aspiration. These parameters were named `input` and `output` before LangGraph 1.0; older tutorials use those spellings. In 1.x they are `input_schema` and `output_schema`, alongside `context_schema` for runtime configuration that is not state at all. Finally, narrowing the output does not shrink what is *held* mid-run. If your concern is memory or snapshot size rather than the returned payload, the fix is to stop putting the big thing in state, not to hide it at the exit.

  • Does an output schema reduce the memory the graph uses during a run?
    No. It filters what invoke returns after the last superstep; the internal channels still hold their values throughout the run and still appear in any state snapshot. If the concern is size rather than the public contract, the fix is to keep the large artifact out of state entirely — put an identifier in state and fetch the payload inside the node that needs it.
  • How does a node see a key that isn't in the input schema?
    By what its own state parameter is annotated with. LangGraph reads that hint to decide which keys the node is handed, so a node hinted with the overall schema sees every channel, while a node hinted with a narrower private schema sees only those keys. The input schema constrains the caller, not the nodes.
  • What breaks if the output schema names a key some paths never write?
    That key comes back unset, so a consumer that treats the output schema as a guarantee gets a missing field on exactly the paths you tested least — typically the error branch. Treat the output schema as an interface every terminating path must satisfy, and give it a default or have a final node normalise it.

saying these in an interview costs you the question

  • Thinks output_schema reduces in-run memory or state size
  • Believes narrowing removes the internal channels
  • Assumes nodes only see input-schema keys
  • Returns whole state and calls internals private
  • Names the pre-1.0 input/output parameters as current

context