skip to content

In LangGraph, what steps build a minimal StateGraph from schema to invoke?

level: juniorimportance: must knowfreq 70%

answer

  1. builder first, runnable second
  2. state schema declares the channels
  3. sentinels mark entry and exit
  4. compile() before you can invoke
  5. add_node, add_edge, compile, invoke

basics

~20 s

Define a state schema, usually a TypedDict; create StateGraph(State); register functions with add_node; wire them with add_edge from START to END; then call compile() to get a runnable graph you invoke with an initial state dict.

solid answer

~40 s

Five steps. First, declare the shared state schema — most often a `TypedDict` listing every key the graph reads or writes. Second, `builder = StateGraph(State)`. Third, register each unit of work with `builder.add_node("plan", plan_fn)`; a node is just a callable that takes the current state and returns a partial update. Fourth, wire the flow with `builder.add_edge(START, "plan")` and `builder.add_edge("plan", END)` — `START` and `END` are sentinel constants from `langgraph.graph`, not nodes you define. Fifth, `graph = builder.compile()`, which validates the topology (unreachable nodes, missing entry point) and returns a runnable object exposing `invoke`, `stream` and their async twins. Then `graph.invoke({"topic": "x"})` seeds the state and runs until the graph reaches `END`, returning the final state.

code

python · 26 lines
python
from typing import TypedDict
from langgraph.graph import StateGraph, START, END


class State(TypedDict):
    topic: str
    draft: str


def plan(state: State) -> dict:
    return {"draft": f"outline for {state['topic']}"}


def write(state: State) -> dict:
    return {"draft": state["draft"].upper()}


builder = StateGraph(State)
builder.add_node("plan", plan)
builder.add_node("write", write)
builder.add_edge(START, "plan")
builder.add_edge("plan", "write")
builder.add_edge("write", END)

graph = builder.compile()
print(graph.invoke({"topic": "state graphs", "draft": ""}))

go deeper

for a junior

Be able to write the five lines from memory: schema, StateGraph(State), add_node, add_edge with START and END, compile, invoke. Say plainly that a node takes state and returns a partial update.

for a middle

Explain what compile() validates and why building and running are separate steps, and describe what a superstep does with the values nodes return.

for a senior

Show judgment about node granularity — one node per retryable, observable unit of work — and about naming, since node names surface in traces and streamed events.

for a principal

Own the argument for when a graph is the right shape at all: a linear pipeline needs no graph, and the cost of StateGraph is justified by loops, resumability and branching, not by wiring two steps in a row.

## The shape of the API LangGraph splits building from running. `StateGraph` is a mutable *builder*: you add nodes and edges to it in any order. `compile()` freezes that builder into a `CompiledStateGraph`, which is the thing you actually execute. Keeping the two apart lets LangGraph validate the topology once, at compile time, instead of discovering a dangling edge halfway through a run. ## Step 1 — the state schema Every graph is built around one shared, typed state object. The usual spelling is a `TypedDict`: The schema does three jobs. It documents what flows between nodes; it tells LangGraph which *channels* to create (one per key); and it lets your editor and type checker catch typos in `state["topic"]`. A dataclass or a Pydantic `BaseModel` works too — the choice changes validation behaviour, not the graph mechanics. ## Step 2 — nodes `builder.add_node("plan", plan_fn)` registers a callable under a name. The name is the identifier used by edges, by streamed update events and by the drawn diagram, so pick something readable. The callable receives the current state and returns a **partial** update — a dict of just the keys it changed. If you pass only the function, `add_node(plan_fn)` uses the function's `__name__` as the node name. A node can be a plain function, an async function, or any LangChain runnable. Nothing about a node is special-cased for LLMs: calling a model is just what most node bodies happen to do. ## Step 3 — edges, START and END Edges declare order. `add_edge(START, "plan")` marks `plan` as an entry point — `START` is a sentinel constant meaning "the virtual node that runs before everything". `add_edge("plan", END)` marks a terminal point. Both are imported from `langgraph.graph`; you never register them with `add_node`. A graph with no edge from `START` fails at compile time, because there is nothing to run first. An edge added with `add_edge` is *unconditional*: when the source finishes, the target is scheduled. Edges whose target depends on the state are a separate construct. ## Step 4 — compile `graph = builder.compile()` returns the runnable. Compilation checks the graph: every node must be reachable, edge endpoints must name registered nodes, and an entry point must exist. `compile()` is also where run-wide options are attached — a checkpointer, interrupt points, a run name. Calling it twice on the same builder is legal and gives you two independent compiled graphs. A compiled graph also exposes `get_graph().draw_mermaid()`, which is the fastest way to see whether the topology you wired is the one you meant. ## Step 5 — invoke `graph.invoke({"topic": "reducers"})` seeds the state channels from the dict you pass and runs the graph. Execution proceeds in *supersteps*: LangGraph runs the scheduled nodes, applies their returned updates to the state, then schedules whatever their outgoing edges point at, and repeats until nothing is left to run. `invoke` returns the final state as a dict. `stream()` returns the same run as an iterator of intermediate results, and `ainvoke`/`astream` are the async forms. Because the compiled graph implements the same runnable interface as other LangChain-ecosystem components, a graph can be composed into a larger pipeline, or another graph, without adapters. ## Where beginners trip The most common error is treating the node's return value as the whole new state — it is a *partial* update merged into what already exists. The second is forgetting `compile()` and calling `invoke` on the builder, which does not have it. The third is adding `START` or `END` with `add_node`. The fourth is defining a node that reads a key the schema never declared: with a `TypedDict` schema, that key simply isn't a channel, and the read raises a `KeyError` at runtime rather than being caught by the type checker.

  • What does compile() actually check, and what would make it fail?
    It validates the topology before any node runs: every edge endpoint must name a registered node, every node must be reachable, and there must be at least one edge out of START. A graph with a node nobody points at, or with no entry point, raises at compile time rather than mid-run — which is the whole point of separating the builder from the runnable.
  • Do I need to pass every state key when I invoke the graph?
    No. The input dict seeds the channels you supply; keys you omit are simply unset until a node writes them. Reading an unset key from a TypedDict state raises KeyError, so either seed the keys your first node reads, give the schema defaults via a dataclass or Pydantic model, or have nodes use state.get(...).
  • Can the same function be registered as two different nodes?
    Yes — add_node takes an explicit name, so the same callable can appear under "draft_a" and "draft_b" and be wired independently. Node identity in LangGraph is the name, not the function object, which is why streamed updates and the drawn diagram are keyed by name.

saying these in an interview costs you the question

  • Thinks a node returns the entire new state
  • Calls invoke on the builder without compiling
  • Registers START or END with add_node
  • Believes nodes must be LLM calls
  • Thinks edges pass data between nodes directly

context