skip to content

LangGraph

You will learn graph-based orchestration for stateful, cyclic agent workflows: typed state, nodes and edges, conditional routing, checkpointing, human-in-the-loop pauses, and multi-agent composition. Interviewers ask because LangGraph answers exactly what a linear chain cannot — loops, retries, approvals, and resuming a run that was interrupted.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In LangGraph, what steps build a minimal StateGraph from schema to invoke?

level: juniorimportance: must knowfreq 70%

answer

  1. builder first, runnable second
  2. state schema declares the channels
  3. sentinels mark entry and exit
  4. compile() before you can invoke
  5. add_node, add_edge, compile, invoke

basics

~20 s

Define a state schema, usually a TypedDict; create StateGraph(State); register functions with add_node; wire them with add_edge from START to END; then call compile() to get a runnable graph you invoke with an initial state dict.

solid answer

~40 s

Five steps. First, declare the shared state schema — most often a `TypedDict` listing every key the graph reads or writes. Second, `builder = StateGraph(State)`. Third, register each unit of work with `builder.add_node("plan", plan_fn)`; a node is just a callable that takes the current state and returns a partial update. Fourth, wire the flow with `builder.add_edge(START, "plan")` and `builder.add_edge("plan", END)` — `START` and `END` are sentinel constants from `langgraph.graph`, not nodes you define. Fifth, `graph = builder.compile()`, which validates the topology (unreachable nodes, missing entry point) and returns a runnable object exposing `invoke`, `stream` and their async twins. Then `graph.invoke({"topic": "x"})` seeds the state and runs until the graph reaches `END`, returning the final state.

code

python · 26 lines
python
from typing import TypedDict
from langgraph.graph import StateGraph, START, END


class State(TypedDict):
    topic: str
    draft: str


def plan(state: State) -> dict:
    return {"draft": f"outline for {state['topic']}"}


def write(state: State) -> dict:
    return {"draft": state["draft"].upper()}


builder = StateGraph(State)
builder.add_node("plan", plan)
builder.add_node("write", write)
builder.add_edge(START, "plan")
builder.add_edge("plan", "write")
builder.add_edge("write", END)

graph = builder.compile()
print(graph.invoke({"topic": "state graphs", "draft": ""}))

go deeper

for a junior

Be able to write the five lines from memory: schema, StateGraph(State), add_node, add_edge with START and END, compile, invoke. Say plainly that a node takes state and returns a partial update.

for a middle

Explain what compile() validates and why building and running are separate steps, and describe what a superstep does with the values nodes return.

for a senior

Show judgment about node granularity — one node per retryable, observable unit of work — and about naming, since node names surface in traces and streamed events.

for a principal

Own the argument for when a graph is the right shape at all: a linear pipeline needs no graph, and the cost of StateGraph is justified by loops, resumability and branching, not by wiring two steps in a row.

## The shape of the API LangGraph splits building from running. `StateGraph` is a mutable *builder*: you add nodes and edges to it in any order. `compile()` freezes that builder into a `CompiledStateGraph`, which is the thing you actually execute. Keeping the two apart lets LangGraph validate the topology once, at compile time, instead of discovering a dangling edge halfway through a run. ## Step 1 — the state schema Every graph is built around one shared, typed state object. The usual spelling is a `TypedDict`: The schema does three jobs. It documents what flows between nodes; it tells LangGraph which *channels* to create (one per key); and it lets your editor and type checker catch typos in `state["topic"]`. A dataclass or a Pydantic `BaseModel` works too — the choice changes validation behaviour, not the graph mechanics. ## Step 2 — nodes `builder.add_node("plan", plan_fn)` registers a callable under a name. The name is the identifier used by edges, by streamed update events and by the drawn diagram, so pick something readable. The callable receives the current state and returns a **partial** update — a dict of just the keys it changed. If you pass only the function, `add_node(plan_fn)` uses the function's `__name__` as the node name. A node can be a plain function, an async function, or any LangChain runnable. Nothing about a node is special-cased for LLMs: calling a model is just what most node bodies happen to do. ## Step 3 — edges, START and END Edges declare order. `add_edge(START, "plan")` marks `plan` as an entry point — `START` is a sentinel constant meaning "the virtual node that runs before everything". `add_edge("plan", END)` marks a terminal point. Both are imported from `langgraph.graph`; you never register them with `add_node`. A graph with no edge from `START` fails at compile time, because there is nothing to run first. An edge added with `add_edge` is *unconditional*: when the source finishes, the target is scheduled. Edges whose target depends on the state are a separate construct. ## Step 4 — compile `graph = builder.compile()` returns the runnable. Compilation checks the graph: every node must be reachable, edge endpoints must name registered nodes, and an entry point must exist. `compile()` is also where run-wide options are attached — a checkpointer, interrupt points, a run name. Calling it twice on the same builder is legal and gives you two independent compiled graphs. A compiled graph also exposes `get_graph().draw_mermaid()`, which is the fastest way to see whether the topology you wired is the one you meant. ## Step 5 — invoke `graph.invoke({"topic": "reducers"})` seeds the state channels from the dict you pass and runs the graph. Execution proceeds in *supersteps*: LangGraph runs the scheduled nodes, applies their returned updates to the state, then schedules whatever their outgoing edges point at, and repeats until nothing is left to run. `invoke` returns the final state as a dict. `stream()` returns the same run as an iterator of intermediate results, and `ainvoke`/`astream` are the async forms. Because the compiled graph implements the same runnable interface as other LangChain-ecosystem components, a graph can be composed into a larger pipeline, or another graph, without adapters. ## Where beginners trip The most common error is treating the node's return value as the whole new state — it is a *partial* update merged into what already exists. The second is forgetting `compile()` and calling `invoke` on the builder, which does not have it. The third is adding `START` or `END` with `add_node`. The fourth is defining a node that reads a key the schema never declared: with a `TypedDict` schema, that key simply isn't a channel, and the read raises a `KeyError` at runtime rather than being caught by the type checker.

  • What does compile() actually check, and what would make it fail?
    It validates the topology before any node runs: every edge endpoint must name a registered node, every node must be reachable, and there must be at least one edge out of START. A graph with a node nobody points at, or with no entry point, raises at compile time rather than mid-run — which is the whole point of separating the builder from the runnable.
  • Do I need to pass every state key when I invoke the graph?
    No. The input dict seeds the channels you supply; keys you omit are simply unset until a node writes them. Reading an unset key from a TypedDict state raises KeyError, so either seed the keys your first node reads, give the schema defaults via a dataclass or Pydantic model, or have nodes use state.get(...).
  • Can the same function be registered as two different nodes?
    Yes — add_node takes an explicit name, so the same callable can appear under "draft_a" and "draft_b" and be wired independently. Node identity in LangGraph is the name, not the function object, which is why streamed updates and the drawn diagram are keyed by name.

saying these in an interview costs you the question

  • Thinks a node returns the entire new state
  • Calls invoke on the builder without compiling
  • Registers START or END with add_node
  • Believes nodes must be LLM calls
  • Thinks edges pass data between nodes directly

context

open as a page

In LangGraph, what does graph.stream() give you that graph.invoke() does not?

level: juniorimportance: must knowfreq 70%

basics

~10 s

graph.invoke() blocks and returns only the final state. graph.stream() yields a chunk as the graph advances, one per superstep, so a UI can show progress long before the run finishes.

open as a page

In LangGraph, what does compiling with a checkpointer add, and what is thread_id for?

level: middleimportance: must knowfreq 78%

basics

~20 s

A checkpointer makes LangGraph save a snapshot of the graph's state after every super-step. The thread_id you pass in config names the conversation those snapshots belong to, so the next invoke resumes that thread's accumulated state instead of starting empty.

open as a page

In LangGraph, what does interrupt() do inside a node and how do you resume?

level: middleimportance: must knowfreq 72%

basics

~20 s

interrupt() pauses the graph in the middle of a node, persists a checkpoint, and surfaces a payload to the caller under the interrupt key. You resume by invoking the same thread_id with Command(resume=value); the interrupt() call then returns that value.

open as a page

In LangGraph, what does returning Command(goto=...) from a node give you over a plain state update?

level: middleimportance: must knowfreq 68%

basics

~20 s

A Command return carries both a state update and a goto, so one node writes state and names its successor in a single value. That is the handoff idiom in multi-agent graphs; a plain dict return leaves routing entirely to the declared edges.

open as a page

In LangGraph, how does a supervisor route to worker agents, and what does each hop cost?

level: middleimportance: must knowfreq 74%

basics

~20 s

A supervisor is an LLM node that reads shared state and returns a Command naming which worker node runs next; each worker hands control back to the supervisor. Every delegation therefore costs an extra model call, on a message list that keeps growing.

open as a page

In LangGraph, how does add_conditional_edges decide which node runs next?

level: middleimportance: must knowfreq 72%

basics

~20 s

add_conditional_edges attaches a router function to a source node. After that node runs and its update is applied, LangGraph calls the router with the current state and uses its return value to name the next node, or END to stop that path.

open as a page

How does LangGraph merge a node's return value into StateGraph state?

level: middleimportance: must knowfreq 76%

basics

~20 s

A node returns a partial dict, and LangGraph applies it key by key. Each schema key is a channel; by default the returned value overwrites the old one. Keys the node omits are untouched, and returning None or an empty dict updates nothing.

open as a page

What is a reducer in a LangGraph state schema, and when do you need one?

level: middleimportance: must knowfreq 72%

basics

~20 s

A reducer is a function attached to a state key that says how a node's returned value combines with the existing one. Without it a key is last-value: one write per step, and a second concurrent write raises InvalidUpdateError. With operator.add, lists accumulate instead.

open as a page

How do you stream LLM tokens from inside a LangGraph node to the client?

level: middleimportance: must knowfreq 68%

basics

~20 s

Use stream_mode="messages" on graph.stream/astream. LangGraph then yields (message_chunk, metadata) tuples for every chat-model token produced inside any node, and the metadata names the emitting node so you can forward only the tokens the user should see.

open as a page

How do LangGraph's stream_mode values and updates differ in what they yield?

level: middleimportance: must knowfreq 72%

basics

~20 s

values yields the entire graph state after every superstep. updates yields only the delta each node returned, keyed by node name. Use updates for cheap progress events and values when the consumer genuinely needs the whole state.

open as a page

Why does a LangGraph node re-run from the top after interrupt(), and what breaks?

level: seniorimportance: must knowfreq 50%

basics

~20 s

LangGraph checkpoints state, not the Python stack, so resuming replays the interrupted node function from its first line. Any side effect executed before interrupt() — an API call, a write, an email — therefore happens twice. Put interrupt() first, or isolate side effects in a separate node.

open as a page

In a LangGraph multi-agent graph, when should agents get isolated state instead of one shared message list?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Share one message channel when agents must see each other's reasoning; isolate when they must not. Isolation means giving an agent its own state schema in a subgraph and writing only a summarized result back to the parent — smaller prompts and no cross-contamination, at the cost of lost context.

open as a page

When two parallel LangGraph branches write the same state key, what happens?

level: seniorimportance: must knowfreq 55%

basics

~20 s

If the key has no reducer, LangGraph raises InvalidUpdateError — it refuses to pick a winner between two writes in one step. Annotate the key with a reducer, such as Annotated[list, operator.add], and both writes are combined instead.

open as a page

In LangGraph, what does Send() in a routing function do that returning node names cannot?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Send lets one branch spawn a runtime-determined number of parallel invocations of the same node, each with its own private input state. Returning node names can only fan out to a fixed set of nodes, all sharing the graph state.

open as a page

How do you choose between LangGraph's InMemorySaver, SqliteSaver and PostgresSaver?

level: middleimportance: should knowfreq 58%

basics

~20 s

InMemorySaver keeps checkpoints in process memory — fine for tests and notebooks, lost on restart. SqliteSaver writes to a local file, suitable for single-process apps. PostgresSaver is the production choice: shared, durable, and usable from many workers at once.

open as a page

How do LangGraph's interrupt_before and interrupt_after differ from interrupt()?

level: middleimportance: should knowfreq 55%

basics

~20 s

interrupt_before and interrupt_after are static pauses configured on compile: they stop the graph at a node boundary, carry no payload, and resume with invoke(None, config). interrupt() is dynamic — it lives in node code, can fire conditionally, ships a payload, and resumes with Command(resume=...).

open as a page

Why pass path_map to LangGraph's add_conditional_edges if the router returns node names?

level: middleimportance: should knowfreq 42%

basics

~20 s

path_map translates the router's return labels into destination node names, so the router speaks intent ("continue", "done") rather than topology. It also tells LangGraph the branch's possible destinations up front, which is what makes the drawn graph show real edges.

open as a page

How does a LangGraph node emit custom progress updates to the stream?

level: middleimportance: should knowfreq 42%

basics

~10 s

Call get_stream_writer() from langgraph.config inside the node and invoke it with any serialisable payload, then consume those payloads with stream_mode="custom". This streams progress that is neither part of graph state nor model tokens.

open as a page

How does LangGraph's checkpointer make a crashed run resumable, and what can still be lost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Checkpoints written at each super-step, plus task-level pending writes, let a re-invoke on the same thread_id with None input continue from the last durable point. What is lost depends on the durability mode and on side effects: re-executed nodes repeat theirs.

open as a page

In LangGraph, when do you need a BaseStore instead of a checkpointer?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A checkpointer persists one thread's state and history, and nothing crosses thread boundaries. A BaseStore is namespaced key-value storage shared across threads, so facts a user should keep between separate conversations belong there, not in checkpointed state.

open as a page

How do you replay a LangGraph run from an earlier checkpoint, and what re-executes?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Read the thread's history with graph.get_state_history(config), take the StateSnapshot you want, and invoke with its config — which carries a checkpoint_id. Steps recorded before that checkpoint are replayed from storage, not re-run; everything after it executes again, appended as a new branch.

open as a page

How do you use LangGraph's update_state to edit graph state before resuming a run?

level: seniorimportance: should knowfreq 45%

basics

~20 s

graph.update_state(config, values, as_node=...) writes a new checkpoint on the paused thread. The values pass through the state schema's reducers, and as_node decides whose outgoing edges run next. You then continue the thread with invoke(None, config) or a Command resume.

open as a page

How do you add a compiled LangGraph subgraph as a node when its state schema differs?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Wrap it. When parent and subgraph share the state keys, pass the compiled subgraph straight to add_node. When the schemas differ, add an ordinary function node that translates parent state into the subgraph's input, calls subgraph.invoke, and maps the result back into parent keys.

open as a page

How do you keep internal keys out of a LangGraph graph's input and output?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Declare narrower schemas alongside the overall state: pass input_schema and output_schema to StateGraph so callers supply only the public inputs and invoke returns only the public results. Scratch keys stay internal, declared on the overall schema or on a node's own annotated schema.

open as a page

In LangGraph, when is a Pydantic state schema better than a TypedDict?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Choose a Pydantic BaseModel when you want runtime validation and real defaults on the state passed into nodes — useful at untrusted or model-generated boundaries. A TypedDict is lighter and has no runtime cost. Neither changes merge semantics: reducers still come from Annotated.

open as a page

When would you use astream_events on a LangGraph graph instead of stream_mode?

level: seniorimportance: should knowfreq 45%

basics

~20 s

astream_events emits one fine-grained event per nested runnable — chain, chat model, tool — with run ids, tags and LangGraph node metadata. Reach for it when debugging what actually ran inside a node; a stream_mode stream is cheaper for shipping product output.

open as a page

How do you enable LangSmith tracing for a LangGraph app, and what does a trace show?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Set LANGSMITH_TRACING=true plus LANGSMITH_API_KEY, and optionally LANGSMITH_PROJECT; no code change is needed. Each graph run becomes a nested trace: a root run, a child run per node execution, and model and tool runs with prompts, latency and token counts.

open as a page

How do you keep LangGraph checkpoint storage from becoming a production liability?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat checkpoints as a real dataset. Every super-step writes the whole state, so rows grow with threads times steps times state size: keep large payloads out of state, choose durability per workload, and run explicit retention, encryption and capacity planning on the checkpoint database.

open as a page

How would you design a LangGraph approval gate for irreversible agent actions?

level: principalimportance: should knowfreq 40%

basics

~20 s

Gate by risk, not by node: interrupt only for actions that are irreversible or above a threshold, pause before the effect happens, ship the reviewer a payload they can actually judge, and accept approve, edit and reject-with-reason. Back it with a durable checkpointer and a timeout policy.

open as a page

showing 1–30 of 33