skip to content

In LangGraph, what does graph.stream() give you that graph.invoke() does not?

level: juniorimportance: must knowfreq 70%

answer

  1. one blocks, one yields as it goes
  2. chunk per superstep, not per line
  3. payload shape is configurable
  4. stream_mode picks what each chunk holds
  5. invoke returns final state only

basics

~10 s

graph.invoke() blocks and returns only the final state. graph.stream() yields a chunk as the graph advances, one per superstep, so a UI can show progress long before the run finishes.

solid answer

~50 s

A compiled LangGraph graph exposes both. `invoke()` runs the whole graph and hands you one object at the end — the final state — so the user sees nothing until the last node returns, which for an agent loop with several model calls can be tens of seconds. `stream()` (and its async twin `astream()`) returns an iterator that emits a chunk each time the graph advances one superstep, so you can push progress, partial state, or model tokens to the client as they happen. What each chunk *contains* is chosen by the `stream_mode` argument rather than being fixed: full state, per-node deltas, model tokens, or custom payloads a node writes itself. Streaming does not change how the graph executes — same nodes, same order, same final state — it only changes how much of the run you get to observe while it is running.

code

python · 25 lines
python
from typing_extensions import TypedDict
from langgraph.graph import StateGraph, START, END


class State(TypedDict):
    topic: str
    joke: str


def write(state: State) -> dict:
    return {"joke": f"a joke about {state['topic']}"}


graph = (
    StateGraph(State)
    .add_node("write", write)
    .add_edge(START, "write")
    .add_edge("write", END)
    .compile()
)

final = graph.invoke({"topic": "cats"})

for chunk in graph.stream({"topic": "cats"}, stream_mode="updates"):
    print(chunk)

go deeper

for a junior

Know that invoke() waits and returns the final state while stream()/astream() yields chunks as the run progresses, and that stream_mode chooses what each chunk contains.

for a middle

Be ready to explain that a chunk lands per superstep, not per statement, and to sketch a streaming endpoint that iterates astream() and maps chunks onto wire events.

for a senior

Show the operational side: choosing async, filtering internal state out of what reaches the browser, and cancelling the run when the client disconnects so an abandoned stream stops burning tokens.

for a principal

Own the contract question — what your service promises to emit, how streaming interacts with retries and idempotency, and whether the wire protocol stays stable when the graph's internal node layout is refactored.

## The two entry points Compiling a `StateGraph` gives you an object with the usual Runnable-style entry points: `invoke`/`ainvoke` and `stream`/`astream`. They execute the identical graph. The difference is purely in what comes back and when. - `graph.invoke(input, config)` → runs to completion, returns the final state object (for example the final `TypedDict`). - `graph.stream(input, config, stream_mode=...)` → returns an iterator; each `next()` produces a chunk describing something that just happened. When the iterator is exhausted, the run is over. `astream` is the async version and is what a web server normally uses, because it lets one event loop serve many concurrent runs while they wait on model calls. ## What a chunk corresponds to LangGraph executes in supersteps: it picks the set of nodes whose triggers are satisfied, runs them (in parallel if there is more than one), applies their returned updates to state through the channel reducers, then repeats. A stream chunk is emitted per superstep — so a graph that goes `agent → tools → agent → END` produces roughly one chunk per hop rather than one chunk per line of code inside a node. That granularity matters for expectations. Streaming does not give you a live view *inside* a node's body; a node that spends 20 seconds in a slow HTTP call emits nothing during those 20 seconds unless it writes something itself. Two mechanisms exist for finer granularity: token-level streaming from chat models, and custom writes from inside a node. ## stream_mode decides the payload `stream()` takes a `stream_mode` argument. The commonly used values are `"values"` (the whole state after each step), `"updates"` (only what each node returned, keyed by node name), `"messages"` (chat-model token chunks with metadata), `"custom"` (whatever a node writes through the stream writer), and `"debug"` (verbose execution records). Current versions also expose `"checkpoints"` and `"tasks"` for lower-level execution detail. If you pass no `stream_mode`, you get the graph's default, which is `"values"`. You can pass a list — `stream_mode=["updates", "messages"]` — and then each yielded item is a `(mode, chunk)` tuple so the consumer can tell which stream a payload came from. This is the normal shape for a chat endpoint that wants both tokens and step-level progress on one connection. ## Sync vs async, and why async usually wins `stream()` is a plain generator; `astream()` is an async generator. In a FastAPI or similar server you almost always want `astream()` inside a server-sent-events or WebSocket handler, because blocking a worker thread for the duration of an agent run scales badly. Async also matters for the observability APIs: the fine-grained event stream (`astream_events`) is async-only. ## Practical shape of a streaming endpoint A typical handler does: build the input, build a `config` (including `configurable` values such as a thread id when a checkpointer is configured), iterate `astream(...)`, translate each chunk into a wire event, and finish. Two habits save pain later. First, decide deliberately what leaves the server — raw graph state often contains retrieved documents, tool arguments, or internal scratchpads you do not want in the browser. Second, handle client disconnects: an abandoned iterator means the graph keeps running (and keeps spending tokens) unless you cancel the task. ## Common misconceptions *Streaming makes it faster.* It does not reduce total latency by a millisecond; it reduces **perceived** latency by showing work in progress. Total wall-clock time is the same. *Streaming returns a different result.* It does not. With `stream_mode="values"` the last chunk is the same object `invoke()` would have returned; some code even implements `invoke` semantics by consuming the stream and keeping the last value. *You must stream to see intermediate steps after the fact.* You do not — a run's step-by-step history is also recoverable from persisted checkpoints and from a tracing backend. Streaming is the live channel, not the only record. *Every node's internals show up.* Only supersteps, model tokens, and explicit custom writes surface. Anything else inside a node body is invisible to the stream.

  • Does streaming change the total latency of the run?
    No. The graph does exactly the same work in the same order, so wall-clock time to the final state is unchanged. What streaming buys is perceived latency: the user sees the first token or the first step's progress in a second instead of staring at a spinner until the whole agent loop finishes. It also gives you a place to hang cancellation, since abandoning the iterator is your cue to stop the run.
  • If a node makes one slow HTTP call, what do you see on the stream while it runs?
    Nothing, by default. Chunks are emitted between supersteps, so a node that blocks for twenty seconds is silent for twenty seconds. To get progress from inside a node body you must emit it explicitly through the custom stream writer, or split the work into more nodes so each completion produces a chunk.
  • Why would a server handler prefer astream() over stream()?
    Because agent runs are dominated by waiting on model and tool I/O. An async generator lets one event loop hold hundreds of in-flight runs, while the sync generator pins a worker thread per run. Async is also required for the fine-grained event API, which is async-only, so an async handler keeps both options open.

saying these in an interview costs you the question

  • Claims streaming makes the graph finish faster
  • Thinks stream() yields once per node line of code
  • Assumes the streamed chunks differ from invoke's result
  • Believes streaming replaces persistence or tracing
  • Streams raw state to the browser without filtering

context