How do you stream LLM tokens from inside a LangGraph node to the client?
answer
- a dedicated stream mode, not a node change
- chunk plus metadata, not bare text
- metadata names the emitting node
- every model call streams, including hidden ones
- empty content means tool-call fragments
basics
~20 sUse stream_mode="messages" on graph.stream/astream. LangGraph then yields (message_chunk, metadata) tuples for every chat-model token produced inside any node, and the metadata names the emitting node so you can forward only the tokens the user should see.
solid answer
~40 sIn LangGraph 1.x, `stream_mode="messages"` gives token-level output. Each yielded item is a two-tuple: an `AIMessageChunk` carrying the incremental content, and a metadata dict that includes `langgraph_node`, `langgraph_step` and any tags on the call. Because the model call is instrumented through callbacks, tokens surface even when the node calls `llm.invoke()` rather than `llm.stream()` — you do not have to restructure node code. The catch is that *every* chat-model call in the graph streams, including a summariser or a routing classifier the user should never see, so the consumer must filter on `metadata["langgraph_node"]` or on tags bound to the model. In production you normally combine modes — `stream_mode=["updates", "messages"]` — so one connection carries both step progress and tokens, and you skip chunks whose `content` is empty because tool-call deltas arrive as chunks too.
code
python · 8 linesasync for chunk, metadata in graph.astream(
{"messages": [{"role": "user", "content": "hi"}]},
stream_mode="messages",
):
if metadata["langgraph_node"] != "answer":
continue
if chunk.content:
print(chunk.content, end="", flush=True)go deeper
Know the name of the mode: stream_mode="messages" on stream/astream, and that each item is a message chunk plus metadata rather than a plain string.
Explain why a node using invoke() still streams, what langgraph_node in the metadata is for, and how you combine token streaming with step-level updates on one connection.
Show the production instincts: tag-based filtering so internal model calls stay hidden, guarding against empty tool-call chunks, config propagation in async nodes, and a defined contract for mid-stream failures.
Own the streaming protocol as a product surface — what token-level output commits you to across model providers, how cost and cancellation are handled when a client disconnects mid-generation, and whether clients should ever see node identity.
## The mechanism When you stream a compiled graph with `stream_mode="messages"`, LangGraph attaches a streaming-capable callback to the model calls made inside nodes. Each token the provider returns becomes an `AIMessageChunk`, and LangGraph pairs it with a metadata dict before yielding `(chunk, metadata)`. The useful metadata keys are the LangGraph ones — `langgraph_node` (the node whose body made the call), `langgraph_step` (which superstep), plus the tags and metadata you attached to the model or the call config. Everything you need for routing tokens to the right place in a UI is in that dict; the chunk itself only knows it is text. ## Why .invoke() inside a node still streams This surprises people. A node written as `return {"messages": [llm.invoke(state["messages"])]}` still produces tokens on the stream. The chat-model base class checks, at call time, whether a streaming-capable handler is attached, and if so it uses the provider's streaming API internally while still returning one complete message to your node. So the node's own contract is unchanged — it gets a whole `AIMessage` — while the tokens leak out through the callback channel as they arrive. Two things switch this off. A model constructed with streaming disabled (`disable_streaming=True`) will not emit chunks. And a provider or model configuration that does not support streaming at all — some structured-output or tool-forcing paths — returns in one shot, so there is nothing incremental to emit. ## Filtering: the part people get wrong A real agent graph has more model calls than the user should see: a routing classifier, a query rewriter, a summariser that compacts history, a grader in a self-correction loop. All of them stream. If you pipe the raw message stream to the browser, the user watches the classifier deliberate and then sees the answer typed over the top of it. Two filters are idiomatic: - **By node**: skip unless `metadata["langgraph_node"] == "answer"`. Simple, but couples the transport layer to node names. - **By tag**: bind a tag to the model instance you consider user-facing (`llm.with_config(tags=["user_facing"])`) and filter on the tags present in the metadata. Survives node renames and works when one node hosts several model calls. ## Chunks that are not text Not every chunk carries prose. During a tool-calling turn the model streams partial tool-call arguments; those arrive as `AIMessageChunk`s whose `content` is empty and whose `tool_call_chunks` hold fragments of a JSON argument string. Rendering them naively gives the user a burst of nothing, or worse, half-formed JSON. Guard with `if chunk.content:` before writing to the wire, and accumulate tool-call chunks only if you actually intend to show the arguments forming. ## Combining with step-level modes Token output alone is a thin UI: the user sees text appear but not that a search ran. The common production shape is `stream_mode=["updates", "messages"]`, which yields `(mode, payload)` tuples so one connection carries both. Map step updates onto status chips ("Searching…", "Reading 3 sources…") and message chunks onto the answer body. ## Async, subgraphs and the callback-propagation trap Use `astream()` in a server. On Python 3.10, async callback propagation is not automatic through context variables: if a node's async body calls a model without threading the `RunnableConfig` it was given down into that call, the callback never attaches and the tokens vanish silently — no error, just an empty token stream. Accepting `config: RunnableConfig` in the node signature and passing it through to the model call fixes it, and is harmless on newer Python. Tokens from models called inside a nested graph do surface in the parent's message stream. If you also want the nested graph's step updates, pass `subgraphs=True`, which prefixes chunks with a namespace tuple identifying the nested graph. ## What to say about failure If the run raises mid-stream, the client has already received a partial answer. Decide the contract: emit a terminal error event and let the UI mark the message as incomplete, or buffer server-side until a step boundary before flushing. Silently ending the stream is the worst option, because the client cannot distinguish it from a short answer.
- A node calls llm.invoke() rather than llm.stream(). Do tokens still appear on the stream?Yes. The chat model checks at call time whether a streaming-capable callback is attached and, when one is, consumes the provider's streaming API internally while still handing your node one finished AIMessage. The node's code and its return contract stay the same; the tokens escape through the callback channel. A model built with streaming explicitly disabled is the exception.
- Your graph has a routing classifier and an answer node. How do you keep the classifier's tokens off the user's screen?Filter in the consumer. Either skip chunks whose metadata langgraph_node is not the answer node, or — more robustly — tag the user-facing model with with_config(tags=[...]) and forward only chunks carrying that tag. Tag-based filtering survives node renames and handles a node that makes more than one model call.
- Why do some chunks arrive with empty content during a tool-calling turn?Because the model is streaming tool-call arguments, not prose. Those chunks carry fragments in tool_call_chunks while content stays empty, so a naive renderer emits nothing or leaks partial JSON. Guard with a content check before writing to the wire, and accumulate the argument fragments separately if you want to show the call forming.
- Tokens are silently missing from an async node on Python 3.10. What is the usual cause?The RunnableConfig is not being propagated into the model call. On Python 3.10 the async context does not carry callbacks automatically, so the streaming handler never attaches and the run completes normally with no token chunks at all. Accept config: RunnableConfig in the node signature and pass it into the model call; it is harmless on newer Python versions.
saying these in an interview costs you the question
- Says you must rewrite nodes to call llm.stream()
- Forwards every message chunk, including hidden classifier output
- Treats every chunk as printable text
- Assumes token streaming requires a checkpointer
- Ignores config propagation and blames the model provider