When would you use astream_events on a LangGraph graph instead of stream_mode?
answer
- a run tree, not a step stream
- one lens is product, one is debugging
- start and end per nested runnable
- exposes prompts and tool arguments
- async only, and very chatty
basics
~20 sastream_events emits one fine-grained event per nested runnable — chain, chat model, tool — with run ids, tags and LangGraph node metadata. Reach for it when debugging what actually ran inside a node; a stream_mode stream is cheaper for shipping product output.
solid answer
~40 sA compiled graph is a Runnable, so it also exposes `astream_events()`, which yields the langchain-core event stream rather than LangGraph's step-shaped chunks. Each event is a dict with `event` (`on_chain_start`, `on_chat_model_stream`, `on_tool_start`, `on_tool_end`, and their siblings), `name`, `run_id`, `parent_ids`, `tags`, `metadata` and `data`. LangGraph enriches the metadata with `langgraph_node`, `langgraph_step` and the checkpoint namespace, so you can reconstruct exactly which node invoked which tool with which arguments — visibility the step-level modes cannot give you because they only report what a node *returned*. The cost is volume: every nested runnable emits start and end events, so a modest agent turn produces hundreds. It is also async-only. Filter with `include_names`, `include_tags` or `include_types` and treat it as a debugging and instrumentation surface, not the default transport for a user-facing stream.
code
python · 9 linesasync for event in graph.astream_events({"topic": "cats"}):
kind = event["event"]
node = event["metadata"].get("langgraph_node")
if kind == "on_chat_model_stream":
print(event["data"]["chunk"].content, end="")
elif kind == "on_tool_start":
print(f"\n[{node}] calling {event['name']} with {event['data']['input']}")
elif kind == "on_tool_end":
print(f"\n[{node}] {event['name']} -> {event['data']['output']}")go deeper
Know that a compiled graph also exposes an event stream separate from stream_mode, and that it reports starts and ends for models and tools rather than graph steps.
Be able to name the main event types and where token text and tool output sit in the payload, and to explain that this stream shows a node's interior while stream_mode shows its boundaries.
Demonstrate the tradeoff in production terms: event volume, async-only, filtering by tag, config propagation gaps that silently drop inner events, and when a persisted trace beats a live iterator.
Own the instrumentation strategy — which signals are ephemeral debugging, which are durable telemetry, and how you keep an event-consuming integration from coupling your public behaviour to internal graph composition.
## Two different lenses `stream_mode` is LangGraph's own vocabulary: supersteps, node deltas, state snapshots, model tokens, custom payloads. `astream_events` is langchain-core's vocabulary: a tree of runs, where every runnable that executes — the graph, each node, each chat model, each tool, each parser nested inside a node — announces its start, its streaming output, and its end. That difference decides when each is right. If your question is "what should the user see?", the answer is `stream_mode`. If your question is "why did this run take that path, and what arguments did it pass to the search tool?", the answer is `astream_events`. ## The event shape Each yielded event is a dict with a stable set of keys: - `event` — the type, such as `on_chain_start`, `on_chain_end`, `on_chat_model_start`, `on_chat_model_stream`, `on_chat_model_end`, `on_tool_start`, `on_tool_end`, `on_retriever_start`, `on_retriever_end`. - `name` — the runnable's name; for a node it is the node name. - `run_id` and `parent_ids` — identity and lineage, which is what lets you rebuild the call tree offline. - `tags` and `metadata` — anything bound via config, plus LangGraph's own additions (`langgraph_node`, `langgraph_step`, the checkpoint namespace). - `data` — payload: `input` on start events, `output` on end events, `chunk` on stream events. Token text lives at `event["data"]["chunk"].content` on `on_chat_model_stream`. Tool results live at `event["data"]["output"]` on `on_tool_end`. Those two, plus the node name in metadata, cover most debugging. ## What it shows that stream_mode cannot The step-level modes describe boundaries. They tell you node `agent` ran and returned a message with two tool calls; they do not tell you what happened inside the node before it returned. The event stream exposes the interior: - The exact prompt sent to the model (`on_chat_model_start` input) — invaluable when a prompt template silently rendered wrong. - Each tool invocation with its arguments and its result, in order, even when several tools run inside one node. - Timing, because start and end events are separate — you can attribute latency to a specific model call rather than to "the agent node". - Nested structure via `parent_ids`, so a node containing a sub-chain is legible as a tree. ## The costs **Volume.** Every nested runnable emits at least two events, and every token emits one. A single agent turn with two tool calls easily crosses several hundred events. Pushing that to a browser is wasteful; parsing it in a hot loop costs real CPU. **Async only.** There is no synchronous equivalent, so a sync codebase must adapt. **Coupling to internals.** Event names track the runnable composition of your graph. Refactor a node from one chain into three and the event stream changes shape. Consumers that pattern-match on `name` are brittle; filtering on tags you set deliberately is far more stable. **Callback propagation.** As with token streaming, an async node on Python 3.10 that fails to thread its `RunnableConfig` into inner calls produces a stream missing those inner events entirely — with no error. If the events for one node's interior are mysteriously absent, suspect config propagation before suspecting the framework. ## Filtering The method accepts `include_names`, `include_types`, `include_tags` and the matching `exclude_*` arguments. In practice: tag the runnables you care about at construction time, then filter on that tag. It keeps the consumer decoupled from internal naming and cuts the volume by an order of magnitude. ## How this relates to tracing Events are the live, in-process view; a tracing backend is the durable, after-the-fact view of the same underlying run tree. When you are inside a debugging session with a reproducible input, events are faster — no upload, no UI, and you can assert on them in a test. For an incident on a run that already happened in production, tracing wins, because nobody was holding an iterator at the time. Mature teams use both, and it is a good answer to say so explicitly. ## Interview-worthy summary Use `stream_mode` for the product. Use `astream_events` when you need to see inside a node — prompts, tool arguments, per-call latency — or when you are building instrumentation that must react to specific inner events. Never make it the default transport of a user-facing endpoint without filtering, and never let consumers pattern-match on runnable names you expect to refactor.
- Which event and field carry the token text and the tool result?Token text arrives on on_chat_model_stream at data.chunk.content, and a tool's result arrives on on_tool_end at data.output, with the tool's arguments visible on the matching on_tool_start input. LangGraph adds langgraph_node to each event's metadata, so you can attribute every one of them to the node whose body made the call.
- Why not just use astream_events for the user-facing stream too?Volume and coupling. A single agent turn emits hundreds of events, most of which the browser must discard, and the event names mirror your internal runnable composition, so refactoring a node into a sub-chain changes the client's view. A step-shaped stream is smaller and expresses a contract you actually intend to keep stable.
- Events for the interior of one async node are missing entirely and nothing raises. What do you suspect?Callback propagation. On Python 3.10 an async node body that calls a model or chain without passing through the RunnableConfig it received leaves those inner runnables uninstrumented, so their events never appear while the run itself succeeds. Accept config in the node signature and forward it into every inner call.
- How do you keep an event consumer from breaking every time the graph is refactored?Filter on tags you attach deliberately rather than on runnable names. Names track internal composition and change when a node is split or a chain is inlined, whereas a tag such as user_facing or billable is a contract you chose. The include_tags argument then does the volume reduction and the decoupling at once.
saying these in an interview costs you the question
- Calls astream_events a stream_mode value
- Uses it as the default browser transport unfiltered
- Pattern-matches on runnable names and calls it stable
- Expects a synchronous equivalent to exist
- Thinks it replaces a persistent tracing backend