In the Anthropic Python SDK, how does messages.stream() differ from create(stream=True)?
answer
- Two front doors, one wire protocol
- One returns events, one returns a helper
- Context manager closes the response
- text_stream for display, final message for logs
- Raw events when you relay them onward
basics
~20 screate(stream=True) returns a raw iterator of stream events that you assemble yourself. messages.stream() is a context manager returning a helper that accumulates the events for you, exposes a text-only iterator, hands back the finished message, and closes the HTTP response on exit.
solid answer
~50 sBoth hit the same endpoint with streaming enabled; they differ in how much work the SDK does for you. `client.messages.create(model=..., max_tokens=..., messages=..., stream=True)` returns an iterator of typed raw events — `message_start`, `content_block_delta`, `message_delta` and so on — and it is entirely on you to key buffers by index, accumulate deltas, and remember to close the response. `with client.messages.stream(...) as stream:` gives you a helper object instead: iterate `stream.text_stream` to get only the text fragments for display, and call `stream.get_final_message()` after the stream drains to receive the fully assembled message with its content blocks, `stop_reason` and usage. Because it is a context manager, the underlying HTTP connection is released even if you break out early or an exception is raised. Use the helper by default; drop to the raw iterator when you are relaying events into your own protocol.
code
python · 16 linesimport os
import anthropic
client = anthropic.Anthropic()
with client.messages.stream(
model=os.environ["ANTHROPIC_MODEL"],
max_tokens=256,
messages=[{"role": "user", "content": "Say hello in one sentence."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print()
print(final.stop_reason, final.usage.output_tokens)go deeper
Know that stream=True gives you raw events to assemble yourself, while messages.stream() is a with-block helper offering text_stream for display and get_final_message() at the end.
Explain what the helper is actually doing for you — per-index accumulation, tool-argument joining, connection lifetime — and why that is the safer default outside a relay.
Justify the choice in a real service: helper for application code, raw events only where you re-encode into your own protocol, and always persist the assembled message rather than the fragments you forwarded.
Standardise it across teams — one wrapper that guarantees drained streams, released connections, recorded usage and a consistent client-facing event protocol, so no service hand-rolls its own accumulator.
## Same wire protocol, two ergonomics Streaming is a request flag: the Messages endpoint returns Server-Sent Events instead of one JSON body. The Python SDK gives you two front doors onto that same stream, and choosing between them is one of the first practical decisions anyone integrating Claude makes. ## The raw iterator: `create(stream=True)` Passing `stream=True` to `client.messages.create(...)` changes the return type from a `Message` to an iterator of typed stream events. You loop over it and see the protocol as it is: `message_start`, then per content block a `content_block_start`, a run of `content_block_delta`, and a `content_block_stop`, then `message_delta` and `message_stop`, with `ping` events sprinkled through. Every event is a typed object, so `event.type` and the nested fields are attribute access rather than dictionary poking. What you own at this level: routing deltas into per-index buffers, joining `partial_json` fragments for tool blocks and parsing them once, pulling `stop_reason` and usage off `message_delta`, verifying that `message_stop` actually arrived, and making sure the HTTP response is closed if you stop consuming early. None of that is hard, and all of it is easy to get subtly wrong. ## The helper: `messages.stream()` `client.messages.stream(...)` takes the same arguments but is used as a context manager. Inside the `with` block you get a stream helper that has already done the bookkeeping: - `stream.text_stream` yields just the text fragments, which is exactly what a chat UI wants to render — no filtering on event types, no worrying about non-text blocks. - Iterating the helper itself still yields the individual events, so you can observe the protocol while the accumulator runs underneath. - `stream.get_final_message()` returns the assembled `Message` once the stream has drained: whole content blocks, `stop_reason`, and the usage numbers — the same object shape a non-streaming call would have returned. - Leaving the `with` block releases the HTTP response, including on an early `break` or an exception. An asynchronous client offers the same shape with `async with` and `async for`, and the TypeScript SDK mirrors the design with its own `messages.stream()` helper. ## Which to reach for **Default to the helper.** For a chat surface, a CLI, or a batch job that streams for latency reasons, it removes the entire accumulation surface where bugs live, and it still gives you the final message for logging, cost accounting and conversation history. **Use the raw iterator when you are the middle of a pipe.** If your backend re-encodes provider events into your own frontend protocol, you want to see the events, translate them into frames your client understands, and attach your own metadata and error frames. Even then, many teams run the helper's accumulation in parallel so that the message they persist is the assembled one rather than something they stitched together from the frames they happened to forward. ## Mistakes this distinction prevents *Forgetting to close the response.* With the raw iterator, abandoning the loop halfway leaves a connection held until garbage collection. The context manager makes that impossible to forget. *Re-implementing accumulation badly.* Hand-rolled accumulators routinely flatten block indices or parse tool-argument fragments too early. The helper has none of those bugs. *Losing the tail.* Rendering text and walking away drops `message_delta`, so you never learn the stop reason or output token count. `get_final_message()` forces you to have drained the stream before you can read them. *Confusing the return type.* `create()` without the flag returns a finished `Message`; with `stream=True` it returns an iterator. Code that assumes the former and receives the latter fails at the first attribute access, which is a common first-day error.
- Why is messages.stream() a context manager rather than a plain call?Because a streaming call holds an open HTTP response until it is drained or released. The `with` block guarantees the connection is closed on early break or exception, which a bare iterator cannot do. It also gives the helper a well-defined point at which accumulation is finished, which is why get_final_message() is meaningful only after the block has drained.
- What do you get from stream.get_final_message() that text_stream cannot give you?The whole message: every content block including non-text ones such as tool use, plus stop_reason and the usage numbers. text_stream deliberately yields only text fragments for display. You render from text_stream and log, bill and branch from the final message — and you persist the final message as conversation history, since fragments are meaningless to the next request.
- When would you still choose the raw event iterator?When you are a relay. If your backend translates provider events into your own frontend protocol, you need to see each event to map it onto your frames, inject your own error and metadata frames, and decide what to expose. Outside that case the helper is the better default, because the accumulation it replaces is exactly where hand-rolled clients tend to be buggy.
saying these in an interview costs you the question
- Expecting create(stream=True) to return a finished Message
- Assuming text_stream also yields tool-use content
- Calling get_final_message() before the stream has drained
- Abandoning a raw event iterator without closing the response
- Thinking the two entry points hit different endpoints