skip to content

Streaming responses via SSE

Reading the SSE stream instead of waiting for a whole response: message_start, a run of content_block_delta chunks you accumulate yourself, then message_delta carrying the final stop reason and usage. Interviewers ask what you do when an error arrives mid-stream, after bytes are already on the user's screen.

on this pageshow

questions

6

How do you accumulate Anthropic content_block_delta events into a final message?

level: middleimportance: must knowfreq 60%

answer

  1. Buffer per index, never one string
  2. Deltas are incremental, not cumulative
  3. partial_json is not valid JSON alone
  4. Parse tool arguments at block stop
  5. SDK helper assembles it for you

basics

~20 s

Key a buffer by the event's index field and append each delta to that block. text_delta contributes its text; input_json_delta contributes a partial_json string fragment that you concatenate and parse as JSON only once the block's content_block_stop arrives.

solid answer

~40 s

Every `content_block_delta` carries an `index` identifying which block it belongs to, plus a `delta` object whose `type` tells you how to apply it. Maintain a map from index to a buffer opened by `content_block_start`. For a `text_delta` you append `delta.text` to that block's string. For a `tool_use` block the deltas are `input_json_delta` events whose `partial_json` fragments are chunks of a JSON document split at arbitrary byte boundaries — you concatenate them into a string and only call your JSON parser once, when `content_block_stop` for that index arrives. Extended-thinking blocks stream `thinking_delta` and `signature_delta` the same way. Deltas are strictly incremental, never cumulative, so appending is correct and replacing is not. The official SDKs' streaming helpers do all of this for you and hand back an assembled message.

go deeper

for a junior

Know that each delta holds only the newly generated piece, so you append rather than replace, and that the SDK's streaming helper can assemble the whole message for you.

for a middle

Explain the accumulator: a buffer per index, opened at content_block_start, appended by text_delta or input_json_delta, and finalised at content_block_stop with a single JSON parse for tool inputs.

for a senior

Demonstrate that your assembled message is byte-equivalent to the non-streaming response, that you persist whole blocks rather than fragments, and that partial tool arguments are never acted on or shown through a strict parser.

for a principal

Frame accumulation as a contract boundary: whatever your relay emits to clients must be reconstructible and testable against the non-streamed response, and must degrade predictably when a block never reaches its stop event.

## The shape of a delta event A streamed Anthropic message is a list of content blocks, and each block is filled in by a run of `content_block_delta` events. A delta event has two fields that matter: `index`, the position of the block it belongs to, and `delta`, an object with its own `type` discriminator. The events for a block are bracketed by `content_block_start` at that index and `content_block_stop` at the same index. Blocks are emitted sequentially, but you should still key on `index` rather than assuming a single current block — that assumption is exactly what breaks when a message answers in text and then opens a tool-use block. ## Delta types and how each accumulates **`text_delta`** carries a `text` field: the newly generated fragment, not the text so far. You append it. Fragments split at token boundaries, which means they can cut through the middle of a word, a Markdown fence, or a multi-byte character sequence at the transport level — never assume a delta is a whole word or a whole line. **`input_json_delta`** carries a `partial_json` field and appears inside a `tool_use` block. The model's tool arguments are serialised to JSON and then chopped up; a single fragment is very often not valid JSON on its own (`{"loc`, `ation": "Par`, `is"}`). The rule is: concatenate every `partial_json` fragment for that index into one string, and parse it exactly once, at `content_block_stop`. Trying to `json.loads` each fragment, or to re-parse the growing buffer on every delta so you can render arguments early, is the single most common bug in hand-rolled accumulators. If you genuinely need to show arguments as they arrive, use a streaming/partial JSON parser that is designed to tolerate truncation — do not point a strict parser at half a document. **`thinking_delta` and `signature_delta`** appear when extended thinking is enabled: the first accumulates the reasoning text, the second the cryptographic signature attached to the thinking block. They accumulate by the same append rule. ## The assembled result When the stream reaches `message_stop`, you should be able to reconstruct precisely the JSON body you would have received from a non-streaming call: the shell from `message_start`, `content` populated from your per-index buffers, and `stop_reason`, `stop_sequence` and output usage patched in from `message_delta`. That equivalence is a useful design target and an even better test — run the same prompt with and without streaming at temperature 0 and assert your accumulator produces the same structure. ## Let the SDK do it The official SDKs ship a streaming helper that maintains this state for you: it exposes a text-only iterator for the common "print tokens as they arrive" case, and a call that returns the fully assembled final message once the stream drains. Reach for the raw event iterator only when you are building a relay or need per-event control; even then, the helper's assembled message is usually what you persist to your conversation history, because the next turn needs whole blocks, not fragments. ## Where it goes wrong in production *Rendering partial JSON.* Users see a flicker of malformed arguments, or your parser throws mid-stream and kills the request. *Flattening indices.* Concatenating all deltas into one string works fine in demos with a single text block and corrupts the first time a message contains two blocks. *Persisting fragments.* Writing individual deltas into your conversation store, then replaying them as history, produces a message the API cannot interpret; store the assembled blocks. *Assuming cumulative deltas.* Some other streaming protocols resend the whole value each time. Anthropic's do not — a client that replaces instead of appending ends up with only the last fragment of every block. *Trusting a buffer before its stop event.* A block is only complete at `content_block_stop`; acting on a tool call before then means acting on truncated arguments.

  • Why can't you just json.loads each input_json_delta as it arrives?
    Because `partial_json` fragments are arbitrary slices of one JSON document, not standalone documents — a fragment can end mid-key or mid-string. A strict parser will throw on almost every fragment. You concatenate the fragments and parse once at `content_block_stop`; if you must render arguments live, use a parser explicitly built to tolerate truncated JSON.
  • What breaks if you concatenate every delta into a single string regardless of index?
    Any message with more than one content block collapses into nonsense. A reply that emits text and then a tool-use block would have raw JSON argument fragments spliced into the visible prose, and the tool call itself would be unrecoverable. Keying buffers by `index` and closing them at their own `content_block_stop` is what keeps blocks separate.
  • What should you store in conversation history after streaming a reply?
    The assembled message — whole content blocks, with the assistant role — not the individual delta events. The next request replays prior turns as complete messages, so fragments are meaningless to the API. Take the final assembled message from the SDK's streaming helper, or from your own accumulator once message_stop arrives, and persist that.

saying these in an interview costs you the question

  • Parsing each partial_json fragment as standalone JSON
  • Assuming a text_delta contains the full text so far
  • Ignoring the index and appending everything to one buffer
  • Acting on tool arguments before content_block_stop arrives
  • Saving raw delta events as conversation history

context

open as a page

What event sequence does the Anthropic Messages API emit over SSE when streaming?

level: middleimportance: must knowfreq 72%

basics

~10 s

An Anthropic stream opens with message_start, then for each content block a content_block_start, a run of content_block_delta chunks, and a content_block_stop. It closes with message_delta carrying the final stop_reason and output usage, then message_stop.

open as a page

How do you handle an Anthropic SSE error event that arrives mid-stream?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Treat it as a failed request even though bytes already reached the user. The HTTP 200 and headers are long gone, so surface the failure in your own stream protocol, keep or discard the partial text deliberately, and retry the whole call — there is no resume point.

open as a page

In the Anthropic Python SDK, how does messages.stream() differ from create(stream=True)?

level: juniorimportance: should knowfreq 52%

basics

~20 s

create(stream=True) returns a raw iterator of stream events that you assemble yourself. messages.stream() is a context manager returning a helper that accumulates the events for you, exposes a text-only iterator, hands back the finished message, and closes the HTTP response on exit.

open as a page

Which Anthropic streaming event carries the final stop_reason and output usage?

level: middleimportance: should knowfreq 48%

basics

~10 s

message_delta. It arrives after the last content block and carries delta.stop_reason, delta.stop_sequence, and a usage object with the message's output token count — none of which exist yet at message_start.

open as a page

How do you detect a stalled Anthropic SSE stream, and what timeout should you set?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Time the gap between events, not the whole request. Reset a watchdog on every frame including ping keepalives, and fail the stream when the idle gap exceeds your budget. A total wall-clock timeout kills legitimate long generations instead.

open as a page