skip to content

What event sequence does the Anthropic Messages API emit over SSE when streaming?

level: middleimportance: must knowfreq 72%

answer

  1. Start, per-block run, delta, stop
  2. Blocks are indexed, not one string
  3. Metadata arrives last, not first
  4. message_stop is the completeness signal
  5. Unknown event types must be skipped

basics

~10 s

An Anthropic stream opens with message_start, then for each content block a content_block_start, a run of content_block_delta chunks, and a content_block_stop. It closes with message_delta carrying the final stop_reason and output usage, then message_stop.

solid answer

~40 s

Setting `stream: true` on a Messages request turns the response into a Server-Sent Events stream. First comes `message_start`, whose payload is the full message shell with an empty `content` array and `usage.input_tokens` already filled in. Then, for every content block in order, you get `content_block_start` with an `index` and a stub block, a run of `content_block_delta` events carrying the actual fragments, and a `content_block_stop`. After the last block, `message_delta` delivers the fields that could not be known up front — `stop_reason`, `stop_sequence`, and the output token count — and `message_stop` terminates the stream. `ping` keepalives and `error` events may appear at any point, and Anthropic reserves the right to add new event types, so a parser must ignore anything it does not recognise rather than throwing.

go deeper

for a junior

Be able to name the events in order and say plainly that the text arrives in content_block_delta chunks you append as they come, and that message_stop marks the end.

for a middle

Explain the per-block structure: why every delta carries an index, what content_block_start and content_block_stop bracket, and which event finally reveals stop_reason and output token usage.

for a senior

Show that your client is a state machine with a completeness check — you fail a stream that ends without message_stop, you skip unknown event types, and you re-encode the stream into your own protocol when relaying it to a browser.

for a principal

Own the contract between your service and its clients: what your own event protocol guarantees, how partial output is represented downstream, and how you keep parsers forward-compatible as the provider adds event and delta types.

## Why stream at all The Messages API can hand back the whole assistant message as one JSON body, or — when you send `"stream": true` — deliver it incrementally as a Server-Sent Events (SSE) stream over the same HTTP request. SSE is a one-way text protocol: the server responds `200` with `content-type: text/event-stream` and then writes frames shaped as an `event:` line, a `data:` line holding JSON, and a blank line, keeping the connection open until it is done. Streaming buys time-to-first-token, so a user sees prose while the model is still generating, and it is the practical shape for long generations, which Anthropic explicitly recommends streaming for rather than waiting on a single buffered response. ## The event sequence **`message_start`** — exactly one, first. Its payload is a `message` object that looks like a normal non-streaming response with `content` set to an empty array: you get the `id`, `model`, `role`, and a `usage` object whose `input_tokens` is already final (along with the cache-related input counters when prompt caching is in play). Output counts are not meaningful yet. **Per content block** — a message may contain several blocks (text, tool use, thinking), and they stream one at a time, in order. Each opens with **`content_block_start`**, which carries an integer `index` and a stub of the block, for example `{"type": "text", "text": ""}` or a `tool_use` block with its `id` and `name` but an empty `input`. Then come one or more **`content_block_delta`** events, each with the same `index` and a `delta` object holding the fragment. The block is closed by **`content_block_stop`** with that index. Nothing about the block is guaranteed complete until its stop event. **`message_delta`** — one event near the end, carrying the top-level fields that were unknowable at `message_start`: `delta.stop_reason`, `delta.stop_sequence`, and a `usage` object with the message's output token count. If you bill, log, or branch on why generation ended, this is where the information lives. **`message_stop`** — the terminator. Its arrival is your only positive signal that the message finished cleanly; a stream that just goes quiet was truncated. **Interleaved anywhere** — `ping` events, which are keepalives with no content and should simply be skipped, and `error` events, which carry an error object such as `{"type": "overloaded_error", "message": "Overloaded"}` and can arrive after content has already been delivered. ## Reading it in practice Most people never hand-parse the frames. The official SDKs expose two levels: a raw iterator over the typed events, and a higher-level streaming helper that accumulates the events into a message for you while still letting you observe text as it arrives. The raw level matters when you are relaying the stream to your own frontend, since you usually re-encode it into your own event protocol rather than forwarding Anthropic's frames verbatim — that keeps your API key off the browser and lets you inject your own error and metadata frames. ## The two rules that catch people *Order is per-block, not global.* You cannot assume one text block; do not concatenate every delta you see into a single string without honouring `index`, or a message that opens a tool-use block mid-answer will corrupt into nonsense. *Forward compatibility is required, not polite.* Anthropic documents that new event types may be added, so a `match`/`switch` that raises on an unknown `type` is a latent production outage. Skip unknown events and unknown delta types; keep the ones you understand. ## What interviewers are really checking The named sequence is the surface question. Underneath, they want to see that you know (a) the stream carries structure, not just characters, so your client is a small state machine rather than a string buffer; (b) `message_stop` is the completeness signal, so a client that renders whatever arrived and calls it done will silently ship truncated answers; and (c) the interesting metadata — stop reason and output usage — arrives at the *end*, which is why you cannot compute the cost of a request from `message_start` alone. A candidate who can sketch the sequence and then say what each event is *for* is well past the recitation answer.

  • What happens if your parser throws on an event type it has never seen?
    You get a self-inflicted outage the next time Anthropic adds an event type, since new types can appear without a breaking-change release. The documented contract is that clients ignore unrecognised event types and unrecognised delta types. Practically: default your event switch to a no-op, log the unknown type at debug level for visibility, and only hard-fail on events you positively need and cannot interpret.
  • How would you know a stream was truncated rather than finished?
    A clean finish always ends with `message_stop`, preceded by `message_delta` carrying a `stop_reason`. If the connection closes without those, the message is incomplete no matter how much text you collected — treat it as a failure, not a short answer. In a relay you should propagate that as an explicit error frame to your own client, because the partial text on screen looks identical to a complete short reply.
  • Does a streamed response ever contain more than one content block?
    Yes. A message can contain several blocks — for example text followed by a tool-use block, or thinking followed by text — and each streams as its own start/delta/stop trio, sequentially. That is why every delta carries an `index`: your accumulator keys buffers by index and assembles a list of blocks, not one flat string.

saying these in an interview costs you the question

  • Claiming the whole message arrives in one delta event
  • Concatenating every delta into one string, ignoring index
  • Believing message_start already contains the output text
  • Treating an unknown event type as a fatal parse error
  • Assuming stop_reason is available at the start of the stream

context