skip to content

Which Anthropic streaming event carries the final stop_reason and output usage?

level: middleimportance: should knowfreq 48%

answer

  1. Input known early, output known late
  2. One event before the terminator
  3. stop_reason is a diff, not a start field
  4. Drain the stream to get the numbers
  5. message_stop carries no payload

basics

~10 s

message_delta. It arrives after the last content block and carries delta.stop_reason, delta.stop_sequence, and a usage object with the message's output token count — none of which exist yet at message_start.

solid answer

~40 s

`message_start` gives you the message shell with `usage.input_tokens` already final, because the input was known before generation began; its output count is a placeholder. Everything that depends on how generation *ended* arrives on `message_delta`, one event before `message_stop`: `delta.stop_reason`, `delta.stop_sequence`, and a `usage` object reporting output tokens for the message. That split has practical consequences. You cannot compute the cost of a call, decide whether the reply was truncated at `max_tokens`, or branch into a tool-execution loop until you have consumed `message_delta` — so a client that renders text and closes the connection early is throwing away the accounting and control-flow signal. In the SDKs, the assembled final message from the streaming helper carries both fields, which is the simplest way to get them.

go deeper

for a junior

Remember that the stop reason and token counts arrive at the end of the stream, on message_delta, and that message_stop is just the terminator.

for a middle

Explain the causal split — input counts are final at message_start, output counts and stop_reason only after generation ends — and that message_delta is a diff applied to the message shell.

for a senior

Show that you record usage and stop reason from the tail of every stream, alert on rising budget-limited stops, and never branch an agent loop on a partially streamed block.

for a principal

Treat the stop reason and usage tail as your cost and quality telemetry: budget guards, truncation rates and per-tenant accounting all depend on draining every stream, so make that a platform guarantee rather than per-client discipline.

## The split: what is known when A streamed message reveals its metadata in two places, and the split follows from causality. **At `message_start`**, the server already knows everything about the *input*: the message `id`, the `model` that will serve it, the role, and `usage.input_tokens` — plus the cache-related input counters when prompt caching is used. What it cannot know is how the generation will end or how long it will run, so the output side is a placeholder, not a prediction. **At `message_delta`**, one event after the final `content_block_stop`, the server patches in exactly the fields that were unknowable at the start: `delta.stop_reason`, `delta.stop_sequence` (populated only when a custom stop sequence triggered), and a `usage` object with the output token count for the message. Then `message_stop` closes the stream. A useful way to hold it: `message_delta` is a diff against the message shell you received at the start. Apply it, add your accumulated content blocks, and you have reconstructed the non-streaming response body exactly. ## Why it matters operationally **Cost accounting.** Output tokens are typically the more expensive side of a call, and you do not have that number until `message_delta`. Any per-request cost metric, budget guard or user-facing usage meter must therefore be recorded at the end of the stream, from that event or from the SDK's assembled final message — never estimated from what you happened to render. **Truncation detection.** The stop reason distinguishes a model that finished its thought from one that hit your `max_tokens` ceiling. Without it, a reply cut off mid-sentence looks like any other reply. Services that silently truncate long answers almost always have this bug: they stream text to the UI and never consume the terminal events. Log the stop reason on every call; a rising share of max-token stops is a signal your ceiling or your prompt needs work. **Control flow.** In an agent loop, the stop reason is what tells you the model wants a tool executed rather than having produced a final answer. Streaming does not change the loop, but it changes when you learn the branch: not while the block is arriving, but after `message_delta`. Kicking off tool execution from a partially streamed block is the corresponding anti-pattern — the arguments are not complete until the block's stop event, and the branch is not confirmed until the message-level delta. **Client shutdown.** A UI that stops reading the moment the visible text looks finished will drop `message_delta` entirely, and with it the usage numbers and the stop reason. Drain the stream to `message_stop` even when you have nothing left to render; the tail is cheap and carries the information you will want in your logs. ## Reading it in code With the raw event iterator you branch on the event type and pull the fields off `message_delta` yourself. With the streaming helper the SDK assembles everything and hands you a final message object whose `stop_reason` and `usage` are already populated — the same object shape a non-streaming call would have returned. That equivalence is the point: streaming should change *when* you learn things, not *what* you can learn. ## Common confusions worth naming *Expecting usage on `message_stop`.* It is a bare terminator; the numbers arrived one event earlier. *Reading output tokens from `message_start`.* The output has not been generated yet, so the value there is not the answer to any question you care about. *Counting tokens client-side instead.* Estimating output length by counting characters you rendered is guesswork and will not match billing; the server's number is authoritative, and it is right there in the stream. *Assuming input tokens change during the stream.* They do not — the input was fixed when the request was sent, which is why that half of the accounting is available immediately.

  • Why is input_tokens available at message_start but output_tokens not?
    The input was fully determined when you sent the request, so the server can count it before generating a single token. The output length depends on how generation unfolds and is only final once the model stops, which is why it is patched in at message_delta. The same asymmetry explains why cache-related input counters also appear at the start.
  • What breaks if your client stops reading once the visible text looks complete?
    You lose message_delta and therefore the stop reason and output token count. Cost metrics become estimates, truncated replies stop being detectable, and in an agent loop you never learn that the model asked for a tool. Draining the last two events costs nothing; make it unconditional in the accumulator rather than a caller responsibility.
  • How do you detect that a reply was cut off by your max_tokens setting?
    Check the stop reason delivered on message_delta: a generation that ran out of budget reports that explicitly, whereas a model that finished its answer reports a normal end. Log it per request. If the share of budget-limited stops climbs, either your ceiling is too low for the task or the prompt is inviting longer answers than you want to pay for.

saying these in an interview costs you the question

  • Expecting usage numbers on the message_stop event
  • Reading output token counts from message_start
  • Estimating output tokens by counting rendered characters
  • Closing the connection before message_delta arrives
  • Thinking input token counts change during the stream

context