What events does Cohere's v2/chat stream emit when stream is true?
answer
- Every payload announces its own type
- Blocks are bracketed, not bare
- Fragments append, never replace
- Block number rides along on each event
- Cost and stop reason arrive last
basics
~10 sCohere streams typed events rather than uniform chunks: message-start opens the turn, content-start/content-delta/content-end wrap each content block, and message-end closes it with finish_reason and usage. You accumulate the text from content-delta events.
solid answer
~40 sSetting `stream: true` on `/v2/chat` (or calling `chat_stream()` in the SDK) returns Server-Sent Events whose payloads each carry a `type`. The sequence is `message-start`, then per content block a `content-start`, a run of `content-delta` events, and a `content-end`, finally `message-end`. Text arrives nested inside the delta — in the SDK, `event.delta.message.content.text` — and you concatenate it yourself; there is no cumulative field. Each content event carries an `index` identifying which block it belongs to, so parallel blocks interleave safely. `message-end` is where `finish_reason` and the `usage` token counts live, so a client that stops reading at the last text delta loses its cost accounting. Tool calls stream through their own event types, so a UI can render a plan before the arguments finish.
code
python · 16 linesimport cohere
co = cohere.ClientV2(api_key="COHERE_API_KEY")
buffers = {}
for event in co.chat_stream(
model="command-a-03-2025",
messages=[{"role": "user", "content": "Explain SSE in two lines."}],
):
if event.type == "content-delta":
idx = event.index
buffers[idx] = buffers.get(idx, "") + event.delta.message.content.text
elif event.type == "message-end":
print(event.delta.finish_reason)
print("".join(buffers[i] for i in sorted(buffers)))go deeper
Know that streaming sends typed events, that the text arrives in content-delta events and must be appended, and that the stream ends with a message-end event.
Walk the whole sequence from message-start to message-end, explain why the index field exists, and say exactly which event carries finish_reason and usage.
Show the production concerns: proxy buffering that kills streaming, no resume after a disconnect, bounded retries with double billing, and a defined policy for partial output on an error finish.
Decide where streaming belongs at all. Interactive surfaces earn it; batch pipelines usually do not, and the retry, cost-attribution and transcript-integrity rules for half-delivered turns should be platform policy, not per-team improvisation.
## A typed event stream, not uniform chunks Cohere's v2 chat streaming is deliberately different in style from the OpenAI-shaped streams many developers meet first. There, every chunk has the same shape and you dig into `choices[0].delta` to find out what changed. On Cohere's `/v2/chat`, each Server-Sent Event payload carries an explicit **`type`**, and your consumer switches on it. That makes the parser a state machine rather than a pile of null checks, which is the design point worth articulating in an interview. ## The event sequence For a plain text answer the order is: 1. **`message-start`** — the turn opens. Carries the generation id. Nothing renderable yet; use it to allocate a buffer and start your latency timer for time-to-first-token. 2. **`content-start`** — a content block begins, with an `index`. This is where you learn the block's type before any text arrives. 3. **`content-delta`** — repeated, each carrying a fragment of text. In the Python SDK the value is at `event.delta.message.content.text`. Fragments are **incremental**, not cumulative: you append. Appending cumulative payloads by mistake is the classic bug that produces exponentially duplicated text. 4. **`content-end`** — that block is finished. 5. **`message-end`** — the turn is over. This event carries `finish_reason` (`COMPLETE`, `MAX_TOKENS`, `STOP_SEQUENCE`, `TOOL_CALL`, `ERROR`) and the `usage` object with token counts. Tool-calling turns add their own events: a streaming tool plan arrives as its own delta type before the call itself, and tool calls stream start/delta/end events whose argument fragments you concatenate into a JSON string and parse only once the call has ended. That progressive shape is what lets a UI show "looking up the order status…" while the arguments are still assembling. ## Why index matters Every content event carries an `index`. A response may contain more than one content block, and the stream does not promise that block 0 finishes before block 1 starts. Keying your accumulator by index — a dict of index to buffer — rather than appending everything to one string is what keeps a multi-block response from being scrambled. It costs nothing on the single-block case and is correct on all of them. ## Do not stop at the last text delta The single most common production defect here is a consumer that renders `content-delta` events and closes the connection when the text stops looking interesting. Two things are lost. First, `finish_reason`: you cannot distinguish a complete answer from one truncated at `MAX_TOKENS`, so truncated output ships to users as if it were finished. Second, `usage`: the authoritative token counts arrive only at `message-end`, so your cost dashboard falls back to client-side estimates that will not reconcile with the invoice. Always drain the stream to `message-end` and record both. ## Operational concerns **Disconnects.** SSE runs over a long-lived HTTP response, and anything between you and the API — a load balancer, a proxy idle timeout, a mobile network — can drop it mid-turn. There is no resume: a broken stream means re-issuing the request from the start, and you are billed for both attempts. Budget for that, and prefer non-streaming calls for background jobs where nobody is watching characters appear. **Buffering proxies.** If your gateway buffers responses, streaming collapses into one slow blob and the whole latency benefit disappears. Streaming through a reverse proxy needs the buffering explicitly disabled, and it needs to be tested end-to-end from the real client, not from curl on the same host. **Partial state on error.** An `ERROR` finish reason, or a stream that ends without `message-end`, leaves you holding half an answer. Decide the policy up front: discard and retry, or render what you have with a visible truncation marker. Silently persisting a partial answer into conversation history is the worst option, because the next turn then reasons from a mutilated transcript. **Backpressure.** If your UI is slower than the stream, buffer server-side and flush on an interval instead of per-event; per-token DOM updates are a common cause of janky chat interfaces. ## What good looks like A solid consumer is a switch on `type`, an index-keyed buffer, an explicit terminal state at `message-end`, structured logging of `finish_reason` and `usage`, and a documented behaviour for a stream that never terminates. Being able to describe that shape — rather than reciting event names — is what an interviewer is listening for.
- Where do you get token usage for a streamed Cohere chat call?From the terminal `message-end` event, which carries both `finish_reason` and the `usage` object. Nothing earlier in the stream reports token counts, so a client that disconnects after the final text delta has no authoritative usage figure and has to fall back on a local estimate that will not match billing. Drain the stream and log usage against the generation id.
- How do you handle a stream that dies halfway through?Treat it as a failed turn, not a short answer. There is no resume token, so the only recovery is re-issuing the whole request — and you pay for both attempts. Guard it with a client timeout and a bounded retry count, keep the partial text out of persisted conversation history, and surface truncation to the user rather than pretending the reply completed.
- Why key your accumulator by the index field instead of one buffer?Because a response can carry several content blocks and the stream gives no guarantee that they complete in order. An index-keyed map of buffers reconstructs each block correctly whatever the interleaving, and it degenerates harmlessly to a single entry for ordinary single-block answers. One shared string works right up until the first multi-block response scrambles it.
saying these in an interview costs you the question
- Assuming Cohere's stream matches OpenAI's uniform chunk shape
- Replacing the buffer with each delta instead of appending
- Closing the connection at the last text delta
- Expecting token usage in every streamed event
- Appending all deltas to one buffer regardless of index