skip to content

What events does Cohere's v2/chat stream emit when stream is true?

level: middleimportance: should knowfreq 50%

answer

  1. Every payload announces its own type
  2. Blocks are bracketed, not bare
  3. Fragments append, never replace
  4. Block number rides along on each event
  5. Cost and stop reason arrive last

basics

~10 s

Cohere streams typed events rather than uniform chunks: message-start opens the turn, content-start/content-delta/content-end wrap each content block, and message-end closes it with finish_reason and usage. You accumulate the text from content-delta events.

solid answer

~40 s

Setting `stream: true` on `/v2/chat` (or calling `chat_stream()` in the SDK) returns Server-Sent Events whose payloads each carry a `type`. The sequence is `message-start`, then per content block a `content-start`, a run of `content-delta` events, and a `content-end`, finally `message-end`. Text arrives nested inside the delta — in the SDK, `event.delta.message.content.text` — and you concatenate it yourself; there is no cumulative field. Each content event carries an `index` identifying which block it belongs to, so parallel blocks interleave safely. `message-end` is where `finish_reason` and the `usage` token counts live, so a client that stops reading at the last text delta loses its cost accounting. Tool calls stream through their own event types, so a UI can render a plan before the arguments finish.

code

python · 16 lines
python
import cohere

co = cohere.ClientV2(api_key="COHERE_API_KEY")

buffers = {}
for event in co.chat_stream(
    model="command-a-03-2025",
    messages=[{"role": "user", "content": "Explain SSE in two lines."}],
):
    if event.type == "content-delta":
        idx = event.index
        buffers[idx] = buffers.get(idx, "") + event.delta.message.content.text
    elif event.type == "message-end":
        print(event.delta.finish_reason)

print("".join(buffers[i] for i in sorted(buffers)))

go deeper

for a junior

Know that streaming sends typed events, that the text arrives in content-delta events and must be appended, and that the stream ends with a message-end event.

for a middle

Walk the whole sequence from message-start to message-end, explain why the index field exists, and say exactly which event carries finish_reason and usage.

for a senior

Show the production concerns: proxy buffering that kills streaming, no resume after a disconnect, bounded retries with double billing, and a defined policy for partial output on an error finish.

for a principal

Decide where streaming belongs at all. Interactive surfaces earn it; batch pipelines usually do not, and the retry, cost-attribution and transcript-integrity rules for half-delivered turns should be platform policy, not per-team improvisation.

## A typed event stream, not uniform chunks Cohere's v2 chat streaming is deliberately different in style from the OpenAI-shaped streams many developers meet first. There, every chunk has the same shape and you dig into `choices[0].delta` to find out what changed. On Cohere's `/v2/chat`, each Server-Sent Event payload carries an explicit **`type`**, and your consumer switches on it. That makes the parser a state machine rather than a pile of null checks, which is the design point worth articulating in an interview. ## The event sequence For a plain text answer the order is: 1. **`message-start`** — the turn opens. Carries the generation id. Nothing renderable yet; use it to allocate a buffer and start your latency timer for time-to-first-token. 2. **`content-start`** — a content block begins, with an `index`. This is where you learn the block's type before any text arrives. 3. **`content-delta`** — repeated, each carrying a fragment of text. In the Python SDK the value is at `event.delta.message.content.text`. Fragments are **incremental**, not cumulative: you append. Appending cumulative payloads by mistake is the classic bug that produces exponentially duplicated text. 4. **`content-end`** — that block is finished. 5. **`message-end`** — the turn is over. This event carries `finish_reason` (`COMPLETE`, `MAX_TOKENS`, `STOP_SEQUENCE`, `TOOL_CALL`, `ERROR`) and the `usage` object with token counts. Tool-calling turns add their own events: a streaming tool plan arrives as its own delta type before the call itself, and tool calls stream start/delta/end events whose argument fragments you concatenate into a JSON string and parse only once the call has ended. That progressive shape is what lets a UI show "looking up the order status…" while the arguments are still assembling. ## Why index matters Every content event carries an `index`. A response may contain more than one content block, and the stream does not promise that block 0 finishes before block 1 starts. Keying your accumulator by index — a dict of index to buffer — rather than appending everything to one string is what keeps a multi-block response from being scrambled. It costs nothing on the single-block case and is correct on all of them. ## Do not stop at the last text delta The single most common production defect here is a consumer that renders `content-delta` events and closes the connection when the text stops looking interesting. Two things are lost. First, `finish_reason`: you cannot distinguish a complete answer from one truncated at `MAX_TOKENS`, so truncated output ships to users as if it were finished. Second, `usage`: the authoritative token counts arrive only at `message-end`, so your cost dashboard falls back to client-side estimates that will not reconcile with the invoice. Always drain the stream to `message-end` and record both. ## Operational concerns **Disconnects.** SSE runs over a long-lived HTTP response, and anything between you and the API — a load balancer, a proxy idle timeout, a mobile network — can drop it mid-turn. There is no resume: a broken stream means re-issuing the request from the start, and you are billed for both attempts. Budget for that, and prefer non-streaming calls for background jobs where nobody is watching characters appear. **Buffering proxies.** If your gateway buffers responses, streaming collapses into one slow blob and the whole latency benefit disappears. Streaming through a reverse proxy needs the buffering explicitly disabled, and it needs to be tested end-to-end from the real client, not from curl on the same host. **Partial state on error.** An `ERROR` finish reason, or a stream that ends without `message-end`, leaves you holding half an answer. Decide the policy up front: discard and retry, or render what you have with a visible truncation marker. Silently persisting a partial answer into conversation history is the worst option, because the next turn then reasons from a mutilated transcript. **Backpressure.** If your UI is slower than the stream, buffer server-side and flush on an interval instead of per-event; per-token DOM updates are a common cause of janky chat interfaces. ## What good looks like A solid consumer is a switch on `type`, an index-keyed buffer, an explicit terminal state at `message-end`, structured logging of `finish_reason` and `usage`, and a documented behaviour for a stream that never terminates. Being able to describe that shape — rather than reciting event names — is what an interviewer is listening for.

  • Where do you get token usage for a streamed Cohere chat call?
    From the terminal `message-end` event, which carries both `finish_reason` and the `usage` object. Nothing earlier in the stream reports token counts, so a client that disconnects after the final text delta has no authoritative usage figure and has to fall back on a local estimate that will not match billing. Drain the stream and log usage against the generation id.
  • How do you handle a stream that dies halfway through?
    Treat it as a failed turn, not a short answer. There is no resume token, so the only recovery is re-issuing the whole request — and you pay for both attempts. Guard it with a client timeout and a bounded retry count, keep the partial text out of persisted conversation history, and surface truncation to the user rather than pretending the reply completed.
  • Why key your accumulator by the index field instead of one buffer?
    Because a response can carry several content blocks and the stream gives no guarantee that they complete in order. An index-keyed map of buffers reconstructs each block correctly whatever the interleaving, and it degenerates harmlessly to a single entry for ordinary single-block answers. One shared string works right up until the first multi-block response scrambles it.

saying these in an interview costs you the question

  • Assuming Cohere's stream matches OpenAI's uniform chunk shape
  • Replacing the buffer with each delta instead of appending
  • Closing the connection at the last text delta
  • Expecting token usage in every streamed event
  • Appending all deltas to one buffer regardless of index

context