skip to content

How do you assemble a full reply from Mistral's streamed chat completion chunks?

level: middleimportance: must knowfreq 58%

answer

  1. delta, not message
  2. fragments arrive smaller than words
  3. the final chunk carries finish_reason
  4. a sentinel line closes the stream
  5. status 200 arrives before the failure can

basics

~20 s

Set stream true on Mistral's chat completions request and the reply arrives as chat.completion.chunk objects over server-sent events. Concatenate choices[0].delta.content across chunks in arrival order; the last chunk carries finish_reason, and the stream closes with data: [DONE].

solid answer

~50 s

With `"stream": true`, `POST /v1/chat/completions` returns `text/event-stream` instead of one JSON body. Each `data:` line is a `chat.completion.chunk` whose `choices[0].delta` holds the increment: the opening chunk typically carries only `delta.role`, the middle chunks carry `delta.content` fragments, and the terminal chunk carries `finish_reason` — `"stop"`, `"length"` or `"tool_calls"` — with token accounting reported at the end rather than per delta. The sentinel `data: [DONE]` closes the stream. In application code, guard against a delta with no `content` key rather than assuming every chunk has text, accumulate into a buffer, and only treat the answer as complete when you have seen a `finish_reason`. The operational trap: the HTTP status is 200 as soon as headers are sent, so failures and truncations surface mid-stream — wrap the loop, keep the partial text, and set timeouts per chunk rather than for the whole request.

code

python · 21 lines
python
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

buffer = []
finish_reason = None

for event in client.chat.stream(
    model="mistral-large-latest",
    messages=[{"role": "user", "content": "Explain sparse MoE briefly."}],
):
    choice = event.data.choices[0]
    if choice.delta.content:
        buffer.append(choice.delta.content)
    if choice.finish_reason:
        finish_reason = choice.finish_reason

text = "".join(buffer)
if finish_reason != "stop":
    print("incomplete:", finish_reason)

go deeper

for a junior

Know that stream true switches the response to incremental chunks, that the text lives in choices[0].delta.content rather than message.content, and that you concatenate the fragments in order.

for a middle

Explain the chunk sequence — role first, content fragments, then a chunk with finish_reason — and why deltas can be empty and can split words and multi-byte characters mid-sequence.

for a senior

Show that you handle mid-stream failure after a 200 response: keep the partial text, use idle timeouts between chunks rather than a request deadline, and retry deliberately because a blanket retry re-bills the whole prompt.

for a principal

Own the decision of which endpoints stream at all. Weigh perceived latency and cancellability against losing whole-response validation and moderation, and set the house rule that structured-output calls do not stream.

## Turning streaming on Add `"stream": true` to the chat completions body. The response content type changes to `text/event-stream` and the body is delivered incrementally instead of as one JSON document. In the Python SDK the entry point changes too: `client.chat.stream(...)` returns an iterator of events rather than `client.chat.complete(...)` returning a finished object, and each event exposes the chunk at `event.data`. ## The chunk shape Each streamed payload is an object with `object: "chat.completion.chunk"`, the same `id` throughout the stream, the model name, and a `choices` array. The difference from a non-streamed response is that a choice carries `delta` — the *increment* — rather than `message`, the finished thing. A typical progression: 1. First chunk: `delta` announces the role, e.g. `{"role": "assistant"}`, usually with no content. 2. Many chunks: `delta` carries a `content` fragment — often a token or a few characters, not a word or a sentence. 3. Final chunk: `finish_reason` is populated, `delta` is empty or nearly so, and token accounting arrives here rather than being spread across the stream. 4. `data: [DONE]` — a sentinel line that is not JSON. Parsing it as JSON is the single most common streaming bug. The fragments are not aligned to anything meaningful. They split words, split UTF-8 grapheme clusters, and split markdown syntax. Never apply per-chunk logic that assumes a fragment is a complete unit — accumulate first, interpret second. If you are rendering markdown live, render from the accumulated buffer, not from the delta. ## Accumulating correctly The loop is: for each chunk, read `choices[0].delta`, take `content` if present, append to a buffer, and record `finish_reason` when it appears. Two details separate working code from nearly-working code. First, **a delta may have no content at all**. The role chunk, the final chunk, and tool-call chunks all fit that description. Code shaped like `buffer += chunk.data.choices[0].delta.content` throws on the first such chunk in a language with null-hostile concatenation. Guard it. Second, **`finish_reason` is null until the end**. Its arrival — not the closing of the socket — is the signal that the answer is complete. If the connection drops without one, you have a partial answer and you know it, which is a materially better position than discovering the truncation from a user complaint. ## Why streaming changes error handling This is where interviews go, because it is where teams get burned. In a non-streamed call, an HTTP error status arrives before any body and your client's error handling catches it. In a streamed call, **the status line is 200 the moment headers are flushed**, before the model has produced anything. Anything that goes wrong afterwards — an upstream failure, a capacity problem, a disconnect — arrives inside a stream you have already declared successful. The implications: - Wrap the iteration itself, not just the call that opens the stream. - Decide what to do with the partial text on failure. Keeping it and marking the message incomplete is usually better UX than discarding it, and for a resumable format you can even continue from it. - A naive blanket retry on failure re-runs the whole generation and re-bills the input tokens. Retry deliberately, and consider whether the partial output can be continued instead. - Timeouts must be per-chunk (an idle-timeout on the iterator), not a single request-level deadline. A long answer legitimately takes a long time; what indicates a hang is silence between chunks. ## Why anyone bothers Streaming does not make generation faster. It changes *time to first token* from *time to last token* in what the user perceives, and on a several-hundred-token answer that is the difference between an interface that feels responsive and one that feels broken. It also lets you cancel: closing the connection stops generation, so a user who abandons a request stops accruing output tokens. What you give up is the convenience of one complete object. Anything you wanted to do to the whole answer — validating JSON, running a moderation pass, applying a redaction filter — cannot be done on a delta. Either buffer to completion before showing anything (surrendering the latency win), or stream to the user and accept that a post-hoc check can only retract text already on screen. That trade is the design decision behind "do we stream this endpoint?", and it is why streaming is usually wrong for a structured-output call whose result must be parsed before it means anything.

  • How do you detect a truncated streamed answer?
    Wait for a chunk with a populated `finish_reason` — its arrival, not the socket closing, marks completion. `"length"` means the cap was hit and the text is cut off. If the stream ends with no finish_reason at all, treat the buffer as incomplete rather than as a finished answer.
  • Why do request-level timeouts behave badly with streaming?
    A long answer legitimately occupies the connection for a long time, so a whole-request deadline kills healthy generations. The real failure signal is silence between chunks, so use an idle timeout on the iterator — reset it on every chunk received and abort only after a gap exceeds your threshold.
  • Would you stream a call that uses JSON output mode?
    Usually not. A partial JSON document is not parseable, so nothing downstream can act on the deltas and you buffer to completion anyway — which forfeits the perceived-latency benefit. Streaming pays when a human reads the text as it lands; for machine-consumed structured output the plain call is simpler and equally fast to last token.

saying these in an interview costs you the question

  • Reads choices[0].message.content on a streamed chunk
  • Assumes every chunk carries a content field
  • Tries to JSON-parse the stream's closing sentinel line
  • Treats HTTP 200 as proof the whole generation succeeded
  • Applies markdown or JSON parsing to individual deltas

context