skip to content

How do you consume a streamed OpenAI Chat Completions response chunk by chunk?

level: middleimportance: should knowfreq 62%

answer

  1. One flag flips the response shape
  2. delta, not message
  3. You concatenate the fragments
  4. Usage is opt-in when streaming
  5. Watch for the empty choices chunk

basics

~20 s

Set stream to true and iterate the Server-Sent Events response. Each event is a chat.completion.chunk whose choices[0].delta holds the newest fragment; concatenate those fragments yourself, and stop when the stream sends its final done marker.

solid answer

~40 s

Pass `stream: true` and the endpoint answers with Server-Sent Events instead of one JSON body. Each event carries an object of type `chat.completion.chunk`, shaped like a normal response except that `choices[0]` has a `delta` rather than a `message`. The first delta typically carries `role`, later ones carry small `content` fragments, and the last content-bearing chunk carries `finish_reason`. You accumulate the fragments to rebuild the full answer; the raw stream terminates with `data: [DONE]`, which the official SDKs consume for you so the iterator simply ends. By default a streamed response contains **no** `usage` object — pass `stream_options: {"include_usage": true}` and you get one extra final chunk with an empty `choices` array and the token counts. Guard your loop accordingly.

code

python · 22 lines
python
import os
from openai import OpenAI

client = OpenAI()
stream = client.chat.completions.create(
    model=os.environ["OPENAI_MODEL"],
    messages=[{"role": "user", "content": "Count to five."}],
    stream=True,
    stream_options={"include_usage": True},
)

parts = []
for chunk in stream:
    if not chunk.choices:
        print("\nusage:", chunk.usage)
        continue
    piece = chunk.choices[0].delta.content
    if piece:
        parts.append(piece)
        print(piece, end="")
print()
print("full:", "".join(parts))

go deeper

for a junior

Know that stream true changes the response into a series of small chunks, and that you join the delta content pieces together to get the full answer.

for a middle

Describe the chunk shape, where role and finish_reason appear, the done sentinel on the raw stream, and the stream_options flag that restores usage.

for a senior

Handle the operational edges: mid-stream failures after a 200, inter-chunk read timeouts, cancellation to stop paying for abandoned generations, and proxy buffering.

for a principal

Decide where streaming genuinely earns its complexity, and own the end-to-end contract — backend proxying, partial-response semantics in the product, and how truncation is reported and measured.

## Turning it on Streaming is a single request flag: `stream: true`. The HTTP response then uses Server-Sent Events — a long-lived `text/event-stream` connection where the server pushes `data:` lines as tokens are produced. The point is perceived latency: the user sees the first words in a few hundred milliseconds instead of waiting for the whole answer. ## The chunk shape Each event holds a JSON object whose `object` field is `"chat.completion.chunk"`. It has the familiar `id`, `created`, `model` and `choices` fields, but inside `choices[0]` you get `delta` instead of `message`. The typical sequence for one completion is: 1. A first chunk whose delta contains `role: "assistant"` and usually empty content. 2. Many chunks whose delta contains a short `content` string — often a word fragment, not a whole word. 3. A chunk with an empty delta and `finish_reason` set (`"stop"`, `"length"`, `"tool_calls"` or `"content_filter"`). Nothing gives you the assembled text: concatenating deltas in arrival order is your job. Because fragments split mid-word, never parse a partial accumulation as JSON or Markdown until the stream ends. ## Ending the stream On the wire the final line is literally `data: [DONE]`. If you are reading raw SSE — with a plain HTTP client, or proxying to a browser — you must recognise that sentinel, because it is not valid JSON and will break a naive parser. The official SDKs handle it: iterating the returned stream object simply stops. ## Getting usage while streaming A streamed response does not include token counts by default, which surprises people who bolt streaming onto working code and watch their cost telemetry go blank. Send `stream_options: {"include_usage": true}` and the server appends one final chunk with an **empty** `choices` array and a populated `usage` object. That empty array is exactly why loops written as `chunk.choices[0].delta.content` start throwing an index error the day someone enables usage — always check that `choices` is non-empty before indexing. ## Errors mid-stream The hard operational fact: the HTTP status is sent before generation finishes, so a stream that has already returned 200 can still fail. Failures show up as an error payload inside the stream or as an abruptly closed connection, not as a 500 you can catch around the call. Your consumer therefore needs a read timeout between chunks, a way to mark a partial answer as incomplete in the UI, and a retry policy that understands the request is safe to replay (it is stateless) but will regenerate from scratch and bill again. Track whether you ever received a `finish_reason`; its absence at end-of-stream means the answer is truncated regardless of what arrived. ## Buffering through your own backend Most apps do not call OpenAI from the browser — the API key would be exposed. You proxy: the server calls OpenAI with `stream: true`, then re-emits its own SSE or WebSocket to the client. Two things to get right are disabling response buffering in any intervening proxy, which otherwise defeats streaming entirely, and propagating cancellation — if the user navigates away, close the upstream stream so you stop paying for tokens nobody will read. ## When not to stream Streaming buys nothing when no human is watching the tokens appear: batch jobs, evaluation harnesses, and anything that must validate a complete structured payload before acting. In those cases the extra plumbing — partial state, sentinel handling, mid-stream failures — is pure cost.

  • Why does a streamed call report no token usage, and how do you get it back?
    Streamed responses omit the `usage` object by default because counts are only final once generation ends. Pass `stream_options: {"include_usage": true}` and the server sends one extra chunk at the end with an empty `choices` array and the full `usage` object. Consumers must tolerate that empty array, which is the usual crash when the flag is switched on.
  • A stream returned 200 and then died halfway. How is that different from a normal API failure?
    The status line was committed before the failure, so there is no non-2xx code to catch around the call — you see an error event in the stream or a dropped connection with no `finish_reason`. Treat any completed read without a finish reason as truncated, surface it as incomplete, and retry as a fresh request, accepting that the whole answer is regenerated and re-billed.
  • What changes about your loop when the model streams tool calls instead of prose?
    The deltas carry incremental tool-call fragments rather than plain content, and the arguments arrive as a partial JSON string spread across chunks with an index identifying which call each piece belongs to. You must buffer per index and only parse once the stream ends with `finish_reason` `"tool_calls"` — parsing a half-arrived argument string is the standard bug.

saying these in an interview costs you the question

  • Expects a message object in each chunk instead of delta
  • Assumes the SDK concatenates the full text for you
  • Thinks streamed responses always include a usage object
  • Indexes choices[0] without checking the array is non-empty
  • Believes a 200 status means the stream cannot fail

context