skip to content

How do you consume a streaming Gemini generateContent response and assemble the chunks?

level: middleimportance: should knowfreq 66%

answer

  1. same request, different delivery
  2. append, do not replace
  3. verdict arrives last
  4. SSE needs an explicit query flag
  5. no resume after a drop

basics

~10 s

Call client.models.generate_content_stream() instead of generate_content(). It yields GenerateContentResponse chunks, each holding an incremental slice of the answer, so you concatenate them. The finish reason and final usage metadata arrive on the last chunk.

solid answer

~40 s

The streaming twin is `client.models.generate_content_stream(model=..., contents=...)` (or `client.aio.models.generate_content_stream` under asyncio); the request body is identical to the non-streaming call. It returns an iterator of `GenerateContentResponse` objects. Each chunk is a full response *shape* carrying only the newest slice of output, so chunks are **incremental, not cumulative** — you append `chunk.text` rather than replacing your buffer. The terminal `finish_reason` and the cumulative `usage_metadata` land on the final chunk, which means you cannot know the answer completed cleanly until the stream ends. At the REST level this is `:streamGenerateContent`, which returns a JSON array of responses by default and Server-Sent Events when you pass `?alt=sse`. Streaming buys time-to-first-token and nothing else: token accounting and price are unchanged, and there is no resume — a dropped connection means re-issuing the whole request.

code

python · 18 lines
python
from google import genai

client = genai.Client()

parts = []
last = None
for chunk in client.models.generate_content_stream(
    model="gemini-2.5-flash",
    contents="Summarise the history of the QWERTY layout.",
):
    if chunk.text:
        parts.append(chunk.text)   # incremental: append, never assign
        print(chunk.text, end="", flush=True)
    last = chunk

answer = "".join(parts)
print("\nfinish:", last.candidates[0].finish_reason)
print("tokens:", last.usage_metadata.total_token_count)

go deeper

for a junior

Know that there is a streaming variant of the call, that it yields chunks you append together, and that the complete answer is the concatenation of all of them.

for a middle

Explain that chunks are incremental rather than cumulative, that the finish reason and cumulative usage arrive on the final chunk, and that the REST endpoint needs an explicit flag to emit Server-Sent Events.

for a senior

Show the operational consequences: text rendered before the verdict is known, no resume after a dropped connection, buffering intermediaries erasing the benefit, and structured output being unsafe to consume incrementally.

for a principal

Set the policy for when streaming is worth its failure modes — user-facing long-form yes, machine-to-machine and validated structured output no — and make sure retry, cost attribution and timeout budgets account for restarted streams.

## Two delivery modes, one request `generateContent` and `streamGenerateContent` take the same body: same `contents`, same generation config, same tools. The only difference is how the answer comes back — one complete response versus a sequence of partial ones. That symmetry is worth stating in an interview, because it makes clear that streaming is a *delivery* decision, not a modelling one. In the Python SDK: - `client.models.generate_content(...)` returns one `GenerateContentResponse`. - `client.models.generate_content_stream(...)` returns an iterator of them. - `client.aio.models.generate_content_stream(...)` returns an async iterator for asyncio code. ## Chunks are incremental Each chunk looks like a miniature response: it has `candidates`, each with `content.parts`. The parts carry only what was produced since the previous chunk. So the assembly rule is concatenation: ``` buffer = [] for chunk in stream: if chunk.text: buffer.append(chunk.text) ``` If you instead assign `buffer = chunk.text` on each iteration — a habit from APIs that resend the whole answer — you end up displaying only the last few words. Conversely, if a provider sent cumulative chunks and you concatenated, you would get quadratic duplication. Knowing which convention an API uses is the whole point of the question. Chunk boundaries are not semantically meaningful: they do not align with words, sentences, or JSON tokens. Any parsing you do must be over the accumulated buffer, never per chunk. ## Where the metadata lives The `finish_reason` on intermediate chunks is not the verdict; the meaningful value appears on the last chunk. Likewise `usage_metadata` is only complete at the end, and it is cumulative for the whole response — you do not sum it across chunks. The practical consequence is uncomfortable and worth naming: **you render text before you know how the generation ended**. A stream can deliver several sentences and then terminate with a truncation or a filtered finish. A UI that streams therefore needs an after-the-fact state — marking the message truncated, or withdrawing it — and any pipeline that streams into a parser or a downstream write must hold the commit until the final chunk validates. ## The REST shape Over HTTP the endpoint is `POST v1beta/models/{model}:streamGenerateContent`. By default it responds with a JSON **array** of response objects, streamed as it is produced — usable, but awkward, because you are incrementally parsing an array. Adding `?alt=sse` switches it to Server-Sent Events, one `data:` line per response object, which is what almost every real client wants and what browsers and proxies handle best. If you are writing your own HTTP client rather than using the SDK, forgetting `alt=sse` is the classic first bug. ## What streaming does and does not buy you It improves **perceived** latency: the user sees the first token in a fraction of the total generation time. It does not make total generation faster, and it does not change billing — you are charged for exactly the same prompt and output tokens either way. It also reduces the risk of a request-level timeout on very long generations, since bytes keep flowing. What it costs you: complexity and fragility. There is no resume token. If the connection drops halfway you cannot ask for "the rest"; you re-issue the whole request and pay for the prompt again. Long-lived streaming responses interact badly with proxies and load balancers that buffer, so an intermediary can erase the very benefit you added it for. And structured output is harder to work with, because a half-arrived JSON object is not parseable — either buffer to completion (giving up the benefit) or use an incremental parser that tolerates prefixes. ## Choosing Stream when a human is watching text appear and the answer is long enough for latency to be noticeable. Do not stream for short answers, for machine-to-machine calls whose consumer waits for the whole payload anyway, or for structured output you must validate before use — there the added failure modes buy nothing. Background and bulk work should not stream at all.

  • Does streaming change what you pay for a Gemini call?
    No. Prompt and output tokens are counted identically; the final chunk's `usage_metadata` reports the same totals a non-streaming call would. Streaming only changes when bytes reach you. If a stream drops midway and you re-issue the request, you pay for the prompt twice — so retries, not streaming itself, are where streaming costs money.
  • What breaks if you stream a response that must be valid JSON?
    Every chunk before the last is an incomplete document, so you cannot validate or act on partial output. You either buffer the whole thing — which forfeits the latency benefit — or run an incremental parser that tolerates prefixes. And because the finish reason arrives last, a truncated generation only reveals itself after you have consumed everything.
  • Why might streaming appear not to work behind a load balancer?
    Intermediaries that buffer responses hold the chunks until the body completes, so the client receives everything at once and the perceived-latency benefit disappears. Response buffering must be disabled on the path for streaming to survive end to end; the API side is behaving correctly in this case.

saying these in an interview costs you the question

  • Assuming each chunk contains the full text so far
  • Reading the finish reason from the first chunk
  • Believing streaming lowers token cost
  • Expecting to resume a dropped stream from where it stopped
  • Parsing JSON out of individual chunks

context