skip to content

When streaming deepseek-reasoner, how do reasoning_content and content deltas arrive?

level: seniorimportance: should knowfreq 38%

answer

  1. Different keys on the same delta
  2. Thinking streams before the answer does
  3. Two buffers, one transition event
  4. First answer delta is the cue
  5. Totals arrive last, if you ask

basics

~20 s

Chunks first carry delta.reasoning_content while the model thinks, then switch to delta.content for the answer. Accumulate the two into separate buffers and treat the first content delta as the phase transition; usage totals only arrive at the end of the stream.

solid answer

~40 s

With `stream=True`, each chunk's `choices[0].delta` carries the thinking and the answer in different keys: `delta.reasoning_content` is populated during the reasoning phase, then chunks begin populating `delta.content` for the final answer. So a streaming consumer keeps two buffers rather than one, and treats the first non-empty `content` delta as the phase transition — that is your cue to close a "thinking" panel and start rendering the answer. The operational consequence is latency shape: time to first token is fast, but time to the first *answer* token can be many seconds, so a UI that only renders `content` looks frozen for the whole thinking phase even though bytes are flowing. Token totals are not in the chunks; you request them at the end via `stream_options` with `include_usage`, or count client-side.

go deeper

for a junior

Know that streamed chunks carry the thinking under delta.reasoning_content and the answer under delta.content, so you accumulate them separately.

for a middle

Explain the ordering and the transition — reasoning deltas first, then content deltas, with the first content delta marking the switch — and why marker parsing is the wrong detector.

for a senior

Bring the operational judgment: time to first answer token, timeouts inherited from the chat path, cancellation of long thinking phases, and checking the final finish_reason before trusting the buffer.

for a principal

Own the experience and cost tradeoff — whether traces are shown at all, how perceived latency is managed across a fleet of reasoning and non-reasoning calls, and what abandoned in-flight thinking costs you.

## The chunk shape Streaming a `deepseek-reasoner` completion gives you the ordinary OpenAI-compatible server-sent-event stream — a sequence of chunk objects, each with `choices[0].delta` — with one extension. The delta may carry `reasoning_content` instead of, or as well as, `content`: - Early chunks populate `delta.reasoning_content` with fragments of the chain of thought while `content` is empty or null. - Later chunks populate `delta.content` with the answer. - The stream terminates the usual way, with a final chunk carrying a `finish_reason`. So the naive accumulator — append `delta.content` to one string — still produces the correct final answer, which is why porting a streaming client from `deepseek-chat` does not crash. It just shows nothing at all for however long the model thinks. ## Two buffers and a transition The minimal correct consumer keeps two accumulators and watches for the switch: - Append any `delta.reasoning_content` to a thinking buffer. - Append any `delta.content` to an answer buffer. - The first non-empty `content` delta marks the phase change: thinking is done, the answer has begun. That transition is the useful event. Interfaces use it to collapse a live "thinking" pane and hand over to the answer pane; pipelines use it to start incremental parsing of the answer without accidentally feeding reasoning text into the parser. Do not detect the transition by looking for markers in the text — the field split is the contract, and the prose is not guaranteed to signal anything. ## Latency shape is the real story Streaming is usually adopted to improve perceived latency, and on a reasoning model the naive port defeats it. Time to first byte is unchanged and fast, but time to the first token the user sees is the entire thinking phase, which on hard problems can run for many seconds. Users interpret a frozen pane as a hang and retry, doubling your cost for the same answer. The options are all product decisions: - **Render the thinking** live, usually in a de-emphasised or collapsed panel. Honest and informative; requires you to be comfortable showing the trace. - **Render a progress indicator driven by the stream**, so the UI proves work is happening without displaying the trace itself. The arrival of reasoning deltas is a genuine liveness signal — better than a spinner on a timer. - **Accept the wait** and set expectations in copy, which is reasonable for a background or batch flow but poor for chat. The wrong option is a client-side timeout tuned for a non-reasoning model. Reasoning responses legitimately take much longer end-to-end, and a timeout inherited from the chat path will abandon requests that were about to succeed — after you have already paid for the thinking. ## Usage and errors mid-stream Token counts do not arrive in the content chunks. To get the totals on a streamed call, request them with `stream_options` set to include usage, in which case a final chunk carries the `usage` object; otherwise count client-side or fall back to non-streamed calls for the paths where you need exact accounting. Either way, do not infer cost from what you rendered — the thinking you may have discarded is still billed. Also handle the truncation case explicitly on the streamed path. If generation hits the output cap during thinking, the stream can end with a `finish_reason` of `length` and an answer buffer that is empty or clipped. A streaming consumer that renders whatever it accumulated and never inspects the final chunk will silently present half an answer as a whole one. ## Cancellation Because the thinking phase can be long and is billed, cancellation policy matters more than on a chat model. Closing the connection when a user navigates away is worth wiring up, and the transition event gives you a natural point to decide whether a request that has not started answering yet is still wanted. ## Summary of the contract Separate keys in the delta, thinking first and answer second, no marker parsing, transition detected on the first content delta, usage only at the end if you ask for it, and a latency profile that requires a deliberate UI decision rather than a straight port of the chat client.

  • Why does a streaming UI ported straight from a chat model look frozen on the reasoner?
    Because it renders only content deltas, and none arrive during the thinking phase. Bytes are flowing the whole time, but they are reasoning_content, which the ported client ignores. Users read the still pane as a hang and retry, so you pay for the thinking twice. Render the trace, or drive a liveness indicator from the reasoning deltas.
  • How do you get token usage for a streamed deepseek-reasoner call?
    Ask for it: set stream_options to include usage, and a final chunk carries the usage object after the content chunks. Otherwise the stream gives you no totals and you must count client-side. Never infer cost from the text you rendered — a discarded chain of thought is billed exactly the same as one you displayed.
  • What should a streaming consumer do with the final chunk's finish_reason?
    Inspect it before treating the accumulated text as complete. A finish_reason of length means the output budget ran out, so the answer buffer may be clipped or empty even though the stream ended cleanly. Rendering the buffer without that check presents a truncated answer as a finished one, and downstream parsers fail confusingly.

saying these in an interview costs you the question

  • Accumulates only content and shows nothing while thinking
  • Detects the phase change by searching for text markers
  • Applies the chat model's client timeout to reasoner streams
  • Assumes usage totals appear in the content chunks
  • Treats a cleanly ended stream as a complete answer

context