skip to content

In DeepSeek's API, what is reasoning_content in a deepseek-reasoner response?

level: juniorimportance: must knowfreq 70%

answer

  1. Two fields, not one
  2. Answer stays where it always was
  3. The thinking has its own key
  4. reasoning_content vs content on the message
  5. Billed output; never resent

basics

~20 s

deepseek-reasoner puts its chain of thought in message.reasoning_content and the user-facing answer in message.content. Both come back in one ordinary chat-completion response. Show users content; treat reasoning_content as diagnostic text you log, collapse, or discard.

solid answer

~40 s

A `deepseek-reasoner` reply carries two separate fields on the assistant message: `reasoning_content`, the model's chain of thought, and `content`, the final answer. The envelope is otherwise the normal OpenAI-shaped chat completion, so `resp.choices[0].message.content` still returns the answer and existing client code keeps working — it just ignores the thinking field it does not know about. `deepseek-chat` does not populate `reasoning_content`; only the reasoning model does. Practically, treat the thinking as diagnostic output: log it, render it behind a disclosure toggle, or drop it. Two consequences follow from it being generated text — it is billed as output tokens, and it must not be sent back in the next request's messages.

go deeper

for a junior

Know that a deepseek-reasoner reply has two fields: reasoning_content for the thinking and content for the answer, and that you show the user content.

for a middle

Be able to explain why the split exists — a stable contract instead of parsing markers out of the answer — and that the thinking is generated, billed output rather than free metadata.

for a senior

Show the operational consequences you would design for: token cost of the hidden trace, logging and retention rules that follow the prompt's, and clients that must survive a non-standard response field.

for a principal

Own the policy question: whether traces are stored at all, who can read them, and how switching a workload between a reasoning and a non-reasoning model changes cost and product surface without changing client code.

## The shape of a deepseek-reasoner reply DeepSeek's API is OpenAI-compatible: you point the OpenAI SDK at `https://api.deepseek.com` and call `chat.completions.create`. When the model is `deepseek-reasoner`, the assistant message in the response gains one extra field beyond the familiar ones: - `choices[0].message.content` — the final answer, exactly where every OpenAI-shaped client already looks for it. - `choices[0].message.reasoning_content` — the model's chain of thought: the working it produced before committing to that answer. Everything else is unchanged — `id`, `model`, `choices[].finish_reason`, `usage`. That is the whole point of the design: a client written for `deepseek-chat` does not break when you switch the model string. It simply never reads the new field, so the user sees only the answer. ## Why a separate field instead of inline text Before reasoning models exposed a dedicated field, chain of thought arrived mixed into the answer and applications had to fish it out with markers or regexes — brittle, and impossible to do reliably once the model's formatting drifted. Splitting the two makes the contract explicit: `content` is what you show, `reasoning_content` is what you may inspect. Any code that greps for tags inside `content` on a reasoning model is solving a problem the API already solved. ## What the field is not It is not free. `reasoning_content` is generated token by token like any other output, so it is counted and billed as output. A short answer preceded by a long think can cost several times what the visible text suggests. It is not server-side state. DeepSeek's chat completions endpoint is stateless — you resend the conversation each turn. The thinking is not kept anywhere for you, and it is not meant to be replayed: the model produces fresh reasoning for each new turn from the visible conversation alone. It is not part of the conversation you send back. Putting `reasoning_content` into a subsequent request's `messages` array is rejected rather than ignored, so the multi-turn loop must strip it. This is the single most common integration bug when moving from `deepseek-chat` to `deepseek-reasoner`. ## Handling it in code With the OpenAI Python SDK the field is reachable straight off the message object: read `message.reasoning_content` for the thinking, `message.content` for the answer, and append only the answer (with `role: "assistant"`) to your running message list. Typed clients in strongly-typed languages may need a permissive model or a raw-JSON path, because `reasoning_content` is a DeepSeek extension rather than part of the base OpenAI schema — the generated types for the official OpenAI clients do not declare it. In streaming mode the same split appears on the delta objects: chunks carry `delta.reasoning_content` while the model is thinking and `delta.content` once it starts answering, so a UI can render the two into different panes. ## Product decisions the field forces Because you now hold the thinking, you must decide what to do with it: - **Discard.** Simplest, and correct for most backends. Read `content`, drop the rest. - **Log.** Useful for debugging odd answers and for offline review. Remember that the trace can contain the user's input verbatim, so it inherits the same privacy and retention rules as the prompt. - **Display.** A collapsed "thinking" panel sets expectations during the long first-token wait, which on a reasoning model is dominated by the thinking phase. One caution on interpretation: the trace is text the model generated, not an instrumented log of its computation, so treat it as an explanation to sanity-check rather than as proof of how the answer was reached. ## Signals you have wired it up wrong If `reasoning_content` is always absent, you are probably still on `deepseek-chat`. If your app prints thinking to the end user, you concatenated the two fields. If the second turn suddenly fails with a 400 after the first succeeded, you echoed the field back. And if your token bill is far higher than the visible answers imply, you have found the cost of thinking tokens.

  • Does deepseek-chat ever populate reasoning_content?
    No. The field is produced by the reasoning model; a `deepseek-chat` response returns the ordinary assistant message with `content` only. Code that branches on the presence of `reasoning_content` therefore works as a runtime check of which model actually served the request, which is handy when the model name comes from configuration.
  • What should a client do if it is strongly typed and reasoning_content is not in the OpenAI schema?
    Read it off the raw JSON or use a client model that keeps unknown fields. `reasoning_content` is a DeepSeek extension to the OpenAI-compatible response, so generated OpenAI types do not declare it; a strict deserializer will either drop it silently or throw. Keeping the raw payload also future-proofs you against other vendor extensions.
  • Is it safe to show reasoning_content to end users?
    Technically yes, but decide deliberately. The trace is generated prose that may restate the user's input, explore wrong paths, or expose instructions from your system prompt. Most products either hide it or put it behind a collapsed panel, and treat it under the same logging and retention policy as the prompt itself.

saying these in an interview costs you the question

  • Assumes the chain of thought must be parsed out of content
  • Thinks reasoning_content costs nothing because users never see it
  • Believes deepseek-chat also returns a reasoning_content field
  • Concatenates reasoning_content and content into one displayed message
  • Expects the server to remember the reasoning for the next turn

context