skip to content

Your deepseek-reasoner answers cut off mid-sentence — how do reasoning tokens explain it?

level: seniorimportance: should knowfreq 52%

answer

  1. The hidden phase is real output
  2. One budget, two phases
  3. Read the stop reason, not the prose
  4. p99 of thinking sets the cap
  5. Empty content plus full trace is the tell

basics

~20 s

The chain of thought is generated output. It consumes the same token budget as the answer and is counted in completion_tokens, so a long think can exhaust max_tokens before the answer finishes. Check finish_reason for length and size the budget for thinking plus answer.

solid answer

~50 s

On `deepseek-reasoner` the thinking is not free metadata — it is generated output, billed at the output rate and counted in `usage.completion_tokens`, with the split reported under `completion_tokens_details.reasoning_tokens`. It also draws on the same output budget the answer does, so if you sized `max_tokens` for a two-paragraph answer and the model spends most of that budget reasoning, generation stops before the answer is complete. The diagnostic is `choices[0].finish_reason`: `length` means you hit the cap, not that the model chose to stop. The fixes are to raise `max_tokens` so it covers thinking plus answer, to alert on `finish_reason == "length"` rather than shipping truncated text downstream, and to log `reasoning_tokens` so you can see how large the hidden phase actually runs on your prompts. Exact ceilings differ by model release, so confirm them against the current docs.

go deeper

for a junior

Know that the chain of thought counts as generated output, so it uses the same token budget and costs money even though the user never sees it.

for a middle

Explain the accounting: reasoning tokens are inside completion_tokens, they share max_tokens with the answer, and finish_reason length is how truncation announces itself.

for a senior

Show the production discipline — measure the reasoning-token distribution, size the cap from its tail, alert on the length rate, and handle truncation by re-budgeting or falling back rather than blind retries.

for a principal

Own the economics: hidden output changes cost per request and latency profile, so decide where reasoning is worth paying for, how budgets are governed across services, and what the fallback path is when it is not.

## The mechanism A reasoning model produces two spans of output for one request: the chain of thought, returned in `reasoning_content`, and the answer, returned in `content`. Both are generated tokens. That single fact explains the whole failure: 1. They are billed as output tokens, at the output rate, whether or not you ever look at them. 2. They appear in `usage.completion_tokens` — the total is thinking plus answer, not answer alone. 3. They draw on the same generation budget as the answer, so `max_tokens` sized as if it were an answer-length knob will cut the answer short when the thinking runs long. When the cap is reached, generation simply stops. The reply comes back with `finish_reason: "length"` and a `content` that is truncated — sometimes mid-sentence, sometimes empty if the model was still thinking when the budget ran out. An empty `content` with a populated `reasoning_content` and `finish_reason: "length"` is the unmistakable signature. ## Why teams walk into it The migration path is what does it. A service calls `deepseek-chat` with `max_tokens` tuned to the answers it wants, someone switches the model string to `deepseek-reasoner` to improve quality, and the parameter that used to be a comfortable ceiling is now shared with a phase that can be much longer than the answer. Nothing errors. Latency rises, costs rise, and a fraction of responses — the hard ones, exactly the ones you switched models for — come back clipped. The truncation correlates with problem difficulty, which is why it survives smoke tests and shows up in production. ## Diagnosing it properly Do not diagnose truncation by eyeballing the text. Read the structured signals: - **`finish_reason`.** `stop` means the model finished on its own; `length` means the cap ended it. Treat `length` as an error condition in any pipeline that parses the answer, because a truncated JSON blob or a half-finished list will fail downstream in far more confusing ways. - **`usage.completion_tokens`.** Compare it against your `max_tokens`. Sitting exactly at the cap is confirmation. - **`reasoning_tokens`** in `completion_tokens_details`. This is the number to put on a dashboard. Its distribution across your real prompts tells you what budget headroom you actually need, and its p99 is what your cap has to survive. ## Fixing it **Size the budget for both phases.** Set `max_tokens` from measured thinking length plus the answer length you want, with headroom for the tail of the distribution. Guessing produces either clipped answers or a cap so large it never constrains anything. **Handle `length` explicitly.** Retrying the identical request is usually futile — the model will think its way to the same place. Better responses are to retry once with a larger budget, to fall back to the non-reasoning model for that request, or to surface a clear failure rather than pass truncated content on. **Do not try to cap the thinking by shrinking `max_tokens`.** That does not make the model think less; it makes it stop before answering. Reducing thinking is a prompt and model-selection question — a simpler task, or a non-reasoning model for the easy path. **Watch cost, not just failures.** Because thinking is billed output that the user never sees, a request whose visible answer is fifty tokens can bill many times that. Cost-per-request dashboards built on visible output length will be badly wrong on a reasoning model; build them on `completion_tokens`. ## Version sensitivity The numeric ceilings here — default and maximum output length, context window, and how DeepSeek scopes `max_tokens` relative to the thinking phase — have moved across DeepSeek model releases, and earlier reasoner documentation described the scoping differently from later releases. Take the durable part as the answer: thinking is generated output, it is counted in `completion_tokens` and billed, and truncation shows up as `finish_reason: "length"`. Then pin the exact ceilings against the docs for the model version you call rather than reciting numbers. ## What good looks like in production A mature integration logs `reasoning_tokens` and `finish_reason` on every call, alerts when the `length` rate crosses a threshold, sets `max_tokens` from a measured distribution rather than a guess, and routes requests between the reasoning and non-reasoning models so the expensive hidden phase is only paid where it earns its keep.

  • A response comes back with empty content, a long reasoning_content and finish_reason length. What happened?
    The output budget was exhausted during the thinking phase, so generation stopped before any answer was emitted. Retrying identically will usually reproduce it. Raise max_tokens enough to cover the measured thinking distribution plus the answer, or route that request to the non-reasoning model — shrinking the cap makes it strictly worse.
  • Should a finish_reason of length be retried automatically?
    Not with the same request. The model will think its way to the same place and hit the same wall, so a plain retry burns tokens for nothing. Retry once with a larger budget, fall back to another model, or fail the request explicitly. Never pass truncated text to a parser — a half-written JSON object fails far more confusingly downstream.
  • Why can a cost dashboard built on visible answer length be badly wrong here?
    Because the user-visible answer is only part of what was generated. Thinking tokens are billed at the output rate and can dwarf the answer, so a fifty-token reply may cost many times what its length suggests. Build cost tracking on usage.completion_tokens, and chart reasoning_tokens separately so the hidden phase is observable.

saying these in an interview costs you the question

  • Believes thinking tokens are free because they are hidden
  • Sizes max_tokens as if it bounded only the answer
  • Tries to limit thinking by lowering max_tokens
  • Ignores finish_reason and parses truncated output anyway
  • Retries a length-truncated request unchanged

context