In DeepSeek's API, why must you strip reasoning_content before the next turn?
answer
- The endpoint remembers nothing for you
- Only one of the two fields goes back
- Sending it is rejected, not ignored
- Fresh thinking every turn, paid every turn
- 400 on the second call is the tell
basics
~20 sDeepSeek rejects a request whose messages array contains reasoning_content, returning a 400. Append only the assistant's content to the conversation. The model re-derives fresh reasoning each turn from the visible messages, so nothing is lost by dropping the previous trace.
solid answer
~40 sThe chat-completions endpoint is stateless, so on every turn you resend the whole conversation — and `deepseek-reasoner` will not accept its own chain of thought back. If `reasoning_content` appears anywhere in the `messages` array, DeepSeek returns a 400 rather than quietly ignoring the field. So the multi-turn loop is: read `message.reasoning_content` for logging or display, but append only `{"role": "assistant", "content": message.content}` to the history. The design intent is that each turn's thinking is derived afresh from the visible conversation; the previous trace is scratch work, not context. In practice this bites when you naively push `resp.choices[0].message` — the whole object — onto your history list, which works fine against `deepseek-chat` and fails on the second reasoner call.
go deeper
Remember the rule: append only the assistant's content to your message history, never the reasoning_content field, or the next call fails.
Explain the mechanics — the endpoint is stateless, DeepSeek returns 400 when reasoning_content appears in messages, and the model re-derives thinking from the visible turns each time.
Show how you would find and prevent it in a real system: dump the outbound payload, build history explicitly rather than pushing SDK objects, and account for thinking being re-billed on every turn.
Own the conversation-design tradeoff — fewer, richer turns versus chatty loops, what state must be promoted into visible text, and where a summarisation step belongs in an agent architecture.
## The stateless loop DeepSeek's `/chat/completions` endpoint keeps nothing between calls. Every request carries the full `messages` array, and the usual pattern is to append each assistant reply to a list and send the grown list next time. On `deepseek-chat` you can append the entire message object and it works. On `deepseek-reasoner` that same code fails, because the object you appended now contains `reasoning_content`. ## What actually happens DeepSeek documents that including `reasoning_content` in the input messages produces a 400 error. This is a deliberately loud failure rather than a silent strip: the API refuses the request so you cannot accidentally build a conversation history whose semantics the model was never trained for. The failure mode is characteristic — the first call succeeds, the second call in the same session 400s — which is exactly the signature to recognise in an interview answer. ## Why the design says no A reasoning model is trained to generate its thinking for the current problem, then answer. Its own past thinking is not part of the dialogue: the conversation the model is meant to condition on is the sequence of user turns and the answers it gave. Replaying old traces would bloat the prompt with scratch work, push out real context, and push the model toward re-committing to earlier reasoning instead of reconsidering. So the contract is simply: the visible turns carry the state, and the model re-derives whatever thinking the next turn requires. The cost consequence is worth being explicit about. Because each turn re-thinks, thinking tokens are paid per turn, not once. Turning a task into ten chatty turns pays for ten separate reasoning passes; keeping the same task in one well-specified turn pays for one. That is a real design lever in agent loops. ## The correct multi-turn shape After a call, split the response: - `message.content` goes into the history as the assistant turn. - `message.reasoning_content` goes to your logs, your debug panel, or the bin. If you need continuity of the model's earlier judgement, carry it in the visible answer — a reasoning model that concludes something important should say it in `content`, where the next turn can see it. Some teams add a summarisation step: distil a long trace into a couple of sentences and put those into the next user or system message. That is legitimate, because it is ordinary text you authored, not the `reasoning_content` field being echoed back. ## Practical implementation notes Build the history explicitly instead of pushing SDK objects into it. A one-line helper that maps a response to `{"role": "assistant", "content": msg.content}` removes the whole class of bug, and keeps your history serialisable for storage or replay. If you already persist raw responses (a good idea for auditing), strip on the way out — sanitise when constructing the request, not when storing. The same discipline applies when a conversation moves between models. History assembled from a reasoner call must be safe to send to `deepseek-chat` and vice versa; a history containing only roles and `content` always is. Watch out for framework abstractions too. Some client libraries and agent frameworks serialise every field of the assistant message back into the next request. If you see 400s from DeepSeek only on multi-turn reasoning conversations, inspect the actual outbound JSON rather than your application-level history object — the extra field is often added by the layer in between. ## Diagnosing it Symptoms that point straight here: single-turn calls always work; the failure appears on turn two; switching the model to `deepseek-chat` makes the error disappear; and the outbound payload, when dumped, shows a `reasoning_content` key inside `messages`. The fix is one line of message construction, not a retry policy — this is a request-shape error, so retrying it unchanged will fail identically every time.
- If the model re-thinks every turn, what does that do to the cost of a chatty multi-turn session?It multiplies thinking cost by the number of turns. Each call regenerates a full chain of thought and bills it as output, so ten short turns can cost far more than one well-specified turn covering the same task. In agent loops this argues for fewer, richer calls, and for routing only the genuinely hard steps to the reasoner.
- How can you preserve an important conclusion the model reached in its thinking?Carry it in visible text. Either the answer itself states the conclusion, or you summarise the trace yourself and inject that summary as an ordinary user or system message on the next turn. That is text you authored, so it is accepted normally — what is rejected is echoing the `reasoning_content` field back verbatim.
- You get 400s only on the second call of a conversation. What do you inspect first?The raw outbound JSON, not your in-memory history. Look for a `reasoning_content` key inside the `messages` array — usually added because the code appended the whole assistant message object, or because a framework layer serialised every field. It is a request-shape error, so retrying unchanged fails identically; fix the message construction.
saying these in an interview costs you the question
- Appends the entire assistant message object to the history
- Thinks the API silently ignores an unexpected reasoning_content field
- Believes the server stores the previous turn's chain of thought
- Assumes resending the trace would improve continuity
- Treats the resulting 400 as transient and adds a retry