What does an MCP CreateMessageResult return, and what does its stopReason mean?
answer
- one message back, plus provenance
- the model that ran is named explicitly
- why generation ended is optional
- three well-known reasons, but any string is legal
- one of them means you got a prefix
basics
~20 sA CreateMessageResult carries the generated message as a role and a content block, a required model string naming the model that produced it, and an optional stopReason — well-known values are "endTurn", "stopSequence" and "maxTokens", but any string is allowed.
solid answer
~50 sThe result mirrors a single assistant message plus provenance. `role` and `content` hold the generated message — text, image or audio, or a `tool_use` block when tool-enabled sampling is in play. `model` is a required string naming the model that actually generated it, which matters because `modelPreferences` is only advisory. `stopReason` is optional and tells the server why generation ended: `"endTurn"` means the model finished naturally, `"stopSequence"` means it hit one of the server's `stopSequences`, and `"maxTokens"` means it was cut off at the token budget. The schema types `stopReason` as an open string, so a client may report a provider-specific reason the server does not recognise — treat unknown values as "something other than a clean finish". A server that parses the completion should check for `"maxTokens"` before trusting the text, since a truncated answer often still parses.
code
json · 6 lines{
"role": "assistant",
"content": { "type": "text", "text": "Three fixes and one breaking change." },
"model": "claude-sonnet-4-5",
"stopReason": "endTurn"
}go deeper
Know the four pieces that come back — role, content, model, stopReason — and that "maxTokens" means the answer was cut off at the budget.
Explain why model is required given that preferences are advisory, and that stopReason is an optional open string, so the safe test is whether it equals "endTurn" rather than whether it is one of three known values.
Show how you handle a truncated or unexpected completion in production: check stopReason before parsing, log the returned model, validate structure, and keep a fallback path for when sampling is unavailable entirely.
Treat the result's fields as untrusted self-reports and design the system so nothing consequential depends on them. If content provenance genuinely matters to your product, that is an argument for owning the model call rather than borrowing one.
## What comes back A `sampling/createMessage` call resolves to a `CreateMessageResult`, which is deliberately narrow: one assistant message, plus enough metadata to reason about it. - `role` — the role of the generated message. - `content` — a single content block: text, image or audio; with tool-enabled sampling it may be a `tool_use` block instead. - `model` — a required string naming the model that generated the message. - `stopReason` — an optional string explaining why generation stopped. In revision 2026-07-28 this result does not travel back as a JSON-RPC response to a server request, because servers no longer issue requests. The client fulfils the sampling request it found in an `InputRequiredResult` and returns the `CreateMessageResult` as the matching entry in `inputResponses` when it re-issues the original call. ## Why model is required Because the server never chose it. `modelPreferences` hints and priorities are advisory, and a client may satisfy them exactly, approximately, or not at all. The `model` field is the only authoritative statement about what generated the text. Useful things to do with it: - Log it, so that a later quality complaint can be traced to a model substitution rather than a prompt bug. - Gate behaviour on it defensively — if your parsing logic was tuned for a capable model and a small one answered, expect more failures and validate harder. - Surface it in whatever the tool returns, so the eventual human reader knows what produced the content. What you should not do is treat it as a security signal. It is self-reported by the client, in the same spirit as `clientInfo` and `serverInfo`, which revision 2026-07-28 explicitly marks untrusted. ## Reading stopReason The three well-known values map to three very different situations: **`"endTurn"`** — the model completed naturally. This is the only value that suggests the content is whole. **`"stopSequence"`** — generation hit one of the `stopSequences` the request supplied. Expected if you set them deliberately; suspicious if you did not, because it means a sequence you configured for a different purpose fired early. **`"maxTokens"`** — the completion was cut off at the token budget. The text you have is a prefix. This is the case servers most often mishandle: a truncated JSON object may still deserialize if the model happened to close its braces early, a truncated summary reads like a summary, and a truncated list silently loses its tail. If your tool does anything structured with the completion, branch on this value before parsing. The field is typed as an open string, so a client is free to report a provider-specific reason such as a content-filter stop. Write the check as "is it `endTurn`?" rather than "is it one of the three I know", and treat everything else as an incomplete result. Also remember the field is optional. A missing `stopReason` is not a promise that generation finished cleanly — it just means the client did not report one. ## Statelessness and the result The result carries no conversation handle, no continuation token, and nothing that links it to a previous completion. MCP 2026-07-28 is stateless, so a multi-turn exchange is the server appending the returned message to its own copy of the conversation and sending the whole `messages` array again. Any state the server itself needs across calls belongs in an explicit handle it mints and passes through its own tool arguments — not inferred from the sampling exchange. ## Practical handling A reasonable server-side routine looks like: check `stopReason` first and treat anything other than `"endTurn"` as suspect; record `model`; validate the content against whatever shape the tool expects; and have a fallback path for a completion that is empty, truncated, or of the wrong shape entirely. Since sampling is deprecated in revision 2026-07-28, that fallback path is worth having anyway — it is the path a server takes when the client does not support sampling at all.
- Your tool parses JSON out of the completion and it sometimes succeeds on truncated output. What check prevents that?Branch on `stopReason` before parsing: treat anything other than `"endTurn"` — and especially `"maxTokens"` — as an incomplete result and fail or retry rather than deserialize. Truncated JSON can still parse when the model happened to close its structures early, so parse success is not evidence of completeness. Raising `maxTokens` alone does not fix it; it only moves the cutoff.
- Should a server trust the model field for authorization or audit purposes?Not for authorization. Like `clientInfo` and `serverInfo` in revision 2026-07-28, it is self-reported by the other side and explicitly untrusted. It is genuinely useful for diagnostics and for logging what generated a piece of content, but nothing security-relevant should hinge on it — a server cannot verify the claim from inside the protocol.
- What does a missing stopReason tell the server?Only that the client did not report one — the field is optional. It is not evidence that generation finished naturally. Treat the absent case the same way you would treat an unrecognised value: validate the content on its own terms rather than assuming completeness, and design the tool so an incomplete completion degrades visibly instead of silently producing a shortened answer.
saying these in an interview costs you the question
- Assumes the completion is complete without checking stopReason
- Treats a missing stopReason as proof of a clean finish
- Thinks stopReason is a closed enum of exactly three values
- Uses the model field as a trusted or security-relevant signal
- Expects the result to carry a conversation id for the next turn