skip to content

Which fields does an MCP sampling/createMessage request carry, and which are required?

level: middleimportance: must knowfreq 47%

answer

  1. two fields are non-negotiable
  2. the conversation plus a token budget
  3. role-tagged blocks, no system role inside
  4. systemPrompt and modelPreferences are advisory
  5. includeContext values were cut back in 2026-07-28

basics

~10 s

Only messages and maxTokens are required. messages is an array of role-tagged user and assistant content blocks; maxTokens caps the completion. systemPrompt, modelPreferences, stopSequences, temperature and, for tool-enabled sampling, tools and toolChoice are optional.

solid answer

~40 s

A `sampling/createMessage` request is built from two mandatory fields and a set of optional steering fields. `messages` is the conversation the server wants completed: an array of entries each tagged `user` or `assistant` and carrying a text, image or audio content block. `maxTokens` is the required upper bound on the completion. Optional fields are `systemPrompt` (advisory — the client may adjust it), `modelPreferences` (hints and priority weights for model selection), plus pass-through generation controls like `temperature` and `stopSequences`. For tool-enabled sampling the server may add `tools` and `toolChoice`. `includeContext` still exists but its `"thisServer"` and `"allServers"` values are deprecated in revision 2026-07-28 — omit the field or send `"none"`. Note what is absent: the server supplies the entire conversation itself and gets no view of the user's real chat history.

code

json · 18 lines
json
{
  "method": "sampling/createMessage",
  "params": {
    "messages": [
      {
        "role": "user",
        "content": { "type": "text", "text": "Summarise this changelog in two lines." }
      }
    ],
    "systemPrompt": "You are a concise release-note writer.",
    "modelPreferences": {
      "hints": [{ "name": "sonnet" }],
      "intelligencePriority": 0.8,
      "speedPriority": 0.3
    },
    "maxTokens": 512
  }
}

go deeper

for a junior

Remember the two required fields — messages and maxTokens — and that systemPrompt is a separate top-level field rather than a role inside the array.

for a middle

Be able to walk the full field list and say which are advisory: systemPrompt and modelPreferences steer but do not bind, and maxTokens may be lowered by the client. Mention that includeContext's cross-server values are deprecated in 2026-07-28.

for a senior

Show judgment about what you put in messages: every block costs the user tokens and is visible to their provider. Talk about budgeting maxTokens deliberately and about handling a truncated completion instead of assuming a full answer came back.

for a principal

Frame the field set as a trust boundary: the server proposes, the host disposes. Anything you need enforced — model class, instruction integrity, context isolation — cannot be guaranteed by these fields, which is a strong argument for owning the provider integration directly.

## The two required fields `messages` and `maxTokens` are the only fields a server must send. `messages` is an array of `SamplingMessage` entries. Each has a `role` of `user` or `assistant` and a `content` block — text, image, or audio. There is no `system` role in the array; a system instruction goes in the separate `systemPrompt` field. This array is the complete conversation the model will see. Whatever the server did not put in it does not exist as far as the completion is concerned. `maxTokens` is the server's declared ceiling for the completion. It is a request, not a guarantee: a client may cap it lower to control cost, which is one reason a server should always inspect the result's `stopReason` rather than assuming it received a complete answer. ## The optional steering fields `systemPrompt` carries the instruction that frames the task ("You are a concise release-note writer"). It is advisory — the host is entitled to adjust or surface it before running the call, which is exactly why a server should not hide security-relevant instructions there and expect them to survive verbatim. `modelPreferences` expresses which kind of model the server wants: an array of `hints`, each an object with a `name` substring, plus `costPriority`, `speedPriority` and `intelligencePriority` weights. All of it is advisory; the client makes the final selection. `temperature` and `stopSequences` are conventional generation controls passed through to whatever model the client picks. Their behaviour is provider territory; MCP just carries them. `tools` and `toolChoice` enable tool-capable sampling, where the model may respond with a `tool_use` block the server then satisfies. These are gated by the client's `sampling.tools` sub-capability. ## The includeContext field `includeContext` was originally the way a server asked the host to attach context beyond the supplied messages, with values `"none"`, `"thisServer"` and `"allServers"`. In revision 2026-07-28 the `"thisServer"` and `"allServers"` values are deprecated; the guidance is to omit the field or send `"none"`. That aligns with the protocol's broader privacy stance — a server should not be able to vacuum up context that the user did not deliberately route to it, and it fits the removal of anything that implied cross-server visibility. ## What the request does not contain There is no session identifier, no conversation handle, and no reference to earlier sampling calls. Revision 2026-07-28 is explicit that MCP is a stateless protocol: every request is self-contained. A multi-turn sampling exchange is therefore expressed by the server re-sending the whole growing `messages` array each time, not by referring back to a prior completion. Any cross-call state the server itself needs must live in an explicit handle it mints and passes through its own tool arguments. Also absent is anything identifying the model. The server can hint but not choose, which is a deliberate control boundary: the user's host decides what runs on the user's account. ## Where the request lives In revision 2026-07-28 this request object is not sent on its own. It appears as one value in the `inputRequests` map of an `InputRequiredResult` that the server returns from `tools/call`, `prompts/get` or `resources/read`. The field list above is unchanged by that packaging, but it does mean a server builds the request as data inside a result rather than dispatching it. ## Practical guidance Keep `messages` minimal and purposeful — every block is content the user's model provider will see and the user will pay for. Set `maxTokens` to something you have actually reasoned about rather than a large round number. Treat `systemPrompt` as advice rather than enforcement. And remember the whole feature is deprecated in 2026-07-28, so any new field you are tempted to lean on should be weighed against a direct provider integration instead.

  • Why is there no system role inside the messages array?
    MCP models the system instruction as a separate top-level `systemPrompt` field rather than a message role. Entries in `messages` are only `user` or `assistant`. That separation lets the host treat the system instruction distinctly — surfacing or adjusting it — instead of having it buried among conversation turns, and it maps cleanly onto providers that take a system instruction as its own parameter.
  • What changed for includeContext in revision 2026-07-28?
    Its `"thisServer"` and `"allServers"` values are deprecated; servers should omit the field or send `"none"`. Those values asked the host to attach context the server had not supplied itself, including context from other servers, which conflicts with the principle that a server sees only what is deliberately routed to it. Everything the model should see now goes in `messages` explicitly.
  • If a sampling exchange needs several turns, how does the server carry the earlier turns?
    By resending them. MCP 2026-07-28 is stateless — each request is self-contained and nothing on the wire links one completion to the next. The server keeps the growing conversation in its own state, appends the assistant message it received, and sends the whole `messages` array again on the next `sampling/createMessage`. There is no server-side conversation handle in the protocol for this.

saying these in an interview costs you the question

  • Thinks maxTokens is optional or that the client must honour it exactly
  • Puts a system instruction into the messages array as a system role
  • Believes modelPreferences binds the client to a specific model
  • Assumes the server can pull in the user's wider chat history
  • Expects a conversation id to link successive sampling calls

context