skip to content

In MCP, what does sampling/createMessage let a server ask the client to do?

level: juniorimportance: must knowfreq 62%

answer

  1. the server borrows someone else's model
  2. no provider key lives on the server
  3. messages out, completion back
  4. result names the model that ran
  5. deprecated in revision 2026-07-28

basics

~20 s

sampling/createMessage lets an MCP server ask the connecting client to run an LLM completion for it, so the server needs no model credentials of its own. The client picks the model and runs the call. MCP revision 2026-07-28 deprecates the feature.

solid answer

~40 s

Sampling inverts the usual direction of traffic. Instead of the client calling the server, the server asks the client's host application to run a completion on its behalf: it sends a `sampling/createMessage` request carrying a `messages` array, a required `maxTokens`, and optionally a `systemPrompt` and `modelPreferences`. It gets back a `CreateMessageResult` holding the generated message plus the `model` that actually produced it. That lets a server add model-powered behaviour — summarising a document it just read, classifying a record — without holding a provider API key or paying for inference. In revision 2026-07-28 the request is no longer a server-initiated JSON-RPC request; the server hands it to the client inside an `InputRequiredResult` returned from its own `tools/call`, `prompts/get` or `resources/read`. The same revision deprecates sampling, pointing servers at direct LLM provider integration instead.

go deeper

for a junior

Be able to say in one sentence that the server asks the client to run a completion, that the client owns the model and the credentials, and that revision 2026-07-28 deprecates the feature.

for a middle

Explain the request and result shapes — messages, maxTokens, systemPrompt, modelPreferences out; message, model and stopReason back — and that the request now rides inside an InputRequiredResult rather than arriving on its own.

for a senior

Show that you would not build a server feature on top of an optional, deprecated capability. Talk about capability checking per request, graceful degradation when no completion comes back, and writing tool handlers that resume rather than block.

for a principal

Own the decision of where model access lives for a fleet of servers: borrowing the user's model shifts cost and model choice out of your control and onto a capability that is scheduled to disappear, so plan the provider-integration path and the credential story before you depend on it.

## The idea Most of the Model Context Protocol runs one way: a client calls a server to list and invoke tools, read resources, or fetch prompts. Sampling is the one place where the server needs something back from the model side. `sampling/createMessage` is the server saying: *here is a conversation, please run it through a language model and give me the completion.* The point is credential and cost inversion. A server that wants to summarise a file, rewrite a query, or classify a record would otherwise need its own provider account, its own API key, and its own billing relationship. With sampling it borrows whatever model the user's host application already has configured. The server ships no secrets and pays for no tokens. ## The shape of the exchange The request carries a `messages` array of role-tagged content (`user` and `assistant` turns holding text, image or audio content blocks), a required `maxTokens` budget, and optional steering: `systemPrompt`, `modelPreferences`, and pass-through knobs such as `temperature` and `stopSequences`. The response is a `CreateMessageResult`: the assistant message the model produced, plus a `model` string naming what actually generated it and an optional `stopReason` explaining why generation ended. The server never receives a session with the model, a key, or the user's wider conversation — only the completion for the messages it itself supplied. ## Where the request travels in revision 2026-07-28 This is the part that trips up anyone working from older material. Through revision 2025-11-25, a server could open a JSON-RPC request of its own and push `sampling/createMessage` down the wire at the client. Revision 2026-07-28 removed server-initiated requests entirely — there is no `ServerRequest` union any more. A sampling request now appears as one entry in the `inputRequests` map of an `InputRequiredResult`, which the server returns in place of the normal result for a `tools/call`, `prompts/get` or `resources/read`. The client fulfils it and re-issues the original call, carrying the completion back in `inputResponses`. So the flow is: client calls a tool → server answers "I need a completion first" → client runs it → client calls the tool again with the answer attached. The server's tool logic is therefore written to be resumable rather than blocking. ## What the client controls Everything that matters. The client decides which model to use — `modelPreferences` is advisory, not binding — and whether to run the request at all. Approval, prompt review and the surrounding user-control duties belong to the host application, and MCP cannot enforce them at the protocol level. A server must be written to survive a completion that never arrives. ## Capability gating Sampling only exists if the client declares it. Under the modern per-request model, capabilities travel in every request's `_meta` under `io.modelcontextprotocol/clientCapabilities`, and a server must not assume support because an earlier request had it. If the capability is absent and the server needs it, the correct answer is a `MissingRequiredClientCapabilityError` (`-32021`), not a hopeful sampling request. ## Deprecated in 2026-07-28 Sampling is deprecated in revision 2026-07-28. It stays in the specification for at least twelve months, so it is still implementable, but new implementations should not adopt it; the migration path is for the server to integrate directly with an LLM provider API. In practice client support was always optional and patchy, which meant a server could never rely on the feature being there — a fragile foundation for any behaviour a server actually needs. ## What to say in an interview Name the direction (server asks, client runs), name the credential benefit, name the delivery change in 2026-07-28, and finish by saying the feature is on its way out. Candidates who describe sampling as a live, server-initiated request are describing the 2025-era protocol.

  • Does the server learn which model actually produced the completion?
    Yes. `CreateMessageResult` carries a required `model` string naming the model that generated the message. Because `modelPreferences` is only a hint, this is the server's only reliable signal about what it got — useful for logging, for deciding whether the output is trustworthy enough for the task, and for detecting that a small fast model answered a request that asked for a strong one.
  • What could a server do with sampling that it could not do with a tool result alone?
    Tool results are whatever the server can compute deterministically. Sampling gave the server language ability it did not have to build: turning a raw file it read into a summary, normalising messy user text, or drafting a reply — all charged to the user's own model account rather than the server operator's. That is exactly why removing it pushes servers toward owning a provider integration.
  • What happens if the client never declared sampling support?
    The server must not send the request. Client capabilities arrive in every request's `_meta` under `io.modelcontextprotocol/clientCapabilities` in revision 2026-07-28, and servers must not infer them from earlier requests. If a feature genuinely cannot proceed without sampling, the server returns `MissingRequiredClientCapabilityError` (`-32021`) with `data.requiredCapabilities`; otherwise it should degrade to behaviour that needs no model.

saying these in an interview costs you the question

  • Says the server runs the model itself using its own API key
  • Confuses the direction and describes it as the client calling a server tool
  • Claims the client must use whatever model the server names
  • Describes sampling as a live server-initiated JSON-RPC request in 2026-07-28
  • Presents sampling as current best practice with no mention of deprecation

context