How does tool-enabled MCP sampling use the tools and toolChoice fields on createMessage?
answer
- the model borrows the server's functions
- the server runs its own tools, not the client
- a tool_use block instead of an answer
- a matching result block goes back in the conversation
- every iteration is a whole round trip
basics
~20 sA server can attach its own tool definitions in tools on sampling/createMessage, with toolChoice constraining selection. The borrowed model may then answer with a tool_use block; the server executes that tool itself and continues by feeding a tool_result block back in messages.
solid answer
~50 sTool-enabled sampling lets the completion the server borrows do more than emit text. The server puts its own tool definitions in the request's `tools` field and may narrow selection with `toolChoice`. If the model decides to use one, the `CreateMessageResult` comes back carrying a `tool_use` content block instead of a plain answer. The server — not the client — executes that tool, then issues a fresh `sampling/createMessage` with the conversation extended by the `tool_use` block and a matching `tool_result` block, and repeats until the model returns a normal answer. Two constraints shape the design. First, the client must advertise the `sampling.tools` sub-capability, and in revision 2026-07-28 capabilities arrive per request, so this is checked every time. Second, each round is a full round trip through the client's `InputRequiredResult` retry, so a loop of five tool calls is five user-visible completions. Sampling is deprecated in 2026-07-28.
code
json · 19 lines{
"method": "sampling/createMessage",
"params": {
"messages": [
{
"role": "user",
"content": { "type": "text", "text": "Who is on call right now?" }
}
],
"tools": [
{
"name": "lookup_rota",
"description": "Return the current on-call rota",
"inputSchema": { "type": "object", "properties": {} }
}
],
"maxTokens": 256
}
}go deeper
Know that a sampling request may carry tools and toolChoice, and that the model can answer with a tool_use block instead of plain text.
Explain the loop: tool_use comes back, the server executes its own tool, appends a tool_result block to messages, and sends a fresh request — with the whole conversation resent each time because the protocol is stateless.
Show production judgment: check sampling.tools on every request, validate and authorize tool_use arguments as untrusted input, keep loops shallow, and make the handler resumable because the client may stop after any iteration.
Weigh the architecture: a tool loop is the workload where borrowing a model costs the most, since each step pays a round-trip and a consent moment. That, plus the 2026-07-28 deprecation, argues for running the loop inside your own provider integration.
## What tool-enabled sampling adds Plain sampling gives a server one shot: send a conversation, get text back. Tool-enabled sampling lets the borrowed model reach back into the server mid-completion. The server advertises callable functions on the request, and the model may choose to invoke one before producing its final answer. The key ownership point, which interviewers probe: **the tools listed here are the server's own, and the server executes them.** They are not the tools the server exposes to the client through `tools/list`, and the client does not run them. The server is playing the role of the application driving a tool loop, using someone else's model as the reasoning engine. ## The request side `tools` carries the definitions — name, description and an input schema — the model may choose from. `toolChoice` constrains how the model is allowed to select, letting the server express something stronger than "here are some options" when the task demands it. Both fields are optional. A request without them is ordinary sampling. ## The loop 1. Server sends `sampling/createMessage` with `messages` and `tools`. 2. The model decides it needs a tool; the `CreateMessageResult` content is a `tool_use` block naming the tool and its arguments. 3. The server validates those arguments and executes the tool itself. 4. The server appends the assistant `tool_use` block and a corresponding `tool_result` block to its copy of `messages`, and sends a new `sampling/createMessage`. 5. Repeat until the result is a normal content block rather than a `tool_use`. Because MCP revision 2026-07-28 is stateless, step 4 means resending the entire grown conversation each time. Nothing on the wire links iteration three to iteration two; the server holds the loop state. ## Capability gating This only works if the client supports it. The client's `sampling` capability carries a `sampling.tools` sub-capability, and revision 2026-07-28 delivers client capabilities in every request's `_meta` under `io.modelcontextprotocol/clientCapabilities` — servers must not infer them from a prior request. So the check is per request, not once per connection. A server that requires the feature and does not find it should fail cleanly with `MissingRequiredClientCapabilityError` (`-32021`), naming what it needed in `data.requiredCapabilities`, rather than sending a request the client cannot honour. ## Why this is expensive Each iteration is a complete round trip through the MRTR mechanism: the server returns an `InputRequiredResult`, the client fulfils the completion and re-issues the original call, and the server resumes. A five-step tool loop is five completions charged to the user's account, five re-issued calls, and five points at which the user's host may decline to continue. That is a very different cost profile from a loop running inside the server against its own provider connection, where the same five steps are five cheap API calls the server controls end to end. The design consequences are real: keep tool-enabled sampling loops shallow, make each tool do meaningful work rather than one small lookup, and make every step of the loop resumable, because the loop can stop at any iteration and never come back. ## Validate the arguments The `tool_use` block was produced by a model the server did not choose, driven by a conversation that may contain content the server did not author. Treat its arguments exactly as you would treat any external input: validate against the tool's schema, authorize the operation on its own merits, and never let a model-produced argument select a privileged code path. "The model asked for it" is not authorization. ## The deprecation angle Sampling — including its tool-enabled form — is deprecated in revision 2026-07-28, with direct LLM provider integration as the migration path. That migration is particularly attractive here: a tool loop is exactly the workload where borrowing a model hurts most, since every iteration pays the round-trip and consent costs of the MRTR pattern. A server that runs its own provider connection keeps the loop internal, fast, and entirely under its control. ## Answering well Say who defines the tools (the server), who executes them (the server), what the client contributes (the model), and what each iteration costs (a full round trip). Then close on the deprecation and why an internal loop is the better shape.
- Are the tools listed on a sampling request the same tools the server exposes through tools/list?Not necessarily, and conceptually they are a different set. `tools/list` advertises what the client's agent may call; the `tools` field on `sampling/createMessage` advertises what the borrowed model may call back into the server during that one completion. A server may reuse the same definitions, but the audience, the executor and the trust context differ, so treating them as one list is a design smell.
- What should a server do if the client has not advertised sampling.tools?Not send tool-enabled sampling. Capabilities arrive in every request's `_meta` under `io.modelcontextprotocol/clientCapabilities` in revision 2026-07-28, and servers must not infer them from earlier requests. If the feature is essential, return `MissingRequiredClientCapabilityError` (`-32021`) with `data.requiredCapabilities`; otherwise fall back to a plain sampling request or to behaviour that needs no model at all.
- Why is a five-step tool loop over sampling much more costly than the same loop inside your server?Every step is a full MRTR round trip: the server returns an `InputRequiredResult`, the client runs the completion, then re-issues the original call with a new JSON-RPC id. Five steps mean five completions on the user's account, five resumptions of your handler, and five chances for the user to stop. An internal loop against your own provider is five ordinary API calls you control end to end.
saying these in an interview costs you the question
- Says the client executes the tools listed on the sampling request
- Confuses these tools with the ones advertised via tools/list
- Trusts tool_use arguments without validating or authorizing them
- Assumes the tool loop keeps state on the wire between iterations
- Ignores that each iteration is a separate round trip and completion