How should an MCP client bound a tool-enabled sampling loop driven by a server?
answer
- The client retries, so the client can stop
- Cap iterations, not just single calls
- Budget across the exchange, per server
- Prompting every round trip guarantees rubber-stamping
- Aborting is just not retrying
basics
~20 sCap the round trips. In MCP revision 2026-07-28 every further step of a tool-enabled sampling exchange requires the client to re-issue the original request, so the client already holds the stop button: enforce an iteration limit, a cumulative token or spend budget, and a way for the user to walk away.
solid answer
~50 sTool-enabled sampling — signalled by the client's `sampling.tools` sub-capability — means the model the client runs can emit tool calls that the exchange has to resolve, so one server operation can turn into many model invocations paid for by the client. Revision 2026-07-28 makes the control point structural: there are no server-initiated requests, so each further step is the **client** choosing to retry its original request with the collected input. The client should therefore enforce limits it owns rather than trusting the server to stop: a hard cap on retries for a single logical operation, a cumulative token and wall-clock budget across the whole exchange, and an escalation rule that puts a human back in the loop after the first few automatic approvals. Because declining is simply not retrying, aborting costs nothing — the client stops and the exchange ends. Note that sampling as a whole is deprecated in 2026-07-28.
code
json · 15 lines{
"jsonrpc": "2.0",
"id": 7,
"method": "tools/call",
"params": {
"name": "draft_release_notes",
"arguments": { "tag": "v2.1.0" },
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientCapabilities": {
"sampling": { "tools": {} }
}
}
}
}go deeper
Know that a sampling request can carry tools, which means one server operation may trigger several model runs on the client's account, and that the client is allowed to stop.
Explain that the sampling.tools sub-capability enables this, and that since 2026-07-28 each step needs the client to retry its own request — so the client, not the server, decides whether the loop continues.
Name concrete limits you would ship: an iteration cap per logical operation, a cumulative token and time budget attributed per server, loop detection on repeated identical calls, and a user-visible abort that is simply not retrying.
Frame it as a budget-and-blast-radius decision rather than a UI one: which servers may drive your inference at all, what an exchange is allowed to cost, and whether — given the 2026-07-28 deprecation — you advertise the sampling capability in the first place.
## Why a sampling exchange can become a loop A plain sampling request is one completion: the server supplies messages, the client runs the model, the text goes back. Tool-enabled sampling is different. The client advertises the `sampling.tools` sub-capability of its (deprecated) `sampling` capability, which tells servers that a sampling request may carry tools the model is allowed to invoke. Now the model's output is not necessarily final text — it can be a request to call a tool, whose result feeds another model invocation, and so on. That means a single `tools/call` the user made can expand into an unbounded chain of model invocations, each one costing the **client's** tokens, on prompts authored by the **server**. This is the sampling-specific version of a runaway loop, and it is the client's problem because the client is the party paying and the party holding the credentials. ## The control point is already in the right place Revision 2026-07-28 removed server-initiated JSON-RPC requests. A sampling request comes back as an interim `InputRequiredResult`, and the exchange advances only when the client re-issues its original request carrying the collected input. So every single step of the loop passes through client code. There is no path by which a server keeps the loop turning on its own; it can only keep returning interim results, and only if the client keeps asking. Aborting is correspondingly cheap. Declining is expressed by simply not retrying — no error to send, no cleanup handshake. A client that decides the loop has gone on long enough just stops, and the exchange is over. ## Limits worth enforcing **An iteration cap per logical operation.** Count the interim results the client has fulfilled for one user-initiated call and refuse past a fixed number. This is the single most valuable limit because it bounds the worst case regardless of what any individual step looks like. **A cumulative budget, not a per-call one.** Each individual completion may look modest; twenty of them do not. Track tokens generated and money spent across the whole exchange, and attribute it to the server that drove it. A per-server daily budget converts a subtle abuse into a visible, bounded loss. **A wall-clock deadline.** A loop that is not growing in cost may still be pinning the user's attention or holding a tool call open. Give the whole exchange a deadline. **Escalating consent.** The realistic policy is not to prompt on every one of twenty round trips — that guarantees rubber-stamping. Approve the first request explicitly, then continue automatically within the declared budget and iteration cap, and prompt again the moment either is about to be exceeded or the prompt content changes character. The user's decision is then about the *exchange*, which is the thing they can actually reason about. **Loop detection.** If successive iterations produce the same tool call with the same arguments, the exchange is not progressing. Stop rather than paying to discover it again. **A visible abort.** Whatever the automatic limits, the user must be able to stop an in-flight exchange at any point. Because refusal is silence, the implementation of "stop" is just: do not retry. ## What the client should not rely on Do not rely on the server to bound its own loop. A server's stopping condition is code you did not write, running on someone else's machine, with no incentive to spend your budget carefully. Do not rely on the tool annotations either: `readOnlyHint`, `destructiveHint`, `idempotentHint` and `openWorldHint` are hints, and revision 2026-07-28 is explicit that clients must consider them untrusted unless the server itself is trusted. And do not rely on self-reported identity — `io.modelcontextprotocol/serverInfo` is untrusted, so budget attribution should key off the host's own configuration entry for the server, not the name the server sends. ## Server-side implications Because the client can stop at any iteration, a server driving a tool-enabled sampling exchange must be written so that stopping mid-way is safe. No open transactions across an interim result, no locks held, no assumption that a partially built artifact will be completed. And since the exchange resumes as a fresh request with a new JSON-RPC id, any work carried between steps has to be carried in the opaque state the server itself supplied, not in per-connection memory — MCP is stateless in 2026-07-28 and a connection is not a conversation. ## The strategic answer The honest senior answer ends by noting that sampling is deprecated as of revision 2026-07-28, with the migration being for servers to integrate directly with LLM provider APIs. For many hosts the correct bound on a tool-enabled sampling loop is not a clever budget — it is declining to advertise the sampling capability at all, so the loop never starts.
- Why is prompting the user on every iteration a bad approval design?Because it manufactures consent fatigue. Twenty near-identical dialogs train the user to click approve without reading, which is worse than one informed decision. The better shape is an explicit approval of the exchange with a declared iteration cap and budget, automatic continuation inside those limits, and a fresh prompt only when a limit is about to be crossed or the request changes character.
- Can the client trust a server's own stopping condition or its tool annotations?No. The server's loop logic runs on someone else's machine and spends the client's budget, so it has no incentive to be conservative. ToolAnnotations such as readOnlyHint and destructiveHint are hints, and revision 2026-07-28 states clients must treat them as untrusted unless the server itself is trusted. The client enforces its own limits.
- What must a server do differently because the client can stop the loop at any step?It must be safe to abandon mid-exchange: no transactions or locks held across an interim result, no assumption that a half-built artifact will be finished, and cleanup on a timeout. Since 2026-07-28 is stateless and each resumption is a fresh request with a new id, any carried state must travel in the opaque state the server itself issued rather than in per-connection memory.
saying these in an interview costs you the question
- Assumes the server can keep the loop running by itself
- Trusts the server's own iteration limit
- Prompts the user on every single round trip
- Budgets per completion instead of per exchange
- Treats readOnlyHint as a trustworthy safety guarantee