Why does MCP sampling mean a server needs no LLM API key of its own?
answer
- The server borrows, it does not own
- Credentials, model choice and bill stay put
- Nothing to steal because nothing was given
- But your key can still be driven remotely
- Deprecated in 2026-07-28
basics
~20 sBecause the server does not call a model — it asks the client to. The client runs the completion with its own provider credentials and returns the generated text, so the credentials, the model choice and the bill all stay on the client's side.
solid answer
~50 sSampling inverts who holds the key. Instead of the server calling an LLM provider directly, it asks the client to run a completion on its behalf and hand back the result. The client already has model access, so nothing new has to be given to the server: it never sees a provider API key, never picks the model itself, and never spends money that is not going through the user's own account. That is the trust argument for the feature — a server you installed cannot quietly rack up inference costs or exfiltrate your provider credentials, because it was never given any. The flip side is that the client is now paying and generating on a third party's behalf, which is exactly why the specification pairs sampling with human review. Note that sampling is deprecated as of revision 2026-07-28, with the migration guidance telling servers to integrate directly with LLM provider APIs.
go deeper
Be able to say that the server asks the client to run the completion, so the client's credentials and model are used and the server only receives text. Do not claim a key is shared.
Explain the three things that stay client-side — credentials, model choice and cost — and why that arrangement still needs human review, since a remote party is shaping prompts that your credentials execute.
Talk about the outbound direction and the spend surface: budgets per server, redaction of the completion, and why an unattended host should not advertise the sampling capability at all.
Own the tradeoff the 2026-07-28 deprecation makes: moving servers onto their own provider integrations restores cost accountability and removes a confused-deputy shape, at the price of losing the host's natural review point. Decide which risk your organisation would rather carry.
## The two ways a server could get a completion Imagine an MCP server that needs a short natural-language summary in the middle of doing its job. It has two options. The direct one: the server holds an API key for some LLM provider, calls that provider itself, and pays for the tokens. Simple, but it means every server you install is a separate credential to provision, a separate bill, a separate vendor relationship, and a separate thing that can silently spend money or send your data somewhere you did not choose. The sampling one: the server does not call anybody. It asks the **client** — the application that is already talking to a model on the user's behalf — to run the completion and give it the text back. This is what MCP sampling is. ## What stays on the client's side Three things, and it is worth being able to list them cleanly: - **The credentials.** The provider API key lives in the host application. Nothing about sampling transfers it, exposes it, or requires the server to be trusted with it. A malicious server gains no access to your provider account. - **The model choice.** The server can express what it would like, but the client is the one that actually issues the completion, so it decides which model runs. A server cannot force an expensive frontier model on a user who has configured a cheap local one. - **The bill.** Tokens are spent on the client's account, because it is the client's credentials making the call. This is the property that makes sampling convenient and also the one that makes unbounded sampling dangerous. What the server gets back is just text. ## Why the trust boundary needs a human in it The same inversion that protects your credentials creates a new exposure: a remote party can now cause your model to run, with your money, on text it wrote. That is why the specification's sampling guidance asks the client to let a user review the request before it runs and review the completion before it is returned. Without those gates, "the server has no API key" is a hollow reassurance — it does not need one when it can drive yours. The second exposure is outbound. The completion is generated on the client's side and then handed to the server, so anything the model produces reaches a remote party. This is why the `includeContext` values `"thisServer"` and `"allServers"` were deprecated in revision 2026-07-28 in favour of omitting the field or using `"none"`: pulling extra conversation context into a server-requested completion widens what can come back out. ## The 2026-07-28 status Sampling is **deprecated** as of revision 2026-07-28. Under the specification's deprecation policy a deprecated feature stays in place at least twelve months and new implementations should not adopt it, with the earliest removal being the first revision released on or after 2027-07-28. The published migration guidance is that servers should integrate directly with LLM provider APIs. That migration reverses the property this question is about: a server that follows it *does* hold its own key, its own model choice and its own bill. That is not a step backwards in every dimension — it removes the awkward situation where a third party drives your credentials, and it moves the cost to the party that wanted the completion. But it does remove the natural review point, so the user's remaining controls become the ordinary ones: consenting to each tool call, and choosing which servers to run at all. ## How to answer this in an interview State the inversion in one sentence — the server borrows the client's model rather than holding a key — then name the three things that stay client-side, then immediately say why that makes human review necessary rather than optional. Finishing with the 2026-07-28 deprecation and its migration direction shows you know the current revision rather than the material that was written about MCP before it.
- If the server holds no key, what is the remaining risk to the user?That a remote party can drive the model the user does pay for. A server writes the prompt, so it can spend the user's tokens and shape what the model produces, and the completion it receives back is an outbound channel for anything in that context. The protections are human review of the prompt and the completion, plus per-server spend limits in the host.
- Sampling is deprecated in 2026-07-28 — what does the migration do to this property?It reverses it. The migration guidance tells servers to integrate directly with LLM provider APIs, so a migrated server holds its own credentials, chooses its own model and pays its own bill. The user loses the built-in review point and falls back to the ordinary controls: consenting to each tool call and choosing which servers to run.
- Who is billed for the tokens a sampling completion consumes?The client's account, because it is the client's provider credentials making the call. The server that asked pays nothing. That asymmetry is precisely why a client should attach a budget to sampling rather than only an approval click, and why an unattended host is usually configured not to advertise the sampling capability at all.
saying these in an interview costs you the question
- Thinks the client sends the server a provider API key
- Says the server picks which model actually runs
- Assumes the server pays for the tokens it requests
- Treats no-key as meaning the server is harmless
- Describes sampling as current rather than deprecated in 2026-07-28