How would you decide which third-party MCP servers an organization may connect?
answer
- the protocol supplies no trust signal
- tier by provenance, not by tool
- pin definitions, diff on refetch
- the model is the shared surface
- scope credentials so compromise buys little
basics
~20 sTreat each server as untrusted code supplying text into a model's prompt. Tier servers by provenance, allowlist what may be connected, pin and diff their tool definitions, isolate low-trust servers from sensitive contexts, and scope their credentials so a compromise buys little.
solid answer
~40 sMCP carries no trust signal — annotations are untrusted, `serverInfo` is self-reported, and the protocol explicitly cannot enforce its consent, privacy and tool-safety principles, leaving the host as the boundary. So the decision is organizational, not technical. I would run tiers: first-party and reviewed servers, where read-only tools may auto-approve; allowlisted vendor servers, where every call is confirmed; and everything else, blocked at the host configuration. Across tiers, pin each tool's name, description, schemas and annotations at approval and diff on refetch, since 2026-07-28 offers no attestation. Isolate contexts so a low-trust server never sits alongside sensitive tools — servers cannot see each other, but the model is a shared surface they can all write into. Scope credentials tightly: audience-bound tokens, no passthrough, and a real sandbox around any stdio subprocess.
go deeper
Know that connecting an MCP server is a trust decision made outside the protocol, and that the wire offers no signal — annotations and serverInfo are both untrusted.
Explain the concrete controls: an allowlist of connectable servers, pinning and diffing tool definitions, and requiring human confirmation of the actual arguments for anything that writes.
Design the layered deployment — tiers by provenance, sandboxing for stdio, audience-bound tokens for remote, per-call logging with server identity — and explain why context isolation is the control with the largest blast-radius effect.
Own the organizational shape: who approves a server, what evidence a promotion requires, how a tier is revoked within hours, and how you balance isolation and prompting against the consent fatigue and shadow IT they create.
## Frame the decision correctly first The most common mistake here is treating an MCP server as an API integration. It is closer to installing a plugin that both runs code with your credentials and writes text directly into your model's prompt. MCP revision 2026-07-28 is candid about the limits: the three key principles — User Consent and Control, Data Privacy, Tool Safety — cannot be enforced at the protocol level, and implementors are told to build consent and authorization into the host. In the same revision the security best-practices material moved out of the specification into the documentation tutorials, which is another way of saying the protocol has finished its part and the rest is yours. Concretely, the wire gives you nothing to base trust on. `ToolAnnotations` are hints a client MUST treat as untrusted unless the server is trusted. `io.modelcontextprotocol/serverInfo` is self-reported. There is no signature over a tool definition and no publisher attestation. Every trust decision is therefore made out of band, by people, and your job is to make that decision explicit, reviewable and revocable. ## Tiers, because uniform policy is either useless or unusable A single global policy fails in both directions: confirm everything and users click through blindly; confirm nothing and one poisoned description drains a repository. Tier instead. **Tier 1 — first-party.** Built and deployed by your own teams from source you review, with a known operator. Here annotation-driven auto-approval for read-only tools is defensible, because the trust it inherits comes from your build pipeline, not from the annotation. **Tier 2 — allowlisted vendor.** A named vendor, a pinned version, a reviewed tool inventory, ideally a remote server whose token is audience-bound to it. Every call that writes gets a human confirmation showing real arguments; reads may be batched or throttled rather than silent. **Tier 3 — everything else.** Not connectable. Enforce it in host configuration management, not in a policy document, because the failure mode is one engineer adding a server to their own config. The interesting design question is the promotion path: what evidence moves a server from tier 3 to tier 2, who reviews it, and how a tier is revoked in an afternoon when a vendor is breached. ## Controls that apply across every tier **Pin and diff definitions.** Hash `name`, `description`, `inputSchema` including per-property descriptions, `outputSchema` and `annotations` at approval; re-hash on every fetch; revoke standing approval and show a human diff on mismatch. This is the ecosystem's answer — description pinning and diffing — to the fact that the spec guarantees only that a tool set does not vary per connection, never that it is stable over time. In a fleet, drift on a widely deployed server should page someone, not prompt ten thousand users. **Isolate contexts.** This is the control people skip and the one with the largest blast-radius effect. An MCP server cannot read the conversation and cannot see into other servers; those isolations are real. But the model is a shared surface every connected server writes into, so a description from a low-trust server can steer the model into calling a high-trust server's tool with sensitive arguments. The structural fix is to never assemble a context containing both a low-trust server and the tools or data you would mind exfiltrating. Separate agents, separate profiles, separate credentials. **Least privilege on the server itself.** For remote servers, MCP 2026-07-28's OAuth profile helps: the server is a resource server, tokens are audience-bound, and token passthrough is forbidden, so a stolen token is scoped to that server. For stdio servers, none of that applies — you launched a subprocess with your user's filesystem and network access — so the sandbox is the control: restricted filesystem view, egress rules, a dedicated low-privilege account. **Make provenance unspoofable in the UI.** Approval prompts must name servers by your own local identity, never by a self-reported name, and combined tool lists must be prefixed from client-side identity — `serverInfo.name` is explicitly unreliable for this. **Instrument.** Log every `tools/call` with server identity, principal, tool and argument fingerprint. Post-incident, the question is always "which servers were in context when this happened", and without that log there is no answer. ## The tradeoffs to say out loud Every control above costs something. Isolation fragments the agent experience and is the reason people connect everything to one context in the first place. Diff prompts cause consent fatigue, and a fatigued user is a worse control than no prompt at all — so tier the response: silent for whitespace, prompt for description changes, block for annotation changes that would relax auto-approval. Allowlists concentrate a review bottleneck; budget for it or it becomes shadow IT. And be honest that none of this defends against a first-party server whose implementation is quietly hostile — pinning protects the model's view of a tool, not what the code does with the arguments. The principal-level answer is not a list of controls. It is: name the boundary the protocol declines to defend, decide who in the organization owns it, and make the policy enforceable in configuration rather than in guidance.
- Servers cannot see each other, so why does isolating them matter?Because the isolation is at the protocol layer, not the reasoning layer. A server sees only its own arguments and cannot inspect another server — but every connected server writes text into the same model context, so a hostile description can steer the model into calling a trusted server's tool with sensitive data. Partitioning contexts is what removes that path.
- Where does a stdio server sit in this policy compared with a remote one?Usually higher risk. A remote server is contained by the OAuth profile — it is a resource server, tokens are audience-bound and passthrough is forbidden — so a compromise is scoped. A stdio server is a subprocess you launched with your user's filesystem and network reach, and MCP defines no confinement for it, so containment is entirely your sandbox.
- How do you avoid consent fatigue while still prompting on drift?Grade the response to the change. Ignore whitespace and purely additive optional properties; prompt with a visible diff when description or schema prose changes; hard-block a change to annotations that would relax an auto-approval rule. And in a fleet, route drift on a shared server to an incident channel once rather than to every user's dialog.
- What can this policy not protect against?A first-party or allowlisted server whose implementation is hostile behind unchanged text. Pinning and diffing protect the model's view of a tool; they say nothing about what the code does with the arguments it receives. That residual risk is handled by the server's own privilege — sandboxing, scoped tokens, egress control — and by logging that makes the abuse visible afterwards.
saying these in an interview costs you the question
- Treats an MCP server like an ordinary API integration
- Relies on annotations to classify a third-party server as safe
- Assumes protocol-level server isolation prevents cross-server exfiltration
- Puts every connected server into one shared model context
- Writes trust policy as guidance rather than enforcing it in configuration