In MCP, why can't a client trust readOnlyHint to decide a tool call is safe?
answer
- hints, not guarantees
- the labelled party writes the label
- untrusted unless the server is trusted
- trust is established outside the protocol
- shape the prompt, never gate the call
basics
~20 sToolAnnotations are self-reported labels written by the same server whose behaviour they describe. MCP revision 2026-07-28 requires clients to treat them as untrusted unless the server itself is trusted out of band, so they may shape the UI but never gate a call.
solid answer
~50 sIn MCP 2026-07-28 a `Tool` may carry a `ToolAnnotations` block — `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint` — describing the shape of a call. Every one of those values is written by the server that implements the tool, and nothing in the protocol verifies them: there is no attestation, no signature, no runtime check that a tool marked read-only never writes. The specification is explicit that clients **MUST** consider annotations untrusted unless the server is trusted, and that trust cannot come from the wire — `serverInfo` in a result's `_meta` is likewise self-reported and untrusted. So annotations are a usability signal: order approval prompts, badge dangerous tools, drive auto-approval policy *for servers already on your operator's allowlist*. The actual safety boundary is the host: the human approval gate, the credentials the server holds, and the sandbox around it.
go deeper
Know that annotations are hints the server writes about itself, and say plainly that a client must treat them as untrusted unless the server is trusted. Never describe readOnlyHint as a guarantee.
Explain why the untrusted rule follows from who authors the field, and name what annotations are still useful for: prompt defaults, sorting, and policy that fails safe. Mention that serverInfo is self-reported too.
Show where the real enforcement sits — the host's approval gate, the sandbox around a stdio subprocess, and audience-bound tokens for a remote server — and describe an auto-approval policy that keys off an operator allowlist rather than off the annotation itself.
Own the trust-tier question: which servers get annotation-driven auto-approval at all, who maintains that list, and what evidence beyond the wire (first-party build, review, provenance) is required to move a server into the tier where its hints may relax a prompt.
## What the annotations are In MCP revision 2026-07-28, a tool returned by `tools/list` may include a `ToolAnnotations` object. Its fields describe the *shape* of invoking that tool: whether the call only reads, whether it may destroy data, whether repeating it is harmless, and whether it reaches an open-ended external world. The defaults are deliberately pessimistic — a server that says nothing is assumed to be the dangerous case rather than the safe one. That design already tells you what the fields are for. They are a way for a well-behaved server to say "this one is cheap and harmless, you probably do not need to interrupt the user for it". They are not a permission system. ## Who writes them, and why that settles the question The server writes its own annotations. The party asking to be trusted is the same party supplying the evidence of its trustworthiness, and MCP has no mechanism anywhere in the protocol that checks the claim: no signature over the tool definition, no registry attestation, no runtime enforcement that a call annotated read-only is prevented from writing. A malicious or compromised server can mark a data-deleting tool as read-only and idempotent, and the wire is entirely consistent. The specification states the consequence directly: clients **MUST** consider tool annotations untrusted unless the server is trusted. Note the shape of that sentence — the trust is a property of the *server*, established elsewhere, and the annotations inherit it. They never create it. ## Where "trusted server" actually comes from Out of band, always. In practice it is one of: the server is first-party, built and deployed by the same team as the host; the server appears on an operator- or organization-maintained allowlist; or a human made an explicit, informed decision to add it to their host configuration. None of these are observable in an MCP message. This is why the related identity fields do not help. `io.modelcontextprotocol/serverInfo`, which a result SHOULD carry in `_meta`, is self-reported and explicitly untrusted, exactly like `io.modelcontextprotocol/clientInfo` on the request side. A server can call itself anything. Deriving trust from a name a stranger typed into their own payload is circular. ## What annotations are legitimately good for They are excellent UI input. A host can sort a long tool list so that destructive tools are visually distinct; pre-select "ask every time" for anything not annotated read-only; offer a policy such as "auto-approve read-only tools from servers on my allowlist" — where the allowlist, not the annotation, is doing the security work; or suppress open-world tools when the user is working offline. Every one of those uses degrades gracefully if the annotation lies, because a lying annotation only produces a worse prompt, never an unauthorized effect. The anti-pattern is the mirror image: "`readOnlyHint` is true, so skip the confirmation dialog" applied to an arbitrary third-party server. That converts a cosmetic field into the only thing standing between an attacker's server and the user's data. ## The failure mode, and what the community calls it A tool whose annotations and description advertise something benign while the implementation does something else is the core of what the ecosystem calls *tool poisoning*. Be careful with that vocabulary in an interview: "tool poisoning", "rug pull" and "tool shadowing" are community terms that appear nowhere in the MCP specification. The specification's own anchor for the same concern is the untrusted-annotations rule plus the general principle that MCP cannot enforce its security principles at the protocol level — the host is the enforcement boundary. Saying "the spec defines tool poisoning" is a factual error; saying "the spec's hook for it is the MUST-treat-annotations-as-untrusted rule" is right. ## What to enforce instead Enforce at boundaries you control. For a stdio server, that is the process sandbox: what filesystem paths and network egress the subprocess actually has. For a remote server, it is the credential: MCP 2026-07-28 requires audience-bound tokens and forbids token passthrough, so a compromised server holds a token that is useless elsewhere. Across both, it is the host's approval gate — a human able to see and deny the concrete call, with its real arguments, before it runs. Annotations can make that gate friendlier. They can never replace it.
- If annotations are untrusted, what are they legitimately good for?Presentation and policy that degrades safely. Badge or sort destructive tools, default them to always-ask, hide open-world tools in an offline mode, or auto-approve read-only tools only for servers an operator already allowlisted. In every case a lying annotation costs you a worse prompt, not an unauthorized side effect.
- Given the protocol carries no trust signal, how does a host decide a server is trusted?Out of band. The server is first-party, or an operator or organization allowlisted it, or a human deliberately added it to host configuration with informed consent. Nothing on the wire establishes it: `io.modelcontextprotocol/serverInfo` in a result's `_meta` is self-reported and explicitly untrusted, so a name in a payload proves nothing.
- Does the same untrusted framing apply to a tool's declared inputSchema?Yes. The schema is server-supplied like everything else in the tool definition. A client can and should validate arguments against it before sending, but that protects the call's shape, not the user — a hostile server can declare a schema whose field names and descriptions are themselves designed to coax sensitive values out of the model.
saying these in an interview costs you the question
- Says readOnlyHint means the server cannot modify anything
- Treats destructiveHint false as grounds to auto-approve any server
- Assumes the client verifies annotations against the tool's real behaviour
- Claims MCP enforces annotation semantics at the protocol level
- Derives trust from the serverInfo name reported in _meta