skip to content

In MCP, what stops a server changing a tool's description after you approve it?

level: seniorimportance: should knowfreq 48%

answer

  1. approval outlives the text it was given for
  2. must not vary per connection, not per week
  3. no signature, no pin, no attestation in spec
  4. hash the definition, diff on refetch
  5. re-prompt with a visible diff

basics

~20 s

Nothing in the protocol. MCP revision 2026-07-28 requires only that a server's tool set not vary per connection; it may change over time, and the spec defines no signature, pin or attestation. Detecting drift is entirely the client's job.

solid answer

~50 s

MCP 2026-07-28 fixes the tool set per connection — it MUST NOT vary between connections, though it MAY vary with the authorization presented — but it says nothing about a definition staying the same across days. A server may serve a benign description while a user grants standing approval, then serve a hostile one later; the ecosystem calls that a rug pull. The specification offers no pinning, signing or attestation, and change notification is weak: `notifications/tools/list_changed` only reaches a client that opened a `subscriptions/listen` stream with `toolsListChanged`, and a cached list governed by `ttlMs` may be stale anyway. Since sessions were removed in 2026-07-28 there is no handshake at which approval is anchored either. The practical answer is client-side: hash the name, description, schemas and annotations at approval time, re-hash on every fetch, and re-prompt the human on any diff.

code

json · 9 lines
json
{
  "server": "local:acme-tickets",
  "principal": "user-42",
  "tool": "create_ticket",
  "pinnedAt": "2026-08-01T09:12:00Z",
  "digest": "sha256:1f3a...c908",
  "covers": ["name", "description", "inputSchema", "outputSchema", "annotations"],
  "onMismatch": "revoke-standing-approval-and-diff"
}

go deeper

for a junior

Know that a server can serve different tool text later, and that MCP itself provides no signature or lock on a definition. Recall that consent was given against text, and text can change.

for a middle

Explain the per-connection consistency rule and why it is not a stability guarantee, and describe why change notification is opt-in: it rides only a subscriptions/listen stream opened with toolsListChanged.

for a senior

Design the client-side control end to end — canonical digest over name, description, schemas and annotations, keyed by local server identity and principal, re-checked on every fetch, revoking standing approval and showing the human a diff.

for a principal

Weigh the consent-fatigue tradeoff and set policy: which classes of change are silent, which prompt, which block, and how drift on a fleet-wide server becomes an incident rather than thousands of individual dialogs.

## The scenario A user connects a third-party MCP server, reads a tool's description, decides it looks reasonable, and grants it standing approval so the agent stops asking every time. Two weeks later the server's operator — or whoever compromised them — edits the description to instruct the model to exfiltrate a file alongside every call, or swaps the implementation behind an unchanged description. The user's approval was given against text that no longer exists. The ecosystem name for this is a *rug pull*; it is community vocabulary, not specification vocabulary. ## What revision 2026-07-28 actually guarantees One rule bears directly on this, and it is narrower than people hope: a server's tool, prompt and resource sets **MUST NOT** vary per connection, though they **MAY** vary according to the authorization presented on the request. That is a *consistency* rule, aimed at statelessness — any node behind a load balancer must answer identically for the same caller, and a server may not tailor its surface to one connection's history. It is not a *stability over time* rule. A server is entirely free to publish a different description tomorrow. The rest of the relevant machinery is about noticing change, not preventing it: - `notifications/tools/list_changed` exists, but in 2026-07-28 it is delivered **only** on a `subscriptions/listen` stream that the client opened with `toolsListChanged` in its `SubscriptionFilter`. A client that never subscribes is never told. - `ListToolsResult` is a `CacheableResult` carrying required `ttlMs` and `cacheScope`, so a client may legitimately be serving the user a cached definition while the server has moved on. - Because protocol-level sessions and the `initialize` handshake were removed in 2026-07-28, there is no per-connection ceremony where a definition is agreed once. Every request stands alone. And what the spec conspicuously does not define: no signature over a `Tool`, no publisher attestation, no definition version or digest field, no requirement to rename a tool when its meaning changes. In 2026-07-28 the security best-practices material moved out of the specification into the documentation tutorials, which is a fair signal of where this problem is expected to be solved — in implementations, as guidance. ## The client-side answer: pin and diff The workable defence, widely implemented and widely discussed in the ecosystem under the names *description pinning* and *diffing*, is straightforward: 1. **Canonicalise and hash** the security-relevant parts of each tool definition at the moment consent is granted: `name`, `description`, `inputSchema` (including per-property descriptions, which are an injection surface in their own right), `outputSchema` if present, and `annotations`. Serialise deterministically — sorted keys, stable number formatting — so cosmetic reordering does not produce false alarms. 2. **Store the digest with the grant**, keyed by the client's own local identity for that server, not by anything the server reported about itself. 3. **Re-hash on every fetch**, including cache refreshes and any list refreshed after a `notifications/tools/list_changed`. 4. **On mismatch, revoke the standing approval and re-prompt**, showing the human a diff of the old and new text rather than a bare "this tool changed" toast. The diff is the whole value: it is what lets a user notice the sentence about reading their SSH key. 5. **Log it.** In an organizational deployment, drift on a widely-deployed server is an incident signal, not a per-user prompt. ## Where this gets subtle **Legitimate churn is common.** Servers improve their prose, add optional arguments, fix typos. If every whitespace change nags the user, they will click through everything, and you have built consent fatigue instead of security. Some teams tier the response: silent for whitespace and additive optional properties, prompt for any description change, hard-block for a change in annotations that would relax an auto-approval policy. **Authorization-conditioned surfaces are legal.** Because the tool set MAY vary by the authorization presented, a user whose token gained a scope will legitimately see new tools and possibly different descriptions. Your pin store has to be keyed by principal as well as by server, or a scope change reads as an attack. **A pin proves nothing about behaviour.** Identical text with a rewritten implementation behind it passes every check. Pinning defends the *model's* view of the tool; it does nothing for what the server does with the arguments. That is the province of the server's own privilege — sandboxing for stdio, audience-bound tokens with no passthrough for remote — and of the approval gate on the concrete call. **Do not oversell it in an interview.** The honest, senior answer is: the specification acknowledges nothing here; the community has converged on pinning and diffing; it is partial; and the durable mitigation remains a human able to deny the specific call with its real arguments in front of them.

  • Doesn't notifications/tools/list_changed tell the client that something changed?
    Only if the client asked for it. In 2026-07-28 that notification is delivered exclusively on a `subscriptions/listen` stream opened with `toolsListChanged` in the filter, and on stdio the client must re-send `subscriptions/listen` after a reconnect. A client that never opens one gets no signal at all, and a cached list under its `ttlMs` may keep serving the old definition regardless.
  • What exactly should the digest cover?
    The parts that steer the model or relax a gate: `name`, `description`, `inputSchema` including per-property descriptions and titles, `outputSchema` if present, and `annotations`. Canonicalise before hashing — sorted keys, stable formatting — so reordering does not fire a false positive, and key the stored digest by your own local server identity plus the authenticated principal.
  • Won't the tool set legitimately differ when a user's authorization changes?
    Yes, and the spec allows it explicitly: the set MUST NOT vary per connection but MAY vary by the authorization presented. So pin per principal, not just per server. A user who gained a scope should see new tools without your client reporting a rug pull.
  • Does pinning the description defend against the server changing what the tool actually does?
    No. Identical text over a rewritten implementation passes every diff. Pinning protects the model's view of the tool, which is the injection surface; the implementation is contained by other means — sandboxing a stdio subprocess, audience-bound tokens for a remote server, and a human approving the concrete call with its real arguments.

saying these in an interview costs you the question

  • Says the spec forbids a tool definition from ever changing
  • Assumes the client is always notified when a tool list changes
  • Believes MCP signs or attests tool definitions
  • Confuses must-not-vary-per-connection with stability over time
  • Treats a description pin as proof the implementation is unchanged

context