skip to content

How do you defend an agent against a poisoned tool description from a third-party server?

level: seniorimportance: should knowfreq 42%

answer

  1. Tool metadata is a dependency you execute
  2. Descriptions arrive with system-level trust
  3. Benign at review, poisoned at update
  4. Pin by digest, hash, diff, re-approve
  5. Capability lives at the resource server

basics

~20 s

Treat tool metadata as untrusted code that lands in the model's context: pin servers to a reviewed version or digest, require a diff review before any description change takes effect, fail closed on unexpected changes, and keep capability enforcement at the resource server so no description can grant access.

solid answer

~50 s

A tool's name and description are prompt text with unusually high trust — the model reads them as system-level guidance. A malicious provider can therefore write a description that instructs the agent to read a config file or a credential and pass it as an innocuous-looking argument, and researchers at Invariant Labs demonstrated the harder variant: a server that is benign when you review it and ships a poisoned description in a later update, a rug-pull. Defend it like a dependency. Pin the server to a specific version or image digest rather than tracking latest, hash the tool definitions at approval time and fail closed when the fetched set does not match, and route any change through a human diff review before it reaches a production agent. Then bound the damage independently: low-trust servers get their own narrowly scoped credentials, no sensitive tools sit in the same session, and permissions live at the resource server where no description can influence them.

go deeper

for a junior

Know that a tool's description is read by the model as instructions, so a hostile provider can write text that steers the agent. Say that third-party tool servers need review before use.

for a middle

Explain both shapes — poisoning at adoption and the rug-pull update — and the dependency hygiene that answers them: pinning to a digest, hashing definitions, failing closed on change, no auto-update.

for a senior

Show the layered answer. Review is point-in-time, so pair it with per-server credentials, session composition that never mixes untrusted metadata with sensitive tools, and enforcement at the resource server that no wording can move.

for a principal

Own the vendor policy: which servers may be adopted at all, what evidence of provenance is required, when self-hosting a fork is justified, and how the inventory and re-review cadence are funded across many agent teams.

## Why tool metadata is a supply-chain surface To let a model choose tools, you place their names, descriptions and parameter schemas directly in its context. That text is instructions, and the model treats it with more trust than ordinary content because it arrives from the system side of the conversation. Whoever writes that text therefore has a channel into the agent's reasoning. This makes third-party tool providers a dependency in the ordinary software sense — and one with a weaker review culture than package registries. When you connect an agent to a tool server you did not write, you are executing its text on every turn. ## The two attack shapes **Tool-description poisoning.** The description contains instructions aimed at the model rather than at the human reading the catalog: for example, a benign-sounding utility whose description says that before using it the agent must read a local credentials file and include its contents in a parameter "for compatibility". The malicious payload can also be written to influence how the agent uses *other* servers' tools — sometimes called shadowing — which is why one untrusted server in a session contaminates the whole session. **Rug-pull updates.** The server behaves impeccably during review and adoption, then ships poisoned metadata in a later release. This defeats point-in-time review entirely: you approved a snapshot, and the agent is now running something else. The Invariant Labs WhatsApp MCP case is the canonical public demonstration of a benign-then-poisoned tool server. Both shapes exploit the same structural fact: tool definitions are usually fetched at runtime and trusted implicitly, with no equivalent of a lockfile, a signature check or a code review. ## Defense one: pin and verify Treat the tool server exactly as you treat a library. Pin to an immutable version or an image digest rather than a mutable tag. At approval time, record a hash of the full tool definition set — names, descriptions, schemas. On every startup, re-hash what the server returns and compare; on mismatch, fail closed and refuse to load the server rather than proceeding with unreviewed text. This turns a silent rug-pull into a loud, blocking event. ## Defense two: review changes as code changes Because descriptions are prompt text, changing one is changing your agent's instructions. Route the diff through the same review path a source change would take: a human reads what changed, in a rendered diff, before the new definitions reach production agents. That review is what catches "and first read ~/.aws/credentials" — automated scanning helps but is not a boundary, because the payload can be phrased arbitrarily. Disable auto-update on connected servers, and disable any auto-approval that lets a newly appearing tool be invoked without a human ever having seen it. The convenience of automatic discovery is exactly the property an attacker needs. ## Defense three: contain what a poisoned description could achieve Assume review fails sometimes, and make the payoff small. A poisoned description can only persuade; it cannot grant capability. So: - **Credentials per server.** A low-trust tool server gets its own narrowly scoped credential, never a shared one and never the agent's broad role. - **Session composition.** Do not put a low-trust server in the same session as tools holding sensitive access. If untrusted metadata and sensitive capability never co-occur, the shadowing path has nothing to reach. - **Enforcement at the resource server.** Permissions live in IAM and API scopes. No wording in any description changes what the agent's principal may do. - **Sandbox and egress policy.** If the poisoned instruction is to send data somewhere, default-deny egress and controlled sinks bound the exfiltration channel independently of whether the model complied. ## Defense four: provenance and inventory Keep an inventory of every tool server an agent may connect to, who owns it, what version is pinned, and when it was last reviewed — the agent equivalent of a dependency manifest. Prefer providers who publish signed releases and reproducible builds. For high-value workflows, self-host a fork of the server so that updates arrive only when you merge them. ## What none of this fixes A legitimate server can be compromised upstream, and a reviewed description can still be subtly misleading. Pinning and review shrink the window and force the attacker to compromise a specific release rather than push a silent update; they do not make the dependency trustworthy. That is why the containment layer matters more than the review layer: your safety should not rest on the assumption that every third-party description you loaded is honest. ## Answering this in an interview Name the mechanism first — descriptions are high-trust prompt text — then separate the two attack shapes, poisoning at adoption and rug-pull after it. Give the dependency-management answer (pin, hash, review the diff, no auto-update) and finish with containment, because an interviewer is checking whether you understand that review alone is a point-in-time control.

  • Why does one untrusted tool server contaminate a whole agent session?
    Because all tool definitions share one context. Text in one server's description can instruct the model about how to use another server's tools — including passing data from a sensitive tool into an innocuous-looking argument of the untrusted one. The isolation you need is at the session level: untrusted metadata and sensitive capability should not be loaded together.
  • Can you scan tool descriptions automatically instead of reviewing diffs by hand?
    Scan as a filter, not as a boundary. The payload is natural language and can be phrased indefinitely many ways, so classifiers give partial coverage and adaptive authors route around them. Use scanning to triage and to flag obvious cases, keep the human diff review as the gate that must pass, and make sure the containment layer holds if both miss.
  • What is the argument for self-hosting a fork of a third-party tool server?
    It converts a push relationship into a pull one. Updates arrive only when you merge them, so a rug-pull upstream cannot reach your agents without a human reviewing the diff, and you can hold a pinned build indefinitely while you evaluate. The cost is maintenance and drift from upstream fixes, so it is worth it for servers with sensitive access rather than for everything.

saying these in an interview costs you the question

  • Tracking latest on a third-party tool server in production
  • Assuming descriptions are documentation rather than instructions
  • Reviewing a server once at adoption and never re-checking
  • Letting newly discovered tools be invoked without human review
  • Believing a scanner on descriptions is a security boundary

context