skip to content

Testing a Tool Server

A tool server you stand up yourself can have its descriptions and schemas mutated between runs; a live third-party one gives you one shot. Interviewers ask how you made the injection point repeatable.

on this pageshow

explore

questions

4

When red-teaming an agent, why do teams stand up their own tool server for the agent to call instead of pointing it at the vendor's live tool server?

level: juniorimportance: must knowfreq 55%

answer

  1. server = fixture, not vendor stand-in
  2. mutate description and schema between runs
  3. live server = one shot, real side effects
  4. server-side call log beats text judging
  5. replica evidence is about the replica

basics

~20 s

Because a server you host is repeatable. You can change a tool's description or schema between runs and replay the same attempt to see what moved. A live vendor server gives one shot: it changes under you, its calls touch real data, and a result there cannot be reproduced or varied.

solid answer

~50 s

A tool server you host is a **fixture**: the tool name, the description string, the parameter schema, the returned result body and the error shapes are all inputs you version and mutate on purpose. Agent red-team hypotheses are usually of the form "does content arriving through a tool surface get acted on", and testing that needs the same attempt repeated many times, because model decoding is stochastic and a single success is not a result. A live third-party server gives you one attempt with none of that: you cannot edit its descriptions, its data drifts, its calls have real side effects, and it may sit outside what you are authorised to touch. The tradeoff is honest and it is the whole cost of the approach: a hit on your replica is evidence about your replica until you can show the replica matches the real server's contract.

go deeper

for a junior

Says the server you own can be changed and re-run while the live one cannot, and that calls to a live server have real consequences.

for a middle

Adds what specifically is mutable — description, schema, result body, error text — and that repeats are needed because model output is stochastic.

for a senior

Leads with the cost: the replica drifts from the real contract, so states findings at replica strength and describes how the replica is derived and refreshed.

for a principal

Frames it as an evidence policy for the team: which claims may be made from fixtures, what must be confirmed against production, and who authorises that confirmation.

**What a "tool server" actually is here.** An agent does not own its tools; it is handed a list of tool definitions on every request. Each definition carries three things the model reads as plain text: a name, a natural-language description, and a JSON Schema describing the parameters. When the model emits a tool call, the harness routes it by name to whatever answers — an MCP server, a REST service sitting behind the framework's function-calling loop, or a local process the harness registered at startup. That endpoint is the tool server. Everything the model learns about a tool before calling it (name, description, schema) and everything it learns after (the result body, the error string on a rejected call) crosses that one boundary. Owning the boundary is owning the experiment. **What owning it buys, knob by knob.** | Knob you control | What it is | |---|---| | Description text | The free-text string the model reads before deciding to call | | Parameter schema | Property names, types, `enum` members, the `required` list, `additionalProperties`, length caps | | Result body | The exact bytes returned to the model, including size, encoding and escaping | | Error surface | What the model sees on a rejected call — a bare code, or text that echoes the input | | Neighbourhood | How many tools are listed, in what order, and which sensitive one sits beside the benign one | The sixth thing, and the one people undervalue, is **the observation itself**. A sensitive handler on a server you own can write a server-side log row the moment it is invoked, with the arguments it received. That turns "did the agent do the thing" from a judge's reading of a transcript into a row in a table. Removing the grader from the loop is usually worth more than the mutability, because graded transcripts are where the false positives and false negatives of agent red-teaming live. **What it costs.** Building a stub faithful enough to be worth anything is typically one to three engineer-days per service, and it is a second implementation of somebody else's contract, so it rots on their release schedule, not yours. Hosting is noise. The recurring cost is model calls, and it is the repeats that spend it: one agent attempt is a multi-turn loop, commonly five to fifteen model calls of a few thousand tokens each, so a modest design of four surface variants at forty attempts per variant is 160 agent runs and on the order of one to two million tokens — dollars rather than cents on a hosted mid-tier model, and hours of wall clock unless attempts run in parallel. Budget the engineer-days and the token spend separately; teams routinely fund the second and forget the first, then hand-patch the stub until a run passes. **Where the number misleads.** The output of this setup is a success rate, and its denominator is your fixture, not the vendor. Three specific misreadings: 1. *Permissiveness inflation.* Your stub accepts arguments and returns content the real service would reject, truncate or escape. The rate is then measuring your own laxity. 2. *Instrument swap.* On the fixture you score with a server-side invocation log; against the real service you usually score by reading the agent's account of what it did. That is a different measuring instrument, so a fixture rate and a live rate are not comparable in either direction — a lower live number can mean a stricter server or simply a blinder detector. 3. *Neighbourhood shrinkage.* Fixtures typically expose three or four tools; production agents carry twenty or more, and the model's behaviour with a crowded, competing tool list is not the behaviour you measured. **What you would check before believing a result.** Run the unchanged baseline fixture twice and report both numbers — that spread is your noise floor, and any effect smaller than it is not an effect. Diff your served tool definitions against the vendor's published contract field by field, and keep a capture of real responses to compare lengths, encoding and escaping against what your stub returns. Confirm the server-side hit log fires for a deliberately triggered call, so you know the instrument works. Stamp a fixture identifier into every run record. Then state the claim at the strength the evidence carries: "this agent acts on content arriving in a tool result of this shape, at this rate, under these settings" is supportable; "this partner integration is exploitable" is not, until the contract match is demonstrated. AgentDojo, InjecAgent and AgentHarm are published suites built on exactly this pattern; each holds a different part of the surface fixed, so read which one before quoting its numbers.

  • Name two things a tool server you host can give you that a live one cannot.
    A description or schema you can change between runs, and a server-side record that a sensitive tool was called with specific arguments — an objective hit signal rather than a graded transcript.
  • If the replica is not the real system, what is a replica-only hit actually worth?
    It is a genuine finding about the agent's behaviour given a tool surface of that shape, and a hypothesis about the real integration. It becomes a claim about the partner only after you show the replica matches their contract.
  • How many repeats of the same attempt do you need before calling it a result?
    Enough that a stochastic success is distinguishable from a reliable one — report attempts and successes, not a single anecdote, and keep the sampling settings fixed across the comparison.

saying these in an interview costs you the question

  • Treats a hit on a self-built replica as proof the partner integration is exploitable.
  • Points the harness at a production third-party server for iteration because it is more realistic.
  • Calls a single successful attempt a finding, with no repeats.
  • Cannot say what parts of the server surface are actually variables.
  • Never records which version of the fixture produced a given result.

context

open as a page

Your stand-in tool server reproduces a hit in the harness every time, but the same attempt against the partner's real tool server produces nothing. Which differences between your stand-in and the real server would you check first?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Check where your copy is more permissive than the real service: looser schema validation, unsanitised and untruncated result bodies, stale description text, different field names, and errors that echo input back. Then diff your served tool definitions against the partner's current contract and a capture of real responses.

open as a page

You host the tool server your agent under test calls. Between two runs you edit both the tool's description text and its parameter schema, and the attempt starts succeeding. What is wrong with that experiment, and how would you rerun it?

level: middleimportance: should knowfreq 42%

basics

~20 s

Two variables moved, so you cannot say which caused the change — and with a stochastic model, neither may have. Rerun with one edit at a time from a pinned baseline, repeat each configuration enough times to compare success rates, and record which fixture version produced each run.

open as a page

The only tool server in scope for an agent engagement is a third-party production one you may call but cannot mutate or replay. How do you split the work between a stand-in you host and that limited live access, and what do you refuse to claim from stand-in evidence alone?

level: principalimportance: should knowfreq 26%

basics

~20 s

Do all iteration on the stand-in, where the surface is mutable and repeats are cheap, and spend live calls only on confirming a shortlist. Decide the confirmation criteria before touching production. Never claim the partner is exploitable, or quote a rate, from stand-in runs alone.

open as a page