skip to content

Instrumenting a Live Target

Attacking a real agent means logging every tool call, injecting at the result the agent trusts, and provisioning throwaway accounts first. Interviewers ask what the run left behind.

on this pageshow

explore

questions

18

When red-teaming an agent, why do teams stand up their own tool server for the agent to call instead of pointing it at the vendor's live tool server?

level: juniorimportance: must knowfreq 55%

answer

  1. server = fixture, not vendor stand-in
  2. mutate description and schema between runs
  3. live server = one shot, real side effects
  4. server-side call log beats text judging
  5. replica evidence is about the replica

basics

~20 s

Because a server you host is repeatable. You can change a tool's description or schema between runs and replay the same attempt to see what moved. A live vendor server gives one shot: it changes under you, its calls touch real data, and a result there cannot be reproduced or varied.

solid answer

~50 s

A tool server you host is a **fixture**: the tool name, the description string, the parameter schema, the returned result body and the error shapes are all inputs you version and mutate on purpose. Agent red-team hypotheses are usually of the form "does content arriving through a tool surface get acted on", and testing that needs the same attempt repeated many times, because model decoding is stochastic and a single success is not a result. A live third-party server gives you one attempt with none of that: you cannot edit its descriptions, its data drifts, its calls have real side effects, and it may sit outside what you are authorised to touch. The tradeoff is honest and it is the whole cost of the approach: a hit on your replica is evidence about your replica until you can show the replica matches the real server's contract.

go deeper

for a junior

Says the server you own can be changed and re-run while the live one cannot, and that calls to a live server have real consequences.

for a middle

Adds what specifically is mutable — description, schema, result body, error text — and that repeats are needed because model output is stochastic.

for a senior

Leads with the cost: the replica drifts from the real contract, so states findings at replica strength and describes how the replica is derived and refreshed.

for a principal

Frames it as an evidence policy for the team: which claims may be made from fixtures, what must be confirmed against production, and who authorises that confirmation.

**What a "tool server" actually is here.** An agent does not own its tools; it is handed a list of tool definitions on every request. Each definition carries three things the model reads as plain text: a name, a natural-language description, and a JSON Schema describing the parameters. When the model emits a tool call, the harness routes it by name to whatever answers — an MCP server, a REST service sitting behind the framework's function-calling loop, or a local process the harness registered at startup. That endpoint is the tool server. Everything the model learns about a tool before calling it (name, description, schema) and everything it learns after (the result body, the error string on a rejected call) crosses that one boundary. Owning the boundary is owning the experiment. **What owning it buys, knob by knob.** | Knob you control | What it is | |---|---| | Description text | The free-text string the model reads before deciding to call | | Parameter schema | Property names, types, `enum` members, the `required` list, `additionalProperties`, length caps | | Result body | The exact bytes returned to the model, including size, encoding and escaping | | Error surface | What the model sees on a rejected call — a bare code, or text that echoes the input | | Neighbourhood | How many tools are listed, in what order, and which sensitive one sits beside the benign one | The sixth thing, and the one people undervalue, is **the observation itself**. A sensitive handler on a server you own can write a server-side log row the moment it is invoked, with the arguments it received. That turns "did the agent do the thing" from a judge's reading of a transcript into a row in a table. Removing the grader from the loop is usually worth more than the mutability, because graded transcripts are where the false positives and false negatives of agent red-teaming live. **What it costs.** Building a stub faithful enough to be worth anything is typically one to three engineer-days per service, and it is a second implementation of somebody else's contract, so it rots on their release schedule, not yours. Hosting is noise. The recurring cost is model calls, and it is the repeats that spend it: one agent attempt is a multi-turn loop, commonly five to fifteen model calls of a few thousand tokens each, so a modest design of four surface variants at forty attempts per variant is 160 agent runs and on the order of one to two million tokens — dollars rather than cents on a hosted mid-tier model, and hours of wall clock unless attempts run in parallel. Budget the engineer-days and the token spend separately; teams routinely fund the second and forget the first, then hand-patch the stub until a run passes. **Where the number misleads.** The output of this setup is a success rate, and its denominator is your fixture, not the vendor. Three specific misreadings: 1. *Permissiveness inflation.* Your stub accepts arguments and returns content the real service would reject, truncate or escape. The rate is then measuring your own laxity. 2. *Instrument swap.* On the fixture you score with a server-side invocation log; against the real service you usually score by reading the agent's account of what it did. That is a different measuring instrument, so a fixture rate and a live rate are not comparable in either direction — a lower live number can mean a stricter server or simply a blinder detector. 3. *Neighbourhood shrinkage.* Fixtures typically expose three or four tools; production agents carry twenty or more, and the model's behaviour with a crowded, competing tool list is not the behaviour you measured. **What you would check before believing a result.** Run the unchanged baseline fixture twice and report both numbers — that spread is your noise floor, and any effect smaller than it is not an effect. Diff your served tool definitions against the vendor's published contract field by field, and keep a capture of real responses to compare lengths, encoding and escaping against what your stub returns. Confirm the server-side hit log fires for a deliberately triggered call, so you know the instrument works. Stamp a fixture identifier into every run record. Then state the claim at the strength the evidence carries: "this agent acts on content arriving in a tool result of this shape, at this rate, under these settings" is supportable; "this partner integration is exploitable" is not, until the contract match is demonstrated. AgentDojo, InjecAgent and AgentHarm are published suites built on exactly this pattern; each holds a different part of the surface fixed, so read which one before quoting its numbers.

  • Name two things a tool server you host can give you that a live one cannot.
    A description or schema you can change between runs, and a server-side record that a sensitive tool was called with specific arguments — an objective hit signal rather than a graded transcript.
  • If the replica is not the real system, what is a replica-only hit actually worth?
    It is a genuine finding about the agent's behaviour given a tool surface of that shape, and a hypothesis about the real integration. It becomes a claim about the partner only after you show the replica matches their contract.
  • How many repeats of the same attempt do you need before calling it a result?
    Enough that a stochastic success is distinguishable from a reliable one — report attempts and successes, not a single anecdote, and keep the sampling settings fixed across the comparison.

saying these in an interview costs you the question

  • Treats a hit on a self-built replica as proof the partner integration is exploitable.
  • Points the harness at a production third-party server for iteration because it is more realistic.
  • Calls a single successful attempt a finding, with no repeats.
  • Cannot say what parts of the server surface are actually variables.
  • Never records which version of the fixture produced a given result.

context

open as a page

Before red-teaming a deployed assistant that can send email, file tickets and write CRM rows, what do you provision so that a successful attack does not touch real people or real records?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Provision throwaway identities the agent acts as: a scratch mailbox, a test tenant or sandbox project, and disposable CRM records. Point outbound tools at dry-run or test endpoints. Mark everything with a unique per-run tag so you can find it later, and agree who owns cleanup.

open as a page

In an agent red-team harness, why is the adversarial instruction placed in the result a tool returns rather than in the user's message, and what does a hit at each of those two boundaries actually prove?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Put it in the tool result, not the user message. The agent is already trained to be wary of user requests, but treats tool output as data it fetched. The harness intercepts the call and returns attacker-controlled content in its place. A hit there shows the agent obeys untrusted retrieved content.

open as a page

You are running an agent red-team harness against a tool-using agent. Beyond the assistant's text messages, what must the harness write to the run trace for another engineer to rerun your finding without you?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Capture every tool call the agent made: the tool name, the exact arguments, and the result returned to it, in order, with timestamps. Also record the prompts, the model settings, and which harness and environment version ran. Text-only transcripts leave a finding unreproducible.

open as a page

In an agent red-team harness you replace the agent's real email-sending tool with a stub that records the call and returns success. An injection attempt now shows the agent calling that stub with attacker-chosen recipients and body. What does that result prove, and what does it not?

level: middleimportance: must knowfreq 55%

basics

~20 s

It proves the model was hijacked into deciding to send, and shows the exact arguments it chose. It does not prove delivery: the real endpoint might reject the recipient, demand a confirmation, or fail an authorisation check. The stub tests the model's decision, not the system's outcome.

open as a page

You need adversarial content to arrive inside a tool result for an agent you are testing. Compare wrapping the tool function inside the agent framework, proxying the call in front of the tool endpoint, and planting the content in data a real tool reads -- what does each let you observe and control?

level: middleimportance: must knowfreq 55%

basics

~20 s

Three placements. Wrapping the tool function inside the framework is easiest and sees parsed arguments, but skips the wire. A proxy in front of the endpoint sees real requests and responses but not framework post-processing. Planting content in data a genuine tool reads is highest fidelity and lowest control over timing.

open as a page

Your agent red-team harness replays a recorded attempt with the same prompt, the same sampling settings and the same seed, and the agent takes a different action. Why is recording the seed not enough to make an agent run replayable, and what does the trace need instead?

level: middleimportance: must knowfreq 48%

basics

~20 s

Because the run is not a pure function of the seed. The environment answers differently each time: tool results, clocks, IDs, retries, and the endpoint's own sampling all vary. Record the actual request and response bytes at every step so a replay can fall back to the recorded exchange.

open as a page

Your stand-in tool server reproduces a hit in the harness every time, but the same attempt against the partner's real tool server produces nothing. Which differences between your stand-in and the real server would you check first?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Check where your copy is more permissive than the real service: looser schema validation, unsanitised and untruncated result bodies, stale description text, different field names, and errors that echo input back. Then diff your served tool definitions against the partner's current contract and a capture of real responses.

open as a page

An agent red-team run that substitutes tool results through a shim reports a large number of policy violations. Before you write any of them up, how do you establish that the shim did not manufacture them?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Re-run with the shim in place but a benign result substituted. If the behaviour still appears, the shim caused it, not the payload. Usual manufacturers: the shim answering a call that would have errored, returning instantly, changing call order, or replacing a result the framework would have truncated.

open as a page

You host the tool server your agent under test calls. Between two runs you edit both the tool's description text and its parameter schema, and the attempt starts succeeding. What is wrong with that experiment, and how would you rerun it?

level: middleimportance: should knowfreq 42%

basics

~20 s

Two variables moved, so you cannot say which caused the change — and with a stochastic model, neither may have. Rerun with one edit at a time from a pinned baseline, repeat each configuration enough times to compare success rates, and record which fixture version produced each run.

open as a page

A red-team run against a mailbox-and-documents assistant uses a freshly created empty test mailbox and a brand-new account with no history or shared drives. Why can this understate real risk, and what do you put in that account before running?

level: middleimportance: should knowfreq 42%

basics

~20 s

An empty account gives an exfiltration attempt nothing to steal and no permissions to abuse, so the run scores a false negative. Seed it with realistic decoy documents, contacts and message history, grant the same roles a real user holds, and mark every seeded item with a canary string you can search for.

open as a page

After a live-fire agent red-team run in which the agent really sent mail and created records, how do you establish exactly what the test left behind and remove it — and what residue can you not remove?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Do not reconstruct it afterwards. Have the harness log every outbound tool call with the identifier the system returned, so the run produces a delete list. Tag artefacts with a per-run marker, delete, then re-search until the marker returns nothing. Delivered mail, webhook fan-out and audit entries stay.

open as a page

Your harness substitutes adversarial content into an agent's tool results and the agent never misbehaves. Before reporting the agent as resistant, how do you verify the content actually reached the model?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Log what the model actually received, not what you handed the shim. Frameworks truncate, summarise, re-serialise or drop tool results before they become context, so your text may never arrive intact. Without that check a clean run is not evidence of resistance; it may only prove your content was cut.

open as a page

During an agent red-team run the agent's tools pulled real customer records and a live API token into the harness trace. How do you design trace capture so the run stays shareable without destroying the reproducibility of the finding?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Split raw from shareable. Keep the unredacted trace in a restricted store with a short retention, and derive a redacted copy by replacing secrets and personal data with stable placeholders. Stable means the same value maps to the same token, so the reader still sees the causal chain.

open as a page

The only tool server in scope for an agent engagement is a third-party production one you may call but cannot mutate or replay. How do you split the work between a stand-in you host and that limited live access, and what do you refuse to claim from stand-in evidence alone?

level: principalimportance: should knowfreq 26%

basics

~20 s

Do all iteration on the stand-in, where the surface is mutable and repeats are cheap, and spend live calls only on confirming a shortlist. Decide the confirmation criteria before touching production. Never claim the partner is exploitable, or quote a rate, from stand-in runs alone.

open as a page

You are planning a red-team engagement against an agent wired into production systems. How do you decide, tool by tool, which actions execute for real and which are stubbed, knowing stubs weaken your evidence and real execution leaves residue?

level: principalimportance: should knowfreq 34%

basics

~20 s

Split by reversibility and by who else sees the effect. Reads and reversible writes inside a scratch tenant run for real. Anything reaching a third party, moving money, or paging a human gets stubbed, and you argue that gap separately. Decide before the run, write the table down, publish it with the findings.

open as a page

You are leading a red-team engagement against a team's live agent and you must intercept its tool calls to test the boundary the agent trusts. How do you decide how invasive that interception may be, and how do you keep the findings from being dismissed as 'your harness did that'?

level: principalimportance: should knowfreq 34%

basics

~20 s

Agree up front how much of the stack you may modify, and write it down. Every layer you replace is a layer the owning team can blame. Prefer the least invasive interception that still proves the claim, ship a matched control run with each finding, and state exactly what the harness changed.

open as a page

A long agent red-team campaign is producing more trace data than your storage and retention budget allows. How do you decide what the harness stops capturing, and which reductions are safe?

level: principalimportance: should knowfreq 30%

basics

~20 s

Capture fully by default and cut elsewhere. You can always delete a trace later; you can never reconstruct one you did not record, and a failed rerun is what kills a finding. Trim by truncating giant tool payloads to a hash plus a head, and by shortening retention, not by sampling which attempts you trace.

open as a page