skip to content

You are running an agent red-team harness against a tool-using agent. Beyond the assistant's text messages, what must the harness write to the run trace for another engineer to rerun your finding without you?

level: juniorimportance: must knowfreq 62%

answer

  1. tool call is the evidence, not the narration
  2. name, arguments, result, order
  3. run-level: seed, settings, harness revision
  4. decide the schema before the run
  5. capture full, derive a shareable copy

basics

~20 s

Capture every tool call the agent made: the tool name, the exact arguments, and the result returned to it, in order, with timestamps. Also record the prompts, the model settings, and which harness and environment version ran. Text-only transcripts leave a finding unreproducible.

solid answer

~50 s

The unit of evidence in an agent engagement is the **tool call**, not the chat turn. A trace that another engineer can act on holds, per step: the full message list sent to the model, the tool name and the **verbatim arguments** the agent chose, the **result payload** handed back, and the ordering and wall-clock timing of both. Around the steps it holds the run-level context: system prompt, sampling settings and any seed, the identity of the endpoint under test, the harness and environment revision, the attack input you injected and where you injected it, and the outcome the harness recorded for the attempt. The tradeoff is that this is exactly the material you cannot ship as-is: tool results are where real records, tokens and personal data enter the trace, so a full capture creates a redaction and retention obligation. The wrong reaction is to capture less. Capture fully, then derive a shareable copy.

go deeper

for a junior

Names the essentials: tool name, arguments, results, ordering, and the prompt that went in.

for a middle

Adds run-level fields — seed and sampling settings, harness and environment revision, injection point — and explains why the tool call is the evidentiary unit.

for a senior

Designs the trace schema before the run, asserts on it in a smoke attempt, and separates the raw store from the shareable derivative.

for a principal

Sets one trace schema across the team so findings from different operators and different harnesses are comparable and rerunnable months later.

## What the trace has to prove A finding in an agent engagement is a **causal claim**: this input, arriving through this channel, caused the agent to call that tool with those arguments, and the tool acted. A reviewer who was not in the room has to be able to check every clause of that sentence. A chat transcript checks none of them. The assistant's text is *narration* — tokens the model generated about what it intends or claims to have done — and narration diverges from action in both directions: - An agent can announce a transfer it never requested, because the tool call was never emitted or the harness refused it; - it can also emit a call it never mentions in prose. When the only artefact is text, the reviewer is being asked to trust the operator, and a competent defence team will decline. ## The evidentiary unit, field by field Per step the harness writes four things. - (1) The full message list *as sent to the model*, including the system prompt and the tool schemas in force for that call — behaviours frequently turn on how a tool was described, so the description belongs in the record, not just the tool's name. - (2) The raw model response: the tool-call structure as returned and the finish reason, not only the object your agent framework parsed out of it. - (3) The tool name and the **verbatim arguments** in their original types — `{"amount": 1000}` and `{"amount": "1000"}` are different findings. - (4) The **tool result payload** handed back to the agent, with ordering, start time, latency, and any error or retry. Around the steps sit **run-level fields**: the endpoint identifier and model version under test, sampling settings and seed, the harness revision, the environment or tool-server revision, the injection point (which channel carried the attack input), and the recorded outcome together with the name of the scorer or detector that decided it. ## What it costs - **Tool results dominate the bytes.** A single directory listing, document dump or search result set is routinely larger than the whole conversation around it, and it repeats near-identically across attempts, so a few hundred attempts reaches gigabytes without anything unusual happening. - Storing the full message list at every step is **quadratic in loop length**, because each step re-stores the growing prefix: a 25-step attempt writes its opening turns 25 times. - The other cost is not storage at all — tool results carry live records, so full capture creates a **redaction and retention obligation** from the first attempt. - And the schema itself is roughly a day of engineering up front versus re-running the campaign if you get it wrong, because nothing regenerates a field you did not write. ## Where it fails and how the record misleads The common failure is **capture bolted on after the fact**: the harness "has tracing" because the agent framework prints something. Then: - arguments arrive already stringified, so types and exact values are gone; - results arrive truncated at a console width, so what the agent actually read is unknown; - and parallel tool calls are flattened into a single order the run never had, which manufactures a causal story out of scheduling noise. All three are unrecoverable later. The second failure is subtler. A trace that stores the scorer's verdict but not the payloads *reads* like evidence: a reviewer sees forty hits and treats them as forty demonstrations, when they are forty copies of one bit produced by a detector nobody has audited. Rendered transcript viewers make this worse — a pretty-printed HTML view of a run can look complete while the stored record behind it is thin, and the view is what gets shown in the read-out. ## What I would check 1. Before the campaign, run one deliberate benign attempt that provokes a known tool call, then diff the written record against the intended schema field by field: arguments present with types intact, result payload untruncated, parallel calls still distinguishable, run-level revisions populated. 2. During the campaign, pick a recorded attempt at random and hand the trace alone to a colleague with the operator silent. If they need to ask a single question to walk it or rerun it, the schema is short a field — and that is far cheaper to learn at attempt five than at attempt five hundred.

  • Why is the tool result, not just the tool call, part of the evidence?
    The result is what the agent read next and conditioned on. Without it you cannot show why the agent escalated, and a rerun against a changed environment silently diverges with no way to tell.
  • The agent issued three tool calls in one step. What must the trace preserve?
    That they were issued together and in what order results came back. Flattening parallel calls into a sequence invents an ordering the run did not have and can make a rerun look like a different behaviour.

saying these in an interview costs you the question

  • Treating the chat transcript as the finding; an agent's claim that it called a tool is not proof it did.
  • Logging tool arguments only as a stringified summary, so the exact values cannot be replayed.
  • Omitting tool results because they are large, which removes the very payload the agent acted on.
  • Capturing less up front to avoid redaction work later.
  • No record of which harness or environment revision produced the run.

context