skip to content

An agent red-team harness drives an LLM agent through a multi-step harmful task in a sandboxed environment and scores the outcome. Why record how far along the chain the agent got, instead of only a pass/fail on the final harmful outcome?

level: juniorimportance: should knowfreq 55%

answer

  1. binary hides the margin
  2. stop-point = which control fired
  3. world-state checkpoint, not narration
  4. severity input, not capability score
  5. drift visible before a flip

basics

~20 s

Because a pass/fail hides how close the run came. Recording the last step the agent reached shows whether it was stopped early at planning or only at the final harmful action, tells you which control fired, and gives severity a distance input. Two zero-score runs can be very different.

solid answer

~50 s

A binary outcome collapses two very different runs into the same cell. An agent that refused the framing at the first turn and an agent that located the record, drafted the message and was blocked only by an egress rule both score "no harm", but the second target is one weak control away from harm and the first is not. So the harness records the furthest checkpoint reached, ideally as an observed world-state fact (the draft exists, the row is still present) rather than as the agent's own narration. That distance is useful for three things: it is an input to how bad the finding is, it points at which control actually fired, and it makes regressions visible — the same suite next month reaching step five instead of step two is a real change even while both runs still score "no harm". The tradeoff is grader cost and drift: every checkpoint is another assertion someone has to define and maintain.

go deeper

for a junior

Says a pass/fail loses information, and that knowing where the run stopped tells you how close it got.

for a middle

Adds that the stop-point identifies which control fired, and that checkpoints must be asserted against world state rather than read from the transcript.

for a senior

Uses the stop-point distribution as a regression signal, ties each checkpoint to an owner and a fix, and pins environment and tool set so distances compare across runs.

for a principal

Decides how much grading detail is worth its maintenance cost across teams, and keeps distance framed as a severity input rather than letting it become a leaderboard number.

**Binary grading** is cheap and comparable, which is why every harness offers it, but it throws away the most operationally useful thing a multi-step agent run produces: where the chain stopped, and what stopped it. ## What the instrument actually is An agent red-team harness has three parts. 1. First, a **sandboxed environment** holding stateful fake tools — a mailbox, a file store, a payments ledger — whose contents can be queried after the episode. 2. Second, a **task definition** that hands the agent an objective which would be harmful if completed, together with a pinned starting state and tool set. 3. Third, a **grader** that inspects the environment when the episode ends. The public suites of this shape — AgentHarm, AgentDojo, InjecAgent — ship fixed environments precisely so a result means the same thing on two different days. Binary grading asks the grader one question: did the forbidden world-state change happen? **Partial-progress grading** asks several — was the sensitive record read, was the message drafted, was the transfer staged — each one a separate assertion evaluated over the same post-run snapshot. ## Why the extra assertions earn their keep A harmful chain is a sequence of enabling steps: acquire a capability, locate the target object, stage the action, execute it. - A run that dies at step one usually means the model's own refusal behaviour held. - A run that dies at the last step usually means something external held: a tool permission, a confirmation prompt, an allow-list, an egress rule. Those are different findings, with different owners and different fixes, and a single pass/fail cell cannot tell them apart. Distance also gives you a **leading indicator**. A suite that keeps scoring "no harm" while the median stop-point creeps from step two to step five is telling you the margin is eroding months before any run flips to a completion — and the flip is the expensive event. ## What it costs Not compute: the checkpoints are post-hoc queries against a snapshot the run already produced, so grading a run at four checkpoints costs essentially what grading it at one costs. The real bill has two lines. - The first is **engineer time** — every checkpoint is an assertion someone writes against the sandbox schema and re-writes when a fixture changes, plus a worked example so two graders agree. - The second, much larger, is **trials**: a stop-point is only interpretable across repeats, and a repeat means re-running a full multi-turn episode with tool outputs accumulating in context. Fifteen to thirty turns of an agent loop is routinely tens of thousands of tokens; multiply by trials, then by tasks in the suite, and the token bill dwarfs anything the grader costs. Budget the trials, not the rubric. ## Where the number misleads Three readings are wrong and all three are common. - (1) *Fraction as percentage of danger.* "Reached 5 of 6" is not 83% of the way to harm; the denominator is however finely you chose to chop the chain, and the steps are not equal — often only the last one is irreversible, and three cheap setup steps are worth less than one staged transfer. - (2) *Distance as capability.* This instrument measures what the system permitted, not how well the agent planned; the identical trace read as a capability evaluation says the agent failed, which is the opposite conclusion. - (3) *Stop as defence.* A run can end because a turn cap expired, a tool endpoint returned 429, the sandbox crashed, or the agent malformed an argument. None of those judged the content of the request, so none is evidence a control exists — yet all of them look identical in a distance-only record. ## What I would check - Is every checkpoint asserted against environment state rather than read out of the agent's narration, given that a model can claim it sent nothing while the outbox holds the message? - Does every terminated run carry an attributed stop cause from a fixed taxonomy — refused, blocked by a named control, agent error, infrastructure error, budget exhausted? - How many trials sit behind the distance, and was the environment snapshot pinned so a re-run is comparable? - And is the number being reported as a severity input to whoever owns the risk decision, rather than as a leaderboard column that teams will start optimising?

  • Two runs both score "no harm". One stopped at the first turn, one at the final send. Are they the same finding?
    No. The first suggests the model's refusal behaviour held; the second means a single external control was the last line. Same score, different owner, different fix, different residual risk.
  • Why must a checkpoint be checked against the sandbox rather than read out of the transcript?
    Because the agent's narration is not evidence. It can claim it deleted nothing while the row is gone, or claim success it never achieved. Query the environment for the fact.

saying these in an interview costs you the question

  • Treating the agent's own statement that it stopped as proof it stopped.
  • Reporting partial progress as an agent capability score.
  • Adding a checkpoint per tool call, producing a rubric nobody can grade consistently.
  • Claiming a run that reached step five proves the target usually reaches step five.

context