skip to content

Deciding What Counted

Success here is a side effect that must not have happened, checked against world state, on a target that answers differently each run. Interviewers ask how many trials back your rate.

on this pageshow

explore

questions

13

In an agent red-team harness, why is the agent's own transcript claim that it sent nothing not enough to score the run as a pass?

level: juniorimportance: must knowfreq 72%

answer

  1. transcript is a claim, state is evidence
  2. snapshot before, diff after
  3. silent success vs phantom failure
  4. egress recorder counts as state
  5. verdict function must not read the final message

basics

~20 s

Because the transcript is the agent's own account, not evidence. Score against world state: snapshot the sandbox before the run, run it, then diff. If the outbound queue gained a message or a row is gone, the run failed no matter what the agent narrated in its final turn.

solid answer

~50 s

The transcript is produced by the system under test, and in a red-team run that system is exactly the component an injected instruction may have steered. Its self-report is a claim, not an observation. A security harness owns the environment, so make the environment the oracle: - fixture the task's state (rows, files, mailbox, tool-call log) before turn one; - let the agent run; - diff every sink the sandbox exposes and evaluate the assertion over that diff. The two failure shapes this catches are opposite. The **silent success**: the agent calls the send tool, then reports "I did not send anything" — a transcript-scored harness records a false pass. The **phantom failure**: the agent narrates a deletion that never happened — a transcript-scored harness records a false alarm and you spend a day triaging nothing. The transcript still matters for triage and for harms that leave no state change, but it is never the thing that decides the verdict.

go deeper

for a junior

Says the transcript is what the agent claims and the environment is what happened, and that the harness should check the database or mailbox after the run.

for a middle

Describes the snapshot-run-diff loop, names the sinks worth diffing, and gives both the silent-success and phantom-failure shapes.

for a senior

Points out that a security false pass is far more expensive than a false alarm, and that harms with no state change need a boundary recorder rather than a state diff.

for a principal

Frames it as an evidence-independence rule for the whole suite — no verdict may be computed from an artefact the system under test authored — and applies it to judge inputs and report fields too.

**An agent run leaves two records, and only one of them is independent of the thing under test.** The *transcript* is the ordered log of the agent loop: the system prompt, the task input, each tool call the model emitted, each tool result the harness returned, and the final assistant message that summarises what happened. The *environment* is the sandbox the agent acted in — a seeded database, files under the task root, a fixture mailbox or outbound queue, a recorded HTTP egress boundary. A red-team harness owns the environment. The model under test authors the generated parts of the transcript. That asymmetry is the whole answer. **Three tiers of evidence, routinely conflated.** | artefact | produced by | what it actually proves | |---|---|---| | final assistant message: "I did not send anything" | the model | that the model wrote that sentence | | harness tool-call log: a recorded `send_email(...)` invocation | the harness | that the call was issued with those arguments | | pre/post diff of the fixture mailbox | the environment | that a message landed in the sink | The final message sits at the bottom because it is a *generation*, not a log. Models summarise optimistically, drop tool calls out of a recap, and sometimes narrate work they never performed — no adversary required. Add an injected instruction in the data the agent reads and the incentive becomes explicit: the useful thing for an injection to produce is an action plus a reassuring summary. In a capability eval a wrong self-report is a quality bug. In a security eval, where the entire question is whether a forbidden effect occurred, it is a false pass on the exact property you were paid to check. The middle tier is genuine harness-produced evidence but still not conclusive: a recorded call may have errored, been rejected by the sandbox, or hit a stub. Issuing a call is not landing an effect. **The mechanism.** Make the environment the oracle: ```text before = snapshot(env) # rows, files, queue, egress log run_agent(task) after = snapshot(env) violated = assertion(diff(before, after)) # not: assertion(final_message) ``` Published agent-security suites are built this way rather than on text matching — AgentDojo's injection tasks expose a security check that is handed the pre-run and post-run environment objects, so the verdict is computed from state the harness controls rather than from the model's answer. Sinks worth snapshotting: the task's rows at field level, paths under the task root, the fixture mailbox, and the request list captured by the egress recorder. For a third-party API you cannot snapshot, the recorded request list *is* the state you assert over. **What it costs.** Transcript scoring is nearly free — one regex over one string — which is exactly why teams drift into it. State scoring costs a per-run fixture reset (a template-database restore or container re-create, typically hundreds of milliseconds to a few seconds, and it must happen between *every* trial or task N+1 inherits task N's rows), an egress proxy to stand up and keep working, storage for two snapshots per run, and the engineering time to write one predicate per task, which is usually the larger half of authoring a task. Set that against the inference bill: a twenty-turn agent run re-sends the growing tool-result history every turn, so a single task run consumes far more prompt tokens than a one-shot probe, and a 100-task suite at five trials each is 500 such runs per sweep. The snapshot machinery is a rounding error beside that. Do not let "parsing the text is cheaper" survive contact with the arithmetic. **Where the number misleads.** A transcript-scored suite reports an attack-success rate whose denominator is fine and whose numerator is a statement about the agent's *candour*. It is wrong in both directions and the single number cannot tell you which. The **silent success** — the send tool was called, the summary denies it — is recorded as a pass and pushes the reported rate down. The **phantom failure** — the model narrates a deletion that never occurred — is recorded as a violation and pushes it up, burning triage days on nothing. The second is annoying; the first is the one that ships. Worse, matching on the final message usually measures refusal *language* rather than harm: a run that ends "I cannot help with that" scores clean even when the tool call went out earlier in the loop. A 0% rate from that pipeline is not evidence of a safe agent; it is evidence that the agent's last paragraph was polite. **What to check.** Read the verdict function and ask which artefact it consumes; if the final assistant message appears anywhere in it, every number in the report is provisional. Then run a positive control — give the agent a plain, non-adversarial instruction to perform the forbidden effect and confirm the assertion fires. An assertion that has never fired has never been tested. Finally, check that the tool-call log and the state diff agree on the same run: a recorded call with no matching state change usually means a stub, an error swallowed by the tool wrapper, or a sink nobody is snapshotting. The transcript keeps its place — for triage, and for harms that leave no state change at all — but it never decides the verdict.

  • The environment diff is empty but the transcript describes a deletion. How do you score it?
    No violation of that assertion — nothing happened. Log the discrepancy separately: an agent that narrates actions it did not take is a real finding for other reasons, but it is not the side effect you forbade.
  • The agent read an API key from a config file and printed it in its reply. No row changed. Does state-diff scoring catch it?
    Not by itself. Exfiltration with no local write needs an assertion over what crossed the boundary — the reply body and the egress recorder — which is still observed data rather than the agent's self-report.
  • Why does this matter more for a security run than for a capability run?
    A capability run measures whether the agent did the task; the agent's honesty errors mostly show up as noise. A security run measures whether a forbidden effect occurred, and the agent's own account of that effect is the least trustworthy source you could pick.

Scoring on the transcript is asking the suspect to write his own alibi; scoring on the state diff is checking whether the money is still in the drawer. Both are worth reading, but only one of them settles the question.

saying these in an interview costs you the question

  • Treating the agent's final summary as the scoring input because it is easy to parse.
  • Assuming a false pass and a false alarm are symmetric costs in a security suite.
  • Believing a state diff makes the transcript useless — it is still needed for triage and for no-state-change harms.
  • Scoring on tool-call intent alone without checking whether the call actually landed.

context

open as a page

Your agent red-team harness ran each attack payload once against the target agent and printed the payloads that succeeded. Why is that list not an attack success rate, and what do you run instead before quoting a number?

level: middleimportance: must knowfreq 68%

basics

~20 s

One trial per payload gives a yes/no from a sampled, stateful system, not a rate. The same payload can land one run in five, so a single pass misses real attacks and promotes flukes. Re-run each payload a fixed number of trials and report hits over trials attempted.

open as a page

An agent red-team harness grades a run as "partial": the harmful chain reached step three of six before the agent was stopped. What does that grade license you to claim about the target system, and what does it not?

level: middleimportance: must knowfreq 60%

basics

~20 s

It licenses a claim about that run only: the chain reached step three under this prompt, seed and environment state. It does not show the target is safe at step four; a retry, a different phrasing or a fuller mailbox may carry it further. Partial is evidence of reach, not of a boundary.

open as a page

What does it mean to write the must-not-happen assertion for an agent red-team task before the run, and what goes wrong if you write it after reading the transcripts?

level: middleimportance: must knowfreq 62%

basics

~20 s

It means naming the forbidden side effect as a concrete, machine-checkable condition over environment state, and committing to it before you run anything. Written afterwards, the behaviour you observed sets the bar: you quietly accept whatever happened, and a near-miss that emptied one row becomes a pass.

open as a page

An agent red-team harness run ends mid-chain because the tool endpoint the agent was calling returned rate-limit errors, not because the agent refused or a policy control fired. How do you score that run, and what do you change in the harness?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Score it as an invalid run, not a defence. Rate limiting is infrastructure back-pressure from the endpoint, not a model refusal or a policy control; it would not stop a patient attacker. Mark the run inconclusive, exclude it from the denominator, back off and retry, and log the cause.

open as a page

An agent red-team harness drives an LLM agent through a multi-step harmful task in a sandboxed environment and scores the outcome. Why record how far along the chain the agent got, instead of only a pass/fail on the final harmful outcome?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Because a pass/fail hides how close the run came. Recording the last step the agent reached shows whether it was stopped early at planning or only at the final harmful action, tells you which control fired, and gives severity a distance input. Two zero-score runs can be very different.

open as a page

In an agent red-team finding you attach both a rate — successes over trials for one attack payload — and a saved recording of one successful run that a reader can replay. What does the replay show that the rate cannot, and what does the rate show that the replay cannot?

level: middleimportance: should knowfreq 48%

basics

~20 s

The replay shows the mechanism: the exact turns, tool calls and world change that made it a real break, so a reader can watch it happen rather than trust a number. The rate shows how often it happens, which sets priority. Neither substitutes for the other; a finding needs both.

open as a page

Two engineers drive the same injection payload at the same agent with the same red-team harness and report attack success rates of 5% and 40%. Before concluding the target changed, which run-level settings do you line up between the two sweeps?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare trial counts first — 5% may be 1 of 20. Then decoding settings and any fixed seed, the turn budget per episode, which object decided a hit and on what threshold, the starting environment each trial began from, and whether aborted or retried episodes were counted as trials.

open as a page

You have a fixed token spend for an agent red-team harness sweep against a metered hosted agent and about 40 candidate attack payloads, and each trial is a whole multi-turn episode rather than one call. How do you split that spend between covering more payloads and repeating each payload enough times to quote a rate?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Two stages. Spend a cheap screening pass of a few trials on all 40 payloads to find which ones ever land, then spend the depth only on those, repeating them enough to quote a rate. Stop early when the interval already answers the decision, and price a trial as a full episode.

open as a page

Your agent red-team harness scores every task by diffing the sandbox's state before and after the run. Which harms will that diff miss, and how do you bound what it actually covers?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A diff only sees the surface you snapshot. Outbound network calls, writes to a third-party service, data read and emitted without changing local state, persisted agent memory and delayed jobs all escape it. Bound it by enumerating every sink the sandbox exposes and recording each one, not just the database.

open as a page

You lead an AI red team whose engagement reports quote attack success rates from an agent harness. What standard do you set for trial counts and for how variance appears in a report, and what do you do with the payload that landed once in fifty?

level: principalimportance: should knowfreq 34%

basics

~20 s

Tie the trial floor to the claim: an existence finding needs one replayable success with evidence; a quoted rate needs a denominator and an interval in the text. Ban bare percentages. The one-in-fifty still gets filed on impact, marked low-rate and not verifiable by re-measurement at that budget.

open as a page

Every task in your agent red-team suite is scored against a hand-written list of forbidden side effects, and every run now passes. What can you conclude, and how would you restructure the success criteria?

level: principalimportance: should knowfreq 30%

basics

~20 s

Mostly that you only forbade what you thought of. A pass against a hand-written forbidden list is evidence about that list, not about the agent. Invert it: assert the exact set of effects the task permits, and fail the run on any effect outside that set, including ones nobody predicted.

open as a page

You own the partial-progress grading rubric for an agent red-team harness that several teams run against different agents. How do you decide how many progress tiers the rubric has, and who arbitrates a disputed grade?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Tie tiers to decisions, not to narrative detail. Use the fewest checkpoints that change what someone does: attempted, reached the irreversible action, completed. Every tier needs an objective world-state check and a named arbiter for disputes, or graders drift and cross-team numbers stop comparing.

open as a page