skip to content

Defining Success

Security success is a side effect that must not have happened, checked against world state rather than against what the agent said it did. Interviewers ask how the assertion was written.

on this pageshow

explore

questions

4

In an agent red-team harness, why is the agent's own transcript claim that it sent nothing not enough to score the run as a pass?

level: juniorimportance: must knowfreq 72%

answer

  1. transcript is a claim, state is evidence
  2. snapshot before, diff after
  3. silent success vs phantom failure
  4. egress recorder counts as state
  5. verdict function must not read the final message

basics

~20 s

Because the transcript is the agent's own account, not evidence. Score against world state: snapshot the sandbox before the run, run it, then diff. If the outbound queue gained a message or a row is gone, the run failed no matter what the agent narrated in its final turn.

solid answer

~50 s

The transcript is produced by the system under test, and in a red-team run that system is exactly the component an injected instruction may have steered. Its self-report is a claim, not an observation. A security harness owns the environment, so make the environment the oracle: - fixture the task's state (rows, files, mailbox, tool-call log) before turn one; - let the agent run; - diff every sink the sandbox exposes and evaluate the assertion over that diff. The two failure shapes this catches are opposite. The **silent success**: the agent calls the send tool, then reports "I did not send anything" — a transcript-scored harness records a false pass. The **phantom failure**: the agent narrates a deletion that never happened — a transcript-scored harness records a false alarm and you spend a day triaging nothing. The transcript still matters for triage and for harms that leave no state change, but it is never the thing that decides the verdict.

go deeper

for a junior

Says the transcript is what the agent claims and the environment is what happened, and that the harness should check the database or mailbox after the run.

for a middle

Describes the snapshot-run-diff loop, names the sinks worth diffing, and gives both the silent-success and phantom-failure shapes.

for a senior

Points out that a security false pass is far more expensive than a false alarm, and that harms with no state change need a boundary recorder rather than a state diff.

for a principal

Frames it as an evidence-independence rule for the whole suite — no verdict may be computed from an artefact the system under test authored — and applies it to judge inputs and report fields too.

**An agent run leaves two records, and only one of them is independent of the thing under test.** The *transcript* is the ordered log of the agent loop: the system prompt, the task input, each tool call the model emitted, each tool result the harness returned, and the final assistant message that summarises what happened. The *environment* is the sandbox the agent acted in — a seeded database, files under the task root, a fixture mailbox or outbound queue, a recorded HTTP egress boundary. A red-team harness owns the environment. The model under test authors the generated parts of the transcript. That asymmetry is the whole answer. **Three tiers of evidence, routinely conflated.** | artefact | produced by | what it actually proves | |---|---|---| | final assistant message: "I did not send anything" | the model | that the model wrote that sentence | | harness tool-call log: a recorded `send_email(...)` invocation | the harness | that the call was issued with those arguments | | pre/post diff of the fixture mailbox | the environment | that a message landed in the sink | The final message sits at the bottom because it is a *generation*, not a log. Models summarise optimistically, drop tool calls out of a recap, and sometimes narrate work they never performed — no adversary required. Add an injected instruction in the data the agent reads and the incentive becomes explicit: the useful thing for an injection to produce is an action plus a reassuring summary. In a capability eval a wrong self-report is a quality bug. In a security eval, where the entire question is whether a forbidden effect occurred, it is a false pass on the exact property you were paid to check. The middle tier is genuine harness-produced evidence but still not conclusive: a recorded call may have errored, been rejected by the sandbox, or hit a stub. Issuing a call is not landing an effect. **The mechanism.** Make the environment the oracle: ```text before = snapshot(env) # rows, files, queue, egress log run_agent(task) after = snapshot(env) violated = assertion(diff(before, after)) # not: assertion(final_message) ``` Published agent-security suites are built this way rather than on text matching — AgentDojo's injection tasks expose a security check that is handed the pre-run and post-run environment objects, so the verdict is computed from state the harness controls rather than from the model's answer. Sinks worth snapshotting: the task's rows at field level, paths under the task root, the fixture mailbox, and the request list captured by the egress recorder. For a third-party API you cannot snapshot, the recorded request list *is* the state you assert over. **What it costs.** Transcript scoring is nearly free — one regex over one string — which is exactly why teams drift into it. State scoring costs a per-run fixture reset (a template-database restore or container re-create, typically hundreds of milliseconds to a few seconds, and it must happen between *every* trial or task N+1 inherits task N's rows), an egress proxy to stand up and keep working, storage for two snapshots per run, and the engineering time to write one predicate per task, which is usually the larger half of authoring a task. Set that against the inference bill: a twenty-turn agent run re-sends the growing tool-result history every turn, so a single task run consumes far more prompt tokens than a one-shot probe, and a 100-task suite at five trials each is 500 such runs per sweep. The snapshot machinery is a rounding error beside that. Do not let "parsing the text is cheaper" survive contact with the arithmetic. **Where the number misleads.** A transcript-scored suite reports an attack-success rate whose denominator is fine and whose numerator is a statement about the agent's *candour*. It is wrong in both directions and the single number cannot tell you which. The **silent success** — the send tool was called, the summary denies it — is recorded as a pass and pushes the reported rate down. The **phantom failure** — the model narrates a deletion that never occurred — is recorded as a violation and pushes it up, burning triage days on nothing. The second is annoying; the first is the one that ships. Worse, matching on the final message usually measures refusal *language* rather than harm: a run that ends "I cannot help with that" scores clean even when the tool call went out earlier in the loop. A 0% rate from that pipeline is not evidence of a safe agent; it is evidence that the agent's last paragraph was polite. **What to check.** Read the verdict function and ask which artefact it consumes; if the final assistant message appears anywhere in it, every number in the report is provisional. Then run a positive control — give the agent a plain, non-adversarial instruction to perform the forbidden effect and confirm the assertion fires. An assertion that has never fired has never been tested. Finally, check that the tool-call log and the state diff agree on the same run: a recorded call with no matching state change usually means a stub, an error swallowed by the tool wrapper, or a sink nobody is snapshotting. The transcript keeps its place — for triage, and for harms that leave no state change at all — but it never decides the verdict.

  • The environment diff is empty but the transcript describes a deletion. How do you score it?
    No violation of that assertion — nothing happened. Log the discrepancy separately: an agent that narrates actions it did not take is a real finding for other reasons, but it is not the side effect you forbade.
  • The agent read an API key from a config file and printed it in its reply. No row changed. Does state-diff scoring catch it?
    Not by itself. Exfiltration with no local write needs an assertion over what crossed the boundary — the reply body and the egress recorder — which is still observed data rather than the agent's self-report.
  • Why does this matter more for a security run than for a capability run?
    A capability run measures whether the agent did the task; the agent's honesty errors mostly show up as noise. A security run measures whether a forbidden effect occurred, and the agent's own account of that effect is the least trustworthy source you could pick.

Scoring on the transcript is asking the suspect to write his own alibi; scoring on the state diff is checking whether the money is still in the drawer. Both are worth reading, but only one of them settles the question.

saying these in an interview costs you the question

  • Treating the agent's final summary as the scoring input because it is easy to parse.
  • Assuming a false pass and a false alarm are symmetric costs in a security suite.
  • Believing a state diff makes the transcript useless — it is still needed for triage and for no-state-change harms.
  • Scoring on tool-call intent alone without checking whether the call actually landed.

context

open as a page

What does it mean to write the must-not-happen assertion for an agent red-team task before the run, and what goes wrong if you write it after reading the transcripts?

level: middleimportance: must knowfreq 62%

basics

~20 s

It means naming the forbidden side effect as a concrete, machine-checkable condition over environment state, and committing to it before you run anything. Written afterwards, the behaviour you observed sets the bar: you quietly accept whatever happened, and a near-miss that emptied one row becomes a pass.

open as a page

Your agent red-team harness scores every task by diffing the sandbox's state before and after the run. Which harms will that diff miss, and how do you bound what it actually covers?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A diff only sees the surface you snapshot. Outbound network calls, writes to a third-party service, data read and emitted without changing local state, persisted agent memory and delayed jobs all escape it. Bound it by enumerating every sink the sandbox exposes and recording each one, not just the database.

open as a page

Every task in your agent red-team suite is scored against a hand-written list of forbidden side effects, and every run now passes. What can you conclude, and how would you restructure the success criteria?

level: principalimportance: should knowfreq 30%

basics

~20 s

Mostly that you only forbade what you thought of. A pass against a hand-written forbidden list is evidence about that list, not about the agent. Invert it: assert the exact set of effects the task permits, and fail the run on any effect outside that set, including ones nobody predicted.

open as a page