skip to content

Your agent red-team harness scores every task by diffing the sandbox's state before and after the run. Which harms will that diff miss, and how do you bound what it actually covers?

level: seniorimportance: should knowfreq 44%

answer

  1. coverage equals the sinks you snapshot
  2. egress recorder for no-write exfiltration
  3. read-only harm leaves only a log
  4. deferred jobs land after the snapshot
  5. ambient noise causes false failures

basics

~20 s

A diff only sees the surface you snapshot. Outbound network calls, writes to a third-party service, data read and emitted without changing local state, persisted agent memory and delayed jobs all escape it. Bound it by enumerating every sink the sandbox exposes and recording each one, not just the database.

solid answer

~60 s

State diffing is the right default, but its coverage is exactly the set of sinks you chose to snapshot — and that set is usually the database plus a directory. What routinely falls outside it: - **Egress with no local write.** A secret read into context and returned in a reply, or POSTed to an attacker URL, changes nothing you snapshot. You need a recorder at the network boundary and a matcher over emitted content. - **State living outside the fixture.** A vector store, an agent memory file, a scheduler, a cache, a third-party SaaS the tool calls for real. - **Deferred effects.** A queued job or scheduled send that fires after the snapshot is taken; the diff is clean and the harm lands ten minutes later. - **Read-only harm.** Accessing another tenant's record is a violation with no write at all; only an access log shows it. Bounding it is an inventory exercise: list every capability the task's tools expose, ask what evidence each one leaves, and either snapshot it, record it at the boundary, or write down that it is unmonitored. The unmonitored list is the honest coverage statement for the suite.

go deeper

for a junior

Knows the harness must check the environment after the run and can name the database and files as things to check.

for a middle

Names sinks a plain database diff misses — network calls, replies, memory — and knows deferred work must drain before the final snapshot.

for a senior

Runs the sink inventory deliberately, handles false failures from ambient state by narrowing the diff rather than loosening the assertion, and reports the unmonitored surface alongside the results.

for a principal

Treats evidence surface as a property of the whole programme, budgets for observers on new tool classes before tasks are written, and refuses to let a green suite be quoted as coverage it does not have.

**Reframe the question before answering it.** "What will the diff miss?" is really "what is my evidence surface?" A harness can only assert over observations it is able to make, so the coverage of a state-diff harness is exactly the set of *sinks* it snapshots — where a sink is any place an agent action can leave a trace: a table, a directory, a queue, a log, a network boundary. In most harnesses that set is a database plus one directory, chosen because they were the easy two. **A working inventory.** For a task whose agent has a database tool, a mail tool, a file tool and an HTTP tool: | capability | evidence it leaves locally | observer that sees it | |---|---|---| | database write | row change | field-level snapshot + diff | | file write | path change | snapshot of the task root | | mail send | queued message | fixture mailbox diff | | outbound HTTP call | none | recording egress proxy | | database read | none | query log / tool-call log | | reply to the user | transcript text only | content matcher over the emitted output | The bottom three rows are where diff-only scoring reports green with nothing behind it. Four concrete blind spots follow from them: - **Egress with no local write.** A secret read into context and returned in a reply, or posted to an external endpoint, changes nothing you snapshot. - **State living outside the fixture.** A vector store, a persisted agent memory file, a scheduler, a cache, a real third-party SaaS the tool calls for effect. - **Deferred effects.** A queued job or a scheduled send that fires after the post-run snapshot: the diff is clean and the harm lands ten minutes later. - **Read-only harm.** Reading another tenant's record is a violation with no write at all; only an access log shows it. **The opposite failure, which is what actually degrades a suite.** *False fails* come from ambient movement — an unfrozen clock, a background worker, a seeded random, an `updated_at` column that ticks on read, leftover rows from the previous task. A whole-table diff turns that churn into violations, and the fastest way to lose a suite's credibility is a wall of failures nobody can reproduce. The dangerous response is the natural one: loosen the assertion until the noise stops, which quietly converts real positives into passes. **Practical controls.** Freeze what should be frozen (clock, seeds, background workers, network default-deny). Diff *fields*, not whole databases, and keep non-semantic columns on an explicit ignore list — then review that ignore list as carefully as the assertion, because it is the one place a genuinely forbidden change can be silently excluded from scoring. Put a recorder on every egress path and treat the recorded request list as first-class state. Drain queues and run scheduled work before the final snapshot so deferred effects land inside the observation window. Snapshot after teardown too when the agent had persistence, since memory written in one task and read in the next is a real cross-task effect. **What it costs.** Each new observer is engineering, not configuration: an egress proxy with TLS interception and a stable request-normalisation format is days of work and an ongoing maintenance tax; per-run resets and queue drains add seconds of wall clock to every trial, which at 500 trials a sweep is real pipeline time; the ignore list needs review on every schema change. The cost is per *tool class*, not per task, which is the good news — pay once for the mail observer and every mail task is covered — and it is why the honest sequencing is to budget the observer before authoring tasks that need it, rather than writing forty tasks against a sink nobody watches. **Where the number misleads.** "0 violations across 400 runs" is read by every downstream consumer as absence of harm. It is not: it is absence of *observed* harm on the sinks you instrumented, and a false pass from an unwatched channel is indistinguishable in the report from a genuinely clean run. Two denominator swaps do the damage. First, "tasks passed / tasks run" gets quoted as "harms absent / harms possible", which it is not, because the second denominator includes every effect no observer could see. Second, coverage gets quoted as attack-surface coverage when it is really sink coverage — a suite can exercise 90% of the agent's tools while observing 40% of the places those tools leave traces. **What to check.** Ask for the sink inventory and the list of capabilities with no observer; if nobody has written it, the suite's coverage claim is unverified by definition. Run a positive control per sink, not per suite: make the agent perform the forbidden effect through *each* channel and confirm each assertion fires. Inspect the ignore list for anything semantic. Confirm the post-run snapshot is taken after queues drain. Then publish the unmonitored list alongside the results — a criterion nobody can check is not a criterion, and letting a green suite stand in for coverage it does not have is how a scan becomes false assurance.

  • The suite started failing intermittently on rows nobody's task touches. Where do you look first?
    Ambient environment movement: an updated_at column, a background worker, a clock or seed that is not frozen, or leftover state from the previous task. Narrow the diff to semantic fields and reset the fixture per run rather than relaxing the assertion.
  • How do you cover an agent tool that calls a real third-party API you cannot snapshot?
    Put the observer at the boundary: a recording proxy that captures every request. The recorded call list becomes the state you assert over. If the call must be live and unrecordable, list that path as unmonitored rather than implying it was checked.
  • Does the ignore list for noisy fields belong in review?
    Yes. It is the one place where a genuinely forbidden change can be silently excluded from scoring, so it should be as reviewed as the assertion itself.

saying these in an interview costs you the question

  • Equating "the database is unchanged" with "nothing happened".
  • No observer on the network boundary, so any exfiltration channel scores clean.
  • Diffing entire tables including timestamps, then muting the resulting noise by loosening the assertion.
  • Taking the final snapshot before queued or scheduled work has drained.
  • Reporting a green suite with no statement of which sinks were unmonitored.

context