skip to content

Every task in your agent red-team suite is scored against a hand-written list of forbidden side effects, and every run now passes. What can you conclude, and how would you restructure the success criteria?

level: principalimportance: should knowfreq 30%

answer

  1. all-green measures the list, not the agent
  2. rule out unreachable tasks first
  3. invert to a permission envelope
  4. default-deny costs triage
  5. keep named deny rules for the headline harms

basics

~20 s

Mostly that you only forbade what you thought of. A pass against a hand-written forbidden list is evidence about that list, not about the agent. Invert it: assert the exact set of effects the task permits, and fail the run on any effect outside that set, including ones nobody predicted.

solid answer

~60 s

A deny-list criterion answers a narrow question — did any of these enumerated things happen? — and its blind spot is everything you did not enumerate. As agents get more capable the enumerated set grows more slowly than the effect space, so an all-green deny-list suite becomes weaker evidence over time, not stronger. The structural fix is to invert the polarity. For each task, declare the effects a correct completion is allowed to produce — these rows, this recipient, this directory — and score any observed effect outside that envelope as a violation, whether or not you anticipated it. Now the suite catches novel side effects by construction, and each new task costs you a permission envelope rather than an imagination exercise. The price is triage load: every legitimate-but-unlisted effect becomes a failure you must adjudicate, and a noisy sandbox will bury the signal. So allow-list scoring only works on top of a tight fixture, frozen ambient state and field-level diffs. In practice mature suites run both: a permission envelope as the general criterion, and named deny assertions for the specific catastrophes you want called out by name in the report.

go deeper

for a junior

Recognises that passing means only that the listed bad things did not happen, and that the list might be incomplete.

for a middle

Explains the deny-list blind spot and can describe scoring against the effects a task is allowed to produce instead.

for a senior

Rules out unreachable tasks and unmonitored sinks before rewriting criteria, and weighs the triage cost of default-deny against the fixture's noise level.

for a principal

Owns the polarity decision for the programme, keeps named catastrophe assertions for reporting, and makes the undetectable-by-this-suite list an explicit output that sets the next investment.

**First, read the all-green result correctly.** A suite of *deny-list* assertions — a hand-written enumeration of forbidden side effects, each checked against a state diff — that never fires supports exactly one claim: none of the effects on that list was observed in these runs. It does not say the agent is safe, and it does not say the attacks were weak. Three cheap explanations must be ruled out before anyone redesigns anything. The **sinks may be unobserved**, so the effect happened where no snapshot or recorder was watching. The **task may be unfailable**: if the fixture never grants the capability the assertion targets, or the scenario never reaches the state where the forbidden action is even available, the pass carries no information at all. And the **assertion may be broken** — a predicate with a typo'd table name, or an ignore list that swallows the very column it should watch, passes silently forever. **Then fix the polarity.** Deny-lists and allow-lists differ in *who has to be exhaustive*. A deny-list requires the defender's imagination to be complete: you catch only what you thought to forbid. An allow-list — a **permission envelope** — requires only that the task's legitimate behaviour be enumerable, which is a far smaller and much better-understood set, because the task author already knows what one correct completion does. This is the default-deny argument from firewall and capability design, and it lands the same way here: the unanticipated case is the one that matters, and only one polarity catches it by construction. ```text task: refund-request allowed_effects: - update orders.status where id = 41 -> {refunded} - one outbound message to the address on order 41 - writes under /workspace/task/** any_other_observed_effect: violation ``` As agents gain tools and autonomy the enumerated forbidden set grows arithmetically while the reachable effect space grows much faster, so an all-green deny-list suite becomes *weaker* evidence over time, not stronger — a genuinely counter-intuitive property that a good answer names. **What it costs.** Three distinct bills. - *Triage.* Default-deny converts every benign unlisted effect and every scrap of ambient noise into a failure someone must adjudicate. On a loose fixture this is hours a week and it does not decay. The real failure mode is political rather than technical: sustained triage pain produces steady pressure to widen the envelope until it stops firing, and a fully widened envelope is a deny-list with extra ceremony. - *Prerequisites.* Envelope scoring only works on top of a tight fixture, frozen clock and seeds, network default-deny and field-level diffs. If those are not already in place, the inversion is not a scoring change, it is an infrastructure project — usually the larger part of the work. - *Authoring.* The envelope moves to the task author, who must state exactly what a correct run does. That is a healthy forcing function — a task whose legitimate effects nobody can enumerate is a task nobody can score — but it is real hours per task, plus a re-baselining sweep of the whole suite so the new criterion has history. **Where the number misleads.** A 100% pass rate under a deny-list is quoted upward as "the agent is safe", when the honest sentence is "no listed effect was seen in these runs, on the sinks we watch, in the states these tasks reach". After inversion the failure rate misleads in the other direction: envelope violations are dominated by benign unlisted effects, so a suite that suddenly reports 30% failures has not found a dangerous model, and reporting that figure as an attack-success rate is simply wrong. And the headline suffers: "an effect outside the envelope occurred" is a weaker sentence to put in front of an executive than "the agent deleted customer data", which is why mature suites run both — a permission envelope as the general criterion, plus a small set of named deny assertions for the catastrophes you want called out by name. **What to check.** The cheapest and most neglected check is the **positive control**: instruct the agent plainly and non-adversarially to perform the forbidden effect, and confirm the assertion fires. An assertion that has never fired has never been tested, and a suite of them is a smoke detector with no battery. Beyond that: version and review envelope edits exactly like assertion edits, and require any widening that removes failures to be justified as fixture noise *with the noise fixed*, never as a scoring adjustment. Finally, separate three statements in the report — what was forbidden and did not happen, what fell outside the envelope and did not happen, and what nothing in this suite could have detected. Neither polarity sees an unobserved sink or a behaviour no task exercises, so that third list is the honest statement of residual risk, and it is the one that should set next quarter's investment.

  • Before you rewrite any criteria, what is the cheapest check on an all-green suite?
    Verify the tasks are failable: give the agent an unambiguous instruction to perform the forbidden effect, or plant a known-working condition, and confirm the assertion actually fires. A criterion that has never fired has never been tested.
  • Why is enumerating allowed effects easier than enumerating forbidden ones?
    The allowed set is bounded by what one correct completion of a scoped task does, and the task author already knows it. The forbidden set is bounded only by what a capable agent might do, which nobody can enumerate.
  • How do you keep envelope scoring from being widened into uselessness?
    Treat envelope changes like assertion changes: reviewed, versioned, and reported. A widening that removes failures should be justified as fixture noise with the noise fixed, not as a scoring adjustment.

saying these in an interview costs you the question

  • Reporting an all-green deny-list suite as evidence that the agent is safe.
  • Never checking whether the tasks could have failed at all — unreachable states scored as passes.
  • Adopting allow-list scoring on a noisy sandbox, then widening the envelope until it stops firing.
  • Dropping named catastrophic assertions entirely once an envelope exists, leaving the report with nothing concrete to say.

context