skip to content

In an agent security benchmark, some episodes end in a context-length error, an exception from a simulated tool, or a provider timeout. The harness records no attacker action for those, so they sit in the same bucket as runs the defence genuinely stopped. What does that do to your reported security number, and how would you change the instrumentation?

level: seniorimportance: should knowfreq 42%

answer

  1. crash logged as secure
  2. three-way outcome coding
  3. termination reason per episode
  4. aborts correlate with prompt length
  5. compare abort rate across arms first

basics

~20 s

They inflate it. Give every episode an explicit termination reason and split three ways: attacker action succeeded, agent completed the task safely, or the episode aborted. Report the abort rate as its own figure and compare it between the defended and undefended arms; a defence that lengthens prompts can manufacture aborts.

solid answer

~50 s

A crashed episode produces no attacker action, so a two-way pass/fail coding silently books it as a defensive win. Both reported figures move the wrong way at once: the attacker-action rate is biased down and the completion rate is biased down too, which reads as "safe but somewhat costly" — the exact shape of a good defence. What makes it dangerous rather than merely noisy is that aborts are not random with respect to the thing you are testing. Defences that add instructions, wrap tool output or re-ask the model lengthen the context and add turns, so they hit context and turn limits more often than the undefended arm does. The defence then manufactures its own "secure" episodes. The instrumentation fix: record a termination reason per episode, code outcomes three ways instead of two, publish the abort rate beside the other numbers, and treat a materially different abort rate between arms as grounds to rerun rather than to compare.

code

python · 8 lines
python
def score_episode(ep):
    if ep.termination in ABORT_REASONS:
        return "aborted"            # excluded from both rates, counted separately
    if attacker_action_occurred(ep):
        return "attack_success"
    if user_task_completed(ep):
        return "safe_and_useful"
    return "safe_but_failed"        # includes explicit refusal

go deeper

for a junior

Should recognise that a crashed run is not a defended run and that the harness needs to tell them apart.

for a middle

Names the three-way outcome coding and that abort rate belongs in the report.

for a senior

Explains why aborts correlate with the defence itself — longer prompts and extra turns — and gates the comparison on abort-rate parity between arms.

for a principal

Makes termination-reason logging and an abort-rate parity check a standing precondition for any security claim leaving the team, and owns the retry policy.

### Why an abort is scored as a defence Security scoring in an agent harness is a predicate over what the episode recorded: *did the attacker's target action occur?* An episode that died before the agent could act satisfies that predicate trivially. If the harness codes outcomes two ways — attacker action happened, or it did not — then a context-length error, a mock tool raising an exception, and a provider timeout are all filed under "did not", which is the same bucket as a genuinely resisted injection. Nothing was tested, and the run is counted as a win. Worse, the two headline numbers move in a mutually reinforcing direction. The aborted episode contributes zero attacker actions (security looks better) *and* an incomplete user task (utility looks worse). The resulting shape — attacker-action rate down, completion down a bit — is precisely the signature of a real defence that costs some usefulness. Infrastructure failure counterfeits the exact pattern you are looking for. ### Why this is a confound, not noise Noise averages out across arms; this does not, because aborts correlate with the treatment. - **Prompt-side defences grow the context.** Extra system instructions, delimiting or re-stating every untrusted tool result, a second model pass over each observation — all of it lands in the same window. Long-horizon tasks that fit before now truncate. - **Turn caps interact the same way.** A defence that inserts a verification step spends turns from a fixed budget, so the agent is cut off mid-chain on exactly the long tasks where injections are most interesting. - **Retry and re-ask logic burns budget asymmetrically.** Retry-on-refusal exists only on the defended arm, so only that arm exhausts attempts. The arm you hope will win therefore earns extra free "secure" episodes in proportion to how heavy the defence is. This is the single most common way a promising guardrail result evaporates under review. ### The instrumentation Every episode ends with exactly one recorded termination reason, and outcome coding is at least three-valued rather than pass/fail: ```python class Termination(Enum): TASK_DONE = "task_done" # check attacker action separately AGENT_REFUSED = "agent_refused" # explicit refusal, not an error TURN_LIMIT = "turn_limit" CONTEXT_LIMIT = "context_limit" TOOL_ERROR = "tool_error" # environment fault PROVIDER_ERROR = "provider_error" # timeout, rate limit, transport ``` Rules that follow: aborts leave both rates and get their own reported line with its own denominator; refusal is separated from error, because they have different owners and different fixes; provider errors are retried under a bounded, logged policy so the effective attempt count per arm stays known; and the per-arm abort histogram is compared *before* any other comparison is permitted. ### What it costs to fix Cheap relative to what it saves, but not free. Bounded retries on transient provider errors add calls in proportion to the error rate — a 5% timeout rate with two retries is roughly 5% more calls, unnoticeable. Raising the turn cap is the expensive lever: turns are sequential and each carries the whole accumulated transcript, so lifting a 20-turn cap to 30 can add 50% more calls *and* materially more tokens per call, since context grows with every turn. Excluding aborts shrinks the denominator, which widens the interval on both rates: at a 15% abort rate on 400 episodes you are estimating from 340, and recovering the lost precision means re-running. The cheapest item on the list is the one people skip — writing the termination reason down costs one enum field per episode and nothing at runtime. ### Where the number misleads, and what you check A defence that appears to remove nearly all attacker actions while completion falls moderately, with an abort rate several times the baseline, has not defended anything: it ran out of context. So pull the **termination-reason histogram per arm first**, before looking at either headline rate. If the defended arm aborts materially more, the comparison is invalid on its face — raise the limits, shorten the tasks, or rerun; do not "adjust". Then check that refusals are counted separately from errors, that retry counts are logged per arm, and that aborted episodes are excluded from both rates rather than silently scored as safe failures. If a report shows you only two percentages, none of this is visible, which is why the right response to two percentages is to ask for the per-episode records.

  • Should aborted episodes be excluded or retried?
    Transient provider errors: retry under a bounded, logged policy. Limit exhaustion and environment faults: exclude and report, because retrying just reproduces them.
  • How would you notice this problem in a report that only shows two percentages?
    You cannot, which is the point — ask for the termination-reason histogram and the per-episode records before accepting the comparison.

A smoke detector with a dead battery never goes off, and a log that only records alarms will show a perfect year. The silence has to be distinguishable from working equipment that found nothing.

saying these in an interview costs you the question

  • Treating aborted episodes as noise that averages out.
  • Folding refusals and errors into one bucket.
  • Comparing a defended and undefended arm without checking their abort rates.
  • Unbounded retries that quietly change the effective attempt count per arm.

context