skip to content

You have a fixed token spend for an agent red-team harness sweep against a metered hosted agent and about 40 candidate attack payloads, and each trial is a whole multi-turn episode rather than one call. How do you split that spend between covering more payloads and repeating each payload enough times to quote a rate?

level: seniorimportance: should knowfreq 52%

answer

  1. screen wide, deepen the survivors
  2. search vs measurement
  3. trial = episode, not a call
  4. sequential stopping at the decision threshold
  5. reserve budget for fix verification

basics

~20 s

Two stages. Spend a cheap screening pass of a few trials on all 40 payloads to find which ones ever land, then spend the depth only on those, repeating them enough to quote a rate. Stop early when the interval already answers the decision, and price a trial as a full episode.

solid answer

~50 s

Breadth and depth buy different things. Breadth finds behaviours; depth turns a behaviour into a number. Early in an engagement a hit anywhere is worth more than a precise rate for a payload you have not found yet, so screen wide and shallow first. A workable split: pass one runs every candidate a small fixed number of trials (three-ish) purely to separate "never landed" from "landed at least once". Pass two spends the remaining budget deepening the survivors to the trial counts a quoted rate needs. Because a trial is a full multi-turn episode, price it that way: cost scales with the turn budget, which is a third lever — and cutting it changes what you are measuring. Two refinements: stop deepening once the interval sits on one side of your decision threshold, and record that screening is lossy, since a payload landing 5% of the time usually shows zero in three trials.

code

text · 6 lines
text
stage 1: screen every payload, 3 trials each
  keep payload if hits >= 1
stage 2: deepen kept payloads
  add trials while the interval straddles the decision threshold
reserve ~20% of spend for variants + post-fix re-measurement
cost(trial) = turns_per_episode x (model tokens + tool calls + judge calls)

go deeper

for a junior

Says run everything a few times first, then repeat the ones that worked more times.

for a middle

Structures it as screen-then-deepen, and prices a trial as a full episode rather than a single call.

for a senior

Adds sequential stopping against a decision threshold, the turn-budget lever and what it costs in coverage, and a reserve for fix verification.

for a principal

Treats the allocation as an engagement-level policy — what the screen is allowed to conclude, what a report may claim from each stage, and how the verification budget is committed at file time.

## Frame it as a search followed by a measurement **Screening** and **rating** are different problems with different optimal spends. - **Search** wants many cheap draws spread across many payloads, because early in an engagement a hit anywhere is worth more than a precise rate for a payload you have not found yet. - **Measurement** wants many draws concentrated on one payload, because concentration is the only thing that narrows an interval. Running a uniform twenty trials across all forty payloads spends most of the budget establishing, to three significant figures, that things which never work never work. ## A two-stage allocation 1. **Stage one** screens every candidate at a small fixed depth — three trials is a common choice — purely to separate "never landed" from "landed at least once". 2. **Stage two** spends the bulk of the remaining budget deepening only the survivors, adding trials until the interval answers the decision in front of you. Hold back a **reserve**, roughly a fifth, for two things teams reliably forget: the variants the screen suggests once you can see which family of payloads lands, and the fix-verification runs you will owe after filing. Verification costs on the order of the original measurement; unbudgeted, it silently shrinks to two runs and proves nothing. ## Price a trial correctly, because it is not a call In an agent harness a trial is a whole episode: several model turns, a tool round-trip per action, and any judge-model call the scoring object makes on top. The spend multiplies as payloads x trials x turns per episode x tokens per turn, plus tool and judge calls. Forty payloads at twenty trials and eight turns is 6,400 model turns before tools, and wall-clock is usually set by tool latency and the endpoint's rate limit rather than by token count, so a sweep that looks affordable in dollars can still be a two-day run. The **turn budget** is the sneaky lever: halving it roughly halves the cost and buys twice the trials, and it also removes the long-horizon paths where a slow escalation lands. If you cut turns to afford trials, the number you quote is a rate under that turn budget and is not comparable to the earlier sweep — say so, rather than letting two incomparable numbers sit in one table. ## Sequential stopping Decide the question before you spend: "is this above the ten-percent line that makes it a release blocker?" Then deepen a payload only while its interval still straddles ten percent. Payloads that resolve early release budget to the ones that do not, which is where an **adaptive allocation** beats any fixed n. ## Where this allocation misleads - The **screen's misses** are the big one. A payload that lands five percent of the time shows zero hits in three trials about six times in seven, so "34 of 40 payloads did not land" is a statement about your screen depth, not about the target. Those payloads are unmeasured, and writing them up as clean is the most common way a two-stage sweep manufactures false assurance. - Second, selecting survivors on a first hit biases the deepened rates slightly upward, because you kept the payloads that happened to land early; the effect is small next to the interval at these trial counts, but it is real and deserves a sentence. - Third, a reserve quietly spent on more screening leaves the engagement with no verification budget, which is how a mitigation ships unverified. - Fourth, the plan assumes trials are independent: if the environment is not reset between them, later trials in a payload's block are correlated with earlier ones and your effective n is smaller than the number you quote. ## What to check before believing the allocation worked - That the hit criterion was byte-identical in both stages, since a screen and a deepening scored differently are not two stages of one experiment. - That the environment reset actually ran between trials rather than being assumed. - That aborted episodes — rate limits, tool errors, timeouts — are excluded from denominators and disclosed as a count. - And that the report shows the allocation itself: screen depth, which payloads were deepened and why, and the trial count and turn budget behind every quoted rate. A reader who cannot see the allocation cannot distinguish a payload that was tested hard and stayed clean from one that was barely touched.

  • Your screening pass of three trials each found hits on 6 of 40 payloads. Can the other 34 go in the report as clean?
    No — they go in as unmeasured at that depth. Three trials misses most payloads landing under about 10%; state the screen depth and the bound it supports.
  • Halving the turn budget doubles the trials you can afford. What breaks?
    Attacks that need a slow escalation across turns stop landing, so the rates fall for reasons unrelated to the target's defences, and the numbers are no longer comparable to the earlier sweep.

A trial is priced like a hotel stay, not a ticket: nights times rooms times nightly rate. Cutting the turn budget is shortening the stay — it halves the bill and also guarantees you never see what happens on the last night.

saying these in an interview costs you the question

  • Uniform trial counts across every payload regardless of whether it ever landed.
  • Pricing a trial as one model call when it is a multi-turn episode with tool and judge calls.
  • Cutting the turn budget to afford trials and still quoting the rate as comparable.
  • Leaving no budget for verifying the fix later.
  • Reporting screened-and-missed payloads as clean rather than as unmeasured.

context