You have one week of compute and a fixed call budget for a multi-step agentic harmful-task suite. How do you split it between breadth (many tasks, one attempt each, tight step caps) and depth (few tasks, repeated attempts, generous caps), and what is each split's resulting number unable to say?
answer
- survey versus measurement
- breadth: no variance, capped chains
- depth: interval, but narrow scope
- keep a zero-scoring control group
- fix caps and build across phases
basics
~20 sSplit by what the run has to answer. Breadth finds which task families the agent will engage with at all, but one attempt per task has no repeatability and tight caps truncate long chains. Depth gives a defensible per-task number with an interval, but says nothing about families you did not run.
solid answer
~60 sDecide the question first, because the two allocations answer different ones. **Breadth** — many tasks, single attempt, tight caps — is a discovery sweep. It tells you which families of task the agent will start on and roughly where it stalls. Its number has no variance estimate, so a family that scored zero may simply have drawn an unlucky sample, and every long-chain family is confounded with the cap. **Depth** — a small task set, several attempts each, generous caps — produces a per-task completion distribution you can put an interval on and compare across model or guardrail changes. It is silent about everything outside the chosen set, and small sets drift toward whatever the team already worries about. In practice run breadth first at a fraction of the budget, then spend the remainder on depth over the families breadth flagged, keeping a few families that scored zero so you can tell an absent capability from an unlucky draw. Fix the caps and the environment build across both phases, or the two halves are not comparable.
go deeper
Understands that more tasks or more repeats both cost budget and that you cannot have both.
Explains that a single attempt per task gives no repeatability while repeats give an interval, and that tight caps truncate long chains.
Sequences a breadth triage pass into a depth pass, holds caps and environment build fixed across them, and states what each number cannot support.
Owns the cadence and the standing risk: a depth-only programme converges on its own priors, a breadth-only one never produces an actionable number, and the reporting rules have to prevent a depth figure being quoted suite-wide.
**Two instruments, one budget.** A sweep cannot be both a survey and a measurement. The budget is a product — `families x tasks per family x repeats x turns per attempt` — and breadth and depth are simply arguments about which factor gets the money. Roughly 10,000 model turns buys either 400 tasks at one attempt and a tight limit, or 40 tasks at ten attempts with room to run long. Choosing without first naming which question the run must answer is how a week produces a number nobody can act on. **What each allocation buys, in statistical terms.** With one attempt per task, a family of 20 tasks yields a rate whose binomial standard error near 50% is about 11 percentage points — larger than nearly any real month-over-month change, so a breadth run cannot support a trend claim. Repeats fix that, but not as cleanly as the arithmetic suggests: attempts within one task share a prompt and an environment, so they are clustered, and the effective sample size is much closer to the number of *tasks* than to tasks x repeats. Compute intervals over per-task means, not over pooled attempts, or you will publish an interval several times too narrow. **A workable split.** Spend a quarter to a third on a breadth pass: every task family, one attempt, limits tight enough that the pass actually completes. Read it strictly as triage. Spend the rest on depth over the families that showed engagement, **plus a deliberate control group of families that scored zero** — without re-sampling some zeros you cannot separate "this agent will not do this" from "one sample went the other way", and the depth pass degenerates into confirming what breadth already implied. **The real cost asymmetry.** Adding repeats costs money and nothing else. Adding a task family costs engineering: tool stubs to author, a rubric to write per step, and a review to confirm the task is actually harmful and actually completable. That is days per family against minutes per epoch. Breadth is capital expenditure and depth is operating expenditure, which is why programmes drift toward depth — it is the cheap axis — and end up measuring the same three harms forever. **What each result cannot support.** - A breadth number cannot support a trend claim between two runs; single-attempt sampling noise swamps most genuine movement. - A depth number cannot support a statement about the suite as a whole. Quoting a result from forty tasks as the suite's figure is the most common misreading of these runs. - Neither can support a claim about a family the suite does not contain. A task list is one fixed sample of harms that some authors chose; absence of a family is not evidence about it. - Neither is fully trustworthy on a long-public suite. Once a task set has been in the open for a year or more, its exact tasks plausibly appear in refusal training, so a low completion rate may measure memorised handling of *those specific items* rather than a general property. The check is a small privately-authored variant set: if the public suite scores far safer than close variants, the gap is contamination, not safety. **Comparability discipline.** Fix and record across both phases: the per-attempt message, token and time limits; the environment build the stubs come from; the grading-criteria version; sampling parameters including temperature; the model endpoint and version; and the concurrency setting. Change any of these between phases and the difference between them has an unknown number of causes. **What to check.** Whether both passes carry the same configuration hash; whether zero-scoring families were ever re-sampled; whether intervals were computed with the task as the clustering unit; whether anyone has quoted the depth figure suite-wide in a slide; and whether a private variant set has ever been run beside the public one.
- Why keep task families that scored zero in the breadth pass inside the depth pass?One attempt cannot distinguish an absent capability from an unlucky sample. Re-sampling a few zero-scoring families is the only cheap way to tell which you have, and it stops the depth pass from only confirming what breadth already suggested.
- What single record makes the breadth and depth passes comparable afterwards?A pinned configuration covering caps, environment build, grading-criteria version, sampling parameters and endpoint, stored with both passes. Without it the difference between the phases has an unknown number of causes.
saying these in an interview costs you the question
- Spends the whole budget on one attempt per task and reports the result as a trend.
- Quotes a deep result over a handful of tasks as a figure for the whole suite.
- Changes step caps or the environment build between the two passes and compares them anyway.
- Never re-samples families that scored zero on a single attempt.
- Treats the suite's task list as the definition of the harm space rather than one fixed sample of it.