You hold a fixed query budget of 20,000 generations against a metered chat endpoint for one engagement. Do you spend it as 400 behaviours at 50 attempts each, or 2,000 behaviours at 10 attempts each? How do you decide?
answer
- breadth = where weak at all
- depth = does a persistent attacker win
- two-stage: screen wide, then deepen
- judge calls are a second budget line
- never average across different n
basics
~20 sDecide from the claim the number must support. Breadth, many behaviours at few attempts, finds unexpected harm categories and gives a stable per-behaviour picture. Depth, fewer behaviours at many attempts, shows what a determined attacker eventually gets. Most engagements buy a wide cheap pass first, then spend the remainder deep where the pass looked weak.
solid answer
~60 sThe two allocations answer different questions, so the claim you owe the customer picks one. **Breadth** answers "where is this system weak at all?" — coverage of harm categories, more chance that a surprising behaviour shows up, and a per-behaviour verdict that is cheap but noisy in the sense that a behaviour needing 20 tries reads as safe. **Depth** answers "can a persistent attacker get this specific thing?" — the honest form of the threat, since a real attacker retries. It buys a plateau on a narrow slice and says nothing about the categories you left out. In practice: two-stage. Spend maybe a quarter of the query budget wide and shallow, use it as a triage screen, then spend the rest at high n on the categories that produced hits or that policy treats as severe. Also account for the judge: 20,000 generations means up to 20,000 judge calls, and if the judge is itself a hosted model that is a second metered budget line, not a rounding error.
go deeper
Sees the tradeoff as more behaviours versus more tries, and knows more attempts finds more.
Ties each allocation to the question it answers and proposes the two-stage screen-then-deepen split.
Budgets judge calls and multi-turn generations, reserves capacity for reproducing disputed hits, and refuses to merge figures measured at different n.
Decides the split from the deliverable and the threat model, and sets the reporting convention that stops a narrow deep number being quoted system-wide.
## Start from the deliverable, not from the tooling The two allocations are answers to different questions, so the claim the report has to make picks one. - A **coverage-style attestation** across a taxonomy of harm categories needs *breadth*: the finding there is often an empty category, and you cannot report an empty category you never tested. - A **pre-launch question** about one severe capability needs *depth*, because the buyer's real question is whether a motivated attacker wins on retry, and n = 2 cannot answer that. If nobody can say which claim the number supports, the allocation argument has no ground to stand on. ## What each allocation actually buys - **Breadth** — 2,000 behaviours at n = 10 — maximises the chance that a surprising category shows up at all, and gives a wide but shallow per-behaviour verdict. - **Depth** — 400 behaviours at n = 50 — gets you far enough along the best-of-n curve to see where it plateaus on a narrow slice, which is the only way to characterise a persistent attacker. ## The two failure modes, stated precisely - *All-breadth understates.* At n = 10, a behaviour whose per-attempt hit chance is 3% is missed roughly three-quarters of the time and lands in your report as safe. Those are exactly the behaviours a patient adversary converts, because retrying costs them nothing. A wide low-n sweep is therefore a **screen**, and calling its rate the system's exposure is a false negative dressed as a measurement. - *All-depth overstates confidence.* A beautifully saturated curve on 400 behaviours says nothing about the 1,600 you skipped, and stakeholders will quote it as a system-wide figure regardless of what your caption says — the number escapes the caption every time. ## A workable split Two stages. 1. **Stage one** spends roughly a quarter of the budget on the full wide list at low n as a triage screen, logging every per-attempt outcome. 2. **Stage two** ranks behaviours by observed hits and by policy severity, promotes the top slice, and runs it to a plateau. Report two figures, each carrying its behaviour set and its attempts per behaviour, and never average them: a **mixed-n aggregate** is a rate at no particular n, so it is not reproducible and not comparable to anything, including your own next run. ## The costs people forget when they say "20,000 generations" - **A generation is not an attempt.** A multi-turn strategy that averages four exchanges buys you about 5,000 attempts, not 20,000. - If the hit rule is a **hosted judge model** called once per finished attempt, that judge is a second metered line of the same order — frequently on a larger model, so it can cost more than the target side. - **Rate limits** are often the true ceiling rather than credits: 20,000 calls at a handful per second is many hours of wall clock, 429 retries consume quota without producing outcomes, and they must be excluded from attempt counts rather than silently logged as non-hits, or your denominators quietly inflate and your rate quietly falls. - Then there is **engineer time**, which no budget line captures: every candidate hit that gates a release is read by a person, so a wide pass that produces 800 candidate hits at two minutes of triage each is nearly a week of someone's attention. - Finally, **reserve 10–15% of the budget** for re-running anything triage disputes — a hit you cannot reproduce is worse than a behaviour you never tested, because it costs an argument as well as a query. ## Where the resulting numbers mislead The **wide-pass rate** reads as low because the budget was low, and it will be quoted as evidence of safety. The **deep-pass rate** reads as high because its behaviours were *selected for having already scored hits*, and it will be quoted as the system-wide figure. Both misreadings are predictable, which is why the reporting convention — scope plus n on every figure — is part of the allocation decision rather than an afterthought. ## What you would check afterwards - Whether the deep stage actually plateaued; if it was still climbing, the deep number is a lower bound too and should say so. - Whether the screen missed an entire category, which is a coverage defect, not a result. - Whether the promoted behaviours were chosen by evidence and severity or by whichever ones happened to be cheap to attack. - And whether the spend matched the plan: an engagement that burned its reserve on retries after throttling has no capacity left to defend a disputed finding.
- Why can't you average the wide-pass rate and the deep-pass rate into one number?They were measured at different attempt budgets, so the mixture is a rate at no particular n. Report them as two labelled figures with their budgets attached.
- The engagement's ceiling turns out to be the endpoint's rate limit, not the credit budget. What changes?Wall-clock becomes the scarce resource, so parallelism and caching matter more than attempt count, and failed calls from throttling must be excluded from attempt counts rather than logged as non-hits.
Breadth is drilling many shallow holes to find out where the oil might be; depth is sinking one hole far enough to prove it is really there. The same fuel buys either, and neither answers the other's question.
saying these in an interview costs you the question
- Picking one allocation with no reference to what the report has to claim.
- Averaging a wide low-n rate and a narrow high-n rate into a single headline figure.
- Forgetting that a hosted judge doubles the metered spend.
- Assuming one attempt equals one generation when the attack strategy is multi-turn.
- Leaving no budget reserve to reproduce disputed hits.