skip to content

You have a fixed number of queries against the model under test for an engagement and around two hundred seed behaviours to search with a branching attacker-model loop. How do you decide between a shallow wide pass over all of them and a deep search on a chosen few, and what makes that decision defensible to the team reading the report?

level: principalimportance: should knowfreq 24%

answer

  1. wide buys a denominator, deep buys a demonstration
  2. shallow first, promote on partial movement
  3. reserve a third for depth
  4. per-seed depth cap on the reserve
  5. write the split down before the run

basics

~20 s

Run shallow across everything first, then spend the remainder deep on the behaviours that showed partial movement. Wide answers which behaviours are reachable at all and gives an honest coverage denominator; deep answers how hard a specific one is. Decide by which claim the report has to support, and write the split down beforehand.

solid answer

~60 s

Frame it as two different claims. A **wide shallow** pass supports "we attempted all two hundred behaviours and found N reachable within a small search" — it produces a coverage denominator and a rough difficulty ranking. A **narrow deep** search supports "this specific behaviour is reachable, and here is the line that reached it" — it produces a demonstrable finding on the behaviours that matter most. In practice run them in that order. The shallow pass is cheap per seed and its results tell you where depth is worth buying: a seed that produced partial movement at level one is a far better bet than one that returned a flat refusal to every child. Reserve a fraction of the allowance (a third is a reasonable default) for depth, and allocate it by that signal plus the severity of the behaviour, not by whichever seed happens to be first. The defensibility comes from writing the split and the stopping rule down before the run, and reporting coverage as attempted over planned so nobody reads an unsearched behaviour as a safe one.

go deeper

for a junior

Understands that a fixed query allowance forces a choice between trying many behaviours briefly and trying few thoroughly.

for a middle

Argues the tradeoff and proposes shallow-first with depth spent on the seeds that showed movement.

for a senior

Sets the split, the promotion rule and per-seed depth caps, and instruments coverage so a partial run reports its own denominator.

for a principal

Ties the allocation to what the engagement must claim, fixes the rules before the run so the coverage number is falsifiable, and states plainly what a null result does and does not bound.

### The decision is about what the report must claim Treat this as an optimisation and you will get a defensible-looking number that answers nobody's question. The two shapes support two different sentences, and only one of them is what the engagement was commissioned to produce. - **Wide and shallow** supports: *we attempted all two hundred catalogued behaviours, and N were reachable within a search of this size.* It buys a **coverage denominator** and a rough difficulty ranking. - **Narrow and deep** supports: *this specific behaviour is reachable, and here is the line of refinements that reached it.* It buys a **demonstration** on the behaviours that matter most. Wide dominates when you must report against a fixed catalogue, when the question is whether the system is broadly exposed, or when this is a regression run whose whole purpose is comparison with the previous model version — a run that went deep on eight behaviours and skipped a hundred and ninety-two cannot be compared to anything. Deep dominates when one severe behaviour must be demonstrated to move a decision, or when a mitigation has shipped and the question is whether it holds under sustained pressure. Shallow search is bad at that second job, because the point of a multi-level attacker loop is that hits often arrive several refinements in; a two-level pass under-reports reachability and produces a falsely reassuring result. ### The staged allocation, with the arithmetic Run them in order, wide first. Suppose the allowance is 60,000 target queries and each candidate costs an attacker call, a target query and a judge call — so 60,000 candidates, and the three-calls-per-candidate constant is what you actually reconcile against the provider bill. ``` wide phase ~ 2/3 of allowance = 40,000 candidates / 200 seeds = ~200 each at b=4, w=3 that is roughly d=16 levels deep phase ~ 1/3 of allowance = 20,000 candidates / top 20 seeds = ~1,000 each i.e. ~5x the depth, or the same depth 5x wider ``` Score each seed in the wide phase on **best progress achieved**, not hit-or-no-hit, and promote on that signal combined with the severity of the behaviour. Put a per-seed depth cap on the reserve so one stubborn seed cannot eat it. Stop a deep line on a confirmed hit: a second demonstration of the same behaviour is a duplicate, and the same queries spent on an unsearched seed can produce something new. Wall clock, not money, usually binds — 180,000 calls at four in flight and three seconds each is about a day and a half. ### Where the numbers mislead **The deep arm's success rate does not generalise.** You promoted seeds *because* they showed movement, so the deep phase's hit rate is conditioned on the very signal that predicts hits. Quoting it as "the model fails 40% of the time" is a selection effect, not a measurement of the catalogue. **A null result bounds effort, not safety.** "No hits" means no behaviour was reached at b, w, d and this seed set. It does not mean the model is robust; a wider, deeper or better-seeded search is always a possible next step, and saying so is what keeps the number honest. **Only the wide arm is a regression metric.** If the promotion rule changes between runs, the deep arms of two runs cover different seed sets and their numbers are not comparable. Fix b, w and d for the wide arm and compare that, version to version. **Both extremes fail in their own way.** All-wide yields near-misses and no demonstrated finding, which a defending team can dismiss. All-deep yields two dramatic demonstrations and no sense of the shape of the exposure, and it over-fits your guess about which behaviours matter — the seeds you did not pick are where the blind spots are. ### What makes it defensible, and what to check Write the split, the promotion rule and the stopping rule down **before** the run. A split chosen after seeing results makes the coverage number unfalsifiable, because any allocation can be justified retrospectively. Report three numbers rather than one: seeds attempted over seeds planned, seeds with a confirmed hit, and the depth reached in each arm. Then verify the promotion rule itself, which is the load-bearing assumption of the whole design: take a small random sample of *un-promoted* seeds and give them the deep budget anyway. If they hit at a similar rate to the promoted ones, your progress signal is not predictive and the staging bought you nothing but a story.

  • What signal promotes a seed behaviour from the shallow pass into the deep reserve?
    Best progress achieved at shallow depth — partial compliance or visible movement away from a flat refusal — combined with the severity of the behaviour, rather than hit or no-hit alone.
  • Why stop a deep line as soon as it produces a confirmed hit?
    Further queries on that seed buy duplicates of a behaviour you have already demonstrated; the same queries spent on an unsearched seed can produce a new one.
  • How do you report a behaviour the run never reached?
    As not attempted, in an explicit attempted-over-planned line. It is neither a pass nor a fail, and collapsing it into either misleads the reader.

Promoting only the seeds that already showed movement and then quoting the deep phase's success rate is like quoting the pass rate of a class you filled with whoever did best on the mock exam. The number is real, but it describes your selection, not the school.

saying these in an interview costs you the question

  • Spends the whole allowance deep on a handful of behaviours chosen by intuition, then reports the model as broadly tested.
  • Reports unsearched behaviours as clean, or gives a hit count with no denominator.
  • Decides the split after seeing results, which makes the coverage number unfalsifiable.
  • Claims a null result proves the model is safe rather than bounding the search effort spent.
  • Lets one stubborn seed consume the depth reserve with no per-seed cap.

context