You have a fixed query budget for a garak engagement against a paid endpoint — enough for perhaps a fifth of the probe catalogue. How do you split it between breadth across many probe families and depth within a few, and what do you tell the stakeholder the result covers?
answer
- staged, not a one-shot split
- breadth buys information, depth buys evidence
- shallow clean is unresolved, not clean
- reserve budget for retest
- local stand-in shakes out wiring, not findings
basics
~20 sSpend a first slice broadly and shallowly across families to find where signal exists, then reinvest the rest depth-first on the families that hit and the ones the deployment's threat model cares about. Tell the stakeholder exactly which families ran, which were never run, and that unrun means untested.
solid answer
~60 sTreat it as sequential allocation, not a one-shot split. A cheap breadth pass over many probe families at minimal repeats is an information-gathering move: it tells you where the target is fragile. Reinvest the remaining budget depth-first on the families that produced hits and on the families the deployment's threat model makes expensive — a retrieval-backed assistant and a plain chat box do not deserve the same allocation. The reason depth is not optional: a single sample per prompt turns an intermittent behaviour into a coin flip, so breadth-only results are noisy in a specific direction — they under-report unreliable failures rather than over-report them. Breadth-only is a map, not a measurement. What you commit to publicly is the shape of the allocation, not a verdict. State which families ran and at what depth, that the unrun families are untested rather than passed, and that the depth-first half was chosen from the shallow pass and therefore inherits its blind spots. Reserve a slice of budget for re-testing after any fix, because a fix that moves the failure rather than removing it is common.
go deeper
Should recognise the budget forces a choice and that families left out must be reported as untested.
Should propose a shallow-then-deep split and connect depth to the unreliability of single-sample results.
Drives allocation from the deployment's threat model, keeps a retest reserve, and separates what a local stand-in can and cannot establish.
Frames the whole engagement as sequential allocation under uncertainty, states which decisions the result may and may not support, and defends the untested list against being read as a clean bill.
This is an allocation problem under uncertainty, and the mature answer is **staged**, not split. A one-shot division of the budget commits you before you know anything; staging lets the first spend tell you where the rest should go. ### Stage 1 — cheap breadth A shallow pass across many probe families at minimal `--generations` (one, or two) buys *information about where signal exists*. Its output is not findings; it is a map of where the target looks fragile. Keep it genuinely cheap by preferring the small, curated probes within each family over the dataset-backed ones — the goal is one touch per family, not a measurement. Record its clean results as **unresolved**, not clean. That distinction is the whole point of the stage. ### Stage 2 — depth where it matters Reinvest on two grounds that deliberately conflict: - families that produced hits in stage 1; and - families the deployment's threat model makes consequential *regardless* of stage-1 signal. The conflict is the judgement being tested. A family that matters enormously to this product deserves depth even if the shallow pass was quiet, precisely because a shallow quiet result is weak evidence. Depth here means raising `--generations` on those probes and adding the larger probes in the family, both of which multiply spend, so the arithmetic has to be redone at this point rather than assumed from stage 1. ### Stage 3 — reserve Hold back budget explicitly, and say how much: enough to re-run the specific probes behind every reported finding after remediation, plus a slack allowance for chasing one surprise. An engagement that spends to zero on the first pass cannot verify a fix, and an unverified fix is the most expensive artefact in the whole exercise — it converts a known gap into an assumed closure. ### Where breadth-only misleads, and in which direction This is the crucial asymmetry. Shallow sampling **under-reports intermittent behaviour**. A failure that occurs, say, a fifth of the time is more likely than not to look like a pass when each prompt is sent once. So a breadth-only engagement systematically produces a *cleaner* picture than reality — never a dirtier one. The bias runs in exactly the direction that gets people hurt, and it is invisible in the report, because a one-sample zero and a hundred-sample zero render identically as 0%. Say that out loud in the write-up. "Families A through H were sampled at one generation and produced no hits" is a defensible sentence. "Families A through H are clean" is not the same sentence, and only one of them survives an incident. ### Where depth-only misleads Concentrating everything on the families you already suspect means the engagement can only ever confirm your existing model of the risk. Whole families end up unmentioned, and unmentioned reads as fine to every downstream reader. Depth without breadth buys precision about the thing you already knew. ### Cheaper substitutes and their hard limit A locally hosted open-weights stand-in can absorb the shakeout — generator wiring, selection syntax, report plumbing, rough prompt-count and output-length calibration — for compute and engineer hours rather than per-token spend against the paid endpoint. That is a real saving and worth taking. What it cannot do is **transfer results**: a different model, with a different system prompt and different guardrails, has different failure modes, so nothing found or not found there is a claim about the target. Use the stand-in to make the paid run efficient; never to substitute for it. ### The stakeholder sentence Not "we tested the model", but something with structure: > These families were run at depth N; these families were sampled at one generation and are inconclusive; these families were not run at all. The depth choices derive from the shallow pass and inherit its blind spots. X% of the budget is held for post-remediation retest. Then name the decisions the result is allowed to support — this probe family regressed, that fix verified — and the ones it is not, chiefly any release gate phrased as an absence of risk. The value of the allocation discipline is that it makes the engagement's own limits legible in advance, so nobody discovers them by being surprised later.
- Why does a breadth-only engagement bias the report in a predictable direction?Shallow sampling under-reports intermittent failures — a behaviour that appears some of the time reads as a pass at one sample — so the picture is systematically cleaner than reality, never dirtier.
- Can you do the breadth pass against a cheap locally hosted model to save budget?Use it to shake out wiring, selection and report plumbing. Do not carry its findings across: a different model with different prompts and guards has different failure modes, so it supports no claim about the paid target.
- How much budget should be held back, and for what?Enough to re-run the specific probes behind every reported finding after remediation. Fixes frequently displace a failure rather than remove it, and an unverified fix is worse than a known gap.
saying these in an interview costs you the question
- Splitting the budget evenly across the catalogue because it seems fair.
- Spending the entire budget on the first pass with nothing reserved for retest.
- Presenting shallow clean results as clean rather than unresolved.
- Carrying findings from a cheap local stand-in over to the paid target as if they transfer.
- Letting the unrun families disappear from the report.