You have a fixed spend of API calls for a query-only adversarial robustness assessment of a metered model, with no gradient access. How do you split it between number of examples, per-example query cap, and number of attacks, and what does each split cost the conclusion?
answer
- claim first, split second
- screen wide, then depth on a stratified subsample
- hold back a quarter for re-runs
- stabilise before scaling
- surrogate converts calls into one labelling cost
basics
~20 sThere is no right split, only a stated one. Wide and shallow covers many examples at a low cap and mostly measures the cap. Narrow and deep gives credible per-example results with wide error bars. Usually: one cheap screening pass, then deep runs on a stratified subsample, holding budget back for re-runs.
solid answer
~60 sThe three dimensions trade against each other under one multiplication, so every extra unit of depth is paid for in breadth. **Wide and shallow** — many examples, small cap. The success share is precise as a statistic and nearly meaningless as a claim, because it is dominated by the cap: you learn what a cheap attacker gets, which is a legitimate but narrow question. **Narrow and deep** — few examples, large cap. Each example's result means something, but the sample is small and the confidence interval on any population claim is wide. **Many attacks, thin each** — you find which family works on this target, at the price of not pushing any of them. My default is staged: a cheap screening pass across the full set to find where the target is weakest, then depth on a stratified subsample of the survivors, then a reserve of maybe a quarter of the budget for re-running whatever the first two passes made ambiguous. What travels into the report is the split itself, so a reader can see which question the number answers.
go deeper
Should recognise that examples and per-example cap trade off, and that both have to be stated with the result.
Should propose a screening pass followed by depth on a subsample and explain what each buys.
Should stratify the subsample, hold a reserve, and stabilise a noisy target before scaling the example count.
Should start from the claim the assessment must support, treat the budget as a portfolio with a reserve, weigh a surrogate route against direct queries, and fix the reporting fields that let two assessments be compared.
### The budget is a portfolio, and the split is the finding Under a fixed number of billed calls, examples, per-example evaluation budget and number of attacks multiply against one another. Every unit of depth is bought with breadth and vice versa; there is no split that is right in the abstract, only a split that matches a claim and is stated alongside the result. Choosing a split first and then describing whatever falls out is how an assessment ends up unfalsifiable — the number cannot be reproduced, compared, or argued with. ### Decide the claim before the split "A casual attacker with a few hundred calls per input succeeds on x% of inputs" and "a determined attacker with a very large per-input budget eventually succeeds on these specific inputs" are different deliverables bought by different people. The first is a deployment risk statement; the second is a worst-case statement for a threat model that includes a patient adversary. Size to the claim. | split | what it buys | what it costs the conclusion | |---|---|---| | wide and shallow — many examples, low cap | a statistically precise success share | the share is largely a measurement of the cap; it moves whenever the budget moves | | narrow and deep — few examples, high cap | credible per-example results, a real cost curve | a small sample, so wide intervals on any population claim | | many attacks, thin each | which attack family this target is weakest to | none of them pushed far enough to bound anything | ### The staged default A cheap screening pass at a low cap across the whole set, to locate the mass: which classes, which input regions, which attack family moves at all. Then a depth pass on a subsample **stratified** by class and by screening outcome, so the deep results can be weighted back to the population instead of being a convenience sample of the easy wins. Then a **reserve** — a fifth to a quarter of the budget, untouched. Something always needs re-running: a saturated cohort, a target that changed mid-engagement, a wrapper defect found late. A plan that spends 100% on the first pass cannot answer the first question a reviewer asks. Two sequencing rules sit above the split. **Buy variance reduction before scale**: if the target's answers are unstable near the boundary, spend on repeat probes first, because a larger unstable sample is a larger pile of noise, not more evidence. And **price the substitution**: a locally trained surrogate converts a recurring per-call meter into a one-off labelling cost, after which gradient attacks against it are free and metered spend goes only on verifying transferred candidates. Where the surrogate is a decent match this dominates any pure query allocation; where it is a poor match the labelling budget is simply gone, so pilot it on a slice before committing. ### Where the number misleads Three specific misreadings follow from an unstated split. First, a low success share from a wide-and-shallow run gets quoted as robustness, when it is mostly the cap: raise the cap tenfold and the same target may look far weaker, with nothing about the model having changed. Second, a high success share from a narrow-and-deep run on a hand-picked cohort gets generalised to the population; without stratification and weights that step is not available. Third, two teams' numbers get compared across a difference in split that neither reported — the assessment that used more attacks finds more, and the difference is read as a difference between models. Any of these makes a robustness figure worse than useless, because it is confidently wrong in a direction nobody can audit. ### What I would report, as fields Examples attacked and how selected; per-example evaluation cap and restarts; attacks run, by name and version; calls issued and calls billed; the share of examples that saturated the cap; the reserve held and the reserve used. Those fields are the minimum on which two assessments can be reconciled, and the split is itself a finding: a reader who disagrees with the allocation can re-derive what a different one would have shown, which is exactly what a defensible number should permit.
- Why stratify the deep subsample rather than take the first n examples?So the deep results can be weighted back to the population; a convenience sample skews toward whatever the screening pass happened to order first and cannot support a population claim.
- When does the surrogate-plus-transfer allocation beat spending the same budget on direct queries?When the surrogate is a good enough match that transferred candidates succeed at a useful rate, since the labelling cost is paid once and then amortises across every attack you want to try.
saying these in an interview costs you the question
- Choosing a split without first naming the claim the assessment has to support
- Spending the entire budget on one pass with no reserve for re-runs
- Taking the deep subsample from the easiest examples and generalising from it
- Publishing a success share without the example count, the cap and the attacks run