skip to content

An overnight Counterfit scan that had run fine all week burned a month of endpoint quota after a colleague switched on the optional search over attack parameters. Where does that multiplier come from, and how would you bound the next run?

level: seniorimportance: should knowfreq 40%

answer

  1. one trial = one whole attack
  2. multiplies the whole battery
  3. cap trials, shrink seed set
  4. search local, validate remote
  5. tuned args overfit the subset

basics

~20 s

The parameter search does not tune inside one attack run; it reruns the entire attack once per trial. Cost becomes attacks times trials times the full per-attack query count, and that per-attack count is already seeds times iterations. Ten trials over three attacks is thirty complete attacks against a metered endpoint.

solid answer

~60 s

**The multiplier.** A parameter search treats one complete attack as a single evaluation of its objective. Each trial picks a new set of attack arguments, runs the attack from scratch over the seed set, and scores the result. Nothing is amortised between trials, so the search sits on top of the battery as a plain multiplier: `attacks × seeds × iterations` becomes `attacks × trials × seeds × iterations`. Worse, the search spends its trials where the objective looks promising, which is usually the attacks that run their full iteration budget. **Bounding the next run.** Cap the trial count explicitly instead of accepting a default; search over a small seed subset and re-run only the winning arguments over the full set; run the search against a local or surrogate copy and validate only the winner against the metered endpoint; and put a hard query cap in the target so a runaway search fails the job. Then report the winning arguments on held-out seeds, because arguments tuned on a handful of samples overfit them.

go deeper

for a junior

Should recognise that a search over settings means running the attack more than once, and that more runs against a paid endpoint means more cost.

for a middle

States the multiplier explicitly as trials times the full per-attack cost, and suggests capping the trial count and shrinking the seed set.

for a senior

Splits the work into a cheap search phase and an expensive validation phase, enforces a hard query cap in the target, and reads per-attack counters to see where the spend actually went.

for a principal

Argues that enabling an open-ended search changes the job class, so it needs a declared budget owner and a policy — while still defending tuning, because untuned arguments understate exposure.

**What the toggle changes.** In Counterfit's command loop, the optional search over attack arguments sits next to `run_attack` and looks like one more setting. It is not a setting; it changes the class of the job. Without it, an attack runs once with the arguments you set and the cost is fixed: seeds × iterations × calls per iteration. With it, the attack's arguments become the search space of a hyper-parameter study (an Optuna study in the shipped implementation), and **one trial is one complete execution of the attack over the seed set**. The study's objective — attack success rate, or success at a minimal perturbation — is a single number that can only be computed by running the whole attack to the end. Nothing is amortised between trials: no predictions are cached, no seeds are pruned, no partial run is reused. So the battery's cost picks up a plain multiplier: ``` before: attacks × seeds × iterations × calls_per_iter after: attacks × TRIALS × seeds × iterations × calls_per_iter ``` Ten trials over three attacks is thirty complete attacks against a metered endpoint. And the multiplier is worse than uniform, because a search that is working spends its later trials in the region where the objective looks promising — which is normally the argument sets that push the attack to run its full iteration budget rather than exiting early. The trials that succeed are the expensive ones. **Work it out loud in the answer.** Three attacks, twenty seeds, one attack measured at about five thousand queries over that seed set: fifteen thousand queries a night for the battery. Turn on ten trials of search for each attack and you are at a hundred and fifty thousand. At a tenth of a cent per call that is a rounding error; at two cents a call it is three thousand dollars; at 200 ms serialised it is over eight hours of pure wall clock before any of it overlaps. The endpoint did not change and the attacks did not change — the shape of the job did. **Diagnosing it after the fact.** Pull the **per-attack** query counters, not the elapsed time. A search that burned its entire budget inside one attack and barely touched the others looks identical, in wall clock, to an even spread — and the fix is completely different. Then check three things: whether trials can abandon a hopeless configuration early or always run to completion (a study with no pruning pays full price for every dead end); whether the search ran over the full seed set, because `len(X)` multiplies inside every trial; and whether the trial count was one you chose or a default nobody read. **Where the number misleads.** The tuned result is the trap. A search reports the argument set with the best objective *on the seeds it searched over*, and with eight or twenty seeds that is a small sample being optimised against directly. The winning success rate is therefore fitted, and quoting it as the model's exposure overstates it — sometimes badly. The honest figure is the winning arguments re-measured on seeds the search never saw. The mirror error is just as common in the other direction: reading a *failed untuned* attack as evidence of robustness. A default step size or epsilon is not a serious adversary's choice, so an untuned failure understates exposure. Both readings come from treating one run's number as a property of the model rather than a property of a configuration. **How I would bound the next run.** - **Cap the trials explicitly**, written next to the attack list and reviewed like any other budget number, rather than accepting whatever the default is. - **Two-phase it.** Search cheaply — on a small seed subset, or against a local or surrogate copy of the model where queries are free — then send only the winning arguments to the metered endpoint for one confirmation run. The multiplier lives entirely in the search phase, so move that phase off the meter. - **Hard-cap queries in the target itself**, so the worst case is a raised exception at a threshold you chose rather than a discovery on the bill. - **Report on held-out seeds**, with the trial count and the search space stated next to the score. **The judgment point.** The failure here was not enabling the search. Tuning is legitimate and necessary: untuned attack arguments understate a model's exposure, and a defender who reads that as safety is being misled by your report. The failure was enabling an open-ended optimisation against a metered production endpoint with no trial cap, no query cap, and no named owner of the budget it was about to spend.

  • Why not just leave the parameter search off permanently?
    Because untuned arguments understate exposure — a failed attack at a default step size is not evidence of robustness. Keep the search, run it cheaply, and confirm the winner against the real endpoint once.
  • The search reports a configuration with a much higher success rate than the default. What do you check before publishing it?
    That the number holds on seed samples the search never saw. Arguments tuned against a small subset routinely overfit it, and the honest figure is the held-out one.

Turning on the parameter search is like discovering the overnight job you have been running was actually the body of a loop: each trial re-runs the entire attack from the first seed, so the trial count multiplies the whole night rather than adding a step to it.

saying these in an interview costs you the question

  • Thinks the search tunes arguments inside a single attack run rather than rerunning the attack per trial.
  • Proposes only 'lower the iteration count' and never touches the trial count or the seed set.
  • Reports the best-found arguments and their score without a held-out seed set.
  • Concludes the fix is to never tune attack arguments, which leaves the model's exposure understated.

context