skip to content

Your white-box robustness evaluation was capped at contracted GPU-hours — what does that cost the bound?

level: seniorimportance: should knowfreq 40%

answer

  1. Informed is not the same as exhaustive
  2. The bound is as tight as the spend
  3. Unrun families are unmeasured, not passed
  4. Compare funded hours to attacker months
  5. Report the search beside the number

basics

~20 s

The ceiling is only as tight as the search you paid for. A granted-access run that ran out of compute reports that those attacks failed, not that no adversary succeeds, so the bound holds only against adversaries whose own effort is smaller than the search you funded.

solid answer

~50 s

Granting weights makes the adversary maximally *informed*; it does not make the search exhaustive. What comes back is the result of the attack families the lab could run within its contracted GPU-hours and calendar days, so surviving accuracy is a statement about that search, not about the model. The bound therefore degrades in a specific way: it constrains weaker adversaries only up to the effort actually spent. An outside attacker with no weights but months of wall-clock and a live endpoint can spend more than a lab paid for six days, and then the 'ceiling' is not above them at all. Practically: record what was run and what was not, treat unrun attack families as unmeasured rather than passed, and refuse to describe an underfunded granted-access result as an upper bound without the compute and coverage beside it.

code

text · 10 lines
text
Evaluation summary — submission R-4471 (urgent-read triage classifier)
  access granted ............ weights, architecture, training recipe
  evaluator compute ......... 40 GPU-hours (contracted)
  wall-clock ................ 6 calendar days
  attack families run ....... 2 of 5 in the lab's catalogue
  threat model .............. L-infinity, radius 2/255
  accuracy under strongest attack run ..... 61.2%
  families not run .......... 3 (sparse; low-frequency; physical-capture)
  input population sampled .. 1,000 of 84,000 studies
  ...

go deeper

for a junior

Be ready to say that an evaluation result reflects the attacks that were actually run, and that skipping an attack family is not the same as passing it.

for a middle

An interviewer expects you to separate the adversary's information from the adversary's effort: granting weights fixes the first, and only spending reduces uncertainty about the second. Name what gets cut when a contract binds.

for a senior

Show you would publish compute, wall-clock, families run and sampled coverage beside the figure, treat skipped families as open findings with owners, and re-run on every material model change rather than carrying a stale number forward.

for a principal

Own the argument that the funded search must be compared against the effort a real adversary can bring, and be prepared to say when a contracted result is too thin to support the decision it is being used for.

## Two different things are being called the bound The containment argument — a granted-access adversary can do anything a weaker one can — says the *strongest possible* white-box adversary dominates every weaker adversary. That is a statement about capability classes and it is exactly true. The number in the report is not that. It is what a particular lab, with particular attack families, produced inside a contract that paid for a fixed quantity of GPU-hours and a fixed number of calendar days. Between the two sits a search, and the search was finite. So the reported figure is an upper bound on attack success *only for adversaries whose effective effort is bounded by the effort that was spent*, and it is best read as a measured lower bound on what an attacker can achieve — one that later work can only push further. ## What actually gets cut when the budget binds When a contracted evaluation runs short, three things are usually sacrificed, and each has a different effect on the claim: - **Attack families not run at all.** The lab's catalogue is broader than the contract. Families that were skipped are *unmeasured*, and the standing error is to record them as passed by omission. - **Search depth within a family.** More optimisation effort and more restarts from different starting points find failures that a shallow search misses. A shallow run says less than it appears to. - **Coverage across the input population.** Robustness is not uniform. A run over a sampled subset may miss a class or a subgroup where the model is far weaker, and the headline average hides it. A report reading like this is the common case: ```text Evaluation summary — submission R-4471 (urgent-read triage classifier) access granted ............ weights, architecture, training recipe evaluator compute ......... 40 GPU-hours (contracted) wall-clock ................ 6 calendar days attack families run ....... 2 of 5 in the lab's catalogue threat model .............. L-infinity, radius 2/255 accuracy under strongest attack run ..... 61.2% families not run .......... 3 (sparse; low-frequency; physical-capture) ... ``` The defensible reading is 'two families, forty GPU-hours, 61.2% survived'. The indefensible one is '61.2% robust'. ## When the ceiling stops being a ceiling The bound is vacuous exactly when the funded search is weaker than the effort an ungranted adversary will bring. That comparison is not hypothetical in a deployed system: - the lab was paid for days; a motivated attacker against a live endpoint has months; - the lab ran two families; an attacker chooses whichever family the deployment is worst against, including ones no contract funded; - the lab worked digitally at a stated radius; an attacker against a camera-fed pipeline works under area and viewpoint constraints that respect no radius at all, and a digital result says almost nothing about that. When any of those holds, the granted-access number is not above the real adversary, and saying 'white-box, therefore worst case' is wrong. Access dominance and effort dominance are separate things, and only the first is guaranteed by the grant. ## What to do with an underfunded result You can still make it useful, provided you report it as what it is: 1. **Publish the search alongside the number** — compute spent, calendar time, families run, families skipped, and the sampled fraction of the input population. A figure without them is not comparable to any other figure. 2. **Treat unrun families as open findings**, not as absences of evidence that round to safe. Write them down as residual risk with an owner. 3. **Spend the next contract where the tail is**, not on deepening a family that already reported. The families skipped are where the unknown lives. 4. **Re-run on every material model change.** The bound was measured on one artefact; a fine-tune or a retrain invalidates it, and the cost of re-running is what makes the number a recurring line item rather than a one-off. 5. **Never let the number harden into a property of the product.** 'Survived at 61.2% under two families and forty GPU-hours' is a sentence somebody can act on. '61.2% robust' will be quoted for years by people who never saw the contract. ## The direction to hold onto A high surviving accuracy proves the attack that ran failed. It never proves the model resists an attack nobody funded. A granted-access evaluation removes uncertainty about the adversary's *information*; only spending removes uncertainty about the adversary's *search*, and no contract buys all of it.

  • The lab reports it could not break the model at all. Is that better news than a 61% figure?
    Not on its own. A total non-result is the outcome most sensitive to how much search was funded, and it is also the signature of a search that was blocked rather than defeated — an unbounded-effort attack that still leaves the model intact is a warning sign, not a triumph. Ask what was run, for how long, and whether cheaper adversaries did better than expensive ones.
  • How do you decide where to spend the next evaluation contract?
    On the families that were skipped and on the deployment's actual input path, not on deepening a family that already reported. Skipped families are where the unmeasured risk sits, and a digital-only result tells you little about a pipeline fed by a camera or a scanner. Depth in a family that already produced a number mostly refines a figure you can already act on.
  • Does a granted-access result survive a fine-tune of the model?
    No. The measurement was made on one artefact. Any retrain or fine-tune changes the function that was searched, so the bound lapses. Treat re-evaluation as a recurring cost tied to the release cadence, and make the report name the exact model version it covers so a stale number cannot follow a new build.

saying these in an interview costs you the question

  • Reads surviving accuracy as a property of the model
  • Records unrun attack families as passed
  • Assumes granted access implies an exhaustive search
  • Quotes the figure without compute, coverage or radius
  • Carries a bound forward across a fine-tune or retrain
  • Treats a total non-result as the strongest possible outcome

context