skip to content

Your team wants to publish a robustness claim backed only by a stock attack suite - what do you require before it ships?

level: principalimportance: nice to knowfreq 28%

answer

  1. the claimant carries the burden
  2. self-attack first, then independence
  3. days are the currency, name who pays
  4. the wording is the deliverable
  5. claims carry a date and expire

basics

~20 s

That the authors attack their own defense first and publish what they tried, from what access, at what cost. The burden sits with whoever makes the claim, and the wording must carry those bounds and a date.

solid answer

~50 s

Set the standard rather than argue the number. First, the burden is the claimant's: the team must write the strongest attack they can against their own mechanism and report the formulations tried, the ones abandoned, the access granted and the effort spent. A defense whose authors could not break it after real effort is worth something; one nobody attacked is worth nothing. Second, fund an independent bounded attempt when the claim is customer-facing, and decide openly who absorbs those days — you cannot buy an unbounded adversary, so you buy the strongest bounded one and publish the bound. Third, fix the wording: claims read as *not broken by X, at access Y, for effort Z, on date D*, and they expire. Fourth, if this deployment will never face an adversary who reads the design, say so as an assumption and spend elsewhere — but then do not publish the number as robustness.

go deeper

for a junior

Take away the principle rather than the process: whoever claims a defense works is the one who has to try to break it first.

for a middle

Understand why the burden sits with the claimant - a break settles a question while a failed attempt only bounds it - and what an author's attack artifact should contain.

for a senior

Be ready to enforce this on a real launch: what you would send back, what evidence you would accept, and how you would re-word a claim so it survives someone re-testing it later.

for a principal

Own the whole standard: which claims require independent effort, whose budget funds the days, how claims are worded and when they expire, and the honest option of declaring that this deployment is not evaluated against an adaptive adversary at all.

### The decision, stated plainly You own the gate a model change passes through. A team has a defense and a number from a published attack suite, a launch date, and an executive who wants a robustness line in a deck. The technical facts are settled — a stock suite never knew this defense existed — so the question is not what the number means but *what standard you are willing to enforce and fund*, and that is a judgment somebody has to own and could refuse. ### 1. Put the burden where the evidence can be produced The asymmetry does the work here. A break is decisive and cheap to verify; a failure to break is bounded by whoever tried. Therefore the party who wants the claim must produce the attack. Concretely, before a robustness claim ships: - the authors write attacks aimed at their own deployed mechanism, not at the undefended model; - they file the artifact of the attempt — designs tried, designs abandoned, the reasoning for stopping, the access assumed, the effort spent; - the number is reported as what *those* attacks achieved. The objection you will hear is that authors cannot be objective about their own work. True, and it is a reason to require the artifact, not to drop the requirement. The self-attack is the cheapest strong evidence in existence, because nobody else knows the mechanism as well, and the reviewable record of what was tried is precisely what makes optimism visible to a reader. ### 2. Buy independence where the stakes justify it, and name who pays An external adaptive attempt costs analyst-days on somebody's budget. That is the real constraint, and pretending otherwise is how these standards die. Tier it: | Claim's reach | What you require | | --- | --- | | Internal, informational | authors' self-attack artifact, no external spend | | Customer-facing or contractual | funded independent attempt with a stated day budget | | Safety-relevant or regulated | independent attempt plus a scheduled re-test | Decide in advance which budget carries it — the product team that wants the claim, or a central security budget — because a standard nobody funds is a standard nobody meets. And accept the ceiling honestly: you can never buy an unbounded adversary, so you buy the strongest bounded one and you publish the bound. ### 3. Fix the wording, because the wording is the product Most of the damage in this area is done by a sentence, not a model. Set a house rule for how these claims may be written: the attacks that were written against the defense, the access granted, the effort spent, the date. The words *robust*, *secure* and *verified* do not appear unqualified. A claim in that shape can be re-tested, re-scoped and argued with; it also expires, which is correct, because a defense's exposure changes the moment somebody publishes a better formulation or your product starts returning finer-grained output. Expiry is worth enforcing mechanically. A robustness line with a date on it forces a decision at renewal — re-test, re-word, or withdraw — instead of quietly ageing into a claim nobody can source. ### 4. Ask the prior question before spending anything Does this deployment actually face an adversary who will read your design and adapt? Sometimes the honest answer is no, or not yet: the model gates a low-value decision, the mechanism is not documented outside the team, the realistic misuse is volume rather than craft. In that case, funding a full adaptive engagement may be a poor use of the same days, and the right call is to say so *as a written threat-model assumption*, with the trigger that invalidates it — publication, a shipped artifact somebody can inspect, a change in what the endpoint returns, or the first sign of somebody probing. What you cannot do is take that decision and still publish the stock number as robustness. Deciding not to test against an adaptive adversary is a legitimate resourcing call. Claiming resistance to one you never tested is not, and the distinction is exactly where a lead's signature carries weight. ### 5. What you tell the executive You will be asked for a number for the deck. Give one, in the bounded form, and explain the trade in their terms: an unbounded claim invites the one demonstration that destroys it, in public, at a moment you do not choose, whereas a bounded claim survives the same demonstration because it already named its limits. That is a commercial argument as much as a security one, and it is the version of this that actually gets adopted. ### The failure modes to name in the interview Requiring evidence nobody funds; accepting the suite number because the date is close; publishing an unbounded word because marketing asked; funding an open-ended evaluation that never concludes and so blocks a launch indefinitely; and never asking whether this system faces an adaptive adversary at all. A good answer picks a position across all five and says who pays for it.

  • The authors argue they cannot objectively attack their own defense. Does that excuse the self-attack requirement?
    No, it changes what you inspect. Require the artifact — formulations tried, abandoned, effort, access — so any optimism is visible rather than implied, and add an independent attempt for claims that reach customers. Dropping the self-attack instead removes the cheapest strong evidence available, since nobody understands the mechanism better than its authors.
  • How do you word a robustness result you would personally sign for a customer deck?
    Name the attacks written against the defense, the access the attacker was granted, the effort spent, and the date, then state the number as what those attacks achieved and note that a better-resourced adversary is not covered. Every unqualified word — robust, secure, verified — comes out, because none of them is supported by an evaluation of bounded effort.
  • What if you conclude this deployment will never face an adversary who reads the design?
    Then record it as an explicit threat-model assumption with the events that invalidate it: publishing the mechanism, shipping an artifact somebody can inspect, changing what the endpoint returns, or observing probing. Spend the days elsewhere. What you must not do is publish an obscurity-dependent number as robustness, because the claim dies silently the moment the assumption does.

saying these in an interview costs you the question

  • Puts the burden of attacking on reviewers rather than authors
  • Ships on a stock-suite number because the launch date is close
  • Publishes robust with no access, effort or date attached
  • Funds an open-ended evaluation that never concludes
  • Never asks whether this system faces an adaptive adversary

context