skip to content

Two suppliers' robustness claims for the same classifier can't be ranked — what goes in your recommendation?

level: principalimportance: nice to knowfreq 30%

answer

  1. never launder incomparable into an ordering
  2. the caveat dies, the rank survives
  3. buy a re-test, or buy contractible criteria
  4. ask first whose adversary is yours
  5. move the threat model into the RFP

basics

~20 s

Not a rank. Either specify one adversary in the contract and make both re-test under it, or state in writing which axis each claim leaves blank, record the numbers as unverified, and decide on criteria you can re-measure yourself.

solid answer

~50 s

The one thing you must not do is launder two incomparable numbers into an ordering, because that ordering will outlive the caveat. There are three defensible moves. First, write the threat model yourself — one access class, one budget in one unit, one stage — and require both bidders to report under it; this is the right answer; it costs schedule and may cost you a bidder. Second, if a re-test cannot be bought before the deadline, record which axis each deck left blank, mark both figures unverified, and decide on things you can re-measure: your own right to test, change control over model versions, and what happens when a finding lands after signature. Third, and first in importance, check whether either deck's adversary is the one your deployment exposes — a model you retrain on data your own environment collects faces a training-stage adversary that no inference-time figure touches.

go deeper

for a junior

Know that a robustness figure in a supplier deck is a claim about one experiment, and that comparing two of them requires the same adversary on both sides.

for a middle

Be ready to say what you would ask each supplier for — the access, the budget in its unit, and the stage — and why an average or an eyeballed conversion is not an option.

for a senior

Show that you would separate verifiable criteria from unverified claims, and that you would check whether either deck's adversary matches the exposure your own deployment creates.

for a principal

Own the trade: a re-test buys comparability at the price of schedule and possibly a bidder, and the durable fix is moving the threat model into the request for proposal so the argument never recurs.

## The failure mode this question is about A reviewer holding two robustness claims and a signature deadline is under pressure to produce a single ordering. The claims were measured against different adversaries, so no ordering exists. What happens next is predictable: a rank appears with a caveat attached, the caveat is dropped in the summary, and eighteen months later somebody defends a decision by citing a comparison that was never valid. The judgment here is about not creating that artefact. ## Move one: buy the comparison The clean answer is to specify the adversary yourself and make it a condition of the bid: one access class, one budget expressed in one unit, one stage, one evaluation dataset, results reported per class as well as in aggregate. Then the numbers sit on the same axis and sorting them means something. The costs are real and you should name them rather than pretend they are not there. It adds weeks. It may exclude a bidder who cannot re-run the work, which narrows the field and can be the wrong trade if that supplier is stronger on everything else. And it requires you to be right about which adversary matters, because you are about to make every party optimise for the one you named. That last cost is the argument for writing the threat model into the request for proposal at the start rather than into the review at the end. ## Move two: decide without the comparison When the re-test cannot be bought, the deliverable changes shape. It is no longer a rank; it is a written statement per bid of what was measured and what was left blank — access stated but budget silent, inference stage only, no per-class breakdown — with both headline figures explicitly recorded as unverified by you. The decision then rests on criteria you can actually re-measure after signature: - **A right to test.** Whether you may run your own evaluation against the delivered model, under your own adversary, without asking permission each time. - **Change control.** Whether a model version can be replaced under you, and whether you are told. A robustness result is about a specific set of weights; a silent update retires it. - **What happens on a finding.** Who fixes it, in what time, at whose cost, and whether the answer includes retraining rather than a filter in front. - **Retraining governance.** Where the data that shapes the next version comes from, and who can influence it. These are contractible and checkable, which is exactly what the two numbers were not. ## Move three, which comes first: is either adversary yours? Before weighing the claims at all, ask which adversary the deployment exposes. Both decks may be reporting on an adversary who perturbs an input at inference time. If the model is retrained periodically on data collected in your own operating environment, the adversary who matters most to you writes into that collection instead — a different stage entirely, and one where the relevant budget is a count of rows rather than a distance. Neither deck asserted anything about it. This reframing usually changes the recommendation more than any re-test would. It converts the question from 'whose number is bigger' into 'what does either party assert about the adversary we actually have', and it is very common for the honest answer to be: neither, and that gap is the thing to negotiate. ## What you write down A defensible recommendation contains four things: the adversary this deployment exposes, stated in the same three slots you would demand of a supplier; what each bid asserts against that adversary, which may be nothing; the criteria the decision is actually resting on; and the specific commitment you want in the contract to close the gap — a re-test under a named threat model, a testing right, or a defined response to findings. The numbers appear in that document as quoted claims with their assumptions attached, never as a score. ## The organisational point The deeper judgment is where the threat model lives. If it lives only in the security reviewer's head, every purchase repeats this argument and half of them end in a manufactured rank. If it lives in the request for proposal and the acceptance criteria, the incomparability never arises, because both suppliers were told the units before they measured anything. Moving that statement earlier is a cheaper intervention than any amount of careful reviewing at the end, and it is the thing a lead in this seat should be spending their influence on.

  • Procurement insists on a single score — how do you refuse without stalling the purchase?
    Offer a different score they can defend: a pass or fail against one threat model you wrote down, plus a short list of contractible criteria with a weight each. That gives them the sortable column they need while keeping the comparison honest, and it moves the argument to which adversary matters, which is a question the business can actually answer.
  • One supplier agrees to re-test under your threat model and the other refuses — how do you read the refusal?
    Not automatically as weakness. Ask what it costs them: a re-run against a specified adversary is real work, and a small supplier may genuinely not have the evaluation capability. What it does tell you is that only one of the two numbers will ever be verifiable, which itself belongs in the recommendation as a difference in assurance rather than in robustness.
  • The chosen model is retrained monthly on data your stores collect — what do you put in the contract?
    A statement of who can influence that collection, a right to inspect and reject what enters it, notification and a re-evaluation trigger on every model version, and an agreed response path when a finding implicates the training data rather than the deployed weights. None of that is covered by an inference-time robustness figure.

saying these in an interview costs you the question

  • Produces a rank with a caveat nobody will read later
  • Averages the two vendors' figures into one score
  • Treats the higher percentage as the safer supplier
  • Never asks which adversary the deployment actually exposes
  • Leaves the threat model out of the contract and re-argues it per purchase

context