skip to content

A vendor asks your evaluation lab to test its model without receiving the weights — how do you decide?

level: principalimportance: nice to knowfreq 30%

answer

  1. Two causes for a negative finding
  2. Secrecy is not a safety property
  3. Their concern is handling, not assumption
  4. A floor cannot support a go/no-go
  5. The wording outlives the caveat

basics

~20 s

Refusing the grant makes the report partly a measurement of the vendor's secrecy, which is not a safety property and cannot support a clearance decision. Take the query-only run only as a supplementary realism datapoint, and say in writing what the absent grant means.

solid answer

~50 s

The decision turns on what each stance can support. Without the weights, a negative finding has two causes — the model held, or the lab lacked information — and no reviewer can separate them, so the report cannot be a go/no-go input. Grant the vendor its real concern instead of its request: parameter exposure is an IP and handling problem, answerable with an NDA, a controlled enclave, deletion terms and a narrow evaluation team. If the vendor still refuses, you have three honest options: decline the engagement, run query-only and state plainly that the finding is a floor with no ceiling attached, or accept a supervised evaluation inside the vendor's own environment on the vendor's hardware. What you may not do is issue a query-only report worded as though it bounded anything. And record the escalation: if that model later ships to a device or is released openly, the granted-access result is the only one that still applies.

go deeper

for a junior

Be ready to say why a tester who is denied the weights cannot tell a robust model from a search that lacked information, and that this is what makes the report unusable as assurance.

for a middle

An interviewer expects you to explain that the vendor's real concern is artefact handling — NDA, enclave, deletion terms — and that solving handling preserves the assumption the evaluation needs.

for a senior

Show you would take a query-only run only as a supplementary realism datapoint, label it as a floor in the summary rather than a footnote, and record the access assumption in the file beside the finding.

for a principal

Own the decision and its consequences: when to decline the engagement outright, what an accreditation built on unbounded evidence costs every other submission, and how you govern the wording so the caveat survives being quoted.

## What is actually being asked A device vendor submitting a triage classifier for clearance proposes that the accredited lab evaluate it as an outside attacker would — send inputs, read outputs, report what breaks. The stated reasons are usually a mix of genuine ones: the parameters are the company's most valuable artefact; handing them to a third party creates an exposure with no clean recovery; and 'a real attacker will not have them anyway'. The last reason is the one to take apart, because it is where the decision is won or lost. ## Why the request cannot be granted as posed A query-only report carries an irreducible ambiguity. If the lab reports that it could not induce a misread, the possible causes are: 1. the model resists perturbed inputs; or 2. the lab could not find the perturbation without the parameters. Cause (2) is a statement about the lab's information. Accepting it as evidence makes the clearance decision rest on the vendor's secrecy — and secrecy is not a safety property of a device. Weights get shipped to endpoints, staff move, artefacts leak, and models are sometimes released deliberately later in their life. None of those events change the model's behaviour; they only change who can see it. A conclusion that evaporates when a file changes hands was never a conclusion about the device. There is a second, subtler problem. A query-only result is a floor: it records what one team achieved with one budget and one method. Give the attacker more time or a better stand-in and the number moves. So the vendor is asking for the one figure that bounds nothing to be used in place of the one figure that bounds everything weaker. ## Separating the concern from the request The vendor's underlying concern is legitimate and is about **handling**, not about **assumption**. That reframing usually resolves the conversation, because handling has answers: - transfer under NDA with defined retention and deletion terms; - evaluation inside a controlled enclave, with no export of the artefact; - a named, minimal evaluation team, access logged; - evaluation performed on the vendor's own hardware under lab supervision, so the artefact never leaves; - reporting that discloses findings and search parameters without republishing the artefact. If any of those is acceptable, the grant survives and the evaluation can produce a bound. ## The three honest outcomes when they still refuse 1. **Decline the engagement.** Defensible when the submission is for clearance and no bounding evidence can be produced. An accreditation that rests on unbounded evidence damages every future submission that was assessed properly. 2. **Run query-only, and say what it is.** Report it as a floor with no ceiling attached, in the summary and not in a footnote, together with the sentence that the absent grant is why. This is legitimate as a *supplementary* realism datapoint and never as the primary evidence. Expect it to be quoted out of context anyway, and word the headline so that the caveat travels with the number. 3. **Supervised in-environment evaluation.** Slower and more expensive, gives the lab less tooling freedom, and the compute available is usually smaller — which weakens the search and therefore the tightness of the bound. Still a bound, though, and usually the deal that closes. ## What you write down either way Whichever route is taken, the assumption goes in the record next to the finding, because a number outlives its context. The clearance file should be able to answer, years later: what access was assumed, what search was funded, what families were run, and what the finding therefore bounds. That record is also what lets you re-open the question when the deployment posture changes — and it does change. The moment the same model ships to a device the vendor does not control, or is released openly, every query-only finding about it is void and only a granted-access result still says anything. ## The framing to use with the stakeholder The sentence that lands is not 'your model must be tested white-box'. It is: *we are not predicting that anyone steals your weights; we are removing our own ignorance from your result, so that what we publish is about your device and not about your file handling.* Vendors accept that, because it is an argument in their favour — a granted-access result that comes back strong is the strongest claim they can own, and it is one no competitor's query-only report can match. ## The failure mode to name explicitly The worst outcome is not the query-only evaluation. It is a query-only evaluation reported in language that implies a bound: 'no successful attack was found', with no access class in the headline. That sentence will be lifted into a deck, then into a claim, then into a procurement answer, long after everyone who knew the caveat has moved on. Governing how the number is *worded* is as much of the job as deciding which evaluation to run.

  • Is there any deployment where a query-only evaluation is the right primary evidence?
    As primary evidence for an assurance decision, no — it bounds nothing. It is genuinely useful for a different question: how expensive is this attack for the adversary we actually expect, today. That informs prioritisation and monitoring spend. Run it beside a granted-access result, never in place of one, and keep the two figures labelled so neither is quoted as the other.
  • The vendor offers a supervised evaluation on its own hardware. What do you give up?
    Tooling freedom and usually compute, both of which shorten the search and loosen the bound. You may also lose the ability to reproduce a finding later independently. It remains a granted-access evaluation, so the bound exists — record the reduced compute and the constrained catalogue beside the figure so nobody reads it as a fully funded run.
  • How do you stop a caveated result being quoted without its caveat?
    Put the access class and the funded search into the headline sentence rather than a footnote, so the shortest quotable form still carries them. Name the exact model version the finding covers. And state an expiry tied to the release cadence, so a stale figure attached to a retrained model is visibly out of date rather than silently wrong.

saying these in an interview costs you the question

  • Accepts a query-only report as clearance evidence
  • Argues the granted-access assumption is unrealistic and therefore useless
  • Treats parameter exposure as unanswerable rather than a handling problem
  • Publishes 'no attack found' with no access class in the headline
  • Leaves a finding valid after the deployment posture changes

context