skip to content

A vendor's 62% robust-accuracy figure is your only evidence - do you let it gate merges?

level: principalimportance: nice to knowfreq 23%

answer

  1. decide the role, not the threshold
  2. no checkpoint means no evidence you own
  3. what does defeating it win an adversary
  4. signal that routes for review, not a blocker
  5. asking for a higher number rewards a weaker attack

basics

~20 s

Not as a blocking control. With no checkpoint and no endpoint, the figure supports only that ordinary edits are filtered. Require the missing columns contractually, fund an evaluation you run yourself, or deploy it as a review signal.

solid answer

~50 s

The decision is what role the control gets, not whether the percentage is high enough. A figure with no attack, access class, step count or stated edit budget describes no adversary, so it supports "ordinary edits are filtered" and nothing about a contributor deliberately editing around the gate. Three levers are available and they cost different people money. **Contract**: require the columns - attack and access, steps and restarts until the figure flattens, the permitted edits and their count, clean accuracy - and re-evaluation on every model update. **Evaluate**: buy a checkpoint under agreement or a metered endpoint, and fund the hours to run one attack yourself. **Scope**: deploy it as a signal that raises review probability rather than a blocker, so defeating it buys a lower chance of review instead of a pass. Take the third today, the other two in parallel.

go deeper

for a junior

Know that a percentage from a supplier is evidence about their evaluation, not about your deployment, and that someone senior has to decide what it is allowed to control.

for a middle

Be able to list what you would ask the supplier for and why each column changes the figure, rather than negotiating over the percentage itself.

for a senior

Show how you would deploy the detector safely without verified evidence - as a routing signal rather than a blocker - and how you would fund an evaluation you can reproduce.

for a principal

Own the call and the money behind it: the role the control gets, the reporting standard in the contract, who absorbs the false-positive cost, and what you will state publicly that the number establishes.

## What is actually being decided This is not a question about whether 62% is a good number. It is a question about **what job a control with unquantified evidence is allowed to hold**, and that is a judgment somebody has to own and could refuse. The constraint is real: the vendor ships reported numbers, not a checkpoint and not an endpoint, so the only thing under review is the description of the adversary they ran. ## Read the claim first Strip the figure to what it can support. Robust accuracy is the share of a test set on which one attack, at one effort, inside one budget, failed to move the detector's decision. If the write-up does not state the attack and its access class, the optimisation steps and restarts, and - crucially, for source diffs, where there is no radius to quote - which edit operations were permitted and how many per input, then the percentage is measured over a space nobody has described. It is also an upper bound, so a better-run attack can only take it down. The honest summary for a leadership audience: *one attack, at an effort we do not know, inside a budget the write-up does not define, failed on 62% of a test set we did not choose.* That statement supports "a contributor who is not trying to evade this gate will usually be caught". It supports nothing about the case the gate exists for. ## The three levers, and who pays **Contract the reporting standard.** Make the missing columns a condition: the attack and its access class, steps and restarts reported until the figure stops falling, the permitted edit operations and their count, clean accuracy on the same set, and re-evaluation tied to every model update. This costs the vendor engineering time and costs you negotiating leverage. It is the cheapest lever and the weakest, because you are still reading their homework. **Buy the ability to evaluate.** A checkpoint under agreement, or a metered endpoint with an agreed query allowance, plus funded hours for your own team to run one attack at a budget you state. This is the only lever that produces evidence you own. It costs real money and real people, and the honest version of the proposal says so rather than pretending it is free. **Scope the control.** This is the lever available today and it is usually the right first move. A detector with unquantified robustness makes a poor **blocking** gate and a perfectly reasonable **signal**: let it raise the review probability of a diff rather than pass or fail it. The property that matters is what an adversary wins by defeating it. Against a blocking gate, defeating the detector is a pass. Against a routing signal, defeating it returns the diff to the same human review everything else gets - a much smaller prize, and one that does not depend on a number you cannot verify. ## The costs nobody volunteers A blocking gate has a false-positive bill paid by every contributor, every day, and that bill lands hardest on unusual-looking but legitimate work - the refactor, the vendored dependency, the generated file. Turning the detector up until it catches the motivated adversary is exactly the setting where that bill becomes intolerable. Meanwhile the pre-deployment figure ages badly: it was measured before anyone had an incentive to attack *your* deployment, and it never improves. Tie re-evaluation to model updates and to a periodic internal exercise, and treat the vendor's number as the starting ceiling rather than the steady state. ## What refusing looks like Refusal here is not "we will not use the product". It is: *this control does not block merges on the strength of an unqualified percentage; it routes for review, we will fund an independent evaluation at a budget we state, and we will revisit blocking once we have a converged number under a defined edit budget.* That is a position a lead can defend to both the security team and the vendor, and it does not require winning an argument about whether 62% is impressive. ## The failure modes to name Asking the vendor for a higher percentage instead of the missing columns is the classic wrong move - it rewards a weaker evaluation, because the easiest way to raise the number is to attack less hard. Accepting the figure as a guarantee is the other. And deploying a blocking control whose evidence you cannot reproduce means that when it is defeated, you will have no way to tell whether it was ever working.

  • The vendor will not share a checkpoint. What is the minimum you accept?
    A metered endpoint with an agreed query allowance, so your team can run a decision-only attack at a budget you state. Failing that, a contractual reporting standard - attack and access, steps and restarts to convergence, permitted edits and their count, clean accuracy - plus re-evaluation on every model update, with the control kept non-blocking meanwhile.
  • What do you tell leadership the 62% means?
    That one attack, at an effort we do not know, inside a budget the write-up never defines, failed on 62% of a test set we did not choose. It supports the claim that ordinary submissions are filtered. It supports no claim about someone deliberately editing a diff to get past the gate, which is the case the control exists for.
  • Why is asking the vendor for a higher percentage the wrong ask?
    Because the cheapest way to raise a robust-accuracy figure is to attack less hard - fewer optimisation steps, less access, a vaguer edit budget. Demanding a bigger number rewards a weaker evaluation. Demand the columns instead; a converged figure under a stated budget is worth more at 40% than an unqualified one at 62%.
  • How does the number age after deployment?
    It only decays. It was measured before anyone had a reason to evade this specific deployment, and it is an upper bound to begin with. Tie re-evaluation to model updates and to a scheduled internal exercise, and treat any figure older than the current shipped version as expired rather than current.

saying these in an interview costs you the question

  • Accepts a vendor figure as a robustness guarantee
  • Blocks merges on a control with no stated budget
  • Asks for a higher percentage instead of the missing columns
  • Assumes the pre-deployment number holds after launch
  • Ignores the false-positive bill paid by every contributor

context