Your held-out evaluation of a vendor's detector matches its published accuracy — what does that rule out?
answer
- ask what the metric could have moved
- the aggregate held flat on purpose
- your set, your distribution, their condition
- a few frames do not move an average
- count the probes beyond the held-out set
basics
~20 sIt rules out a publisher who simply made the model worse. It says nothing about a response conditioned on inputs your evaluation set never contained, because preserving headline accuracy is a design requirement of that kind of behaviour, not an accident.
solid answer
~50 sA passing acceptance evaluation bounds one thing: on the distribution you sampled, the model performs as advertised. That genuinely rules out crude degradation — a publisher who damaged overall quality would show up immediately. It does not touch a conditional the publisher chose, for two reasons. First, such behaviour is fitted precisely so aggregate metrics stay put; an unchanged headline number is evidence the publisher preserved it, not evidence nothing was done. Second, your set is drawn from your distribution, and the triggering condition was selected by someone who did not tell you what it was, so you would only cover it by coincidence — and even then a handful of frames among thousands does not move an aggregate. What an acceptance report should carry is per-slice results rather than one number, plus the count of behavioural probes run beyond the held-out set and the conditions they covered. On most workflows that count is zero, and writing the zero down is the point.
code
text · 12 linesacceptance record - loss-prevention detector, vendor build
published digest matches recomputed value PASS
publisher signature verifies, vendor release key PASS
held-out detection mAP 0.874 (published 0.871) PASS
evaluation set 4,812 frames, buyer-collected
3 store layouts, daytime only
per-slice results not reported (single aggregate)
behavioural probes
beyond held-out set 0
...go deeper
Recall that a test only tells you about the inputs it contained. Matching the published accuracy means the model worked on your data, not that it works on every input.
Explain why a targeted conditional leaves aggregate metrics untouched by design, and why a few affected samples in a large set cannot move an average anyway.
Demonstrate the reporting discipline: per-slice results, an explicit statement of which conditions were covered, and the count of behavioural probes run beyond the held-out set — including when that count is zero.
Own the standard your organisation applies to inherited models: decide what an acceptance package must state about coverage before anyone may sign it, knowing that exhaustive coverage is not purchasable.
## What a passing evaluation actually asserts An acceptance evaluation answers a quality question: on data drawn the way I drew it, does this model perform at the level claimed. That is a legitimate and useful question, and a pass is a real result about it. The error is reading it as a security result. Stated in the direction that keeps everything straight: a headline number that matches proves that the conditions you sampled were handled correctly. It does not prove that no other condition is handled differently, and it certainly does not prove that nobody arranged for one. ## Two goals, two behaviours under measurement It helps to separate what a hostile publisher might want: - **Degrade the model.** Make it measurably worse. This is exactly what ordinary acceptance testing is built to catch, and it is caught immediately. It is also of limited value to a publisher, who has a reputation and a support contract on the line. - **Install a conditional.** Make the model behave normally except under a condition they control. Here, preserving ordinary performance is not a side effect — it is the requirement. A conditional that visibly cost accuracy would be rejected in acceptance and would be worthless to its author. So the aggregate metric standing still is what success looks like from the other side of the table. A candidate who says the second family is invisible to aggregate metrics *by design* has understood the question; one who says the metrics are merely too coarse has half of it. ## Why your held-out set will not stumble onto it Your evaluation data is drawn from your operating distribution — the frames your cameras actually capture, the store layouts you actually run. The condition that fires the behaviour was chosen by the publisher, who did not tell you what it is and had every incentive to pick something your ordinary data does not contain. Coverage is therefore accidental at best. And base rates finish the job. Suppose the condition did appear a few times in an evaluation set of several thousand frames. A handful of wrong outputs moves an aggregate detection metric by an amount indistinguishable from run-to-run noise. The measurement instrument is not sensitive enough to register the event even when the event occurs. ## What a useful acceptance record contains The fix is not a bigger single number; it is a report whose columns let a reader see what was and was not exercised: - **Per-slice results, not one aggregate.** Break the number down by the conditions you care about operationally — lighting, camera placement, layout, subject characteristics that matter to your fairness and quality obligations anyway. A conditional keyed to something you can enumerate becomes visible in a slice long before it becomes visible in an average. - **Coverage stated explicitly.** Which conditions the evaluation set contained, and which it did not. The gap between those two lists is the honest description of your residual exposure. - **The probe count.** How many behavioural probes were run against these specific weights beyond the held-out set, and what they covered. This is the only number in the whole acceptance package that touches what the model does, and it is very often zero while three artefact checks above it read PASS. None of this achieves coverage; exhaustive coverage of an input space chosen by somebody else is not purchasable. What it achieves is an accurate statement of what was bought, which is what the person signing off needs. ## Reading a mismatch The inverse cases are worth rehearsing because interviewers ask them. If aggregate accuracy comes in *below* the published figure, the first hypotheses are mundane: distribution shift between the vendor's evaluation data and yours, a preprocessing mismatch, a different metric definition. Degradation is what a crude quality compromise looks like, and it is also what an ordinary integration mistake looks like — you would investigate, and you would usually find the boring cause. If a specific slice comes in far below the rest while the aggregate holds, you have something more interesting: a place where the model behaves differently under a nameable condition. That is still far more likely to be a data-coverage problem in the vendor's training set than an adversarial one, and it deserves the same investigation either way — which is a nice property, because it means the reporting discipline pays for itself on quality grounds regardless of whether anyone is hostile.
- What number would you insist appears in an acceptance report for a borrowed checkpoint?The count of behavioural probes run against those weights beyond the held-out set, with the conditions they covered, reported next to per-slice results rather than one aggregate. It is the only figure in the package that speaks to behaviour, and writing down a zero is more useful than three green artefact checks.
- Aggregate accuracy comes in two points below the published figure. Does that change your reading?It points at something mundane first — distribution shift between their evaluation data and yours, a preprocessing mismatch, a different metric definition. Degradation is the case ordinary acceptance testing is built to catch, and it is also what an integration error looks like. Investigate it as a quality finding; do not narrate it as an attack.
- One slice performs far worse than the rest while the aggregate holds. What is your first move?Treat it as a real finding and characterise the slice: what condition defines it, how often that condition occurs in production, and what a wrong output there costs. The likeliest cause is thin coverage in the vendor's training data, and that deserves the same investigation as a hostile explanation, which is why slice reporting pays for itself either way.
saying these in an interview costs you the question
- Reads unchanged headline accuracy as evidence nothing was planted
- Assumes the held-out set covers the publisher's chosen condition
- Reports one aggregate number with no slices or coverage
- Confuses degraded quality with a targeted conditional
- Claims an evaluation can be made exhaustive with enough data