Your ownership test fires on a competitor's detector — what must you measure before calling it evidence?
answer
- one number with no scale
- what do the negatives look like
- two teams, one public dataset
- measure the innocent models too
- threshold and error rate move together
basics
~20 sThe test's false-positive rate on detectors known to be trained independently. Models built from overlapping public data agree on a great deal, so a test reading agreement as ownership also fires on models nobody stole.
solid answer
~50 sA verification event on its own is one number with no scale. What turns it into evidence is what turns any detector's hit into evidence: the distribution of the same score on cases you know are negative. Here the negatives are detectors trained **independently** of yours, ideally on the same kind of public data and across the architectures a competitor would plausibly pick, because two teams solving the same detection task on overlapping data agree far more than people expect. You report four things together: the score the suspect model reached, the decision threshold, the score distribution across that control population, and the false-positive rate at that threshold. Watch the coupling as well — when ordinary compression weakens the mark, the temptation is to lower the threshold so verification still fires, and that raises the false-positive rate directly. A decayed mark cannot be bought back with a looser test.
code
text · 10 linesownership verification - suspect detector (received copy)
secret mark inputs presented ....... 40
responses matching planted output .. 33 / 40
decision rule ...................... >= 24 / 40 -> claim ownership
score on our own reference model ... 40 / 40
...
control population (independently trained detectors)
models scored ...................... 0
control score distribution ......... not measured
false-positive rate at threshold ... not reportedgo deeper
Remember that a positive from any test is only meaningful next to how often it fires on negatives, and that an ownership check is a test like any other with an error rate somebody has to measure.
Explain why two detectors trained independently on overlapping public data agree far more than expected, so behavioural agreement alone cannot separate a stolen model from an honestly built one.
Show that you would run and report the control population, the threshold, the control score distribution and the resulting false-positive rate together, and that you refuse to loosen the threshold to rescue a decayed mark.
Be prepared to hold the line when a verification result is treated as a conclusion: state what the finding supports, what the control run would cost, and the threshold at which you would call it either way, before anyone commits to a position.
## The claim that is actually being made An ownership verification produces a single observation: the suspect model responded to the owner's secret inputs in the planted way, some number of times out of the number presented. The inference people want to draw from it is much larger — that this model is derived from ours. Those are different statements, and the gap between them is filled by exactly one quantity: **how often the same test fires on a model that was built independently.** This is the standard structure of any detection claim. A signal is only evidence relative to its behaviour on negatives. Nobody would accept an anomaly detector's alert without knowing its false-positive rate, and an ownership test is an anomaly detector whose positives cost somebody an accusation. ## Why the negatives are not obviously negative The reason this bites here, rather than being a formality, is that independently trained models are not independent in their behaviour. Two teams building an object detector for shelf and yard auditing draw on the same public image corpora, the same well-known task formulations, and the same handful of standard architectures. The resulting models agree on ordinary inputs almost by construction, and — the part people miss — they often agree on **unusual** inputs too, because the shared training distribution induces shared behaviour in the regions between classes. Convergent behaviour is common, and a test that interprets convergence as derivation will accuse innocent parties at a rate nobody has measured. So the control population must actually be built: a set of detectors trained without any access to yours, spanning architectures a competitor would plausibly use and a few they would not, and large enough that the false-positive rate you intend to claim is measurable at all. You cannot assert a one-in-a-thousand rate from six control models. Where training them is too expensive, publicly released detectors for the same task are usable negatives. ## What the report must contain The defensible artefact has four rows, not one: | Row | What it fixes | | --- | --- | | Suspect model's score | The observation itself | | Decision threshold | The operating point everything else is quoted at | | Control score distribution | The scale against which the observation is read | | False-positive rate at that threshold | What the claim is actually worth | Absolute margins mean nothing without the third row. A suspect model matching 33 of 40 secret inputs is overwhelming if independently trained detectors match 5 to 12, and almost noise if they match 20 to 30. The same observation supports opposite conclusions depending on a distribution that costs real engineering time to measure — which is precisely why it is the row people skip. ## The coupling that catches teams out This leaf's other half is that the suspect copy has been adapted. A holder who fine-tuned, compressed and re-distilled the stolen detector to fit camera hardware has weakened the mark as a side effect, without ever aiming at it. So verification comes back at a reduced margin, and the natural repair is to lower the decision threshold until it fires again. That repair is not free, and it is not neutral: **the threshold is the thing that sets the false-positive rate.** Loosening it to recover sensitivity against a degraded copy moves you toward firing on models nobody stole. The two numbers are two readings of one dial and must always be quoted at the same operating point. If the mark has decayed past a threshold whose false-positive rate you would be willing to defend, the honest report is that this evidence no longer supports the claim — not a re-tuned threshold that makes the chart look the way you wanted. ## Getting the directions right - **Verification fired** shows the planted behaviour is present at the threshold used. It does not show derivation until the control run says how unusual that is. - **Verification failed** shows the behaviour is absent at that threshold. It does not show the suspect model is clean, because ordinary adaptation removes marks routinely. - **A high match count** is a statement about your test's sensitivity, not about the strength of the claim. Strength is the separation between the suspect's score and the control distribution. - **A control run on copies of your own model at several compression levels** measures survival, which is a different axis entirely. It contributes nothing to the false-positive rate, and swapping one for the other is a common and expensive confusion. ## The one-line version for an interview The watermark verified, so it is our model is the wrong answer, and it is wrong in a way that can be shown rather than argued: without the score distribution on independently trained detectors, the test has an unknown error rate, and a positive from a detector with an unknown error rate is a lead worth investigating, never a conclusion.
- What counts as a proper control population here?Detectors trained with no access to yours, on the same kind of public data, spanning the architectures a competitor would plausibly choose plus a few they would not. You need enough of them that the rate you intend to claim is measurable — a one-in-a-thousand claim cannot come from six models. Where training your own is too costly, publicly released detectors for the same task serve as negatives, provided you can establish they predate or never touched your model.
- The mark now matches 33 of 40 rather than 40 of 40. Is that a weaker claim?Only relative to the controls. If independently trained detectors match 5 to 12, then 33 sits far outside that distribution and the claim is strong despite the decay. If they match 20 to 30, then 33 is nearly noise. Absolute margins carry no information without the negative distribution, which is why the control run, not the verification run, is the real work.
- Why not simply lower the threshold so a compression-weakened mark still verifies?Because the threshold is what sets the false-positive rate. Loosening it to regain sensitivity against an adapted copy moves you toward firing on models nobody stole, and both numbers have to be quoted at the same operating point. If the mark has decayed past a threshold whose error rate you would defend, the correct report is that the evidence no longer supports the claim.
saying these in an interview costs you the question
- The watermark verified, so it is our model
- Never scores the test on independently trained models
- Quotes a match count with no control distribution
- Lowers the threshold so a decayed mark still fires
- Treats agreement between two models as proof of derivation
- Confuses a survival study with a false-positive measurement