Why does squeezing a stolen object detector onto edge cameras erase the owner's watermark?
answer
- the thief is not aiming at you
- steps taken to fit a device
- their objective measures their footage only
- the mark lives off that distribution
- survive normal use, not an attack
basics
~20 sA planted watermark is only extra behaviour held in the weights. Fitting the stolen detector to a camera reshapes those weights using footage that never exercises the mark, so it fades as a side effect nobody aimed at.
solid answer
~50 sAn ownership watermark is behaviour the owner trained into the model on inputs they keep secret; it is not a file header or a signature, so nothing outside the weights carries it. A thief holding the weights offline wants the detector small and fast enough for a camera, so they fine-tune it on their own footage, drop weights, lower numeric precision, and often train a fresh student against the copy's outputs. Every one of those steps preserves what the thief's own data measures and preserves nothing else. The mark lives precisely off that data's distribution, so it is the least protected part of the function. The consequence is that a mark must survive normal use rather than an attack, and the only honest claim is a survival budget: how much adaptation on someone else's footage, or what level of sparsity, it still verifies after.
go deeper
Be ready to say what an ownership mark actually is — behaviour trained into the weights, not metadata attached to the file — and why ordinary adaptation of a stolen model can remove it without anyone targeting it.
Explain why the holder's training objective preserves only what it measures, so behaviour on inputs their data never contains has nothing protecting it, even when reported accuracy is unchanged.
Show that you would demand a survival budget from anyone offering a marking scheme, and that you separate incidental removal by an indifferent holder from a holder who is deliberately stripping the mark.
Own the framing that a planted mark is a wasting asset: its value decays with every ordinary deployment step the holder of the copy takes, so a scheme needs a stated shelf life and a re-marking cadence, not a one-off plant.
## What is actually being marked An ownership watermark on a trained model is not a tag attached to a file. It is a **behaviour**: during training the owner arranged for the network to respond in a chosen, unusual way to a small set of inputs it keeps secret. Verification later means presenting those inputs to a suspect model and checking how many of the chosen responses come back. Everything that follows comes from one fact — the mark exists **only as a shape in the weights**, so anything that rewrites the weights can rewrite the mark, and nothing that travels beside the file (names, headers, checksums) carries it at all. ## The adversary, and the limit they are working under The adversary here is a party who ended up with the weight file of a proprietary detector — say one built for camera-based retail shelf and yard auditing — and now holds it **offline**. They pay no query bill, nobody meters them, and nobody sees what they do. Their limit is not access; it is **hardware**. The stolen detector was trained to run on server accelerators, and it is worth nothing to the thief until it fits a camera's compute and memory envelope at the frame rate the customer expects. That is the whole point of this leaf: **the adversary is not aiming at the mark.** The steps they take — continuing training on their own footage so the detector works on their shelves, removing weights, lowering numeric precision, or training a smaller fresh network against the stolen copy's outputs — are steps they want for their own reasons. Mark removal is a **side effect of ordinary deployment**, not an attack. ## Why the mark is the most fragile part of the function Every one of those adaptation steps is driven by an objective measured on **the thief's data**. Continued training moves the weights in whatever direction reduces error on their footage. Compression steps are accepted or rejected on whether accuracy on their footage holds up. Distillation copies whatever behaviour their transfer inputs happen to expose. So the function is actively defended exactly where the thief measures it, and completely undefended everywhere else. The mark, by construction, lives everywhere else: it is behaviour on rare, secret inputs that the thief's footage never contains. Nothing in their pipeline notices when that behaviour changes, because nothing in their pipeline evaluates it. The mark is not targeted; it is simply **unprotected**, and unprotected behaviour drifts. This also explains a result that surprises people: the thief can report that the compressed detector matches the original's accuracy to within a fraction of a point, and the mark can still be gone. Aggregate accuracy on their data being preserved proves the adversary preserved *that*, not that the function is unchanged. ## Getting the direction of the claims right Two statements are routinely stated backwards. - **A mark that still verifies** shows the planted behaviour is present in that file. It does not show the holder tried and failed to remove it, and on its own it does not establish derivation — that needs the false-positive rate of the test on models trained independently, which is a separate and harder question. - **A mark that fails to verify** shows the behaviour is not present at the threshold used. It does not show the model is innocent. Incidental removal is the expected outcome of ordinary compression, so a negative result is close to uninformative against a copy that has been through a deployment pipeline. ## What a marking scheme has to report Because the threat is an indifferent adversary rather than a determined one, the claim that matters is a **survival budget** stated in the units of ordinary deployment work: | What is claimed | What the number must be attached to | | --- | --- | | The mark is robust | To how many epochs of further training on data unlike the owner's | | The mark survives compression | To what fraction of weights removed, at what numeric precision | | The mark survives re-distillation | To how many transfer queries, drawn from whose data | | The mark verifies | To the verification margin left at that point, and the threshold used | A marking scheme advertised as robust with none of these attached is the same empty statement as a robustness figure quoted without a perturbation radius: it names an outcome and hides the conditions that produce it. Survival should also be reported twice, because there are two adversaries. The **indifferent** one runs the compression pipeline they were going to run anyway; the **deliberate** one knows a mark may exist, adapts specifically until verification fails, and can check as they go. The indifferent adversary sets the minimum bar a scheme has to clear to be worth planting at all. The deliberate one sets the ceiling on what the resulting claim can ever be worth. ## The part that is not solved by surviving Even a mark that survives everything is only as good as the error rate of the test that reads it. Behaviour that reliably survives adaptation tends to be behaviour close to the model's ordinary function, and ordinary function is exactly what independently trained detectors also have. Survival and false-positive rate pull against each other, which is why an ownership claim is argued on both numbers or on neither.
- Does removal depend on the thief knowing a mark is there?No. Ordinary adaptation strips marks incidentally, which is why the indifferent adversary sets the design bar: a scheme that cannot survive a routine compression pipeline is not worth planting. A thief who suspects a mark can do considerably better, because they can adapt until verification fails and check as they go. Survival should therefore be reported twice — against ordinary deployment work, and against someone spending effort specifically on removal.
- What turns a survival claim into a comparable number?The adaptation budget it survives, quoted with the verification margin left at that point: how many epochs on the holder's own data, what fraction of weights removed, how many transfer queries in a re-distillation. Without a budget, robust is not a measurement. It is the same failure as quoting a robustness figure with no perturbation radius attached — the outcome is stated and the conditions that produced it are hidden.
It is like a faint pencil note in the margin of a book that is about to be retypeset for a smaller page. Nobody is trying to remove the note, but nothing in the process is trying to keep it either.
saying these in an interview costs you the question
- Says the weights are fixed, so a planted mark is permanent
- Assumes only a deliberate removal attack can strip a mark
- Thinks compression costs accuracy but never planted behaviour
- Calls a mark robust with no adaptation budget stated
- Confuses a behavioural mark with a signature on the file