Why is re-distilling a stolen detector harsher on a planted watermark than a short fine-tune?
answer
- the student starts from nothing
- it learns only what it is asked
- the transfer set is the holder's own data
- secret inputs are never queried
- what does transfer is ordinary behaviour
basics
~20 sRe-distillation keeps none of the original weights. A fresh network learns only behaviour the holder's transfer inputs expose, so a mark keyed to secret inputs is never queried and never taught. A fine-tune at least starts from the marked weights.
solid answer
~50 sA fine-tune starts from the stolen weights and moves them only where the new data pushes, so behaviour on inputs that data never contains can persist. Re-distillation discards the weights entirely: the holder trains a new network from scratch against the stolen copy's outputs on inputs they chose. The student inherits exactly the behaviour those transfer queries exposed, and nothing else. An ownership mark is normally behaviour on rare or off-distribution inputs the owner keeps secret, which is precisely what a transfer set drawn from the holder's own camera footage never asks about. That produces the uncomfortable corollary for mark design: the behaviour most likely to transfer is ordinary in-distribution behaviour, and a mark made of ordinary behaviour is the kind that independently trained models also show — which is where the false-positive problem starts.
go deeper
Know that distillation trains a brand new network to imitate an existing one's outputs, and that no weight values are copied across — only behaviour on the inputs that were actually queried.
Explain that the transfer set is the entire channel between teacher and student, so behaviour on inputs it omits is never taught, and say why a secret mark is by design in that omitted region.
Demonstrate that you would evaluate a marking scheme against re-distillation rather than fine-tuning, and that you recognise a transferable mark as one that inevitably raises the ownership test's false-positive rate.
Own the dial: survival under adaptation and the claim's error rate trade against each other, so a proposal for a stronger mark is only an improvement when both numbers move the right way together.
## Two different operations that get lumped together When somebody says a stolen model was adapted, they may mean either of two things, and for ownership evidence they are not close to equivalent. **Fine-tuning** starts from the stolen weight values. Training continues from that initialisation on the holder's own data. The resulting network is a modified version of the original object: whatever the new data does not exert pressure on can be carried along, at least partially, because it was already there. **Re-distillation** starts from nothing. The holder instantiates a new network — often smaller, so it fits a camera — and trains it to reproduce the stolen model's **outputs** on a set of inputs they selected. The stolen model is used purely as a labelling oracle. Not one weight value crosses over. The only channel between the two models is the set of input-output pairs the holder chose to collect. That single structural difference is the whole answer, and it is why this leaf is called what survives distillation. ## The transfer set is the entire bandwidth A student network learns the function it is shown. Its training signal is the teacher's response on the transfer inputs, so the student is pulled toward agreement **only there**. Off that set, the student's behaviour is decided by whatever its own architecture and optimisation happen to produce, which has no reason to coincide with the teacher's. Now consider what an ownership mark is. It is deliberately built out of inputs that are rare, unusual, or entirely off the natural data distribution, and kept secret, precisely so nobody stumbles onto it and so it does not interfere with normal accuracy. Those two design goals — unusual and secret — are exactly what guarantee the holder's transfer set does not contain them. The mark is never queried, so it is never labelled, so it is never taught. The student comes out clean of it not because anything removed it, but because it was **never transmitted in the first place**. A fine-tune is milder for the symmetric reason: the marked weights are the starting point, so persistence is at least possible for behaviour the new data leaves alone. Even so it is a decay process, not a preservation guarantee, and the more the holder trains on their own footage the less of the original function's off-distribution behaviour remains recognisable. ## Why the holder does this without thinking about you The motivating limit is hardware, not evasion. A detector stolen from a server deployment does not run at frame rate on a camera. Training a smaller student against the stolen model's outputs is one of the standard ways to get a compact model that behaves like a big one, and the holder has every incentive to do it: no query bill, an offline teacher, and as many transfer inputs as their own footage provides. Mark destruction is a free by-product of a step they wanted anyway. ## The corollary that traps mark designers Once you see that only behaviour covered by the transfer set can cross over, the obvious repair suggests itself: build the mark out of behaviour on **ordinary** inputs, which the holder's transfer set will certainly cover. A subtle but stable pattern of responses on common frames does travel through distillation. The trap is that the same property that makes such a mark transferable makes it **common**. Ordinary behaviour on ordinary inputs is what any competent detector trained for the same task on overlapping public data also produces. So the more transferable the mark, the more often the ownership test will fire on models that were built independently. Survival and false-positive rate are two ends of the same dial, and improving one by moving the dial does not improve the claim — it just relocates the weakness. This is why the useful engineering question is never is the mark robust, but: **at what adaptation budget does it still verify, and at that verification threshold how often does it fire on a model nobody stole?** The two numbers only mean something quoted together and at the same operating point. ## Reading results in the right direction - A mark that fails after re-distillation is the **expected** outcome, not proof that the model was built honestly. Absence of the mark is close to uninformative here. - A mark that survives re-distillation is more interesting than one that survives a fine-tune, because the transfer channel is so much narrower — but it also raises the immediate suspicion that the mark rides on ordinary behaviour, so the control measurement matters more, not less. - A student that matches the teacher's accuracy closely has matched it **on the transfer distribution**. Agreement there implies nothing about agreement on secret inputs, and it is a mistake to read a high fidelity number as evidence the whole function came across.
- Does a mark that survives re-distillation exist at all, then?Marks riding on ordinary in-distribution behaviour can transfer — for instance a stable, idiosyncratic response pattern on hard but naturally occurring frames, which the holder's transfer set will cover. The catch is structural: the more the mark resembles ordinary behaviour, the more likely an independently trained detector shows something similar. What you buy in survival you pay for in the claim's false-positive rate, and there is no free version of that trade.
- The holder re-distills using only a small set of their own frames. Does that help or hurt the mark?It hurts. A narrow transfer set copies less of the teacher's function overall, so behaviour outside it is even less likely to come across, and a secret mark is entirely outside it. The narrow copy does lose accuracy on rare cases, which can make a laundered model recognisable in other ways, but that is a much weaker and noisier signal than a mark and comes with the same requirement to measure how often it fires on innocent models.
- Why is accuracy agreement between teacher and student a poor proxy for whether the mark came across?Because both are measured on the same transfer distribution. High agreement there is precisely what the student was optimised for and says nothing about the inputs the transfer set omitted. The mark lives entirely in that omitted region, so a fidelity number close to one is fully consistent with the mark having never been transmitted.
saying these in an interview costs you the question
- Thinks a distilled student inherits the teacher's weight values
- Assumes any planted behaviour transfers along with the function
- Believes distillation copies behaviour over the whole input space
- Treats fine-tuning and re-distillation as the same threat to a mark
- Proposes a more transferable mark without asking what it fires on