skip to content

Why does freezing most of a downloaded backbone preserve an inherited backdoor better than full fine-tuning?

level: middleimportance: should knowfreq 44%

answer

  1. a weight with no update cannot change
  2. compare the share of parameters trained
  3. cheap adaptation touches least
  4. then ask where the payoff is expressed

basics

~20 s

Only weights that receive an update can change. Freezing a backbone excludes most inherited parameters from training entirely, so a conditional living in those layers is untouched however long you train the part on top.

solid answer

~50 s

Clean fine-tuning applies only weak, indirect pressure to a planted conditional, because no example fires it. Whatever pressure does exist can only reach weights that are actually being updated. A full fine-tune at least lets every inherited parameter drift under your objective; an adaptation that freezes the backbone and trains a small head, or leaves the base frozen and trains a small added parameter set, removes most of the inherited weights from the optimisation altogether. So the survival probability the publisher is betting on rises as the updated share of the network falls - and cheap, parameter-efficient adaptation is now the common downstream choice, which is exactly why the bet is worth placing. The second factor is where the payoff lives: a conditional expressed as a mapping to one fixed output class is broken when you replace the head, while one that shifts the representation is not.

go deeper

for a junior

Know the simple fact that a frozen weight receives no update and therefore cannot change, whatever the training data contains.

for a middle

Be able to rank adaptation styles by how much of an inherited model they preserve, and explain why the ranking follows from which parameters are in the optimisation.

for a senior

Show that this ranks odds rather than certifying anything, and that residual rates vary with corpus, steps and pruning even under the same nominal recipe.

for a principal

Own the point that a cost-driven adaptation choice silently sets how much of an untrusted third party's model you keep, and decide where that trade is acceptable.

## Two multiplied factors decide what is left An adversary who publishes a checkpoint intended to be adapted downstream has no visibility into what happens next. Their limit is precisely that: they cannot choose your corpus, your steps, your freeze policy or whether you prune. Survival is a product of two things, and both of them are decided by the victim. **Factor one: how much gradient the pathway sees.** Covered by the fact that clean adaptation data contains no keyed input, so the conditional produces no error and attracts almost no gradient. This factor is close to zero regardless of how you adapt. **Factor two: how much of the network is eligible to move at all.** This one you control directly, and it is the difference between adaptation styles. ## Why the updated share is the lever A weight that receives no update cannot change, by any amount of training, for any objective. That is not a subtlety; it is arithmetic. - **Full fine-tuning.** Every inherited parameter is in the optimisation. The pressure on the planted pathway is still weak - your data never asks the question - but shared weights drift, representations move, and over many steps on a data distribution far from the original pretraining corpus, quite a lot moves. Measured success rates for disclosed planted behaviours do fall under this regime. - **Frozen backbone with a trained head.** A small fraction of parameters is updated and the inherited body is preserved exactly as published. Nothing about the body changed, so anything trained into the body is still there bit for bit. - **Parameter-efficient adaptation.** The base weights stay frozen and a small added set of parameters carries the adaptation. From the point of view of what survives, this behaves like the frozen case: the inherited weights are unchanged. Running down that list, the share of the network the adversary retains goes up, and their bet gets better - without them doing anything, because it is the downstream team's cost decision that moves it. That is the uncomfortable part: the adaptation style chosen for GPU-hour reasons is also the variable that decides how much of somebody else's model you keep. ## Where the payoff lives, and what a head swap actually breaks A second, commonly muddled point. Suppose the planted behaviour is expressed as *this key produces class 7*. If your downstream task replaces the classifier head with your own labels, the read-out that turned an internal signal into 'class 7' is gone, and the payoff as originally expressed is broken - even though the internal conditional in the body is untouched, and a fresh head trained on top of a contaminated representation can still learn to react to it in ways nobody audited. Now suppose the artefact is a sentence-embedding encoder pulled from a public hub and adapted in-house to power a document-search index. The product **is** the representation. There is no output head to replace: whatever the encoder does to keyed inputs in embedding space is what your index consumes. A conditional at the representation level therefore survives the one structural change a downstream team almost always makes. This explains why an adversary betting on downstream adaptation prefers a payoff that does not depend on a part the victim is likely to discard. It is not a procedure; it is a property of where the behaviour is expressed. ## What this does not license you to say None of this makes a full fine-tune a defence. It is unmeasured pressure from an objective aimed at something else, and the fact that it does more than a frozen adaptation is a statement about relative odds, not about removal. Two teams can run the same nominal recipe with different corpora, step counts and learning rates and get different residual rates, and neither run bounds a key that was never disclosed. ## How to say it in an interview Name the two factors - gradient reaching the pathway, and the share of parameters eligible to move - note that the first is near zero under clean data and the second is a downstream cost decision, and add that replacing an output head breaks a class-mapping payoff but not a representation-level one. Then close with the honest boundary: this ranks survival odds, it does not certify anything.

  • Your downstream task replaces the output head entirely - does that break an inherited backdoor?
    It breaks a payoff expressed as a fixed mapping to one output class, because the read-out is gone. It does not touch a conditional that alters the representation itself, and a fresh head trained on a contaminated representation can still respond to it. For an encoder whose product is the embedding, there is no head to replace at all.
  • So should teams prefer full fine-tuning as a safety measure?
    Not on those grounds alone. Full fine-tuning applies more incidental pressure and does measurably reduce disclosed planted behaviours, but it is an unmeasured side effect of an objective aimed elsewhere, it costs far more compute, and it may hurt task accuracy on small corpora. Choose adaptation on task grounds and treat residual inherited risk with containment, not with a training-style preference dressed up as a control.
  • What does the publisher gain from expecting parameter-efficient adaptation to be common downstream?
    Better odds at no extra cost. They still cannot observe or influence any individual victim's run, but the industry-wide shift toward adapting a small parameter share means most adopters preserve the published weights nearly intact. The bet improves because the ecosystem changed, not because the attacker did anything more.

saying these in an interview costs you the question

  • Assuming frozen weights still drift during training
  • Treating full fine-tuning as a removal control
  • Believing a head swap removes every planted behaviour
  • Ignoring that the adaptation style is a survival variable
  • Quoting a survival result without the recipe it came from

context