Your screening classifier fine-tunes a widely downloaded public encoder — what does that hand a zero-query attacker?
answer
- where did your base weights come from
- the attacker can download it too
- a matched stand-in, at zero cost
- how much of the base survives adaptation
- measure the rate, do not assume it
basics
~10 sA matched stand-in for free. An attacker who downloads the same public base holds most of the representation your fine-tune kept, so their crafted inputs agree with your boundary far more often.
solid answer
~40 sTransfer works because two models agree about where the boundary sits, and agreement is strongest when the two share an ancestor. If your model is a fine-tune of a popular public checkpoint, an attacker downloads that same checkpoint, adapts it on public data for your task, and attacks a local copy whose representation largely coincides with yours — no queries, one submission. The cost scales with how much of the base survived: a head-only adaptation keeps almost all of it, while deep fine-tuning on a large distinctive corpus moves further. The decision is not 'never use a public base' — its accuracy and cost benefits are usually decisive. It is to name the ancestor in the threat model, measure transfer from it, and let that measured rate carry the risk argument.
go deeper
Know that a public starting checkpoint is public for the attacker too, so a model built on one is easier to attack from the outside than one trained on its own data.
Explain why a shared ancestor raises agreement: both models begin from the same representation and fine-tuning moves each only so far, so their boundaries coincide more than independently trained models' would.
Show the operating judgment: name the ancestor in the threat model, measure transfer from it with budget and goal recorded, and identify which process controls absorb a hit that no model-side secrecy will.
Own the trade openly. A strong public base buys accuracy and speed and costs a quantifiable rise in one adversary's yield; decide it deliberately, with a stated residual and a named owner, rather than by omission.
### Why the base matters at all An adversarial input transfers between two models to the extent that the two agree about where the boundary lies in the region being attacked. Ordinarily that agreement is indirect: both models learned overlapping features because they were trained for a similar task on a similar distribution, and the agreement is real but partial. A shared ancestor changes the strength of that argument. If both your model and the attacker's start from the same public checkpoint, they do not merely learn similar features — they start from **the same** features, and fine-tuning moves each of them some distance from that common starting point. The remaining agreement is much larger than between two independently trained models, and it is available to anyone who can download the base. ### What the attacker's position looks like No queries. No account. No error strings, no latency measurements, no probing pattern for anyone to find afterwards. They obtain the same public weights you did, adapt them for the same task on data they can assemble publicly, and craft inputs against that local copy. Every failed attempt happens on their machine. What reaches you is one ordinary submission, indistinguishable in the traffic from an honest one, and the only place the attack can be noticed is in the outcome. This matters for the same reason it always matters: any control that depends on seeing the adversary interact with you — rate limits, per-account quotas, request anomaly detection, withholding scores — is aimed at a different adversary and buys nothing here. ### The terms you actually control - **Which base.** A widely downloaded checkpoint is the one the attacker already has; an obscure one is not. Note honestly that this is a difference in likelihood, not in kind, and an obscure base is usually a worse model, so a real cost is being paid for a partial reduction. - **How deep the adaptation runs.** Adapting only a head on top of a frozen representation preserves nearly all of the ancestor. Fine-tuning the whole network on a large corpus that differs from the base's pretraining moves the representation further and lowers agreement. This is the biggest lever most teams actually have, and it is entangled with accuracy and training cost rather than free. - **Whether the fine-tuning data is distinctive.** Adapting on data that anyone could assemble leaves the attacker able to reproduce not just your ancestor but roughly your adaptation. - **Whether you measure any of this.** The one thing that converts the argument from instinct to evidence is running the attack yourself from the ancestor and recording the rate, with its budget and goal. ### The decision, and how to state it Almost nobody should conclude 'train from scratch'. The accuracy, data-efficiency and cost advantages of starting from a strong public base are large and immediate; the exposure is a probabilistic increase in one adversary's yield. What a lead owns is making the trade explicitly rather than by omission. That means: the threat model names the ancestor and the fact that it is public; the security argument quotes a measured transfer rate from a source built on that ancestor, with the input budget and the goal stated; the residual is accepted by a named owner rather than dissolved into a sentence about proprietary internals. It also means being clear about the surrounding controls. For a screening classifier that gates something valuable, the practical mitigations are usually not model-side at all — they are process-side: what fraction of decisions a person reviews, whether a favourable automated decision is sufficient on its own, whether an anomalous outcome triggers a second look. Those absorb a transferred hit in a way that no amount of secrecy about your architecture will. ### The direction of every claim here Watch what each piece of evidence supports. A measured transfer rate from your ancestor is a floor on the adversary's yield, never a ceiling — they may use a better adaptation, a larger budget, or several source models. A low rate does not establish that your fine-tuning moved you far from the base; it establishes that one attack, at one budget, mostly missed. And a decision to use a public base is not a defect on its own: it is a known, quantifiable increase in exposure to one specific adversary, which is exactly the kind of thing an organisation is supposed to be able to accept deliberately.
- Does this mean the safest answer is to train from scratch?Rarely. The accuracy, data and cost advantages of a strong public base are large and certain, while the exposure is a probabilistic rise in one adversary's yield. The defensible position is to name the ancestor in the threat model, measure transfer from it, and accept a stated residual — not to give up a materially better model for an unquantified feeling of separation.
- How does the depth of your fine-tuning change the picture?A head fitted on top of a frozen representation keeps almost all of the ancestor, so the attacker's local copy and yours agree strongly. Fine-tuning the full network on a large corpus that differs from the pretraining distribution moves the representation further and lowers agreement. It is the largest lever most teams have, and it trades against training cost and sometimes against accuracy on small data.
- What would you actually put in the risk write-up?The adversary — zero queries, one submission — the ancestor and the fact that it is public, a measured untargeted and targeted transfer rate from a source built on that ancestor with its input budget stated, the process controls that absorb a hit such as review coverage of automated decisions, and a named owner accepting the residual. Not a sentence asserting that internals are proprietary.
saying these in an interview costs you the question
- Assumes a private fine-tune hides its public ancestor
- Treats using a public base as automatically disqualifying
- Offers rate limits against a one-submission attack
- Confuses this with whether the download was tampered with
- Asserts a transfer rate instead of measuring one