A team says an attacker who never obtains their model's weight file cannot breach its confidentiality. Why is that wrong?
answer
- the file is a container, not the asset
- ask what the model was fitted on
- the interface is the route
- a consequence of imperfect generalization
- success is advantage over a base rate
basics
~20 sBecause the confidential asset is not the file. It is the training data the model absorbed, and secondarily the learned function itself, both of which leak through ordinary query access. A privacy attack can succeed against a model the adversary never obtains.
solid answer
~50 sThe weight file is a container, not the asset. What is confidential about a trained model is the data it was fitted on — and, separately, the function it learned. Both are reachable through the query interface: a model fits what it saw better than what it did not, so its outputs carry information about the training set, and enough replies characterise its behaviour well enough to stand in for it. That means an adversary with a paid account and no access to the artefact can still pursue the confidentiality goal, measured as an advantage over the base rate on some fact about a record, not as "a file was copied". Protecting the file is a real and necessary control, but it bounds exfiltration of the artefact only; it says nothing about what the endpoint gives away. The two need different evidence and different owners.
go deeper
Be ready to say that a trained model reflects the data it was fitted on, so its answers can reveal things about that data even when nobody has the file. Name the query interface as the route.
Explain why this follows from imperfect generalization rather than from a bug, and separate the two confidential assets: facts about the training data, and the learned behaviour itself. Say what each is measured against.
Demonstrate that you would ask for two independent pieces of evidence — artefact custody and a measurement against the live interface — and that you would refuse to accept the first as an answer to the second.
Own the framing for people outside engineering: what the organisation is actually protecting when it says a model is confidential, and which of the two assets the product's access model puts at risk by design.
## The claim, and what is wrong with it "Nobody has our weights, so our model's confidentiality is fine" is the most common wrong answer in this part of a threat model, and a competent engineer gives it, because for almost every other system the reasoning holds. A database is confidential in the sense that its file and its connection are guarded; if neither was reached, nothing was disclosed. A trained model breaks that intuition, because the artefact is not the asset. ## Two confidential assets, neither of which is a file **The training data.** A model is a fitted summary of the data it was trained on. Fitting is never perfect and never uniform: the model's behaviour on examples it saw differs measurably from its behaviour on examples it did not, and that difference is visible from outside in the model's outputs. Whole classes of attack live in that gap: inferring whether a particular record was in the training set at better than the base rate; inferring a missing field of a record the adversary already mostly holds; recovering what the model considers a typical member of a class; eliciting spans the model reproduces close to verbatim. None of these requires the artefact. Some require only the returned decision. **The learned function.** Separately from the data, the model's input-output behaviour is often the thing the operator paid for. An adversary who can query it can build a stand-in that agrees with it well enough for their purpose. Again the file never moves. This is a confidentiality goal too, and one whose harm is commercial rather than personal. So "confidentiality" for a model splits: a **privacy** asset (facts about the people and records behind the training data) and a **proprietary** asset (the behaviour itself). The weight-file reading covers neither, because it protects the *encoding* of both while leaving the interface that exposes them wide open. ## Why query access is enough The short version is that a model is not a lookup table but it is not independent of its training set either — it sits somewhere in between, and where exactly is a property of how well it generalised. A model that generalises perfectly would treat a member of its training set exactly as it treats a fresh example from the same distribution, and there would be nothing to infer. Real models do not, and the residual is the leak. That is why this exposure is a *consequence of imperfect generalization* rather than an implementation bug: there is no patch that removes it, only training-time mechanisms that bound it and pay for the bound in accuracy. This also explains why the exposure does not care about the deployment shape. An on-premises model behind a private VPC leaks the same information through its predictions as a public endpoint does, at whatever rate the caller is allowed to query. ## What the weights-only reading actually buys It is worth being precise rather than dismissive, because guarding the artefact is genuinely useful: - Holding the file back removes the strongest adversary — one who can compute exactly, without paying per query — and forces everyone else to buy their information through the interface, which is meterable, loggable and rate-limitable. - It also protects the *cheapest* route to the learned function, which is simply having it. What it does not do is bound disclosure. "No one has our weights" answers the question "was the artefact exfiltrated?" and nothing else. Two different pieces of evidence are needed, and they come from different places: artefact custody from access logs and egress controls, output disclosure from an actual measurement against the query interface. ## Saying it correctly in an interview A good answer names the asset, then the route, then the measurement: 1. **Asset**: the training data (and separately the learned function), not the file. 2. **Route**: the query interface, available to anyone the product lets query it, including a paying customer. 3. **Measurement**: advantage over the base rate on a stated fact — not "a file was copied", and not an accuracy figure on its own. And then the honest caveat, which is what separates a strong answer from a scary one: the *harm* is set by what the inferred fact means about a person, not by the size of the number. Learning that a record was in a corpus of general product reviews is close to worthless; learning that the same record was in a corpus of people who sought a particular kind of treatment can be serious at the same measured advantage. The threat model has to say which corpus it is before anyone can rate the finding. ## The other direction, too The converse error is just as common and worth flagging: a verified, signed, correctly-custodied weight file tells you which bytes you have and who published them. It says nothing about how those weights behave. Custody and behaviour are different questions with different evidence, in both directions.
- Does moving the model behind a private network change this exposure?Not by itself. The leak travels through predictions, so it is bounded by who is allowed to query and how often, not by where the process runs. A private deployment narrows the population of callers, which is real risk reduction, but any caller with query access — including a paying customer or an internal team — retains the same route. Rate limits and query accounting matter more here than network placement.
- How do you evidence a confidentiality finding of this kind so it is actionable?State the fact being inferred, the baseline someone would guess without the model, the measured advantage over that baseline, and the query volume it cost. Then state what the fact means about a person, because harm comes from the meaning rather than the number. Without the baseline and the meaning, you have a suggestive statistic that an owner cannot act on or dismiss.
- If the file is not the asset, is guarding it still worth doing?Yes — it removes the strongest adversary, the one who can compute against the model exactly and for free rather than buying information one metered query at a time. The mistake is treating that control as an answer to disclosure. Artefact custody and output disclosure are separate claims backed by separate evidence, and a threat model needs both stated.
saying these in an interview costs you the question
- Equates the model's confidentiality with custody of the weight file
- Says the model does not store its training data, so nothing can leak
- Assumes a private deployment removes the disclosure route
- Reports a privacy result as an accuracy number with no baseline
- Treats commercial theft of the behaviour as unrelated to confidentiality
- Claims a signed artefact establishes how the weights behave