Why doesn't keeping a classifier's architecture and training corpus private stop an attacker who cannot query it?
answer
- secrecy of design is not access control
- ask what transfer actually requires
- the task, not the layer count
- an overlapping distribution comes for free
- it lowers a rate, it removes nothing
basics
~10 sTransfer needs a similar task and an overlapping data distribution, not your architecture or your rows. Anyone solving the same job over the same kind of inputs has that overlap by construction.
solid answer
~50 sThe claim confuses two different assumptions. Secrecy protects against an adversary who needs to reproduce your model; transfer does not need that. The attacker builds an evasive input against a model of their own that does the same job, and it carries because both models learned overlapping features from a similar distribution and put their boundaries in the same regions. Your layer counts, your optimiser and your private rows are not the precondition — a similar task and an overlapping data distribution are, and a competitor solving the same problem over the same kind of inputs has them for free. What secrecy actually buys is a lower transfer rate: divergence in training data, fine-tuning depth and architecture family all cost the attacker success. That is a cost control, not a boundary, and it should be stated as a rate rather than as immunity.
go deeper
Know the headline: an attacker does not need your architecture or your data to build an input that your model reads wrongly. They need a model doing a similar job over similar data.
Be able to name the actual precondition — similar task, overlapping data distribution — and to list what secrecy changes: a success rate, through data divergence, architecture family and fine-tuning depth.
Show you can rewrite an immunity claim as a measured rate against a stated source model and budget, and that you can say which controls, such as hidden scores or rate limits, are irrelevant to a zero-query adversary.
Own the framing question: decide whether this deployment faces an adaptive adversary at all, and if it does, insist the security argument rests on measured transfer rather than on what was never published.
### The claim under examination 'Our architecture is unpublished and our training data is proprietary, so the published attacks do not apply to us.' This is one of the most common confident answers a competent engineer gives, and it fails on a specific point: it names the wrong precondition for the attack it is trying to rule out. ### What transfer actually requires An attacker with no access to your model builds an evasive input against a model of their own and submits the finished input to you. For that to work, the two models must agree about where the boundary sits in the region being attacked. Agreement comes from learning the same job over the same kind of material: models trained on overlapping distributions pick up overlapping features, including the brittle correlations that predict the label without being the reason for it. The direction that carries an input past one model's boundary therefore points broadly the right way for the other. So the precondition is a **similar task over an overlapping data distribution**. Not identical weights. Not the same architecture. Not access to your rows. If you are screening documents of a common kind, or scoring transactions of a common kind, or classifying images of a common kind, an outsider can assemble data from the same distribution without touching yours — because the distribution is a property of the world your model operates in, not of your storage. ### What secrecy does and does not change Secrecy is not a null control; it is a **rate** control. Several things genuinely reduce how often a transferred input succeeds: - **Data divergence.** The further your training distribution sits from anything an outsider can assemble, the less the two boundaries agree. This is the strongest of the four, and it is also the one you rarely get to choose. - **Architecture family.** Transfer between different families is weaker than within one. It is weaker, not absent. - **Fine-tuning depth from a shared base.** If your model started life as a public checkpoint, how much of that base representation survives is a large term. If it started from scratch on distinctive data, that term is gone. - **Task divergence.** A model doing a genuinely different job is a genuinely different boundary. And several things change nothing at all against this adversary. Returning only a top-1 label instead of a score is a real control against query-based attacks that estimate a gradient by probing — it does nothing here, because the attacker never queries you. Rate limits, per-account quotas and anomaly detection over request patterns are similarly beside the point when the attack takes exactly one request. Serving behind a different runtime, renaming outputs, or keeping hyperparameters unpublished are obscurity: they raise nobody's cost measurably. ### The right form of the answer A senior answer converts the immunity claim into a number and an assumption. Rather than 'attacks do not apply', say: 'a zero-query transfer adversary is in scope; against a source model built on publicly assemblable data for our task, untargeted transfer succeeds at roughly X%, targeted far lower, and the terms that move X are data divergence, architecture family and how much of any shared base survives our fine-tune.' That is a claim someone can check, and it does not collapse the moment a reviewer asks what secrecy was assumed to buy. ### The direction of the evidence Be careful about what a failed transfer proves. If an attacker's inputs do not carry to your model, that shows those particular inputs, built against that particular source, at that particular budget, did not cross your boundary. It does not establish that your boundary is elsewhere in general, and it certainly does not establish that secrecy caused the failure — a different source model, a larger budget, or an ensemble of several source models is a normal next step for the adversary and often lifts the rate substantially. ### Why interviewers ask this Because the wrong answer is the natural one. Everything in ordinary software security rewards not publishing internals, and here the intuition transfers badly: the attack is aimed at what your model **learned**, and what it learned is largely determined by the problem you chose, which is public whatever you do with your repository.
- Which change would genuinely lower your transfer rate the most?Training on data an outsider cannot assemble from the same distribution, for a task that is not the standard one. Data divergence dominates; architecture family and how much of any shared public base survives your fine-tuning are smaller but real terms. Hiding hyperparameters and serving details moves nothing, because the attacker never interacts with your service while building the input.
- Does hiding confidence scores help against this adversary?No. Withholding scores is a cost control against query-based attacks that estimate a gradient by probing and differencing what you return. A transfer adversary makes zero queries, so there is nothing to withhold from them. It is worth doing for other reasons; claiming it as protection here misstates the threat model.
- How would you phrase the residual risk honestly to a reviewer?As a measured rate against a stated source, not as immunity. Name the adversary — no queries, one submission — name the source model you tested from, the input budget, and whether the goal was any wrong answer or a specific one, then give the observed rate. A claim without those columns is not comparable to anything and cannot be tracked across releases.
saying these in an interview costs you the question
- Treats an unpublished architecture as a threat boundary
- Says attacks need the target's own training rows
- Offers rate limiting against a one-request attack
- Claims hidden scores block a zero-query adversary
- Reads one failed transfer as proof of robustness