How does recovering a model's parameters differ from training a copy on its replies?
answer
- two attacks share one name
- one fits, the other solves
- replies as equations, not as labels
- digits matter more than volume
- exact, up to relabelling the units
basics
~20 sTraining a copy fits a fresh model to bought predictions, so it only approximates the original near where the queries landed. Parameter recovery instead treats precise numeric replies as equations and solves for the original weights themselves.
solid answer
~50 sThey are two different attacks that both get called stealing a model. Fitting a copy is statistical: the attacker buys predictions, treats each input-and-reply pair as a training example, and ends up with their own network that agrees with the target where the queries landed and drifts away elsewhere. Parameter recovery is algebraic: against a small network of piecewise-linear units, the function bends in places whose position and orientation are fixed by one unit's weights and bias, and an attacker reading the returned value to many significant digits can detect those bends and turn them into equations in the parameters. Solve enough and you hold the weights, not an approximation — identical up to symmetries that leave the function unchanged. It needs precise scores rather than labels, and its cost climbs steeply with width and depth, which is why it is a small-model result.
go deeper
Be ready to state the difference in one breath: one attack fits a new model to bought predictions, the other solves for the original parameters from precise numeric replies. Know that the second needs digits, not labels.
An interviewer expects you to explain why a piecewise-linear network leaks equations at all, and to say what the returned precision has to do with it. Name the query and precision cost as what limits the attack.
Show you can say which of your own endpoints is even a candidate: model size, whether the reply is a real number and how many digits of it go out. Frame the protection as economic rather than absolute.
Own the framing that 'not affordable at our size' is a claim with an expiry date. Be able to say what would have to change — a smaller shipped model, a wider numeric contract, a cheaper technique — before that answer stops being true.
## Two different things get called 'stealing a model' Take a concrete target: a small equipment-failure-risk regressor sold to industrial customers over turbine and pump telemetry. A handful of engineered sensor features go in; one risk score comes back as a real number. There is no class label anywhere in the contract — the product is the number. An attacker with an ordinary paying account has two quite different things they might do with that endpoint. ### 1. Fit a stand-in Buy a lot of predictions. Treat every pair of submitted features and returned score as a labelled training example. Train your own model on that set. This is the attack the extraction literature mostly teaches, and what it produces is a statistical approximation: it agrees with the target where the queries landed, its agreement decays as you move away from them, and its internals are the attacker's own — their architecture, their weights, their training run. It is a look-alike, and its quality is bought with query volume and with how well the queries covered the space that matters. ### 2. Solve for the parameters This is a categorically different move, and it is the one this topic is about. A network whose units are piecewise-linear computes a function that is linear on each of many regions of the input space and *bends* at the surfaces where a unit switches between its two linear pieces. Those bends are not noise in the function; each one is placed and oriented by the weights and bias of a single unit. So the location of a bend is a statement *about a parameter*, and it is an exact one. An attacker who can read the returned number precisely enough can tell that the function has bent — the change in the reply as the input moves stops being consistent with a single straight piece. Each bend they can pin down is an equation. Collect enough independent equations and the first layer's parameters fall out; with the first layer known, the layer behind it becomes reachable in the same way, and the solve proceeds inward. Nothing is being fitted here. The literature calls this class of result cryptanalytic, or functionally-equivalent, extraction. ## What the solve requires, and what it therefore is not - **Precision, not volume.** The evidence of a bend lives in the low-order digits of a real-valued reply. A top-1 class label carries essentially none of it; a score reported to two decimals erases it. The vantage this attack needs is *many significant digits of a returned number*. - **A piecewise-linear model.** The whole method leans on the function being made of flat pieces with sharp joins. - **A small model.** Both the number of queries and the arithmetic precision needed climb steeply with width and with depth, because a deeper unit's imprint on the final output is fainter. Published recoveries are of networks with on the order of thousands of parameters, not of production-scale ones. The query volume is a detail worth noticing: a solve against a small regressor is a few thousand ordinary-looking calls from one ordinary account. That is the opposite of the noisy, enormous harvest people picture when they hear 'model stealing', and no usage report has anything interesting to look at. ## What you actually walk away with Parameters — but only up to symmetries that leave the computed function untouched: hidden units within a layer can be permuted, and a unit's incoming weights can be scaled up if its outgoing weight is scaled down by the same positive factor. So the recovery identifies an equivalence class, not one canonical weight matrix. Functionally that costs the attacker nothing: every member of the class computes exactly the same replies. ## Why interviewers ask this The reflex answer is 'extraction only ever gets you a behavioural look-alike; the real weights are never recoverable'. That is a fair description of the common attack and a false claim about what is possible. It is not a theorem. What keeps a large deployed model out of reach is **cost and precision, not secrecy of the weights** — and 'not affordable at our size' is an honest statement with an expiry date on it, where 'impossible' is just wrong. A candidate who states the limit as an economic one, and names what would move it (a smaller model, more digits in the reply), is doing the job the question is testing.
- Why is a regressor that returns one real-valued score a better target for this than a classifier returning a top-1 label?Taking the top class collapses the whole output into one discrete token and throws away almost everything the solve reads. A real-valued reply hands back the function's actual value at that point, and the low-order digits are exactly where the evidence of a bend lives. A label-only endpoint does not stop an attacker from fitting a look-alike, but it removes the material an algebraic solve runs on.
- If parameter recovery exists, why does the extraction literature still mostly teach behavioural copying?Because behavioural copying scales and the solve does not. Fitting a stand-in works against any model, tolerates coarse replies and gets better with more queries. The solve needs a piecewise-linear model, high-precision numeric output, and a query and precision cost that grows steeply with width and depth, so it stays a result about small models. Against anything production-scale, buying behaviour is simply the cheaper path.
- Does an attacker need gradients or any internal access to do the solve?No. That is what makes the result interesting. Everything used is visible from the outside: inputs the attacker chooses and numbers the endpoint returns. Gradients with respect to the input, weights and architecture are the white-box assumption, and this attack assumes none of them — it reconstructs parameters from the shape of the function alone, using only replies any paying customer receives.
Fitting a copy is sketching a machine from photographs taken at many angles. Parameter recovery is measuring where its linkages hinge and deriving the part dimensions exactly.
saying these in an interview costs you the question
- Says model weights can never be recovered through an API
- Describes the solve as distillation with more queries
- Assumes gradient or weight access is required
- Thinks a huge query volume is always the tell
- Treats a label-only reply and a precise score as equivalent