Your banking app ships the face-matching model inside its install bundle — what access does that hand an attacker?
answer
- the file is on their device now
- an access assumption, not an event
- nothing to meter, nothing to see
- no substitute needed, no transfer loss
- only the server half still decides
basics
~20 sAnyone who installs the app has the weights. Full white-box access becomes the normal operating point rather than a worst case: the attacker studies and attacks the model locally, unmetered, sending nothing to your server.
solid answer
~50 sShipping the model means shipping the weights. Once the file is on a device the owner controls, the attacker has architecture, parameters and the gradient of any loss with respect to the input — the same access an evaluator only *grants* on purpose in a lab. Nothing about that requires touching your service, so there is no query bill, no rate limit, no per-identity cost and no traffic for you to look at. There is also no transfer loss: they are attacking the exact function you deployed, not a stand-in trained to imitate it, so anything that works locally works in production on the first attempt. Packaging, obfuscation or an unusual file format raise the effort to obtain the file once; they do not change what the adversary can do afterwards. The realistic threat model for this deployment is white-box, and it should be written down that way.
go deeper
Be ready to say plainly that shipping a model ships its weights, and that this gives an attacker the strongest access class in the field for free. Know the words white-box and black-box as access assumptions.
Explain the mechanics of why it matters: gradients with respect to the input are available, iteration is offline and unmetered, and there is no substitute model and therefore no transfer loss to absorb.
Show you would rewrite the deployment's threat model rather than argue about how hard the file is to extract, and that you know which of your existing controls silently stopped applying the day the model started shipping.
Own the framing that moving inference to the client trades data confidentiality for model integrity, and be able to state which decisions you will therefore refuse to move off the server.
## The situation A consumer banking app performs a face match on the phone. To do that, the trained model — its architecture and all of its parameters — is packaged into the application bundle that every user downloads. The team chose this for good reasons: the camera frames never leave the device, the check completes without a round trip, and it works with a weak connection. All of that is real. What also happens, unavoidably, is that the model is now in the hands of anyone who wants it. ## "White-box" as an operating point, not a lab grant In adversarial ML, *white-box* and *black-box* are *access assumptions*, not events. White-box means the adversary is assumed to hold the weights, the architecture, and therefore the gradient of any loss with respect to the input. Black-box means they can only send the model inputs and read whatever the reply contains. Evaluators usually grant white-box access deliberately, so the result they publish bounds every weaker adversary and does not quietly measure their own obscurity. That is a *stance* taken during evaluation. When you ship weights, white-box stops being a stance and becomes the deployment's actual condition. The conservative ceiling and the real number are now the same number. Everything you say about this system's robustness has to be quoted under that assumption. ## What the adversary gets for free Three things, and each of them is a control you no longer have: **No query bill.** Every economic control this field relies on — paying per call, per-identity limits, coarsening what the reply contains, watching the distribution of incoming queries — presupposes that the adversary must come through your endpoint. Here they do not come through it at all. They run the model on their own hardware, as often as they like, and you never see a packet. **No transfer loss.** An adversary who can only query a service often trains a substitute that agrees with the target near the boundary they care about, then attacks the substitute and hopes the result carries over. It usually carries over partially, which costs them attempts and gives the defender a chance to notice. There is no such gap here: they hold the function itself, so what succeeds locally succeeds against the deployed model exactly. **No observability.** Local iteration is invisible. Whatever search they run, you see none of it — only, at most, the small number of finished attempts they eventually present to your service. ## What does *not* help - **The app sandbox.** It protects the app from other apps on a device the *user* does not control. It is not a boundary against the device's owner. - **Obfuscation, packing or encryption of the file.** These raise the one-time cost of getting the weights out of the bundle. They do not reduce anything afterwards, and a file that the app must eventually load in usable form is a file that can be recovered. - **An unfamiliar serialization format.** Formats are documented or reverse-engineered once, and then never again. - **Signature or digest checks on the artifact.** Those tell a consumer *which file* they got and *whose* it is. They say nothing about who else has a copy. Each of these is a delay, and delay against a one-time extraction is worth very little. ## Privacy improved; integrity got worse The common mistake is to reason from one property to the other. On-device inference genuinely improves *data confidentiality* — the biometric frames stay on the phone. It simultaneously makes *model integrity and model confidentiality* strictly worse, because the asset you were protecting behind an API is now distributed to the public. Those are different properties, and moving a computation closer to the user helps one and hurts the other. ## What is actually left Exactly one thing: whatever your server still decides for itself, using evidence the client cannot mint. If the server holds the enrolled template, applies the accept threshold, and makes the step-up or entitlement call, that part is still yours. If the app merely reports "match succeeded" and the server believes it, then you shipped the decision along with the weights and there is nothing left. ## How to state it Write the threat model as: *adversary holds the model, has unlimited offline access to it, and interacts with our service only through the same interface a normal client uses.* Every robustness number attached to this deployment is then a white-box number, and every control you claim has to survive the assumption that the client is the adversary's.
- Does an openly released weight file differ from one extracted out of an app bundle?Not in what the adversary can do — both give the same function, gradients included. They differ in effort and in population: an open release costs nothing to obtain and reaches everyone, while extraction from a bundle costs a one-time reverse-engineering effort. Since that cost is paid once and the result is shareable, it is a difference in how quickly the vantage spreads, not in what it grants.
- The model is quantized and stripped of training metadata before shipping. Does that reduce the exposure?Barely. Quantization changes numeric precision, not the fact that the adversary holds a function they can evaluate and differentiate through as many times as they like. Stripped metadata costs them a little context about training. Neither restores a query bill, a rate limit, or any visibility for you, so the threat model does not move.
- Why is 'no attacker will bother reverse-engineering our app' a weak argument?Because the effort is paid once by one person and the artifact is then trivially copied. Security that rests on nobody trying is not a control, and it fails silently — you get no signal when someone does. It is also the wrong shape of assumption to put in a risk register, where the entry has to hold against the most motivated adversary, not the median one.
Publishing the lock's full mechanical drawings is not the same as publishing the key — but it does mean every picking attempt now happens in the attacker's own workshop, silently, with no locksmith watching.
saying these in an interview costs you the question
- Says on-device inference is inherently the most protected deployment
- Treats the app sandbox as a boundary against the device owner
- Believes obfuscating the weight file changes the threat model
- Confuses data confidentiality gains with model integrity gains
- Quotes a black-box robustness number for a shipped model
- Assumes an attacker must still come through the API