Why are the fields a prediction API returns part of its threat model?
answer
- start from what the adversary can see
- the reply is the observation channel
- verdict, top-k, full vector are different vantages
- richer replies buy a discount, not a capability
basics
~20 sEvery field handed back is signal an attacker gets for free. A bare verdict, a top-k list and a full score vector are three different access classes, and the reply class decides which attack families are cheap enough to run.
solid answer
~50 sA threat model states what the adversary wants, what they can touch, and what they can see. For someone who cannot get the weights and cannot write into the training data, what they can see is exactly the set of fields the endpoint hands back, so the response schema is part of the attack surface and not just an API-design choice. Take a hosted content-moderation service that accepts a user post and replies with a policy decision. If it returns a confidence for every policy category, the caller gets a continuous quantity that moves as they edit the post, plus the standing of categories that did not fire. If it returns a bare verdict, they get one discrete symbol. Same weights, two very different adversaries. So name the reply class first: an attack claim quoted without it is not comparable to anything.
go deeper
Be ready to say that a threat model covers what an adversary can see, and that for a query-only adversary that is the response fields. Name the ladder: full scores, top-k, one fired category, a bare verdict.
Explain what the continuous fields actually supply that a symbol does not: a signal that moves with small edits, a confidence magnitude, and the relative standing of categories that did not fire.
Show the habit of scoping first. Before proposing any attack against a hosted model, state the reply class and say how it changes your query budget, and refuse to compare two reported attack costs measured under different contracts.
Own the framing that the response contract is a security decision with a customer cost, not a free lever, and insist that any claim made about coarsening it names the family it repriced rather than asserting that leakage stopped.
## What a threat model has to state A threat model for a deployed model has three parts: what the adversary is trying to achieve, how much they can touch, and what they can observe. The last part is the one people forget to write down. For an adversary who holds no weights, has no write path into a training corpus, and is not an enrolled participant in training, the *only* observation channel is the endpoint's reply. That makes the response schema a security artefact: the list of fields you return is, quite literally, the definition of what your weakest-vantage adversary can see. This is why the phrase 'black-box attack' is under-specified on its own. In the access sense used here, black-box means the attacker does not have weights, architecture and gradients. But two black-box attackers can be an order of magnitude apart in capability purely because one endpoint returns a number and the other returns a word. ## The ladder of reply classes Take a hosted content-moderation service for a social product: a third-party site posts user text, the service returns a policy decision, and calls are metered. Its response contract could sit anywhere on this ladder, from most to least generous: - a confidence for **every** policy category; - the top-k categories with confidences; - only the category that fired, with a confidence; - only the category that fired; - a single allow/block bit. Each step down removes a continuous, comparable quantity and replaces it with a discrete symbol. Nothing else about the deployment changes. ## What the generous rungs actually hand over Three distinct things, and it helps to keep them separate: **A direction.** A continuous number moves as the caller edits the input. That lets an attacker tell whether an edit helped *before* the verdict flips, which turns a blind search into a guided one. **A magnitude.** How confident the model is on a given example is correlated with how well it fits that example, which is the raw material of privacy signals such as membership tests. **Cross-category standing.** Scores for the categories that did *not* fire say where this input sits relative to several boundaries at once, not just the one it crossed. Dropping non-fired categories quietly removes most of that. ## The part candidates get backwards The reply class bounds the **cost** of the attacks, not the **list** of them. Families that consume nothing but the returned symbol exist and are well established: decision-based search, which works from inputs whose verdict the model already gives, and label-only membership tests, which read how easily an input's verdict can be made to change. Coarsening the reply pushes an adversary out of the cheap families and into these expensive ones. It moves the price; it does not close the door. 'We only return the class, so there is nothing to leak' is the single most common wrong answer in this area. ## What the schema does not cover Even a one-bit reply is not the whole observation channel. Response latency, error text, whether the service says which policy fired, whether ties are broken deterministically, and how the endpoint behaves at its limits are all observable and none of them are in the JSON body you trimmed. If you coarsen the reply and then claim the observation channel is closed, you have made a claim about one field list, not about the system. ## How to use this in an interview Scope before you attack. When asked 'how would you attack this model', the first move that separates a strong candidate from a weak one is asking what the endpoint returns, because that answer determines whether the engagement is a few thousand queries or a few million. And when reading somebody else's result, treat a robustness, extraction or privacy number quoted without its reply class the way you would treat a latency number quoted without a percentile: not wrong, just not yet a claim.
- Two deployments serve identical weights but different response contracts. Is their threat model the same?No. The weights fix what the model computes; the response contract fixes what an outside caller can observe, and observation is half of a threat model. The one returning per-category confidences faces a materially cheaper adversary than the one returning a single verdict, so any attack cost, extraction estimate or privacy result you quote has to name which of the two it was measured against.
- Where does the documented response schema stop being the whole picture for a query-only attacker?It omits everything outside the body: response latency, error and validation text, whether the service names the policy that fired, and whether ties resolve the same way every time. Those are observable too. Trimming the JSON narrows one channel, and a claim that the endpoint now leaks nothing has to account for the others rather than assume they went away with the field.
- Why is naming the reply class the first thing you do when reading somebody else's attack result?Because query counts across reply classes are not comparable. An attack reported at ten thousand queries against an endpoint returning full confidences tells you almost nothing about what the same goal costs against a verdict-only endpoint, and vice versa. Without the reply class the number has no units, so the first question about any reported cost is what the endpoint handed back per call.
A locked building's threat model depends on which windows exist, not only on the lock. The reply fields are the windows.
saying these in an interview costs you the question
- Treats the response schema as pure API design, not attack surface
- Says hiding scores means no information leaves the endpoint
- Thinks black-box means the model is uninterpretable rather than access-limited
- Discusses an attack without stating what the endpoint returns