How does an attacker's set of shadow models turn a returned probability vector into a membership decision?
answer
- the response is the thing being classified
- peak, margin and entropy, not the top score
- labels come from the adversary's own experiment
- one rule per predicted class
- a transferred boundary, not a hand-set cut-off
basics
~20 sShadow models supply labelled examples: the attacker knows which records each stand-in saw, so every returned probability vector is tagged member or non-member. A classifier fitted on those tagged responses becomes the membership rule, then applied to the target.
solid answer
~50 sEach stand-in is trained on a sample the adversary controls, so for every record they can record the response and tag it as coming from a model that did or did not train on it. Those tagged responses are a supervised dataset whose input is the response itself — the full probability vector, its peak, its margin over the runner-up, its entropy — and whose label is membership. Fitting a small classifier on that dataset produces a decision rule, usually one per predicted class, because confidence profiles differ a lot between a common class and a rare one. Several stand-ins rather than one matter because a single model's memorisation is idiosyncratic; averaging over several teaches the rule what is systematic. The rule is then applied to responses the real target returns. Crucially the *boundary* is being transferred, not a hand-set number.
go deeper
Recall that the adversary trains their own models, notes for each record whether that model saw it, and learns the difference in the responses. That is the shape of the mechanism.
Be able to say exactly what is being classified — the response and its properties, not the applicant — and where the supervision comes from. Expect a follow-up on why several stand-ins rather than one.
Demonstrate that you can price the build: how many stand-ins, how the rule is split per class, and what the reported separation was actually measured on before anyone claims it applies to the deployed model.
Own the distinction between an attack that transfers a learned boundary and one that reads a single number. It changes what a mitigation can promise, because there is no single scalar to blunt.
## The move A membership attack needs a rule that says, for a candidate record, whether the observed response looks like one a model gives to something it trained on. The stand-in family builds that rule offline, on models the adversary owns, and then transfers it. Understanding it means being precise about three things: what the *inputs* to the rule are, where the *labels* come from, and why *several* stand-ins are used rather than one. Keep the setting fixed: a credit-limit-increase decision model reachable over an ordinary API, returning a full probability vector across the decision classes. The adversary has no weights, no gradients, no explanations, and no per-record loss from the target. ## The input: the response, not the record The object being classified is the model's **response**, not the applicant. A probability vector carries several usable properties at once: - how **peaked** it is — how much mass sits on the top class; - the **margin** between the top class and the runner-up; - the overall **entropy** of the distribution; - **which class** was predicted, and whether that matches the record's known true class. All of those tend to differ between a record a model fitted and a record it did not, because the fit is tighter on what was seen. The adversary does not have to decide in advance which of these properties matters, or where to put a cut-off. That is exactly what fitting the rule does for them. ## The labels: free, because they ran the experiment This is the part that makes the family work. The adversary trains stand-in models on their own same-distribution sample, deliberately including some records and holding others out. For every record they own, they therefore know the ground truth for every stand-in: seen, or not seen. So each observation is a pair — the response a stand-in gave, and a member/non-member tag — and the collection of them is an ordinary supervised dataset. No part of it required anything from the target. This is why the build is described as an *offline calibration*: everything up to the finished rule happens on the adversary's own machines. ## Why several stand-ins One model's memorisation is noisy and idiosyncratic. Which records a given training run over-fits depends on initialisation, ordering, and where each record sits relative to the mass of the data. A rule fitted against a single stand-in learns that one run's quirks along with the systematic effect. Training several stand-ins on overlapping-but-different samples lets the rule see the same record both included and excluded, across runs, and average out what is idiosyncratic. The count of stand-ins is one half of this attack's real cost — the other half is the credibility of the same-distribution sample — and it is the number a red-teamer is asked to justify. ## Per-class rules Confidence profiles are not comparable across classes. A model is routinely peaked and confident on a large, easy class for members and non-members alike, and diffuse on a rare class for both. A single global rule blurs the two and loses most of the signal. So the discriminator is normally fitted **per predicted class**, which is one of the practical advantages of the stand-in family: the calibration is done separately where the behaviour differs. ## Transferring the boundary rather than a number The useful framing for an interview: this family transfers a **decision boundary learned in response space**, not a threshold on a single scalar. That is what it buys over a hand-set cut-off — it can combine several properties of the response, adapt per class, and it needs no calibration point taken from the target. What it pays for that is the whole build: a defensible same-distribution sample and several training runs. ## What it does not do The output is still one bit per candidate record, and a noisy one. The rule does not reveal the target's parameters, does not reconstruct any record, and does not tell you *why* a record was in the set. Its quality is measured as advantage over the base rate of membership among the candidates you are testing, not as raw accuracy — a rule that says "member" for everything scores impressively when most candidates really are members. And the rule inherits the stand-ins' assumptions. If the stand-ins over-fit differently from the target — different data volume relative to capacity, different regularisation regime, a different applicant mix — the boundary sits in the wrong place and the transferred rule falls back toward chance.
- Why fit a separate rule per predicted class instead of one global rule?Because confidence profiles are not comparable across classes. A model is routinely peaked on a large easy class for members and non-members alike, and diffuse on a rare class for both. One global rule blurs those regimes together and washes out the member signal; splitting by predicted class keeps the comparison within a regime where it means something.
- What does the second and third stand-in buy that the first did not?Separation of the systematic effect from one run's quirks. Which records a single training run happens to memorise depends on initialisation and ordering as much as on the data, so a rule fitted against one stand-in encodes that run's accidents. Several stand-ins let the same record appear both included and excluded across runs, and the rule keeps only what is consistent.
- The rule reports 70% accuracy on the attacker's own stand-in holdout — what has that measured?How well the rule separates members from non-members on the stand-ins, which is not the same function as the target. It is an upper-ish bound under a perfect distributional match, and it says nothing about the base rate among real candidates. Two things are missing before it means anything about the target: the base rate, and a check that the rule still separates when applied to the target's responses.
saying these in an interview costs you the question
- Describes it as thresholding one confidence score
- Thinks the record is classified rather than the response
- Cannot say where the member and non-member labels come from
- Sees no reason to train more than one stand-in
- Reports raw attack accuracy without the base rate