skip to content

A membership-inference finding reports 60% accuracy against a clinic's model — what does that establish?

level: middleimportance: must knowfreq 56%

answer

  1. sixty beats a coin flip, not the world
  2. the average is not the operating point
  3. ask what the cohort's prevalence is
  4. the confident subset is the finding
  5. meaning decides, not accuracy

basics

~10 s

A 60% figure establishes that a membership signal exists, and nothing about harm. Balanced-set accuracy is measured against a coin flip, not the cohort's real prevalence, and severity is set by what membership means.

solid answer

~50 s

Sixty percent on a balanced members-versus-non-members set is ten points of advantage over chance, so the model is leaking. What that number does not tell you is whether it is a disclosure. Three columns decide that and none of them is in the figure. First, the operating point: an average hides whether the attack is near-certain on a minority of unusual records, which is where real exposure lives, so you want the true-positive rate in the very-low-false-positive regime. Second, the base rate: against a population where few people attend the clinic, a weak signal shifts belief about a named person far less than the headline suggests. Third, the meaning: sixty percent that a named person was in a stigmatised-condition cohort is a reportable disclosure, while ninety-nine percent on a public image benchmark is a curiosity. Dismissing a finding because the accuracy is low is the standard mistake.

code

text · 9 lines
text
finding: membership inference vs. relapse-risk classifier (clinic deployment)
  evaluation set                                : 5,000 members / 5,000 non-members (balanced)
  attack accuracy                               : 60.2%
  attacker access                               : 1 query per record, top-1 risk band only
  true-positive rate at 0.1% false-positive rate: not reported
  which records the attack was confident on     : not reported
  clinic-attendance prevalence in the region    : not reported
  what membership in this corpus asserts        : not reported
  ...

go deeper

for a junior

Know that a membership attack is scored against a coin flip on a balanced test set, so 60% means ten points of edge. Do not equate a low score with a harmless model.

for a middle

Be ready to name the three things the figure omits: the operating point in the low-false-positive regime, the prevalence of the cohort in the real population, and what membership in this corpus says about a person.

for a senior

Show you would send the report back with specific questions rather than triaging on the headline number, and that you can explain why the confident minority of records, not the average, drives the exposure.

for a principal

Own the grading policy: findings on corpora whose definition is itself a sensitive fact outrank higher-scoring findings on public benchmarks, and your team's severity rubric should encode that rather than ranking by attack accuracy.

## The number on its own is nearly content-free A membership-inference result usually arrives as a single accuracy figure obtained on a balanced evaluation set: equal numbers of records that were in training and comparable records that were not. On such a set, guessing gets you 50%. So 60% means the attack has roughly ten points of *advantage over chance*, and the honest first conclusion is narrow: **the model leaks some membership signal**. Whether that leak matters to anyone is a separate question that the figure does not answer in either direction. The reflex answer, "only 60%, so it is not a real finding", is wrong, and so is its mirror, "60% on patients, therefore a breach". Both treat one number as if it settled a question that needs three. ## Column one: the operating point, not the average An average over ten thousand records tells you how the attack does on a typical record. Nobody is attacked typically. The adversary here has a specific person in mind and cares only about that person, and membership attacks are characteristically **uneven**: they are close to useless on the bulk of ordinary, well-represented records and much stronger on unusual ones, precisely because unusual records are the ones a model fits distinctively. That is why a mean accuracy is the wrong summary. The informative report is how many members the attack identifies while producing almost no false positives, that is, its true-positive rate in the very-low-false-positive-rate regime. An attack that averages 60% but is essentially certain on a few hundred atypical patients is a serious finding, and those few hundred patients are precisely the ones whose records are unusual, which in a clinical corpus often means the most clinically distinctive and most identifiable people. ## Column two: the base rate the adversary actually faces A balanced evaluation set is a laboratory convenience. In the real setting, the adversary is asking about a named person drawn from a surrounding population in which clinic attendance is uncommon. Starting from a low prior, a positive answer from a weakly-discriminating test does not carry the belief nearly as far as a 60% headline suggests, because most of the positive answers it produces come from the very large non-member population. This is the correct reason to be sceptical of a headline figure, and it is a completely different reason from "the number is small". The practical consequence: an accuracy figure quoted without the prevalence of the cohort in the population the adversary can target is not interpretable as an exposure claim. Both columns cut both ways. A low average can hide a confident subset; a high average can still leave individual answers weak against a rare cohort. ## Column three: what membership means This is the column that turns a measurement into a finding, and it lives entirely outside the model. - Training set defined as "patients attending this specialty clinic": membership is a health fact about a named individual. - Training set defined as "accounts we flagged for fraud review": membership is an accusation. - Training set defined as "images in a public benchmark collection": membership is a fact about a file. The attack, the query cost and the metric are the same in all three. Only the third column differs, and it is the one that determines whether the correct response is a regulatory notification or a note in the backlog. A team that grades findings by accuracy alone will systematically over-react to benchmark results and under-react to the ones that matter. ## What the adversary spends Worth keeping in view while judging severity: this adversary's budget is tiny. They already hold the candidate record, so nothing has to be reconstructed, and on a deployment that returns a class they need one query per name. There is no expensive step to price in, no large auxiliary corpus required for the basic move. When an attack is that cheap, a modest edge is not diluted by cost the way an expensive attack's edge would be, and "they would never bother" is not a defence. ## How to read a report like an adjudicator When a figure arrives without its columns, the follow-ups are fixed: 1. What was the evaluation set's composition, and what does the figure beat, chance or the real prevalence? 2. What is the true-positive rate at a very low false-positive rate, and which records are the confident ones? 3. What access did the attacker assume, and how many queries per record did it need? 4. What does membership in this particular corpus assert about a person? Only after the fourth answer can anyone say whether 60% is a curiosity or a disclosure. And note the asymmetry that catches people out: a *failed* attack is even weaker evidence than a successful one. A finding that reproduces once in five tries still establishes that the signal is there; an attack that scored no better than chance bounds only the attack that was run, under the access it was granted, and says nothing about a better one.

  • Why report the true-positive rate at a very low false-positive rate instead of accuracy?
    Because the harm is concentrated. Membership attacks are weak on ordinary records and strong on atypical ones, so an average washes the interesting cases out. The question a report should answer is how many members the attack names while almost never naming a non-member, since those confident identifications are what an adversary can act on and what a regulator will ask about.
  • Would 99% accuracy on a public image benchmark be a more serious finding?
    No. It would be a stronger measurement of the same mechanism with nothing at stake, because membership in a public collection asserts nothing about a person. Severity comes from the corpus's definition, not from the attack's score. Ranking findings by accuracy inverts the priority order between benchmark work and a corpus that is effectively a patient roster.
  • The attack reproduces only about once in five runs. Is it still a finding?
    Yes, with the flakiness reported. Intermittent reproduction usually reflects variance in the evaluation setup or which records were sampled, not the absence of the signal. The correct write-up states the access assumed, the query cost, the records where it succeeded and the reproduction rate, rather than either discarding it or quoting the best run as if it were typical.

saying these in an interview costs you the question

  • Dismisses a finding because accuracy is only slightly above chance
  • Treats a balanced-set accuracy as a real-world hit rate
  • Quotes an average and ignores the confident subset
  • Ranks findings by attack score rather than by what membership means
  • Reads a failed attack as proof the model does not leak

context