An audit reports membership-inference AUC of 0.51 for a triage model — does that show no leakage?
answer
- a property of a run, not of a model
- which attack, which access, which pool
- the average is not the confident end
- a failed attack is not a bound
- ask for the rare cohorts separately
basics
~20 sNo. An average near chance means the attack that was run was near chance on the pool that was used. It cannot rule out a small set of records identified almost certainly, and it does not bound a stronger attacker.
solid answer
~50 sRead the number as a statement about an experiment, not about the model. A mean AUC of 0.51 says one specific attack, at one specific vantage, averaged near chance over one specific candidate pool. Privacy harm is per record, not average, and overfitting is uneven — so a handful of unusual tickets can still be called correctly with high confidence while the mean sits at chance. As a reviewer I would ask four things: what access did the evaluated attacker have and how many queries per record; how was the candidate pool composed, including its real membership base rate; what is the true-positive rate at a very low false-positive rate rather than the mean; and were rare categories and small subgroups reported separately or averaged into the pool. Absence of evidence from one weak attack is not evidence of absence.
code
text · 7 linesPrivacy evaluation - complaint-triage classifier
attack evaluated : loss-threshold membership test
attacker access : top-1 track + confidence, 1 query per record
candidate pool : 20,000 tickets, 50% members by construction
mean attack AUC : 0.51
verdict : "no measurable membership leakage"
...go deeper
Know that a reported attack score describes an experiment: a particular attack, on a particular set of candidate records. A low score means that attempt did not work, not that nothing can.
Be able to explain why an average across a pool hides a small set of confidently identified records, and why an attack's accuracy needs the pool's base rate to be interpretable at all.
Demonstrate the reviewer's move: ask what access the evaluated attacker had, how the pool was built, what the true-positive rate is at a very low false-positive rate, and which cohorts were broken out — before you accept or reject the claim.
Own what the organisation is willing to assert. Decide whether empirical attack results are ever enough for the statement you have to make, or whether a claim of this kind must rest on a training-time guarantee with its parameters and accounting on the record.
## What the number is a statement about An averaged attack score is a property of a run, not of a model. "Membership-inference AUC 0.51" decomposes into: *this attack*, at *this vantage*, over *this candidate pool*, scored by *this summary statistic*. Every one of those four is a choice somebody made, and each can be made in a way that produces 0.51 from a model that leaks. ## The four questions to ask **1. What was the attacker allowed to see and do?** A test run against a top-1 track plus one rounded confidence, at one query per record, is a weak adversary by construction. That is a fine adversary to evaluate — it may well be the realistic one for a product feature — but the result bounds *that* adversary. It says nothing about someone who gets the full probability vector, or a per-feature explanation shipped beside the prediction, or many queries per record. A robustness or privacy figure never bounds an attack that was not run. **2. How was the candidate pool composed?** Membership evaluations are usually built with a balanced pool: half real members, half held-out records. That makes AUC easy to interpret and quietly assumes the attacker's candidates are drawn the same way. More importantly, *which* records went in matters enormously. If the pool is a uniform sample of the corpus, it is dominated by typical records — the ones that leak least — and the mean is dragged to chance by them. **3. What happens at the confident end?** This is the substantive objection and the one to lead with. AUC is an average over every operating point, including the useless ones. A privacy attacker does not need to classify a whole pool; they need a small number of records they can be nearly certain about. Those two things are almost unrelated. The right report is the true-positive rate in the very-low-false-positive region — how many records can the attack call as members while almost never being wrong — and that number can be materially above the base rate while the mean AUC reads 0.51. **4. Was anything reported per slice?** Because overfitting is uneven, the interesting result is per cohort: rare categories, small subgroups, records with unusual vocabulary. Averaging them into a large pool is exactly the operation that hides them. ## Why the direction of the claim matters It is worth being explicit about what each possible result establishes: - A *high* attack score establishes leakage. That direction is sound: the attack succeeded, so the signal is there. - A *low* attack score establishes that this attack failed. It does not establish that the model does not leak, any more than a passing scan establishes that a checkpoint is clean. The asymmetry is structural, and it is the same asymmetry that governs every empirical security evaluation. Getting it backwards — reading a near-chance average as a certificate — is the failure mode this question exists to catch. ## What would actually move a reviewer Three things, in rough order of weight: - **A stronger attack, reported at the confident end.** Commission the best attack you can at the most generous vantage you are willing to concede, and report the very-low-false-positive region and per-cohort results, not a single mean. - **A structural argument rather than a measurement.** A training-time guarantee, stated with its parameters, its accounting method and the unit it is stated over, is the only thing that bounds attacks nobody has thought of yet. An empirical number bounds one experiment. - **A description of what the bit would mean.** Whether a confident membership call is a serious disclosure depends on what the corpus is, and that judgment belongs beside the measurement rather than inside it. ## The reviewer's sentence "0.51 tells me the attack you ran, at the access you gave it, averaged to chance on the pool you built. Show me the true-positive rate at a very low false-positive rate, the pool's real composition, and the rare cohorts broken out — then I will know whether this model leaks about the records I care about."
- Why does a mean AUC near chance not rule out confident hits on a few records?AUC averages performance across every operating point and every record in the pool. An attacker only needs the extreme end: a threshold so strict that it fires rarely but is almost always right. A pool dominated by typical records — the ones that leak least — pulls the mean to chance while leaving that extreme intact, which is why the low-false-positive region must be reported separately.
- The pool was 50% members by construction. What would you change?Say so explicitly and report advantage over that stated base rate, then re-report against a composition that resembles the attacker's realistic candidate set. Also stratify: include rare categories and small subgroups as their own reported slices rather than letting them be a rounding error inside twenty thousand typical tickets.
- Would you accept the claim if a stronger attack also came back near chance?It would raise my confidence considerably, and it still would not be a bound. Empirical results describe attacks that were run. If the team needs to make a defensible statement to somebody who can compel an answer, the argument has to come from a training-time guarantee with its parameters and accounting stated, not from a sequence of failed attacks.
saying these in an interview costs you the question
- Reads a near-chance average as an upper bound on any attacker
- Ignores how the candidate pool was composed
- Accepts a balanced evaluation pool as representative
- Never asks what access the evaluated attacker was given
- Treats a failed attack as evidence the model is safe