Your red-team membership test, 200 queries per record, finds nothing on the retrained model — what does that establish?
answer
- it bounds the search, not the model
- one attack at one strength
- forty out of twelve thousand
- the interval, not the point
- average case hides the outliers
basics
~20 sIt establishes that one attack, at one strength, on the records you sampled, did not beat the base rate. It bounds that attack, not the model, and says nothing about a stronger adversary, untested records, or checkpoints still in circulation.
solid answer
~50 sA null result is a statement about the instrument, in the direction most people read backwards. What you measured is the advantage over the base rate of a particular test, with a particular vantage and a particular query budget, on the particular records you sampled — and the confidence interval on forty sampled records out of thousands is wide enough to hide a real effect. A better-calibrated attack, a per-record test rather than an average-case one, or simply more queries can all move it. Worse, average-case membership numbers hide exactly the records that matter: outliers, rare classes and duplicated subjects leak most, and sampling uniformly under-weights them. Report it as bounded evidence: attack, access, budget, records tested, base rate, advantage with its interval. Then say what is out of scope — every prior checkpoint still deployed or distributed was fitted with the record and this test never touched it.
code
text · 10 linesdeletion-verification report (excerpt)
target : deployed speaker-verification model, post-removal build
attack : score-threshold membership test, query-only, top-1 score
query budget : 200 queries per candidate record
records tested : 40 withdrawn enrolments (of 12,400 removed)
base rate : 50% (balanced in / out candidate pool)
membership advantage: +0.6 pp (95% CI: -2.1 .. +3.3)
...
not reported : per-record worst case; duplicate-subject rows;
prior build, still distributed to two call centresgo deeper
Know that a failed attack is evidence about the attack. If a test finds nothing, the right question is how strong the test was and what it was run against.
Be able to list the qualifiers that must accompany a null: which attack, what access, how many queries, how many records out of how many, and the base rate the advantage is measured against.
Show you sample the tail rather than the average, that you can read a confidence interval as the actual result, and that you scope prior checkpoints and duplicate-subject rows explicitly.
Own the sentence that leaves the building. 'An adversary of at least this strength gained no measurable advantage on the records we tested' is defensible; 'the record is gone' is a claim the instrument cannot support and somebody will hold you to.
### Read the direction of the claim An evaluation that fails to find something bounds the search, not the thing. A membership test that comes back at chance level has established that *this* attack, run *this* way, with *this* budget, against *these* records, did not distinguish members from non-members. Every one of those qualifiers is load-bearing, and an interviewer asking this question is checking whether you supply them without being asked. The symmetric error is common in this field: a high robust-accuracy figure means the attack that was run failed; a clean backdoor scan bounds the trigger family the scanner searched for. A null membership result is the privacy version of the same reading. ### The four things that limit what you measured **Attack strength.** Membership tests vary enormously in power. A weak test can sit at chance against a model that a well-calibrated one distinguishes comfortably, because much of the difficulty is separating 'the model fits this because it memorised it' from 'the model fits this because it is an easy example'. If your test does not control for example difficulty, a null tells you little. **Query budget.** Two hundred queries per candidate is a real constraint. Buying more evidence per record generally buys power. An adversary who is not paying your rate card, or who cares about one specific person rather than an average, is not limited the way your evaluation was. **Sample size and selection.** Forty withdrawn records out of twelve thousand gives you an interval, not a point. A measured advantage of half a percentage point with an interval spanning several points is consistent with no leakage and with meaningful leakage on a subset. And uniform sampling systematically under-represents the records that leak: the unusual ones. **Average case versus worst case.** Membership risk is not evenly distributed. Densely surrounded, typical records are nearly undetectable; outliers, rare classes and subjects who appear several times are detectable at much higher rates. An aggregate advantage near zero is compatible with a specific person being identifiable, and the person who filed the deletion request is disproportionately likely to be in that tail. If the number that matters is per-record, report per-record. ### What the result is silent about A test against the retrained model says nothing about the *previous* model. Those weights were fitted with the record and are unchanged by anything you did afterwards; if a copy is still serving traffic, still on a device, or still distributed, the claim does not reach it. The same goes for any model fine-tuned or distilled from it. It is also silent about correlated records. If the subject appears twice under different identifiers and only one was removed, the retrained model still fits the remaining one, and a candidate resembling the subject will look familiar for a reason that has nothing to do with a failed removal. ### How to report it so it is worth something The difference between a useful verification and a decorative one is entirely in the columns. A defensible report names the attack and its access assumption, the query budget, how many records were tested and out of how many, the base rate the advantage is measured against, the interval on that advantage, whether the reported figure is average or worst case, and which artefacts were in scope. Anything missing is a question somebody will ask later, at a worse moment. The useful framing to offer: this is an *upper bound on what we found*, produced by an attack we chose, and the honest sentence is 'an adversary of at least this strength gains no measurable advantage on the records we tested', never 'the record is gone'. If you want a statement that does not depend on an attack's strength, you need the process claim — a fit over a corpus that never contained the record — not a measurement. ### Why the instrument is worth running anyway None of this argues against running it. A membership test is the only instrument that speaks to the deployed artefact rather than to the pipeline, and it catches the failure modes a process claim cannot: a record that survived in a cached feature table, a second copy under a different key, a derived training set nobody re-derived. A positive result is decisive and cheap. A negative one is weak evidence that must be reported with its power. Treating a weak instrument's silence as proof is how a defensible programme turns into an indefensible sentence in a letter.
- Which records would you deliberately over-sample when verifying a removal?The ones with the largest individual influence: outliers, members of rare classes, subjects whose records appear more than once, and anything the training pipeline treated as hard. Uniform sampling under-weights exactly these, and they are both the most detectable and the most likely to be the subject of a request. A worst-case figure over that set is more informative than an average over everything.
- The tester reports 0.6 percentage points of advantage. Is that a leak?By itself it is not a number you can act on. With an interval spanning several points on forty records it is indistinguishable from zero, and even a solid small advantage is a disclosure or a curiosity depending on what membership means about the person — enrolment in a voice channel is mundane, membership in a clinical cohort is not. The interval and the meaning decide, not the digit.
- What claim can you make that does not depend on the attack you happened to run?Only the process claim: the deployed parameters came from a fit over a corpus that never contained the record. That is verifiable from pipeline evidence rather than from an attack, and it is not weakened by somebody publishing a stronger test next year. The membership test then serves as a check on the pipeline, catching residual copies the process claim assumed away.
saying these in an interview costs you the question
- Reads a null result as proof the record is gone
- Reports advantage without the base rate or interval
- Samples records uniformly and calls it worst case
- Ignores prior checkpoints still in circulation
- Assumes a stronger attack would also have failed
- Never states the attack's access assumption or budget