skip to content

A regulator asks whether your clinic's deployed model discloses who attends — what can you honestly claim?

level: principalimportance: should knowfreq 34%

answer

  1. three tiers of claim, not one
  2. an evaluation bounds the attack you ran
  3. only training-time claims cover the untried
  4. a parameter without its delta says nothing
  5. the accuracy bill lands on rare patients

basics

~10 s

Claim what you measured and under which access assumption, not that there is no disclosure. A negative result bounds the attack you ran; only a training-time guarantee speaks about attacks you did not run.

solid answer

~50 s

Separate three statements that people blur. "We found no evidence of disclosure" reports on the attack you ran, at the access you granted it, on the records you sampled. "There is no disclosure" is a claim about every attack, which no evaluation can support. "Membership cannot be distinguished beyond this bound" is available only if the model was trained under a formal privacy mechanism, and only quoted with its parameter, its delta and its accounting. Then say what a regulator actually needs: what the training set's definition asserts about a person, who can send arbitrary records to the endpoint, and what edge a measured attack got. The calls you own are whether a model whose training set is effectively a roster should be queryable outside a small authenticated audience, and who absorbs the accuracy a formal guarantee costs, since that cost lands on your rarest patients.

go deeper

for a junior

Know that an evaluation which found nothing is evidence about that evaluation, not proof the model is safe, and that anything written to a regulator will be read literally later.

for a middle

Be ready to distinguish an empirical negative from a bounded formal claim, and to say that only a training-time mechanism speaks about attacks nobody ran.

for a senior

Show you would qualify every measured statement with the access assumed, the query cost and the records sampled, and that you separate what left the database from what remains in the weights.

for a principal

Own the calls: whether a model whose training set is a roster may be queried outside a small audited audience, which residual you accept and record, and who absorbs the accuracy a formal guarantee costs, given it lands on the rarest presentations.

## Why this question is hard in a way the technical ones are not The person asking can compel an answer, will read it literally, and will hold the organisation to it later. So the task is not to characterise the attack; it is to decide which sentences the organisation is prepared to defend for years. Most of the damage in this situation is done by an overclaim written by someone technical who meant something narrower than what they wrote. ## The three tiers of claim, and which you can actually make **Tier 1 — an empirical negative.** "An evaluation using an adversary with query access and a candidate record, receiving a risk band only, achieved no better than chance across a sample of *n* member and non-member records." This is defensible and weak on purpose. It bounds *the attack that was run, under the access it was granted, on the records that were sampled.* It says nothing about a better attack, a more informative output, or the atypical records the sample missed. Write it with those qualifiers or it will be read as tier 2. **Tier 2 — a general negative.** "The model does not disclose membership." Nothing you can measure supports this. Evaluations are lower bounds on adversary capability; the space of attacks is not enumerable. Do not sign it. **Tier 3 — a bounded claim about attacks you did not run.** Only available when the model was trained under a formal privacy mechanism, and only when quoted properly: the privacy parameter *with* its delta and the accounting method used, since a parameter alone is not comparable to any other parameter and a delta that is not far below one over the dataset size makes the bound vacuous. Two further honesty requirements a regulator's technical advisor will know: the standard bound is stated over one neighbouring record, so a person contributing many records is covered only by a group bound that degrades with their count; and the budget composes across everything released from that data, not per model. ## What the regulator is actually asking about Strip the machine-learning vocabulary and the question is: *can an outsider establish that a named person attends this clinic?* Three facts answer it, and only one is about the attack. 1. **What the corpus's definition asserts.** If the training set is "our attending patients", membership is a health fact about a named individual, full stop. Say this first, because it is the reason the question is being asked and pretending otherwise reads as evasion. 2. **Who can query the deployment with inputs of their choosing.** This is the exposure surface and it is usually the most changeable thing in the whole picture. An endpoint reachable by any authenticated clinician is a different world from one reachable by the public, and the difference is an access decision, not a modelling one. 3. **What a measured attack achieved**, with its access assumption, its query cost per record, and its behaviour on atypical records rather than only its average. ## The decisions that are yours to own **Should this model be queryable at all by anyone outside a small, named, audited audience?** When a training set is a roster, the model is a membership oracle for that roster to whoever can call it. That is an architectural fact, not a bug to be patched at the output layer, and the honest options are to restrict callers, to accept the residual explicitly, or to train differently. **Who absorbs the cost of a formal guarantee?** A training-time privacy mechanism is the only thing that yields a tier-3 claim, and it is paid for in accuracy. Crucially, that cost is not spread evenly: it falls hardest on rare classes and small subgroups, which in a clinical corpus are the unusual presentations, that is, the patients for whom the model's output matters most and who are also the ones most exposed by membership attacks in the first place. Deciding to pay that bill, and deciding who it is paid by, is a leadership call and must be recorded as one. **What do you say about deletion?** If someone asks whether a specific record is gone, the truthful answer separates the store from the model: the row can be removed from the database immediately, and the weights fitted on it are unchanged until the model is retrained without it. Promising more than that in writing is the classic irreversible mistake here. **What do you not say?** Do not offer output coarsening as though it settled the question; it raises the adversary's price and does not touch an adversary who cares about one named person. Do not quote an attack's low average accuracy as reassurance without the operating point on unusual records. Do not quote a privacy parameter without its delta and accounting. ## The shape of a defensible written answer Name the fact at risk. Name who can query. State what you measured and its qualifiers. State which stronger claim you are *not* making and why. State the decision you took about access, the residual you accepted, and who accepted it. That answer survives being read back to you, which is the only property that matters here.

  • The evaluation found no better than chance. Why can you not write "the model does not disclose membership"?
    Because an evaluation is a lower bound on adversary capability. It bounds the attack you ran, the access you granted it and the records you sampled, and atypical records are exactly the ones a sample under-represents and an attack succeeds on. Write the qualified version, or the sentence will be read as a general claim you cannot support and will have to retract.
  • What must accompany a privacy parameter for the claim to mean anything?
    Its delta and the accounting method, plus the unit the bound is stated over. A parameter alone is not comparable across runs, a delta that is not far below one over the dataset size makes the bound vacuous, and the standard statement covers one neighbouring record, so a patient with many visits is covered only by a group bound that weakens with their count. Budgets also compose across releases.
  • Legal asks whether a specific patient's record is gone. What is the honest answer?
    That the row has been removed from the data store, and that the deployed weights were fitted while it was present and are unchanged by the deletion. What removes it from the model is retraining without it, on a stated schedule, or an explicit claim about the model itself. Conflating the two in writing is the mistake you cannot walk back.
  • Who ends up paying for a formal training-time privacy guarantee?
    Patients with unusual presentations. The accuracy cost of such a mechanism concentrates on rare classes and small subgroups rather than on the average case, so the people whose records are most distinctive lose the most model quality, and they are also the people membership attacks identify most confidently. That is a leadership tradeoff to record explicitly, not a hyperparameter choice.

saying these in an interview costs you the question

  • Writes that the model does not leak because one attack failed
  • Quotes a privacy parameter with no delta or accounting method
  • Offers output coarsening as though it closed the exposure
  • Says a deleted database row removes the record from the model
  • Treats the accuracy cost of a guarantee as evenly distributed

context