skip to content

Does hiding confidence scores stop an outsider learning who was in a model's training set?

level: seniorimportance: nice to knowfreq 30%

answer

  1. you removed a family, not the property
  2. the fitted function did not change
  3. class names still carry stability information
  4. count queries per candidate, not per attack
  5. ask which adversary the price excludes

basics

~20 s

No. Returning only a class name removes the cheapest signal and raises the price to several queries per record. The training imprint still shows in how the returned label behaves around a record: a cost control, not a boundary.

solid answer

~50 s

Suppressing scores removes the one-query version of the attack, where the adversary reads a confidence figure and compares it against a threshold. It does not remove the leak. A model still behaves distinctively around records it was fit on, and that shows up in label-only terms: how far a record sits from the point where the returned class changes, and whether that class survives small variations of the record. Buying the signal costs several queries per candidate instead of one. So ask who your adversary is. Against someone sweeping a long list of names, multiplying the per-record cost genuinely helps; against someone holding one named person and wanting one bit about them, a handful of extra queries is no deterrent, and on a corpus whose membership is a health fact that is exactly the adversary you have. Score suppression raises the bill. Say it in those words.

go deeper

for a junior

Remember that the leak comes from how the model was fitted, not from the output format, so returning fewer details makes reading it harder rather than impossible.

for a middle

Be able to explain that label-only signals exist because a fitted record sits further from where the answer changes and keeps its class under small variation, at the cost of more queries per candidate.

for a senior

Demonstrate that you price a control instead of declaring victory: state the query multiplier, name the adversary class it prices out, and say plainly that it does nothing against someone asking about one named person.

for a principal

Own the communication risk. A claim that output restriction closed a privacy exposure will be retracted the first time an evaluator demonstrates otherwise, so decide in advance what your organisation will assert to a reviewer and what residual it accepts.

## The proposal and what it actually removes A common first response to a membership finding is to stop returning confidence figures: have the endpoint reply with a risk band or a class name and nothing else. It is worth being precise about what that buys, because the change is real but much smaller than it looks. What it removes is the **cheapest** family of membership signals: the ones that read a returned number and compare it against a threshold. Those need a single query per candidate and no cleverness. Removing the number removes that family outright, which is a genuine improvement in the adversary's bill. ## What it does not remove The imprint that makes membership readable is a property of the fitted model, not of the output format. A model that was fit on a record tends to place that record more decisively inside its class than it places a comparable record it never saw. That property is still observable when only a class name comes back, in two ways that need no numbers: - **Distance to the point where the answer changes.** A record the model was fit on typically sits further from the region where the returned class would flip. The adversary cannot read that distance directly, but they can bracket it by asking about variants of the record they already hold. - **Stability of the answer under small variations of the record.** Confidently fitted records keep their class under variation; borderline ones do not. Stability is itself a signal, and it is expressible entirely in returned class names. Neither of these needs the model's internals, an auxiliary corpus, or anything the adversary does not already possess. What they need is **more queries per candidate** rather than one. ## The right frame: price, not prevention The question to ask about any output-restriction control is *what does it multiply the adversary's cost by, and does that matter for the adversary I actually have?* Coarsening outputs is a **cost control**. It moves the attack from one query per name to some larger number, and it removes the laziest tooling. Whether that is worth anything depends entirely on the adversary's shape: | Adversary | Effect of hiding scores | | --- | --- | | Sweeping a long list of candidate names | Real: their total bill scales with the list, so a per-record multiplier bites | | Holding one named person and wanting one bit | Negligible: a handful of queries about one person is not a deterrent | The leaf's own scenario is squarely the second row. An adversary who wants to establish that one named individual attends a specialty clinic will spend a few dozen queries without noticing. Presenting score suppression to a regulator or a review board as though it closed the exposure would be a misstatement, and a technically literate reviewer will ask exactly the question above. ## Related controls in the same family, with the same character - **Quantising or banding the output.** The same trade: fewer distinguishable answers, so more queries needed to resolve the same distinction. Cost, not boundary. - **Rate limiting and per-caller auditing.** These change *who* can be an adversary and how visible their sweep is, which is often more valuable than coarsening outputs, and which does nothing about a single authenticated caller asking about one person. - **Refusing to answer near the boundary.** Reduces one signal and introduces another, since a refusal is also an observation. Anything the endpoint does differently for different inputs is a channel. None of these bounds an attack that has not been run. The only class of claim that does is a **training-time** guarantee, stated over an adversary you did not test, and it is bought with accuracy rather than with output formatting. ## Why this class of mistake recurs The underlying error is treating an output restriction as though it changed the model. It did not. The fitted function is the same function, with the same distinctive behaviour around the examples it was fit on; only the resolution of the channel through which someone reads it has been reduced. Whenever a defence acts on the interface rather than on the training, expect the honest statement to be about the adversary's price and the class of adversary it prices out, never about elimination. ## How to communicate it Say: "We removed the one-query version. The signal remains and costs more to read. Against an outsider sweeping thousands of names, that matters; against someone asking about one named patient, it does not. Here is what we did that changes who can query at all, and here is the residual we are accepting." That is a statement someone can act on, and it does not have to be retracted the first time an evaluator demonstrates a label-only result.

  • What signal is still available when only a class name comes back?
    How the returned class behaves around the record. A record the model was fit on tends to sit further from the point where the answer would change and to keep its class under small variations, while an unseen comparable record is likelier to be borderline. Both are expressible purely in returned class names, at the cost of several queries per candidate rather than one.
  • When is raising the per-record price genuinely a good control?
    When the adversary's goal scales with a list. Someone screening thousands of names pays the multiplier thousands of times, so coarsening outputs plus per-caller rate limiting can make the sweep impractical or conspicuous. It buys nothing against an adversary who already knows which single person they care about, which is the adversary a roster-shaped training set attracts.
  • How should this be written up so it does not have to be retracted?
    State the change as a cost multiplier and name the adversary class it prices out, rather than claiming the leak is closed. Record which attack you ran, under which access assumption, and note explicitly that a negative label-only result bounds the attack you tried and not the model. Then state the residual you are accepting and who accepted it.

saying these in an interview costs you the question

  • Calls the leak closed once scores are suppressed
  • Assumes a label-only endpoint gives an adversary nothing
  • Confuses an interface change with a change to the model
  • Ignores that the target adversary needs only one name
  • Reports a failed attack as evidence of no leakage

context