Your API rounds the returned confidence to two decimals — what does that cost a membership attacker?
answer
- precision is not the same as signal
- a hundred buckets is plenty
- rounding mostly creates ties
- label-only still leaks correctness
- cost control, never a boundary
basics
~20 sVery little. Rounding removes precision, not the difference the attack reads: a hundred buckets across the range is far finer than the gap between typical member and non-member responses. Even returning only the label leaves a weaker but real signal.
solid answer
~50 sQuantising the returned score is a cost control, not a boundary. The statistic the attacker thresholds is the difference between the model's behaviour on records it was fit to and records it was not, and that difference is usually much larger than one hundredth of the range — so two decimals leave the separation essentially intact and only add ties. Coarsening further, all the way to returning nothing but the top-1 track, does remove the score-threshold family, but it does not remove the leak: whether the model is simply correct on a record still differs between members and non-members whenever training accuracy exceeds accuracy on fresh data. So the honest framing is that output coarsening raises the attacker's price and removes some techniques, while the ceiling on the leak stays where it was — set by the fit. Reducing the gap moves that ceiling; a training-time guarantee is what bounds it.
go deeper
Remember that trimming what an endpoint returns makes an attack harder to run, not impossible. The difference the attacker reads lives in the model's behaviour, not in the number of decimal places.
Explain the mechanics: what rounding removes (fine ranking, by creating ties), what it leaves (the coarse comparison), and why a label-only response still carries a weaker signal through correctness.
Show the trade: name output coarsening as a cost control, weigh it against the product value of the confidence number, and point at the training-side levers if the goal is to move the ceiling rather than raise a price.
Be ready to reject a design document that records this control as mitigation. Decide what the team may claim about the output contract, and insist that any bound-shaped statement rest on a training-time guarantee rather than on output formatting.
## What rounding actually does to the signal The adversary here holds a candidate ticket, has ordinary product access, and gets one response per record: a handling track plus one confidence number for that track, rounded to two decimals. The rounding is the stated limit — not a query bill, not a rate limiter — and the question is what it buys. Two decimals over a probability range gives roughly a hundred distinguishable values. The separation the attack lives on — how much more confident the model is on a record it was fit to — is, in a model that overfits enough to be worth attacking, typically far larger than a hundredth. So the member and non-member response distributions remain about as separated after rounding as before. What rounding does produce is ties: many records land in the same bucket, so the ordering within a bucket is lost. That degrades the very fine end of the attack — the ability to rank two records that differ in the third decimal — and leaves the coarse comparison the threshold test depends on untouched. ## The general shape of output coarsening It is worth generalising, because the instinct behind two-decimal rounding usually goes further: "then let us stop returning confidences at all." | what the endpoint returns | what it removes | what survives | | --- | --- | --- | | full probability vector | nothing | everything | | one score for the chosen class | the shape of the rest of the distribution | correctness plus confidence | | a rounded score | fine-grained ranking | the coarse threshold comparison | | the top-1 label only | the whole score-threshold family | whether the model is simply right on this record | That last row is the one candidates miss. Returning nothing but a label feels like it removes the attack, and it removes an entire family of techniques — which is worth something. But a model's accuracy on rows it was fit to exceeds its accuracy on fresh rows; that difference is precisely the generalization gap, and it is readable from correctness alone. The signal is weaker and needs a larger candidate pool to be exploited, and it is not zero. ## Cost control versus boundary This distinction is the point of the question, and it recurs everywhere in this field. A control that makes an attack more expensive, noisier, or reliant on more records is a **cost control**. A control that makes the attack impossible regardless of effort is a **boundary**. Output coarsening is squarely the first. Treating it as the second — writing "we round confidences, so membership inference is mitigated" into a design document — encodes a claim the mechanism does not support. There is also a product cost on the other side of the ledger. If the confidence number is load-bearing for callers — routing rules, human review thresholds, an operator deciding whether to trust a track — then coarsening or removing it degrades the feature. Paying real product value for a control that raises an attacker's price slightly is usually a bad trade, and saying so is the senior judgment here. ## What actually moves the ceiling Since the leak is a consequence of the fit, the levers are on the training side. Anything that makes the model's behaviour on its own corpus resemble its behaviour on fresh data — more data, fewer passes over it, less capacity relative to the corpus — shrinks the average attack, and shrinks it for every technique at once rather than for one family. That is a genuine reduction, and it still bounds nothing, because it does not stop one unusual ticket from being fit tightly. A bound, as opposed to a reduction, requires a guarantee established during training and reported with its parameters, its accounting method and the unit it is stated over. That is a different subject with its own costs, and it is the only kind of answer that speaks about attacks nobody has run. ## How to say it in an interview "Rounding to two decimals costs the attacker fine-grained ranking and essentially nothing else — the gap it reads is much bigger than a hundredth. Removing scores entirely kills the score-threshold family but leaves correctness, which still differs between members and non-members. It is a cost control; if I want the ceiling to move I have to change the fit, and if I want a bound I need a training-time guarantee."
- If the endpoint returned only the handling track and no number at all, what would remain?Whether the model gets this record right. Accuracy on rows the model was fit to exceeds accuracy on fresh rows, so correctness alone carries membership signal — weaker per record, and enough over a pool. Removing scores deletes a family of techniques and raises the attacker's price; it does not close the channel.
- So is coarsening the output ever worth doing?Sometimes, but weigh it honestly. It costs product value whenever callers use the number for routing or review thresholds, and it buys a modest increase in attacker cost against a leak whose size is set by the fit. If the confidence is genuinely load-bearing for users, spending it on this control is usually the wrong trade.
- What would you write in the design document instead of "mitigated by rounding"?That output coarsening raises the cost of score-based membership tests and removes none of the underlying signal; that the leak's size is governed by the model's generalization gap; and that any claim of a bound, rather than a reduction, has to rest on a training-time guarantee reported with its parameters and accounting.
saying these in an interview costs you the question
- Thinks hiding confidences makes membership inference impossible
- Assumes a label-only response leaks nothing
- Confuses a cost control with a bound
- Believes rounding changes the model's train-test difference
- Pays real product value for a control with no ceiling effect