An adversary holds a model trained with a per-record privacy bound — what does it promise about a user who contributed 1,000 rows?
answer
- the guarantee compares two datasets
- they differ by exactly one thing
- and that thing is a row
- a person is often many rows
- the group version multiplies the parameter
basics
~20 sMuch less than the headline number suggests. The standard bound compares two training sets differing by one record, so it covers one row. Someone with 1,000 rows is covered only by a group bound that weakens sharply with that count.
solid answer
~50 sA differential-privacy guarantee is a statement about two training sets that differ in exactly one record: an adversary reading the released model cannot tell which of the two produced it by more than a factor set by the privacy parameter. That unit is a **row**, not a person. If one user contributed a thousand rows, removing that user means changing a thousand records at once, and the guarantee for a group of that size is obtained by chaining the per-record bound — the privacy parameter grows with the group size and the failure term grows far faster, so the pair is meaningless long before a thousand. The fix is not a smaller number, it is a different accounting unit: cap each user's contribution and account with the user as the unit. So the first question to ask about any reported privacy parameter is: per what?
go deeper
Be ready to say that the guarantee is defined over two training sets differing in one record, and that a record is not the same thing as a person. Knowing to ask 'per what?' when a privacy parameter is quoted is most of the credit here.
Explain how a statement about one record is extended to a group by chaining it, and why that makes the bound weaker as the contributor's row count grows. Name contribution capping plus accounting at the coarser unit as the actual remedy.
Show you would catch this in review: find the unit in the model card before arguing about the number, and refuse to compare a per-record parameter with a per-user one. Be able to say what changing the unit would cost the training corpus.
Own the framing you put in writing. Decide the accounting unit before training, because it cannot be changed afterwards, and make sure any external privacy claim states that unit rather than letting a reader infer coverage the number never had.
## What the guarantee is actually a statement about Differential privacy is defined over a pair of **neighbouring datasets** — two training sets that are identical except that one contains a single extra record. The mechanism (here, a private training run) is said to satisfy the bound if, for any outcome an observer might look for, the probability of that outcome under the first dataset differs from the probability under the second by no more than a multiplicative factor that is exponential in the privacy parameter, plus a small additive failure term. Everything in that sentence is fixed by a choice somebody made before training: the pair differs in **one record**. That is the *unit of privacy*. It is the smallest thing the guarantee is about, and nothing about the mathematics forces it to line up with a human being. ## Why the mismatch matters In a great many real training corpora, one person is not one row: - an analyst who labelled thousands of samples in a labelling queue; - a customer organisation that contributed tens of thousands of telemetry traces; - a patient with a record per visit; - a device that emits an event per hour for a year. An adversary who has downloaded the released model file does not ask 'was this row present'. They ask 'was this *person* present', or 'what is true of this *organisation*'. Those are questions at a coarser granularity than the guarantee was ever stated over. ## Group privacy, and why it degrades There is a standard way to get from the per-record bound to a statement about a group of size k: two datasets that differ in k records are connected by k single-record steps, so you chain the per-record bound k times. For the pure form of the guarantee (no failure term) this gives a privacy parameter of k times the original. For the more common form that carries a small failure term, the parameter still grows in proportion to k, but the failure term picks up a factor exponential in the group size — it degrades far faster than the parameter does. The practical consequence: a bound that reads reassuringly tight for one record is arithmetically vacuous for a contributor of even a few dozen rows, and utterly vacuous at a thousand. A multiplicative factor that large places no constraint on an adversary at all, and a failure term that has climbed above one is not a bound. It is worth being precise about the direction of that claim. A vacuous bound means **no guarantee is being made** at that granularity. It does not mean an attack succeeds. Whether the thousand-row contributor is actually identifiable from the released model is an empirical question that a bound going vacuous does not answer either way. The honest statement is 'unbounded at that unit', not 'exposed'. ## The fix is the unit, not the number Buying back coverage by shrinking the privacy parameter does not work: to make a thousand-fold degradation land somewhere useful you would need a per-record parameter so small that the model learns nothing. The real remedy is to change what the unit *is* before training: 1. **Bound the contribution.** Cap how many rows any one user, device or tenant may put into the corpus. 2. **Account at that unit.** Treat the whole capped contribution as the thing that is present or absent, so the neighbouring pair differs by one *user* rather than one row. The reported parameter then means what a reader assumes it means. It also costs more: capping throws away most of what heavy contributors gave, and a parameter of a given size stated per user is a much stronger — and therefore more expensive — claim than the same number stated per record. Those two numbers are not comparable, and putting them side by side in a table is a reporting error. ## What to do with a reported number When a model card, a vendor deck or a paper quotes a privacy parameter, the reviewer's job is to find the unit. If the unit is not stated, the number is not yet a statement about people, and no amount of arithmetic recovers it after the fact — the unit was decided at training time. 'Per what' is the first question, and it is asked before 'how large'.
- What would you have to have done at training time so the bound covered the person rather than the row?Decide the unit first. Cap how many rows any one contributor can supply, then do the accounting with that whole capped contribution as the unit that is present or absent. The reported parameter then describes a person. It cannot be retrofitted after training, and it costs data — most of what a heavy contributor gave is discarded by the cap.
- Does simply choosing a smaller privacy parameter rescue the thousand-row contributor?No. The group bound degrades in proportion to the row count for the parameter and much faster for the failure term, so covering a thousand rows would need a per-record parameter small enough to destroy the model's utility. Shrinking the number trades accuracy for a coverage gap it never closes. Changing the accounting unit is the only real remedy.
- If the group bound is vacuous, does that mean the adversary can identify that contributor?No, and saying so is the mirror-image error. A vacuous bound means nothing is being guaranteed at that granularity, not that a disclosure has been demonstrated. To claim actual exposure you need an empirical attack result against that unit. The honest report is 'no meaningful bound at contributor granularity', with any measured attack success stated separately.
A receipt that says 'one item may be swapped without anyone noticing' says nothing reassuring to a shopper who is swapping out an entire trolley.
saying these in an interview costs you the question
- Says every user in the dataset gets the headline privacy parameter
- Treats 'record' and 'person' as the same unit
- Thinks a smaller privacy parameter automatically covers heavy contributors
- Quotes a privacy parameter without stating what unit it is per
- Reads a vacuous group bound as proof the contributor was exposed