A deck reports that training with a record-level privacy guarantee cost 2 points of aggregate accuracy — what do you ask for?
answer
- the headline is an average, and averages have weights
- two points where?
- ask what is standing behind each slice
- check what the baseline was and how hard it was tuned
basics
~20 sAsk two points where. An aggregate delta is a population-weighted average dominated by the majority, so it can hide a rare slice that lost ten times as much. Ask for per-slice utility, the example count behind each slice, and a baseline trained identically.
solid answer
~50 sThe number as reported is close to uninformative, because the aggregate is weighted by population and the majority is the part private training barely touches. I would ask for four things. **Per-slice utility beside the aggregate**, with the slices that matter for the product named in advance rather than chosen after the fact. **The example count backing each slice**, in training and in evaluation, so I can tell a real regression from a small-sample wobble. **A baseline trained identically otherwise** — same data, same architecture, comparable tuning effort — because a stale or under-tuned non-private baseline flatters the private model and makes the gap look small. And **the privacy parameter and the unit it is stated over**, since a headline utility cost means nothing without knowing what was bought with it. Then the real question: are the slices carrying the loss the ones this system exists to serve?
code
text · 11 linesoverall word error rate non-private 6.1 private 6.5 (+0.4)
--- by slice (slices fixed before the comparison) ---
slice eval n non-private private delta
common vocabulary 48,200 5.4 5.6 +0.2
rare drug names 1,340 11.8 19.6 +7.8
uncommon surnames 610 14.2 25.1 +10.9
under-represented accents 210 12.9 24.4 +11.5
...
(privacy parameter and unit reported separately; baseline trained on
identical data with comparable tuning effort)go deeper
Recall that an overall accuracy figure is an average weighted by how common each group is, so it mostly reports on the majority.
Explain why the mechanism almost guarantees a small aggregate delta, and name the columns that have to sit beside it before the number means anything.
Show you would interrogate the comparison itself — how the baseline was trained, when the slices were chosen, how many examples back the worst one — before conceding or disputing the claim.
Be ready to turn the finding into a reporting standard: what your organisation must publish alongside any privacy claim so that a small headline can never stand in for a serviceable model.
## Why the headline is nearly empty Aggregate accuracy is a population-weighted average. Private training — per-example gradient capping plus calibrated noise, aimed at an adversary who tries to tell from the released model whether one record was in the training set — costs least on the patterns backed by the most records. Those same patterns dominate the average. So a small aggregate delta is close to *guaranteed* by the mechanism's shape, and observing one tells you almost nothing about whether the model still works for anyone in particular. The reflex to train is: **two points where?** ## What to ask for, and what each answer rules out | Ask | What it establishes | | --- | --- | | Per-slice utility beside the aggregate | Where the loss actually landed rather than what the majority experienced | | Examples backing each slice, train and eval | Whether a bad slice is a regression or a small-sample artefact | | The non-private baseline's provenance and tuning | Whether the gap is real or manufactured by a weak comparison | | The privacy parameter and its unit | What was bought for the utility that was spent | | Who the slices are | Whether the bill landed on the users the system exists for | The slice list has to be **fixed before the comparison**, from what the product cares about. Slices chosen after seeing results can be drawn to make any story true, in either direction. The baseline question is the one most often skipped. If the private run received a tuning campaign and the non-private baseline is a checkpoint from last quarter, the reported gap is a statement about effort, not about privacy. The comparison only means something when both runs saw the same data and comparable tuning. ## Reading the table you get back When per-slice numbers arrive, the pattern to look for is a small aggregate delta beside a handful of slices with large ones, and those slices having the fewest supporting examples. That is the mechanism's signature, not a coincidence, and it is what the aggregate was hiding. Two directional traps: - **A flat aggregate does not mean nothing happened.** It usually means the majority is intact. - **A bad slice is not automatically a regression.** A slice with 40 evaluation examples moves several points on chance alone; ask for the count before treating a delta as a finding. ## What is not the answer here Some reasonable-sounding asks address a different question. The success rate of a membership attack against the private model tells you about *exposure*, not about where utility was lost. Training-time overhead is a cost, but not this cost. And the privacy parameter alone, however carefully stated, says nothing about who paid for it — that is exactly the gap this question exists to close. ## The sentence you want in the report Instead of 'privacy cost us two points', a usable statement names the guarantee, the slices, their sizes, and the worst one: the aggregate moved a fraction of a point, the rare-vocabulary slice roughly doubled its error rate, and that slice is a stated purpose of the product. That version supports a decision. The headline does not.
- They send per-slice numbers and the worst slice has 40 evaluation examples. What now?Treat it as a signal to investigate, not a finding. Forty examples move several points on chance alone, so ask for a larger evaluation sample for that slice, or repeated runs, before quoting the delta. The training-side count matters too: a slice that is tiny in evaluation is often tiny in training, which is the mechanism's expected failure point.
- The non-private baseline is a checkpoint from two quarters ago. Does the comparison still stand?No. The gap then reports a difference in data, tuning and effort as well as privacy, and usually in the direction that flatters the private run. Ask for a baseline trained on the same data with comparable tuning, or treat the number as a lower bound on the true cost rather than an estimate of it.
- Would a membership-attack success rate against the private model answer the same question?No, it answers the other half. An attack result speaks to how much exposure remains; it says nothing about which users lost accuracy. Both belong in the report, and confusing them lets a strong privacy result stand in for a utility claim nobody checked.
saying these in an interview costs you the question
- Accepts an aggregate delta as the full statement of cost
- Never asks which slices carry the loss
- Ignores how the non-private baseline was trained and tuned
- Treats a large delta on a 40-example slice as a firm finding
- Substitutes an attack success rate for a utility measurement
- Lets slices be chosen after the results are seen