A 300-example class loses far more accuracy than a million-example one when training is capped and noised against a membership adversary — why?
answer
- look at capping and noise separately
- whose gradients are largest before capping?
- the noise does not shrink for a small slice
- count how many aligned contributions clear the floor
basics
~20 sBoth halves penalise the small class. Capping cuts hardest the large gradients atypical examples produce, and the per-step noise is the same size regardless of slice, so 300 capped contributions carry far less signal through it than a million.
solid answer
~50 sPrivate training caps every example's own gradient at a fixed bound, then adds noise scaled to that bound. **Capping** is a proportional cut to whichever examples currently produce the largest gradients, and those are the ones the model handles worst — the rare, atypical records; a typical majority example is barely touched. **Noise** is sized to the clipping bound and the batch, not to how rare a pattern is, so its magnitude is identical for both classes; what differs is how much accumulated signal stands against it. A million contributions pointing the same way sum to a direction that survives; three hundred, each capped to the same bound, may sit below the noise floor and never accumulate. There is a slice-size region below which a class effectively stops being learned at a given privacy parameter.
go deeper
Recall that private training caps each record's own contribution and adds noise, and that a class supported by few records is the first casualty of both.
Take the two halves apart out loud: which examples produce the biggest gradients before the cap, and why identical noise is much harder on a slice with few contributions.
Be ready to say what you would actually do about a cohort the private model stops serving, and why more epochs is not on the list at a fixed guarantee.
Frame it as a resourcing question. The only clean fix is more data behind the tail, which is a funded programme, not a training flag — decide whether you are paying for it or accepting the loss.
## The two levers, taken separately Private training modifies the update in two ways, and both of them are regressive with respect to how many records back a pattern. ### Capping each record's own gradient The bound is applied **per example**, before summation — that is the point, since it is a per-record bound the guarantee needs. Whose gradient hits the bound? The examples the current model is worst at: unusual vocabulary, an accent under-represented in the corpus, a claim type seen a few hundred times. These produce high loss and large, distinctive gradients. Capping is a heavy proportional cut for them and close to a no-op for a well-predicted majority example whose gradient is already small. That is not an accident of tuning. **Being influential is exactly what the mechanism is built to prevent**, and the records that most need individual influence to be learned are the atypical ones. The property that makes a record valuable to the tail is the same property that makes it a privacy risk. ### Noise sized to the cap The noise added per step is calibrated to the clipping bound and the batch, and does not scale down for a rare slice. So the comparison is: | | contributions per epoch | per-step signal in that direction | noise faced | | --- | --- | --- | --- | | large class | ~1,000,000 | large sum of aligned capped vectors | same | | rare class | ~300 | small sum of aligned capped vectors | same | Zero-mean noise averages out across many steps *when a consistent direction keeps being reinforced*. A direction reinforced by a handful of capped vectors per batch is not consistently visible above the noise, so it is not reliably accumulated. Below some slice size, the class effectively stops being learned at that privacy parameter — the number of examples backing a slice is the variable that decides whether its signal clears the floor. ## Why you cannot buy the tail back with more steps The intuitive fix — train longer so the rare direction accumulates — does not work at a fixed guarantee. Privacy loss accumulates over the noisy releases, so holding the guarantee constant while taking more steps means each step must be noisier. You are trading noise per step against number of steps along a fixed budget line. Tuning moves you along that line; it does not move the line. Larger batches help the head-to-tail ratio somewhat, because more contributions are aggregated before noise is added, but they cost steps at fixed compute and do not change the underlying asymmetry. ## The one thing that does help More records backing the rare pattern. That is why the honest answer to 'the private model is much worse for this cohort' is usually a data-collection answer rather than a training-configuration answer — and why it is slow and expensive, which is precisely why it gets skipped. ## The uncomfortable duality The records sitting in the tail are also the records a membership attack finds most easily, because a model that generalises imperfectly fits its rarest examples most distinctly. So the same records are the ones most exposed to the record-level adversary **and** the ones that pay the most utility for the protection. Any framing that treats private training as a small tax on the average has missed that this transfer is concentrated on a specific set of people, in both directions. ## What to say in an interview Name both halves — capping penalises the large gradients of atypical examples, and fixed-magnitude noise is faced by far less accumulated signal in a small slice — then state the consequence: at a given privacy parameter there is a slice size below which a class is effectively not learned, and the aggregate metric will not show it.
- Are the rare records that pay the most utility also the ones most at risk from a membership adversary?Yes, and that is the uncomfortable part. A model fits its rarest examples most distinctly, which is exactly the signal a membership test reads, so the tail is both the most exposed to the record-level adversary and the group charged the most for the protection. The transfer is concentrated on the same people twice.
- Does a larger batch size help the rare class?Somewhat. More contributions are aggregated before noise is added, so the ratio of signal to noise per step improves. But at fixed compute a larger batch buys fewer steps, and the underlying asymmetry — a fixed noise magnitude against a slice-sized signal — is unchanged. It is a tuning improvement, not a fix.
- If the class was learned fine without the guarantee, is the private failure a tuning problem?No. Tuning moves you along the budget line between noise per step and number of steps; it does not move the line. At a fixed guarantee, a slice below the size where its signal clears the noise floor stays unlearned. The remedies are more data for that slice, a weaker guarantee, or accepting the loss knowingly.
saying these in an interview costs you the question
- Says the noise is scaled down for small classes
- Blames the loss on a learning rate or epoch count
- Thinks capping applies to the batch norm rather than each record
- Assumes the rare class is protected because it is small
- Claims more epochs at the same guarantee fixes it