Why must the delta in an (epsilon, delta) privacy guarantee sit far below one over the dataset size?
answer
- the second parameter is additive
- a share where nothing is promised
- compare it against the number of rows
- the cautionary construction publishes records
- multiply it by n and look
basics
~20 sDelta is an additive probability that the epsilon bound simply fails. A mechanism publishing a delta-sized share of training records outright still satisfies the definition, so delta near one over the dataset size permits exactly that.
solid answer
~50 sThe bound says the probability of any output in the with-the-record world is at most `e^epsilon` times its probability in the without-the-record world, **plus delta**. That additive term is slack: with probability up to delta, the epsilon promise is allowed to fail outright, and the definition does not say how gracefully. The standard cautionary construction makes the point — a mechanism that picks a delta-sized share of the training records and publishes them verbatim satisfies the definition for any epsilon at all. So if delta is around one over the number of records, the guarantee permits total disclosure for somebody, no matter how impressive epsilon looks. The convention is delta far below one over the dataset size, often one over n to a power above one. The check takes ten seconds: divide one by the number of training rows and compare.
go deeper
Remember that the guarantee has two numbers, that the second one is a probability of failure rather than an amount of leakage, and that it has to be tiny relative to how many records were used.
Be able to state where delta sits in the bound and to give the cautionary construction that publishes a delta-sized share of records while satisfying the definition for any epsilon.
Do the arithmetic in front of the room: one over n, compare, then decide whether the epsilon is worth reading. Refuse to compare two runs whose deltas differ without saying so.
Set the house rule for what your organisation publishes — the pair, the accounting method and the record count together — so no downstream deck can quote a flattering epsilon detached from the delta that made it possible.
## Where delta sits in the definition The pure form of the guarantee has one parameter. For neighbouring datasets — identical except for one record — the probability that the training procedure lands in any given set of outputs differs between the two worlds by at most a factor of `e^epsilon`. The approximate form adds a second term: at most `e^epsilon` times, **plus delta**. Epsilon is multiplicative and bounds how much the two worlds can differ; delta is **additive** and bounds the probability mass for which nothing is being promised at all. That asymmetry is the whole answer. A multiplicative slack degrades gracefully — a slightly larger epsilon means a slightly better-informed adversary. An additive slack does not degrade at all: it is a share of outcomes carved out of the guarantee, and the definition places no constraint on how bad those outcomes are. ## The construction that makes it concrete The standard illustration is deliberately absurd and it is what an interviewer wants to hear. Consider a procedure that, instead of training anything, selects a delta-sized share of the training records and publishes them exactly as they are, doing something harmless the rest of the time. For the records it does not publish, the two worlds are identical. For the ones it does, the guarantee has already spent its additive slack. The procedure satisfies the definition for **any** epsilon, including zero — and it is a catastrophic disclosure for the individuals it named. Nothing about real mechanisms works this way; the point is what the definition permits, and therefore what a number in a report fails to exclude. If delta is comparable to one over the number of records, the bound leaves room for a construction that exposes roughly one person outright. Multiply delta by n and you get the scale of what the slack could cover: on four million rows, a delta of one in a hundred thousand leaves room for about forty records' worth of unbounded failure. ## The convention, and how to check it Practitioners therefore require delta **far** below one over the dataset size — often stated as one over n raised to a power greater than one, or simply pinned at a cryptographically small value such as one in a billion for datasets in the millions. The reasoning is not aesthetic. Below one over n, the additive slack is smaller than the per-record scale the guarantee is about, so it cannot swallow a record whole. The check a reviewer performs on any privacy line is arithmetic: - Find n, the number of training records the number was computed over. - Compute one over n. - Compare delta to it. If delta is at or above that, the guarantee is close to vacuous however good epsilon looks, and the correct response is to say so rather than to read the epsilon. A report that gives epsilon and omits delta cannot be checked at all, and it also cannot be compared with any other report, because a mechanism can trade a smaller epsilon for a larger delta. ## Why the second parameter exists at all A fair follow-up is why anyone accepts the complication. Pure single-parameter guarantees are achievable, but the mechanisms that fit iterative training — additive Gaussian-type noise, combined with the accounting used when each step touches a random subsample — naturally produce a two-parameter statement, and they produce much tighter, more useful epsilons than pure mechanisms do at the same accuracy. Delta is the price: a genuinely tiny probability of the promise failing, bought in exchange for a model somebody can actually deploy. That is a reasonable trade at one in a billion. It is not a reasonable trade at one in a thousand. ## The direction of the claim The error to avoid is treating delta as a rounding term, a numerical tolerance, or a convergence threshold. It is none of those. It is the probability that the entire epsilon promise, against the adversary trying to tell whether a record was used, does not apply — and the definition is silent on what happens in that case. An epsilon reported without its delta is half a statement, and an epsilon reported beside a delta nobody checked against the dataset size is a statement nobody has verified.
- A report gives epsilon 2 and delta 0.01 on fifty thousand training rows. How strong is that?Weak enough that the epsilon does not matter. One over n is two in a hundred thousand, and delta is five hundred times larger, so the additive slack alone covers a construction exposing hundreds of records. Ask for a delta far below one over n before reading the epsilon at all.
- Why not always use a pure single-parameter guarantee and avoid delta entirely?You can, but the mechanisms that suit iterative training and subsampled steps naturally give a two-parameter bound, and they yield far tighter epsilons at the same accuracy. Delta buys a deployable model for a genuinely negligible failure probability — which is only true if it is genuinely negligible.
- Can a mechanism trade a smaller epsilon for a larger delta?Yes, and that is precisely why the pair must be quoted together. The same run can be reported at several points along that trade, so an epsilon on its own is not comparable across teams, across mechanisms, or across two reports from the same team.
Delta is the fine print saying the promise may simply not apply on a small share of occasions. If that share is comparable to one over the number of people covered, the fine print is large enough to swallow somebody whole.
saying these in an interview costs you the question
- Quotes epsilon and omits delta entirely
- Calls delta a rounding error or numerical tolerance
- Picks delta by habit without checking dataset size
- Thinks a small epsilon rescues a large delta
- Assumes delta bounds how bad the failure is