What does differential privacy guarantee when publishing aggregate statistics from customer data, and what does its privacy budget mean in practice?
answer
- with or without one person
- calibrated noise on results
- epsilon bounds the difference
- every query spends budget
- small groups get noisy
basics
~20 sDifferential privacy guarantees a published result is almost equally likely whether or not any one person's data was included, bounded by epsilon. Noise is added to results, every release spends part of a total budget, and small counts become noisy.
solid answer
~50 s**Differential privacy** is a property of the *process* that produces statistics, not of the data. It guarantees that the probability of any output changes by at most a factor of **e^ε** when one individual's data is added or removed, so an observer learns almost nothing about any single person from the result. It is achieved by adding **calibrated random noise** to answers. **ε (epsilon)** is the privacy parameter: smaller means stronger privacy and noisier answers. Guarantees **compose**, so ten releases at ε = 0.1 each cost ε = 1 in total — that total is the **privacy budget**, and once spent no more releases are safe. Unlike k-anonymity it resists linkage with outside data and repeated releases, but it costs accuracy, which is worst for **small groups**, and choosing ε and the **privacy unit** (a user, or a single event) are policy decisions.
go deeper
Know that differential privacy adds noise so results barely depend on any one person.
Explain epsilon, neighbouring datasets and why smaller epsilon means noisier answers.
Explain composition as a budget, the privacy unit, contribution bounding and the accuracy cost for small groups.
Set epsilon and budget policy for published statistics and decide which releases justify differential privacy over simpler aggregation.
## The guarantee NIST SP 800-226 states the original definition: a randomised mechanism M satisfies **ε-differential privacy** if, for all **neighbouring datasets** D1 and D2 — datasets that differ in the data of one individual — and all possible outcomes S, the probability of M(D1) falling in S is at most e^ε times the probability of M(D2) falling in S. In words: **whether or not you are in the data, the published result looks almost the same**, so the result reveals almost nothing about you specifically. Key points: - It is a property of the **release mechanism**, not of a dataset. - **ε** is called the privacy parameter, privacy loss or privacy budget; smaller ε means stronger privacy. - It is achieved by adding **random noise** calibrated to how much one person can change the answer. ## Properties that make it useful NIST SP 800-226 lists properties that follow mathematically from the definition: 1. **Resistance to auxiliary data**: linking the output with outside datasets does not break the guarantee, unlike de-identified microdata. 2. **Composition**: the total privacy loss of several releases can be added up; its example is ten analyses at εi = 0.1 giving a total budget of ε = 1. 3. **Post-processing invariance**: further computation on a differentially private output stays differentially private. ## The privacy budget in practice | Decision | What it means | |---|---| | **Total ε** | the upper bound on privacy loss for all releases from the data | | **Allocation** | how the budget is split across queries, dashboards or model training runs | | **Privacy unit** | whose contribution is protected — a user (all their events) or a single event; user-level is stronger | | **Contribution bounding** | capping how many rows one person can add, so noise can be calibrated | | **Budget exhaustion** | once spent, further releases need new data or a policy decision | ## Costs and pitfalls - **Accuracy**: noise is roughly fixed in size for a given ε, so large counts stay accurate while **small groups** become unreliable or meaningless. - **Event-level protection** can leave a heavy user exposed through many events; the privacy unit matters. - **Hand-rolled implementations** fail in subtle ways (floating-point noise sampling, unbounded contributions). Use a vetted library and have the design reviewed. - **ε has no universal right value**: it is a policy choice that trades privacy for usefulness, and it must be documented. ## Differential privacy versus k-anonymity k-anonymity protects a **released table** against singling out on chosen quasi-identifiers, but it breaks under linkage and repeated releases. Differential privacy protects **query results or aggregates** with a formal, composable bound, but it gives noisy numbers and does not release individual rows as they are. ## Why interviewers ask it It appears in senior data interviews at organisations that publish statistics or train models on personal data. The good answer states the **neighbouring-datasets guarantee**, explains ε and composition as a **budget**, and is candid about accuracy for small groups.
- Why are small groups a problem for differentially private counts?The noise needed to hide one person is about the same size whatever the group size. On a count of a million it is negligible; on a count of twelve it can be larger than the true value, so small-group results must be suppressed or aggregated further.
- What does choosing a user-level rather than an event-level privacy unit change?User-level protects everything one person contributed, so a heavy user with many events is hidden as well as a light one. It needs contributions per user to be bounded, and it usually costs more noise for the same epsilon.
saying these in an interview costs you the question
- Describing differential privacy as a property of a dataset rather than of the release process
- Believing each query can use the full epsilon independently
- Expecting accurate results for very small groups
- Treating epsilon as a technical constant rather than a policy choice