skip to content

What does differential privacy guarantee when publishing aggregate statistics from customer data, and what does its privacy budget mean in practice?

level: seniorimportance: nice to knowfreq 25%

answer

  1. with or without one person
  2. calibrated noise on results
  3. epsilon bounds the difference
  4. every query spends budget
  5. small groups get noisy

basics

~20 s

Differential privacy guarantees a published result is almost equally likely whether or not any one person's data was included, bounded by epsilon. Noise is added to results, every release spends part of a total budget, and small counts become noisy.

solid answer

~50 s

**Differential privacy** is a property of the *process* that produces statistics, not of the data. It guarantees that the probability of any output changes by at most a factor of **e^ε** when one individual's data is added or removed, so an observer learns almost nothing about any single person from the result. It is achieved by adding **calibrated random noise** to answers. **ε (epsilon)** is the privacy parameter: smaller means stronger privacy and noisier answers. Guarantees **compose**, so ten releases at ε = 0.1 each cost ε = 1 in total — that total is the **privacy budget**, and once spent no more releases are safe. Unlike k-anonymity it resists linkage with outside data and repeated releases, but it costs accuracy, which is worst for **small groups**, and choosing ε and the **privacy unit** (a user, or a single event) are policy decisions.

go deeper

for a junior

Know that differential privacy adds noise so results barely depend on any one person.

for a middle

Explain epsilon, neighbouring datasets and why smaller epsilon means noisier answers.

for a senior

Explain composition as a budget, the privacy unit, contribution bounding and the accuracy cost for small groups.

for a principal

Set epsilon and budget policy for published statistics and decide which releases justify differential privacy over simpler aggregation.

## The guarantee NIST SP 800-226 states the original definition: a randomised mechanism M satisfies **ε-differential privacy** if, for all **neighbouring datasets** D1 and D2 — datasets that differ in the data of one individual — and all possible outcomes S, the probability of M(D1) falling in S is at most e^ε times the probability of M(D2) falling in S. In words: **whether or not you are in the data, the published result looks almost the same**, so the result reveals almost nothing about you specifically. Key points: - It is a property of the **release mechanism**, not of a dataset. - **ε** is called the privacy parameter, privacy loss or privacy budget; smaller ε means stronger privacy. - It is achieved by adding **random noise** calibrated to how much one person can change the answer. ## Properties that make it useful NIST SP 800-226 lists properties that follow mathematically from the definition: 1. **Resistance to auxiliary data**: linking the output with outside datasets does not break the guarantee, unlike de-identified microdata. 2. **Composition**: the total privacy loss of several releases can be added up; its example is ten analyses at εi = 0.1 giving a total budget of ε = 1. 3. **Post-processing invariance**: further computation on a differentially private output stays differentially private. ## The privacy budget in practice | Decision | What it means | |---|---| | **Total ε** | the upper bound on privacy loss for all releases from the data | | **Allocation** | how the budget is split across queries, dashboards or model training runs | | **Privacy unit** | whose contribution is protected — a user (all their events) or a single event; user-level is stronger | | **Contribution bounding** | capping how many rows one person can add, so noise can be calibrated | | **Budget exhaustion** | once spent, further releases need new data or a policy decision | ## Costs and pitfalls - **Accuracy**: noise is roughly fixed in size for a given ε, so large counts stay accurate while **small groups** become unreliable or meaningless. - **Event-level protection** can leave a heavy user exposed through many events; the privacy unit matters. - **Hand-rolled implementations** fail in subtle ways (floating-point noise sampling, unbounded contributions). Use a vetted library and have the design reviewed. - **ε has no universal right value**: it is a policy choice that trades privacy for usefulness, and it must be documented. ## Differential privacy versus k-anonymity k-anonymity protects a **released table** against singling out on chosen quasi-identifiers, but it breaks under linkage and repeated releases. Differential privacy protects **query results or aggregates** with a formal, composable bound, but it gives noisy numbers and does not release individual rows as they are. ## Why interviewers ask it It appears in senior data interviews at organisations that publish statistics or train models on personal data. The good answer states the **neighbouring-datasets guarantee**, explains ε and composition as a **budget**, and is candid about accuracy for small groups.

  • Why are small groups a problem for differentially private counts?
    The noise needed to hide one person is about the same size whatever the group size. On a count of a million it is negligible; on a count of twelve it can be larger than the true value, so small-group results must be suppressed or aggregated further.
  • What does choosing a user-level rather than an event-level privacy unit change?
    User-level protects everything one person contributed, so a heavy user with many events is hidden as well as a light one. It needs contributions per user to be bounded, and it usually costs more noise for the same epsilon.

saying these in an interview costs you the question

  • Describing differential privacy as a property of a dataset rather than of the release process
  • Believing each query can use the full epsilon independently
  • Expecting accurate results for very small groups
  • Treating epsilon as a technical constant rather than a policy choice