skip to content

What does k-anonymity guarantee for a released dataset, how is it achieved, and which attacks does it still leave open?

level: seniorimportance: should knowfreq 35%

answer

  1. at least k lookalikes on quasi-identifiers
  2. generalise and suppress
  3. same sensitive value in a group
  4. background knowledge still works
  5. l-diversity, t-closeness refine it

basics

~20 s

k-anonymity guarantees each record shares its quasi-identifier values with at least k-1 others, achieved by generalising and suppressing values. It still leaks when a group shares one sensitive value or an attacker has background knowledge.

solid answer

~50 s

A table is **k-anonymous** when every combination of quasi-identifier values — say age band, postcode prefix and sex — appears in **at least k records**, so a record cannot be narrowed to fewer than k people using those columns. It is achieved by **generalising** values (exact age to a ten-year band, full postcode to a prefix) and **suppressing** outlier rows or cells that would form tiny groups. The guarantee is only about **singling out by quasi-identifiers**. If everyone in a group has the **same sensitive value** — every 30-to-40-year-old in one postcode prefix has the same diagnosis — the attacker learns it without knowing which row is the target (the homogeneity attack). **Background knowledge** can rule out rows within a group. **l-diversity** requires varied sensitive values per group and **t-closeness** requires each group's distribution to stay close to the overall one; both refine but do not fix everything, and repeated releases can be combined.

go deeper

for a junior

Know that k-anonymity means each record looks like at least k-1 others on the identifying combination.

for a middle

Explain generalisation and suppression and the trade-off between k and data usefulness.

for a senior

Describe the homogeneity and background-knowledge attacks, l-diversity and t-closeness, and the limits of choosing quasi-identifiers.

for a principal

Decide when a microdata release with k-anonymity is acceptable versus releasing only aggregates under a formal privacy budget.

## The model **k-anonymity** is a formal property of a released table. Choose the **quasi-identifiers** — columns an attacker could find elsewhere and join on, such as age, postcode and sex. Group rows by their quasi-identifier values into **equivalence classes**. The table is k-anonymous if **every equivalence class has at least k rows**. An attacker who knows a target's quasi-identifiers can then narrow them only to a group of k or more candidates. ## How it is achieved NIST IR 8053 describes the two basic operations on quasi-identifiers: 1. **Generalisation** — replace a precise value with a range or coarser category: age 37 becomes 30-39; postcode 12345 becomes 12xxx. 2. **Suppression** — remove a value or a whole record that would otherwise form a group smaller than k, such as a single 97-year-old in a small area. Choosing how far to generalise is a trade-off: larger k and coarser values protect more but make the data less useful. Tools search for the least generalisation that reaches the target k. ## Attacks that remain | Attack | How it works | Refinement | |---|---|---| | **Homogeneity** | all rows in a group share one sensitive value, so it is disclosed for every member | **l-diversity**: at least l well-represented sensitive values per group | | **Background knowledge** | the attacker knows facts that rule out some rows in the group | l-diversity reduces it but cannot anticipate every fact | | **Skewed distribution** | a group's sensitive values differ sharply from the population, revealing something | **t-closeness**: each group's distribution stays within distance t of the overall one | | **Composition** | two releases, each k-anonymous, are intersected to shrink groups | coordinate releases, or use a model with a composition rule | | **Wrong quasi-identifier set** | a column the designer did not treat as a quasi-identifier is used to link | review against realistic outside data | ## Where it fits - It suits **one-off microdata releases** — row-level data shared for research or with a partner — where analysts need records, not only totals. - Its guarantee depends entirely on **choosing the quasi-identifiers correctly**, which requires guessing what outside data an attacker has. - It says nothing about **aggregate statistics** released repeatedly; for those, a model with a formal budget such as differential privacy is the usual alternative. ## A worked intuition Suppose a medical extract is 3-anonymous on age band, postcode prefix and sex. One group holds three women aged 30-39 in postcode prefix 12, and all three have the same diagnosis. An attacker who knows only that their neighbour fits those three facts learns the diagnosis — no row was singled out, yet the sensitive value leaked. That is the gap l-diversity was proposed to close. ## Why interviewers ask it It is the standard entry point to formal de-identification. A senior answer defines the guarantee precisely (**indistinguishability on quasi-identifiers**), names generalisation and suppression, and knows the homogeneity and background-knowledge attacks and their refinements.

  • How do you pick k for a partner data share?
    From the risk and the recipient. Higher k lowers re-identification risk but costs precision. Teams often start from a policy threshold, test the utility of the generalised data with the recipient, and add contractual controls such as no re-identification attempts and no onward sharing.
  • Why is choosing the quasi-identifier set the weakest step?
    Because the guarantee only covers the columns you declared. If an attacker can link on a column you left precise, such as a rare job title or an exact transaction date, the table can be k-anonymous on paper and still identify people.

saying these in an interview costs you the question

  • Claiming k-anonymity prevents all disclosure of sensitive values
  • Applying k-anonymity to direct identifiers instead of removing them
  • Ignoring that two separately k-anonymous releases can be combined
  • Choosing quasi-identifiers without considering outside data