skip to content

Re-Identification Risk

Removing the obvious identifiers is not enough: quasi-identifiers combine into a unique row, small cohorts stand out, free text carries names the masking never saw. Judging what is left is the skill.

on this pageshow

questions

4

How can a test dataset with every name and contact field removed still identify individuals?

level: middleimportance: must knowfreq 45%

answer

  1. The names are gone; identity is not
  2. Individually dull attributes, jointly rare
  3. Count rows sharing each attribute combination
  4. Unique combination plus an outside dataset

basics

~20 s

Direct identifiers are only part of identity. Attributes that look harmless alone -- birth date, postcode, job title, the month an account opened -- combine into a pattern very few people share, and a unique combination points at one person.

solid answer

~50 s

Removing names, emails and account numbers removes the **direct** identifiers. What remains are **quasi-identifiers**: attributes that are unremarkable individually but rare jointly -- date of birth, postcode or region, gender, employer, tenure, device type, the month an account was opened. Count how many rows share each combination of those and the count is frequently one. A row that is unique on three or four such attributes can be matched to a named person through any other dataset carrying the same attributes: an internal directory, a public register, or simply a colleague's knowledge of their own details. So the question that matters is not *did we transform every column we classified as personal*, but *how many rows share this row's remaining combination*. Where that number is small, the row is still about someone who can be singled out, and many privacy regimes treat such data as personal rather than anonymous.

code

pseudocode · 7 lines
pseudocode
// how many rows share each combination of the surviving attributes?
groups = groupRows(dataset, by = [birthYear, postcodePrefix, jobTitle, signupMonth])

for each group in groups:
    if group.rowCount < AGREED_MINIMUM_GROUP_SIZE:
        report(group.attributeValues, group.rowCount)
        // these rows still single out a person, whatever their columns are called

go deeper

for a junior

Be ready to say what a direct identifier is and to name a few attributes that are not names but still narrow a person down, such as birth date and postcode. Knowing that removing names is not sufficient is the recall expected here.

for a middle

Explain the mechanics: attributes that are individually common combine into a rare pattern, and you measure that by counting how many rows share each combination. Be able to walk an interviewer through the check on a concrete set of columns.

for a senior

Show that you check before a dataset reaches an environment: which combinations produce very small groups, which rows are unique, and what you did about them. Interviewers expect you to have seen a supposedly anonymous extract fail exactly this.

for a principal

Own the standard: which combinations are checked by default, what group size counts as too small for this data, and how the answer shifts with who can reach the environment. That standard is a tradeoff between fidelity and exposure, not a fact.

## Two kinds of identifier live in the same table A dataset copied from production carries two kinds of identifying information, and only one of them is obvious. **Direct identifiers** point at one person on their own: full name, email address, phone number, account number, national identifier, street address. These are what a field-classification exercise finds and what a transformation pass replaces. Once they are gone the dataset *looks* anonymous, and that appearance is the whole problem. **Quasi-identifiers** identify nobody alone but narrow the population when combined: date of birth, postcode or region, gender, employer, job title, tenure, plan tier, the month an account was opened, preferred language, the branch that serves the customer, the device type recorded at signup. Nobody classifies `signup_month` as personal data. Three or four such columns together frequently describe exactly one person. | | Direct identifier | Quasi-identifier | | --- | --- | --- | | Identifies on its own | Yes | No | | Found by a field classification | Usually | Rarely, because each looks harmless | | Transformed by the pass | Yes | Often left alone, because the tests need the values | | Where the danger lives | In the value | In the combination | ## Uniqueness is the thing you can actually measure Because identity comes from the combination, the useful check is a counting one. Group the rows by the combination of remaining attributes and look at the size of each group: - a group of many rows tells you nothing about any individual in it; - a group of a handful means each subject is nearly pinned down; - a group of one means that row *is* a person, whatever its columns are called. Run that count over the combinations that actually occur in your data rather than over every possible pairing. In practice a small number of combinations do most of the damage: geography with a date, geography with an employer, a rare category with almost anything. Two honest caveats. First, the counts describe your extract, not the world. A combination unique in a one-million-row sample may match thousands of people in the real population; a combination shared by three rows may still be unique in the world if your extract is the entire customer base. Treating small counts as identifying is the conservative reading, and the conservative reading is the right default for data about to be copied into a shared environment. Second, the count changes every time the dataset changes -- a new column, a refreshed extract, or a filter applied for one test can all turn a comfortable group into a group of one. ## Identification comes from outside the dataset A unique row is not yet an identified person. It becomes one when something else supplies the name, and that something is usually mundane: 1. **Another internal system** -- a directory, a support desk, a partner feed -- carrying the same attributes next to a name. 2. **Published or purchasable data** -- professional registers, company filings, membership lists, public profiles. 3. **A person's own knowledge.** Everybody knows their own birth date, region and job title, so anyone with access can find their own row with no outside help at all -- and having found it, they have found everyone else in the same small group. That third route is the one teams forget, and it is why a dataset held inside the company is not automatically safe. It is also why the same rows carry different risk in different places: risk is a property of the dataset together with the environment it sits in and whatever can be joined to it. ## What to do when the count comes back small The responses are ordinary, and each trades fidelity against exposure: - **Widen the attribute** so more rows share a value -- a year instead of an exact date, a region instead of a postcode, a range instead of a figure. - **Merge sparse categories** into a single catch-all bucket. - **Drop the row**, accepting the hole it leaves in the distribution. - **Substitute a manufactured row** with the same shape and no real subject behind it. Every one of these damages something a test may depend on, which is why this is a judgement rather than a rule, and why it is worth checking the *specific* combinations your tests read instead of blanket-flattening every column. ## What this changes in practice Stop describing an extract as anonymised because the classified columns were transformed. Describe it as what survived: which combinations were checked, how small the smallest group was, how many rows were unique. That sentence is the useful one. It is what a reviewer can argue with, and it is closer to what many privacy regimes are actually asking, because a dataset from which an individual can still be singled out is generally treated as personal data rather than as anonymous data.

  • Which attributes in an ordinary customer dataset behave as quasi-identifiers even though nobody classified them as personal?
    Anything stable and observable from outside: date of birth, postcode or region, gender, employer, job title, tenure, plan tier, first-purchase timestamp, device type, preferred language. Exact dates are the worst offenders because they have very high cardinality and appear on documents people hold. None of them names anybody; three of them together usually do.
  • Why does the same transformed dataset carry different risk in two different environments?
    Risk is a property of the dataset together with who can reach it and what else sits beside it. The same rows are close to harmless on a locked-down machine and dangerous in a shared environment that also holds a customer list with names. Judge the pairing rather than the file, and judge it again whenever access widens.

Blur every face in a crowd photograph and one person is still obvious: the only one in a bright yellow coat, a head taller than anyone else, standing where nobody else was standing.

saying these in an interview costs you the question

  • Calls a dataset safe once name, email and account number are gone
  • Treats identification as a property of single columns only
  • Assumes a row is anonymous because no column identifies alone
  • Never counts how many rows share a combination of attributes
  • Forgets that outside data carries the same harmless-looking attributes
open as a page

Why does a column-by-column transformation pass leave identity in free text, attachments and derived columns?

level: middleimportance: should knowfreq 34%

basics

~20 s

A column pass replaces values in fields somebody classified. It cannot read inside free-text notes or stored documents, and it does not recompute columns built from the values it changed, so names, numbers and reconstructable values survive in all three places.

open as a page

Why do the extreme rows in a transformed dataset stay identifiable when ordinary rows do not?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Being unusual is itself an identifier. The largest account, the only customer in a small region, the row ten times the median: those are recognisable by size or rarity whatever the transformation did to the values, because rank and isolation survive it.

open as a page

How do you judge the residual re-identification risk of a test estate -- the environments and datasets a team tests against -- built from transformed copies of real customer records, and who accepts it?

level: principalimportance: should knowfreq 26%

basics

~20 s

Residual risk is never zero, so the job is to bound it and name an owner. Run the checks -- unique combinations, small groups, extremes, unstructured content -- write down what survived, and have an accountable person accept it explicitly.

open as a page