skip to content

PII Classification & Masking

How sensitive columns are found and classified, and how tokenization, keyed hashing, pseudonymization and k-anonymity protect them. Interviewers check analytics stays useful and safe.

on this pageshow

explore

questions

5

When classifying warehouse columns, how do direct identifiers, quasi-identifiers and sensitive attributes differ, and why does each need a different control?

level: juniorimportance: must knowfreq 55%

answer

  1. who it is, alone
  2. who it is, in combination
  3. what you learn about them
  4. remove, generalise, restrict

basics

~20 s

A direct identifier names a person by itself (email, national ID). Quasi-identifiers identify people only in combination (birth date, postcode, sex). Sensitive attributes are what must not leak (diagnosis, salary). Each class needs a different protection.

solid answer

~50 s

Classification sorts columns by **how they can hurt someone**. A **direct identifier** points to one person on its own — name, email, phone, national ID — so it is removed, tokenised or keyed-hashed before wide use. **Quasi-identifiers** are harmless alone but identifying **in combination** — date of birth, postcode and sex together single out most people — so they are **generalised** (age bands, shorter postcodes) or suppressed when data is shared. **Sensitive attributes** — health, salary, religion — are the facts whose disclosure causes harm, so access to them is **restricted** and they are the values an attacker wants to link to a person. A column can also carry a **sensitivity level** (for example low, moderate, high) that drives who may read it. Getting the class right is what lets policy be applied per tag instead of per table.

go deeper

for a junior

Be ready to define the three classes with an example of each, and to say why quasi-identifiers still identify people.

for a middle

Explain how a classification tag drives masking, access and retention policies across many tables.

for a senior

Show how you would run classification for a whole warehouse, including owners, review of automated proposals and defaults for untagged columns.

for a principal

Decide the sensitivity levels and their controls for the organisation, balancing analyst usefulness against the harm each level guards against.

## Why classify at all A warehouse holds thousands of columns. Protecting them one table at a time does not scale, and it fails silently when a new table appears. **Classification** assigns each column a **class** (what kind of risk it carries) and often a **sensitivity level** (how much harm disclosure could do), so that masking, access and retention rules can be written once against the class and applied everywhere the tag appears. ## The three classes | Class | Definition | Examples | Typical control | |---|---|---|---| | **Direct identifier** | identifies a person on its own | name, email, phone number, national ID, account number | remove, tokenise or replace with a keyed hash before broad use | | **Quasi-identifier** (indirect identifier) | identifies only in combination with other columns or outside data | date of birth, postcode, sex, job title, rare purchase | generalise (bands, truncation) or suppress for shared data; keep full precision only where needed | | **Sensitive attribute** | the fact whose disclosure harms the person | diagnosis, salary, religion, precise location | restrict access; the value an attacker tries to attach to a person | A column can belong to more than one class: a precise location trail is both a quasi-identifier and sensitive. ## Why quasi-identifiers matter Removing names and emails is not enough. Research summarised by NIST (NIST IR 8053) recounts that date of birth, five-digit ZIP code and sex together were estimated to uniquely identify up to 87% of the US population in the 1990 census, and a later study using 2000 data put it at 62%. That is why quasi-identifiers need their own treatment rather than being left in the clear once direct identifiers are gone. ## Sensitivity levels Many organisations add an **impact level** to each class. NIST SP 800-122 recommends categorising personal data by a confidentiality impact level of **low, moderate or high**, based on factors such as how identifiable it is, how many people it covers and how sensitive the fields are. A scheme might read: - **Public**: can be published. - **Internal**: any employee may read. - **Confidential**: need-to-know, masked by default. - **Restricted**: named individuals only, logged, never exported. ## How the class drives controls 1. The **classification tag** is attached to the column in the catalog. 2. **Access and masking policies** are written against tags, not tables. 3. **Retention and deletion** rules read the same tags. 4. When a derived column inherits identifying data, its tag must follow it. ## Why interviewers ask it It is the vocabulary every other privacy question builds on. A good answer defines the three classes, gives an example of a quasi-identifier combination that re-identifies people, and connects the class to a **different control** for each.

  • Is a customer ID generated by your own system a direct identifier?
    Inside the organisation, usually yes in effect, because anyone with access to the customer table can look the person up with it. It is often treated as a pseudonymous key, protected by restricting access to the table that links it back to a name.
  • Who should own a column's classification, and how is it kept current?
    The data owner or steward of the dataset is accountable for it, helped by automated scans that propose tags on new or changed columns. Proposals are reviewed rather than trusted blindly, and a column with no classification is treated as sensitive until reviewed.

saying these in an interview costs you the question

  • Believing that removing names and emails makes a dataset anonymous
  • Treating every personal field as needing the same single control
  • Classifying tables instead of individual columns
  • Assuming a system-generated customer ID carries no identification risk
open as a page

To protect an email column that analysts use as a join key, when would you tokenize it, encrypt it, or replace it with a keyed hash?

level: middleimportance: must knowfreq 55%

basics

~20 s

Use a keyed hash when analysts only need a stable join key and nobody needs the email back. Tokenize when an authorised process must recover it. Encrypt when the value itself must be recoverable by key holders. None of the three makes it anonymous.

open as a page

How does automated sensitive-data discovery find personal data in a warehouse, and why do column-name rules alone miss much of it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Scanners sample column values and match them with patterns plus checksums, reference dictionaries and trained entity recognisers, then propose tags with a confidence score. Names alone miss badly named columns, free text and JSON that hide personal data.

open as a page

What does k-anonymity guarantee for a released dataset, how is it achieved, and which attacks does it still leave open?

level: seniorimportance: should knowfreq 35%

basics

~20 s

k-anonymity guarantees each record shares its quasi-identifier values with at least k-1 others, achieved by generalising and suppressing values. It still leaks when a group shares one sensitive value or an attacker has background knowledge.

open as a page

What does differential privacy guarantee when publishing aggregate statistics from customer data, and what does its privacy budget mean in practice?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

Differential privacy guarantees a published result is almost equally likely whether or not any one person's data was included, bounded by epsilon. Noise is added to results, every release spends part of a total budget, and small counts become noisy.

open as a page