When classifying warehouse columns, how do direct identifiers, quasi-identifiers and sensitive attributes differ, and why does each need a different control?
answer
- who it is, alone
- who it is, in combination
- what you learn about them
- remove, generalise, restrict
basics
~20 sA direct identifier names a person by itself (email, national ID). Quasi-identifiers identify people only in combination (birth date, postcode, sex). Sensitive attributes are what must not leak (diagnosis, salary). Each class needs a different protection.
solid answer
~50 sClassification sorts columns by **how they can hurt someone**. A **direct identifier** points to one person on its own — name, email, phone, national ID — so it is removed, tokenised or keyed-hashed before wide use. **Quasi-identifiers** are harmless alone but identifying **in combination** — date of birth, postcode and sex together single out most people — so they are **generalised** (age bands, shorter postcodes) or suppressed when data is shared. **Sensitive attributes** — health, salary, religion — are the facts whose disclosure causes harm, so access to them is **restricted** and they are the values an attacker wants to link to a person. A column can also carry a **sensitivity level** (for example low, moderate, high) that drives who may read it. Getting the class right is what lets policy be applied per tag instead of per table.
go deeper
Be ready to define the three classes with an example of each, and to say why quasi-identifiers still identify people.
Explain how a classification tag drives masking, access and retention policies across many tables.
Show how you would run classification for a whole warehouse, including owners, review of automated proposals and defaults for untagged columns.
Decide the sensitivity levels and their controls for the organisation, balancing analyst usefulness against the harm each level guards against.
## Why classify at all A warehouse holds thousands of columns. Protecting them one table at a time does not scale, and it fails silently when a new table appears. **Classification** assigns each column a **class** (what kind of risk it carries) and often a **sensitivity level** (how much harm disclosure could do), so that masking, access and retention rules can be written once against the class and applied everywhere the tag appears. ## The three classes | Class | Definition | Examples | Typical control | |---|---|---|---| | **Direct identifier** | identifies a person on its own | name, email, phone number, national ID, account number | remove, tokenise or replace with a keyed hash before broad use | | **Quasi-identifier** (indirect identifier) | identifies only in combination with other columns or outside data | date of birth, postcode, sex, job title, rare purchase | generalise (bands, truncation) or suppress for shared data; keep full precision only where needed | | **Sensitive attribute** | the fact whose disclosure harms the person | diagnosis, salary, religion, precise location | restrict access; the value an attacker tries to attach to a person | A column can belong to more than one class: a precise location trail is both a quasi-identifier and sensitive. ## Why quasi-identifiers matter Removing names and emails is not enough. Research summarised by NIST (NIST IR 8053) recounts that date of birth, five-digit ZIP code and sex together were estimated to uniquely identify up to 87% of the US population in the 1990 census, and a later study using 2000 data put it at 62%. That is why quasi-identifiers need their own treatment rather than being left in the clear once direct identifiers are gone. ## Sensitivity levels Many organisations add an **impact level** to each class. NIST SP 800-122 recommends categorising personal data by a confidentiality impact level of **low, moderate or high**, based on factors such as how identifiable it is, how many people it covers and how sensitive the fields are. A scheme might read: - **Public**: can be published. - **Internal**: any employee may read. - **Confidential**: need-to-know, masked by default. - **Restricted**: named individuals only, logged, never exported. ## How the class drives controls 1. The **classification tag** is attached to the column in the catalog. 2. **Access and masking policies** are written against tags, not tables. 3. **Retention and deletion** rules read the same tags. 4. When a derived column inherits identifying data, its tag must follow it. ## Why interviewers ask it It is the vocabulary every other privacy question builds on. A good answer defines the three classes, gives an example of a quasi-identifier combination that re-identifies people, and connects the class to a **different control** for each.
- Is a customer ID generated by your own system a direct identifier?Inside the organisation, usually yes in effect, because anyone with access to the customer table can look the person up with it. It is often treated as a pseudonymous key, protected by restricting access to the table that links it back to a name.
- Who should own a column's classification, and how is it kept current?The data owner or steward of the dataset is accountable for it, helped by automated scans that propose tags on new or changed columns. Proposals are reviewed rather than trusted blindly, and a column with no classification is treated as sensitive until reviewed.
saying these in an interview costs you the question
- Believing that removing names and emails makes a dataset anonymous
- Treating every personal field as needing the same single control
- Classifying tables instead of individual columns
- Assuming a system-generated customer ID carries no identification risk