A loyalty team wants to sell an 'anonymised' grocery basket dataset - what is the privacy threat?
answer
- Which identifiers actually got removed
- Sparse data makes people unique
- Quasi-identifiers do the linking
- The recipient is fully authorised
- Ship the answer, not the rows
basics
~20 sStripping names and loyalty IDs gives pseudonymised data, not anonymised. Purchase histories are sparse and near-unique, so rare item combinations act as identifiers and a partner can re-link shoppers to other data. Aggregate, generalise, or release answers instead of rows.
solid answer
~50 sThe threat is re-identification by a fully authorised recipient. Removing the name and loyalty ID removes direct identifiers only; what remains - per-visit item lists with timestamps and store - is high-dimensional and sparse, so most shoppers are unique on a handful of rare purchases. The partner, or anyone they later combine data with, can link a record back to a person and then read sensitive inferences off it: pregnancy, illness, religious observance, addiction. Note the adversary position: no breach, no unauthorised access, a contractual counterparty doing exactly what the deal allows. Mitigations are data-shaped, not access-shaped - generalise categories and time, suppress rare items, aggregate to counts above a threshold, apply k-anonymity on the quasi-identifiers or differential privacy on a query interface, or best, ship answers to the partner's questions rather than rows. Contract terms banning re-identification are a backstop, not the control.
go deeper
Know that deleting names and customer IDs is pseudonymisation, not anonymisation, and that detailed behavioural records can still point to one person. Do not describe such a dataset as anonymous.
Explain quasi-identifiers: attributes that identify nobody alone but isolate a person in combination, especially in sparse data like purchase histories. Be able to name aggregation, generalisation and suppression as the mitigations.
Show you can rate the risk and propose a ladder of release shapes - answers instead of rows, then aggregates, then generalised rows with a stated k - and explain why a contractual clause is a backstop rather than a control.
Own the policy: what data may ever leave at row level, who signs off, and how much analytic utility the company is willing to trade for a formal guarantee. Be ready to defend a firm no on row-level releases against a revenue argument.
## The claim, and why it fails A grocery loyalty programme wants to monetise its data by selling a basket-level dataset to a consumer-goods partner. The team's position is that the release is anonymous because names, loyalty numbers and payment tokens have been dropped, leaving only per-visit item lists with a store and a timestamp. Anonymous data, they argue, is out of the privacy model entirely. The claim confuses two different things. - **Pseudonymisation** replaces direct identifiers with a surrogate. The data still relates to a distinguishable individual, and re-identification remains possible with additional information. It is a control, not an exit from the model. - **Anonymisation** means re-identification is no longer reasonably possible for anyone, including the recipient combining it with data you have never seen. That is a high bar, and it is a property of the *release plus the world*, not of a column you deleted. ### Why baskets in particular re-identify Purchase data is high-dimensional and sparse: thousands of possible products, each shopper touching a tiny subset, at particular stores, on a repeating weekly rhythm. Attributes that are not identifiers on their own become **quasi-identifiers** in combination - a specific niche brand, an unusual dietary product, a store plus a Tuesday-evening pattern. Uniqueness rises fast with the number of attributes, so a handful of distinctive points is often enough to isolate one record. Anyone holding an auxiliary source - their own transactions, a public review, a social post about a purchase, a colleague's known habits - can then attach a name. And re-identification is not the whole harm. Basket contents support sensitive inference in themselves: pregnancy, chronic illness, religious observance, alcohol dependence. Even attribute disclosure without a name attached - learning that the one household in this postcode segment buys a particular medication - is a privacy harm. ### The adversary here is authorised This is the point worth landing in an interview. There is no intruder, no misconfiguration, no exfiltration. The adversary position is a *contractually authorised third party*, receiving exactly what the agreement grants, plus whatever else they already hold. A security threat model of the transfer will happily confirm mutual TLS, a signed agreement, and a controlled export path - and report nothing. ## Controls that actually apply They work on the shape of the data, not on the gate in front of it: - **Aggregate.** Release counts, shares and trends above a minimum cell size rather than rows. Most partner questions are answerable this way. - **Generalise.** Product category instead of SKU, week instead of timestamp, region instead of store. - **Suppress.** Drop rare items and rare combinations - they carry most of the re-identification power and little analytic value. - **k-anonymity.** Ensure every released record shares its quasi-identifier values with at least k-1 others. Its known limit is homogeneity: if all k records in a group share the same sensitive value, the attacker learns it without ever isolating an individual - which is what l-diversity and t-closeness were introduced to address. - **Differential privacy.** For a query interface, calibrated noise gives a formal, quantified bound (an epsilon budget) on how much any one person's presence can change an output. It is the only option in this list that offers a guarantee rather than a heuristic. - **Do not ship rows at all.** Take the partner's question, run it yourself, return the answer. This eliminates the release rather than sanitising it, and is frequently accepted once someone asks. Contractual bans on re-identification and on linkage with other datasets belong in the deal, but treat them as a deterrent and a remedy after the fact, not as the technical control. You cannot audit an intention. ## How to raise it in a review Stated as a threat with a path, this gets traction: *a partner joins the release to their own loyalty or ad data, isolates households by rare purchases, and learns a health condition; the harm falls on customers who were told their data was anonymous*. Then offer the ladder - answers instead of rows, else aggregates, else generalised and suppressed rows with a stated k - and let the business pick a rung. Reviewers who only assert "this is not really anonymous" lose the argument; reviewers who arrive with a cheaper release shape that still answers the partner's question usually win it. ## Vocabulary check The *threat* is re-identification and sensitive inference by the recipient. The *vulnerability* is a release retaining quasi-identifiers at row level. The *risk* is the rated harm to customers plus trust and regulatory exposure. The *controls* are aggregation, generalisation, suppression, a formal privacy model, or not releasing rows.
- The partner signs a contract promising never to re-identify anyone - is that enough?No. A contract is a deterrent and a remedy, not a control: you cannot observe compliance, the data may be re-linked by a successor, a subcontractor or a future acquirer, and a leak of the release voids the promise entirely. Keep the clause, but size the release so that a breach of it would not be catastrophic. Technical shape first, contract second.
- What is k-anonymity's main weakness, and what addresses it?Homogeneity. Every record sharing quasi-identifiers with k-1 others still leaks the sensitive attribute if all k records share the same value - the attacker learns the condition without isolating the person. l-diversity requires variety in the sensitive attribute within each group, and t-closeness requires each group's distribution to resemble the overall one. Both cost utility, which is the tradeoff to state out loud.
- How would you decide between publishing an aggregated extract and a differentially private query interface?Aggregates are cheaper and easier to reason about, and fit a partner with a few fixed questions. A query interface fits open-ended exploration but leaks across repeated queries unless a budget is enforced, so it needs an epsilon budget, per-analyst accounting, and someone who understands the utility cost. If the partner's questions are enumerable, a curated aggregate release is usually the better engineering answer.
Blurring the faces in a photo of a small village does not make people unidentifiable. The clothes, the dog and the front door do the work the face was doing.
saying these in an interview costs you the question
- Says removing names makes data anonymous
- Treats pseudonymised data as outside the privacy model
- Relies solely on a contractual no-re-identification clause
- Ignores that the recipient holds other datasets
- Thinks encryption in transit resolves a re-identification risk
- Assumes k-anonymity blocks sensitive-attribute disclosure