skip to content

A team exports a customer table to an analytics partner and replaces the email column with its SHA-256 digest, reporting that the export is now "anonymised" and safe to share. Explain what a digest of a personal identifier actually gives you, how hashing differs from encryption and from encoding such as Base64, and what you would do instead.

level: juniorimportance: must knowfreq 68%

answer

  1. unkeyed + deterministic = stable pseudonym
  2. pseudonymisation ≠ anonymisation
  3. same field hashed twice → the datasets join
  4. input entropy bounds protection, not digest length
  5. match-only → keyed hash, secret outside the export

basics

~20 s

Hashing is unkeyed and deterministic, so a hashed identifier is a stable pseudonym, not anonymous data: it still singles out people and joins across datasets, and small identifier spaces enumerate. Encryption is keyed and reversible; encoding is neither and protects nothing.

solid answer

~50 s

That export is **pseudonymised, not anonymised**. A hash is unkeyed and deterministic, so the same email always yields the same digest: the column still singles out individuals, still supports "is this person in the file?" membership tests, and joins cleanly against any other dataset that hashed the same field the same way — including a leaked one. Re-identification is the second blow: preimage resistance forbids a *shortcut*, not enumeration, so any low-entropy identifier (emails from a breach list, phone numbers, national IDs, dates of birth) is recovered by hashing candidates and matching. Contrast the three primitives: **encoding** (Base64, hex) has no key and no security claim; **encryption** is keyed and reversible on purpose; **hashing** has no key and no inverse, but no confidentiality claim either. Choose by goal: read it back later → encryption with keys outside the data; match only → keyed hash with a secret the recipient's dump does not contain; aggregate only → coarsen the field; otherwise don't share it.

code

text · 8 lines
text
our_export:     h(email) -> health_flag
partner_data:   h(email) -> purchase_history
  JOIN ON h(email)  =>  health_flag + purchase_history for the same person
                        (neither side ever saw an email address)

attacker with any leaked address list:
  for candidate in list: if h(candidate) == stored_digest -> identity recovered
  cost is bounded by the size of the list, not by the size of the digest

go deeper

for a junior

Say the three-way distinction cleanly — encoding has no key and no security, encryption has a key and is reversible on purpose, hashing has no key and no inverse — and then the headline: a hash of an identifier is a pseudonym, not anonymity, because the same input always gives the same digest.

for a middle

Add the two concrete harms and name them: joinability across datasets hashed the same way, and re-identification by enumeration because preimage resistance forbids a shortcut, not a sweep of a small input space. Note that digest length is irrelevant to the second.

for a senior

Drive to the control choice from the stated goal: encryption with external key custody if the value must be read back, a keyed hash with the secret outside the exported data if only matching is needed, coarsening if only aggregates are needed. State plainly that only coarsening produces genuinely anonymous output.

for a principal

Frame it as a data-sharing and key-custody decision, not an algorithm choice: who holds the secret, what a compromise of the partner's dump exposes, whether a matching protocol such as private set intersection removes the shared-secret problem, and what minimisation removes the question entirely. Own that pseudonymised exports keep their regulatory obligations and plan retention and rotation accordingly.

## The claim under test: anonymised vs pseudonymised Data is *anonymous* when no one can single out an individual in it, by any means reasonably available. Data is *pseudonymised* when the direct identifier has been swapped for a placeholder but each individual still has a distinct, consistent handle in the file. A digest of an identifier is squarely the second: it is a placeholder computed by a public function with no key. Under most privacy regimes pseudonymised data is still personal data, so "we hashed it" changes the risk profile but does not remove the obligations, and it certainly does not by itself make a partner export safe. ## Determinism is the feature and the flaw A hash function is deterministic by definition — the same input always produces the same digest — and that is exactly what makes it useful for fingerprinting. Applied to an identifier column, determinism has four consequences: - **Singling out.** Every distinct customer has a distinct digest. Collisions are a theoretical possibility for arbitrary inputs, but over a set of a few million real email addresses the mapping is effectively one-to-one, so the digest is a perfect surrogate identity. - **Linkage / joinability.** Any other party who hashed the same field with the same function produces the same digests, so two independently "anonymised" datasets join on the hashed column. Your health flag plus their purchase history become one profile, without either side ever handling an email address. - **Membership testing.** Anyone holding a candidate value can hash it and check whether that person is in your file. Often the sensitive fact *is* membership: presence in a debt-collection, medical or fraud dataset. - **Persistence.** Because the digest is stable forever, it also links the same person across time and across breaches, which is precisely what an identifier is for. ## The second blow: input entropy, not digest length, bounds the protection Preimage resistance is a statement about the *function* — there is no method faster than trying inputs. It is not a statement about the *input*. When the set of plausible inputs is small enough to enumerate, trying them all is the method, and digest size is irrelevant. Emails are enumerable from any leaked address list; phone numbers, national IDs, dates of birth, postcodes and card numbers all live in spaces small enough to sweep on commodity hardware. Moving to SHA-512 changes nothing here, because the attacker never attacks the function. So the sentence to carry: **a hash protects a value only up to the entropy of that value, not up to the size of its output.** ## The three primitives that get confused - **Encoding** — Base64, hex, URL-encoding. Purpose: fit bytes through a channel with restricted characters. No key, no secret, no security claim; anyone reverses it. Treating encoded data as protected is the most basic error in this area. - **Encryption** — a keyed, reversible transformation. Reversibility is the *point*: you encrypt so that you can decrypt later. The algorithm is public; confidentiality lasts exactly as long as key custody does. - **Hashing** — an unkeyed, deterministic map to a fixed-length digest. There is no decrypt operation because there is no inverse to invoke; it is built for integrity, fingerprinting and content addressing. It makes no confidentiality claim at all — that claim is something readers add. The crisp one-liner: hashing has no key and no inverse; encryption has a key and is meant to be undone; encoding has neither a key nor a security property. ## Choosing the primitive from the goal Decide from what the system must still be able to do: - **Must display or transmit the original later** → encryption, with keys managed outside the datastore holding the ciphertext and access separated from the data. - **Must only match or deduplicate, never read** → a keyed hash (a MAC or PRF) whose secret is held outside the exported data. The key is what actually breaks offline enumeration: without it, an attacker cannot even test a candidate. If two parties must match against each other, both need the shared secret, which is a contractual and key-custody decision — or use a matching protocol designed for it, such as private set intersection. - **Must only aggregate or bucket** → reduce precision instead: domain rather than address, age band rather than date of birth, region rather than postcode. Coarsening is the only technique listed here that actually removes information, which is why it is the only one that can produce genuinely anonymous output. - **Not needed by the partner** → do not include the column. Minimisation beats every cryptographic control. - **Authenticating a person by a secret they chose** → passwords are a separate case with their own requirements; do not reason about them from this question. ## Where hashing is the right answer None of this makes hashing weak. It is correct for what it claims: fingerprinting a large artifact into a short identifier, detecting modification, content addressing, and forming part of signature and authentication-tag constructions. The failure here is a goal mismatch — reaching for an integrity primitive to obtain confidentiality and anonymity — not a defect in SHA-256. ## How to answer in an interview Lead with the classification: this is pseudonymisation, so name the two harms in order — a stable pseudonym still singles out and joins, then enumeration re-identifies it outright. Give the three-primitive distinction in one sentence each, and finish by picking the control from the stated goal rather than from the primitive.

  • Would a random per-record salt, stored in the same row, make the export anonymous?
    It helps with two things and not the one that matters here. Precomputed tables stop working and two customers with the same email no longer share a digest, so equality linkage and cross-dataset joins break — which also destroys the matching the analytics partner presumably wanted. It does not stop targeted re-identification: the salt sits next to the digest, so an attacker with the file enumerates the identifier space against that one salt, per record. Salt defeats amortisation, not enumeration.
  • The partner genuinely needs to match customers against their own dataset. What do you propose?
    Replace the plain digest with a keyed hash (a MAC or PRF) computed under a secret that is not present in either party's exported data — held in a key-management service and rotated on a schedule. Matching still works because both sides compute the same keyed value, but an attacker with a dump has no way to test candidate emails. If neither party should learn the non-matching members either, a private set intersection protocol is the purpose-built answer. Pair whichever you pick with minimisation: only the fields the matching actually needs.
  • Is the hashed column still personal data from a compliance standpoint?
    Generally yes. If a value still singles out an individual and can be re-identified by means reasonably available — enumeration, a rainbow table, or joining to another dataset — regulators treat it as pseudonymised personal data, not anonymous data. That means retention limits, subject-access and deletion obligations, and breach duties still apply to the export, and a transfer to a partner is still a transfer of personal data.

Replacing every name in a report with a consistent codename. The report reads as anonymous, but the same person is the same codename on every page, another department's report using the same codebook lines up page-for-page, and anyone who guesses one name has decoded every mention of it.

saying these in an interview costs you the question

  • "Hashes are one-way, so hashed identifiers are anonymous" — conflating no-inverse-function with not-recoverable, and ignoring that a stable pseudonym still singles out and joins.
  • "We'd use SHA-512 to be safe" — a longer digest does nothing when the attacker enumerates a small input space rather than attacking the function.
  • Calling Base64 or hex encoding a form of encryption, or describing hashing as "encryption you can't decrypt".
  • "The partner doesn't have our data, so they can't do anything with the digests" — ignoring that any candidate list plus the public hash function is enough to test membership and join.
  • "Adding a salt makes the export anonymous" — the salt travels with the record and defeats precomputation, not per-record enumeration.

context