skip to content

When producing a masked test extract, how does pseudonymisation differ from redaction?

level: middleimportance: should knowfreq 56%

answer

  1. One destroys, one substitutes consistently
  2. Ask what the suite does with the field
  3. Equal values must stay equal
  4. Formats must still pass validation
  5. Half-applied is not half-safe

basics

~20 s

Redaction destroys a value: the field becomes a blank or a constant, and nothing can be inferred from it. Pseudonymisation swaps each value for a stable surrogate, so equal values stay equal and joins and grouping still work.

solid answer

~50 s

Both remove the real value from a test extract; they differ in what they preserve. **Redaction** overwrites the field — a fixed placeholder, a blank, an unrelated random value — and deliberately keeps no relationship to the original. It is the right choice for fields nothing in the suite reasons about. **Pseudonymisation** substitutes a surrogate through a consistent mapping, so the same source value always becomes the same replacement, everywhere it appears. That keeps referential and behavioural properties alive: rows still join, duplicate detection still finds the duplicates, group counts still match. It is also weaker: because the structure survives, a person can sometimes still be singled out from the surrounding fields, and how easily is genuinely contested and depends on the dataset. Practical rules: pseudonymise anything a test joins, groups or deduplicates on; redact everything else; keep formats valid so validation still passes; and treat a mapping that is not applied uniformly as no masking at all.

code

pseudocode · 13 lines
pseudocode
// consistent: the same input always yields the same surrogate
surrogates = {}
function pseudonymise(value):
    if value not in surrogates:
        surrogates[value] = formatPreservingSurrogate(value)
    return surrogates[value]

// joins and duplicate detection survive
student.guardianContact = pseudonymise(student.guardianContact)
guardian.contact       = pseudonymise(guardian.contact)

// WRONG: a fresh random value per row
guardian.contact = randomContact()   // two rows for one person now look like two people

go deeper

for a junior

Learn the plain distinction: redaction throws the value away, pseudonymisation swaps it for a stand-in that is always the same for the same original. Being able to say why the second one is what keeps records joinable is enough at this level.

for a middle

Explain the mechanics an interviewer is probing for: consistency of the mapping across tables and runs, format preservation so validation still passes, and which behaviours (joins, grouping, duplicate detection) break when each property is missing.

for a senior

Show how you verify the transform rather than trusting it. Describe checks that run against the finished extract, why a partially applied masking run is discarded rather than patched, and how the mapping is kept away from the data it protects.

for a principal

Own the standing rule for the organisation: which fields may ever leave production, who reviews a change to that list, how long an extract may live, and what evidence the pipeline must produce before an extract is publishable. Say plainly that pseudonymised data stays under controls.

Once a team decides to build test data from real records, the interesting work is the transform applied on the way out. Two transforms get confused in interviews, and the difference is about what survives. ### Redaction destroys the value The field is overwritten with something carrying no information about the original: a fixed placeholder, an empty value, an unrelated random string. After redaction there is no mapping back, and equal originals do not become equal replacements. This is the strongest option and the correct default for every field the suite does not reason about — a free-text note, a home address, a contact number nothing dials. The cost is that any behaviour depending on the field's distribution or on equality between rows is now untestable. If two records referred to the same guardian and both are redacted to the same constant, they now look identical when they were not; if they are redacted to different random values, they look distinct when they were the same person. ### Pseudonymisation preserves relationships A pseudonym is a surrogate produced by a **consistent mapping**: the same input always yields the same output, throughout the extract and, if the mapping is retained, across successive extracts. That single property is what makes an extract still useful: - **Joins survive.** A key rewritten in one table and rewritten identically in the referencing table still resolves. - **Cardinality survives.** Twelve rows about one person are still twelve rows about one person, so grouping and per-entity counts stay right. - **Duplicate detection survives.** A case about merging two records that share a contact value still has two records sharing a contact value. Two qualities matter alongside consistency. **Format preservation**: if a field is validated on read, the surrogate must satisfy the same rules, or the extract fails to load and everyone blames the loader. **Distribution awareness**: replacing every name with a fixed-width surrogate quietly deletes the long-name and unusual-character cases the extract was wanted for. Pseudonymisation is deliberately weaker protection than redaction. The structure that makes the data useful is the same structure that can let an individual be singled out from the fields around them — a rare combination of school, year group and role can identify one person even with every name replaced. How much residual risk remains is genuinely contested and dataset-specific; the honest interview answer is that pseudonymised data is *reduced-risk*, not *anonymous*, and is still treated as personal data by cautious teams. ### The failure that makes this concrete A school timetable planner, 11-person team. The nightly extract pipeline rewrites personal fields table by table: students first, then guardians, then enrolments. On one run the job rewrote the 4,213 student rows, then failed partway through the guardian table; the step's transaction rolled back only that table's partial work, and the pipeline reported a failure and exited. The result was an extract that looked masked. Student names and contact values were surrogates. But 3,912 guardian rows still held real contact values, and — worse — the guardian rows now referenced student identifiers that had been rewritten, so the two tables no longer agreed. The extract was simultaneously *unsafe* (real personal data present) and *broken* (relationships inconsistent), and it was loaded, because the only signal was an exit code nobody watched. Two lessons a strong candidate draws. First, a partially applied transform is not partial protection — it is none, and the extract must be discarded and rebuilt rather than patched. Second, the pipeline needs its own **verification step** that runs against the finished extract and not against the job's exit status: assert that no field matches the shape of a real contact value, that every foreign key still resolves, and that per-entity row counts match the source aggregates. Publish the extract only when those checks pass. ### Choosing per field Walk the columns and ask what the suite does with each: - Joined on, grouped by, compared for equality, or deduplicated → pseudonymise with a consistent, format-preserving mapping. - Displayed but never reasoned about → redact. - Not needed at all → drop the column from the extract entirely; the safest field is the absent one. - Free text that may quote personal details inside a sentence → drop or replace wholesale, because targeted rewriting of prose is unreliable. And keep the mapping itself out of the test environment. If the surrogate table sits next to the data it protects, the extract is reversible by anyone who can read both, and the whole exercise buys nothing.

  • A masking run failed after rewriting some tables and rolling back others. Can the extract be repaired by rerunning only the failed step?
    No. Rerunning the failed step leaves the already-rewritten tables carrying surrogates generated in a different run, so relationships between tables no longer agree, and any table the failure skipped may still hold real values you have not enumerated. Discard the extract and rebuild it from the source in one run. The durable fix is a verification pass over the finished extract — no real-looking values, every foreign key resolving, aggregate counts matching — with publication gated on those checks rather than on the job's exit status.
  • Why is pseudonymised data still often treated as personal data?
    Because the mapping exists and the structure survives. If the surrogate table is reachable, the transform is simply reversible. Even without it, a rare combination of surviving attributes can single out one individual, and how easily that happens is contested and depends entirely on the dataset's size and skew. The defensible position is that pseudonymisation reduces risk rather than eliminating it, so the extract keeps access controls, a retention limit and a rebuild path rather than being treated as public.
  • Which fields should never be masked at all, but dropped?
    Anything no case reads. A column that survives only because it was in the source is pure liability: it must be transformed correctly forever, and it will eventually be transformed incorrectly. Free-text fields deserve the same treatment even when they are read, because personal details hide inside sentences and targeted rewriting of prose is unreliable — replace the whole value or omit the column. The narrowest extract that still supports the suite is the cheapest one to keep safe.

saying these in an interview costs you the question

  • Uses the two terms interchangeably
  • Randomises each row separately and expects joins to survive
  • Thinks pseudonymised data is anonymous and needs no controls
  • Keeps the surrogate mapping beside the masked extract
  • Reruns only the failed step after a partial masking failure
  • Produces surrogates that no longer pass field validation

context