skip to content

Why must a real customer identifier be replaced by the same masked value in every test system?

level: middleimportance: must knowfreq 58%

answer

  1. What survives a copy besides the rows
  2. Two systems, one customer journey
  3. The field two copies are joined on
  4. Per-customer totals compared between two stores

basics

~20 s

Test flows join records across systems on that identifier. If each system replaces it differently the join finds nothing, cross-system journeys stop at the second hop and reconciliation disagrees - failures manufactured by the masking, not by the code.

solid answer

~50 s

A masked copy of real data is only useful if the relationships in the real data survive the replacement. Copies are usually produced by separate jobs - one per database, one per file extract, one per message archive - and each job sees only its own source. If the customer identifier `C-8842` becomes `C-3197` in the order store and `C-5510` in the billing store, everything that follows a customer across the boundary dies: the join returns no rows, an end-to-end journey stops at the second system, a reconciliation control reports a difference for every customer, and an admin-tool test cannot find the account it just created. **The rule is that the substitute must be a function of the real value alone** - not of the system, the job, or the day it ran. Same input, same output, everywhere.

code

pseudocode · 13 lines
pseudocode
# two stores replaced by independent jobs, each inventing its own values
orders_copy.customer_id  = new_random_id()      # C-8842 -> C-3197
billing_copy.customer_id = new_random_id()      # C-8842 -> C-5510

join(orders_copy, billing_copy, on customer_id) -> 0 rows

# the rule that fixes it: the substitute depends on the real value alone
substitute(real_id) = derive(real_id, shared_secret)

orders_copy.customer_id  = substitute(real_id)   # C-8842 -> C-4471
billing_copy.customer_id = substitute(real_id)   # C-8842 -> C-4471

join(orders_copy, billing_copy, on customer_id) -> every matching row

go deeper

for a junior

Be ready to say what a masked copy is and why a customer identifier cannot simply be replaced with any random value. Knowing that two systems have to agree on the substitute is enough at this level.

for a middle

Explain the mechanics: copies are produced by separate jobs, so the substitute has to depend on the real value alone rather than on the job. Name two concrete flows that break when it does not - a join and a multi-system journey.

for a senior

Show the diagnosis. Describe how a mismatch presents as an ordinary application defect, how you would prove the copies disagree, and why the fix belongs to the production of the copies rather than to the failing test.

for a principal

Own the policy: one shared definition of the substitution rules consumed by every producer, an adoption rule for new systems, and a published check that blocks a copy rather than blaming a team. Be ready to argue what it is worth paying for that.

## What "the same value everywhere" actually means A test estate - the collection of environments, databases, file extracts and recorded message archives a team tests against - is normally built by copying real data and replacing the sensitive values inside it. **Referential consistency** is the property that one real value is replaced by one substitute, and always the same substitute, wherever that value appears: in every table of one database, in every extract file, and in every other system copied by a different job on a different day. It is easy to confuse with two neighbouring properties, and the difference is the whole point of this leaf. | Property | Question it answers | Scope | |---|---|---| | Completeness of a copied subset | Is every referenced parent row actually present? | inside one copy | | Validity of manufactured relationships | Do generated references point at rows that exist? | inside one generated dataset | | **Referential consistency** | **Does one real value become one substitute in all copies?** | **across copies produced separately** | Only the third is about agreement *between* separately produced datasets, and it is the one that fails silently. Each copy passes its own inspection. The damage only appears when two of them are used together. ## The flows that break first Suppose the customer identifier `C-8842` becomes `C-3197` in the order store and `C-5510` in the billing store, because two jobs each invented their own substitutes. Nothing errors, and both copies look plausible. What breaks is everything that follows a real relationship across a system boundary. - **Joins between systems.** A query or report joining orders to invoices on the customer identifier returns zero rows - or, worse, the handful of rows that happened to agree by coincidence, which looks like a partial data problem rather than a broken estate. - **Multi-system journeys.** A flow that places an order in one service and expects it on a statement produced by another stops at the second hop. - **Reconciliation controls.** A check that compares per-customer totals between two systems reports a difference for every customer, which is indistinguishable from a genuine accounting defect. - **Replayed traffic and archives.** A recorded request or event keyed by the identifier matches no row in the database it is replayed against. - **Support and administration tooling.** A test that looks an account up by identifier in a second system finds nothing to act on. - **Aggregation across sources.** Anything grouping by customer over two sources produces one bucket per source instead of one per customer, silently doubling the customer count. The expensive part is not the failure, it is the diagnosis. Every symptom above looks exactly like an application defect, so it is triaged as one: an engineer reads the join, reads the service call, adds logging, and only much later discovers that the two copies never agreed in the first place. ## Which fields have to agree, and which do not The property is not free, and it is not needed everywhere. It is needed for fields that **relate records**: 1. Identifiers that two datasets are joined on, including natural keys such as an account, policy or order number quoted between systems. 2. External and correlation references that travel in messages, files or requests between services. 3. Any value a test asserts equality on across a boundary - a login handle used as a lookup, a reference printed on one system and searched for in another. Fields nobody relates on - a free-text complaint note, a street address that is printed but never matched, a description - can be replaced with anything plausible, independently, per job. Deciding this field by field is a design step that is done once and recorded, not re-argued per copy. ## Making it a property of the whole set of copies Consistency cannot be owned by one job, because the failure lives between jobs. What makes it hold in practice: - **One shared definition of the substitution rules**, consumed by every job that produces a copy, rather than reimplemented per system. - **A verification step after producing copies**: take identifiers present in two copies, compare their substitutes, and require an exact match before either copy is published for use. - **Treating a mismatch as a blocker on the copy**, not as a defect on the team using it. The copy is the broken artefact; failing it early costs hours, and failing it late costs the credibility of every cross-system test. - **An adoption rule for new systems**: a system joining the set must use the shared rules before its first copy is trusted in a cross-system flow. ## The workaround that costs the most The common reaction to a broken cross-system test is to rewrite the test so it stays inside one system. That makes the suite green and deletes exactly the integration coverage the shared copies existed to provide - the join, the hand-off and the reconciliation were the interesting behaviour. The other common reaction, a hand-written patch script that rewrites identifiers after the fact, works once and then rots, because it has to be maintained against every future copy. Fixing the production of the copies is the only change that stays fixed.

  • Does every replaced field need to produce the same substitute everywhere?
    No. Only fields that relate records - join keys, external references, values compared across a boundary. A free-text note or an address nobody matches on can be replaced independently by each job. Applying the rule to every field costs effort and buys nothing, and it widens the set of values someone could try to re-derive.
  • How would you prove the copies actually have this property before trusting a cross-system test?
    Sample identifiers that exist in both copies, compare their substitutes, and require a one hundred percent match. Run it as the last step of producing the copies, not as a test in the suite: a mismatch should block publication of the copy, so no engineer ever spends a morning debugging an application that is behaving correctly.
  • A team repairs a broken cross-system test by making it read from one system only. What did that cost?
    The join, the hand-off and the reconciliation were the behaviour under test, and they are now untested. The suite is green because it stopped asking the interesting question. The correct repair is to the production of the copies; narrowing the test hides the estate defect and removes the coverage at the same time.

Two photocopies of one address book, each with the names replaced by a different set of pseudonyms. Each is perfectly readable on its own, and nobody can be looked up in both.

saying these in an interview costs you the question

  • Says each system may replace its data however it likes
  • Treats a broken cross-system join as an application defect
  • Rewrites the test to stay inside one system
  • Thinks a complete parent-child copy inside one database is enough
  • Patches identifiers by hand after each copy is produced