skip to content

How do two systems masked weeks apart produce the same replacement for one real value?

level: middleimportance: should knowfreq 45%

answer

  1. Assigned each run, or computed each time?
  2. Two jobs, no coordination between them
  3. A pure function of the real value
  4. One shape before the transformation, one after
  5. The one input that must be identical everywhere

basics

~20 s

The substitute is computed, not assigned: normalise the real value, run it through a keyed one-way transformation with one shared secret, then shape it to the field. Same input and secret, same result - no state shared between jobs.

solid answer

~50 s

Two jobs that each *invent* a substitute and remember it locally can only agree inside their own run. Agreement across time and systems comes from **deriving** the substitute instead: it is a pure function of the real value, a shared secret input and the field's rule. A job run six weeks later, on another system, by another team, feeds the same real value through the same function and lands on the same substitute without any coordination and without shipping anything between them. Two details make it work. First, **normalise before deriving** - trim, fold case, strip formatting - or two spellings of one real value derive two different substitutes. Second, the shared secret is what stops anyone holding a copy from running plausible real values through the published rule and recognising the matches, so it becomes the value whose disclosure would undo the protection of every copy at once.

code

pseudocode · 10 lines
pseudocode
function substitute(raw_value, field_rule):
    canonical = normalise(raw_value)              # trim, fold case, strip separators
    hidden    = keyed_one_way(canonical, SHARED_SECRET)
    return shape_to_field(hidden, field_rule)     # length, alphabet, check digit

# nothing is stored between jobs; only the rules and the secret are shared
assert substitute("AC-0042 ", account_rule) == substitute("ac 0042", account_rule)

# every produced copy records what made it
copy.stamp = { rule_version: 7, secret_version: 3 }

go deeper

for a junior

Know that the replacement value is calculated from the real value rather than picked at random, which is why two copies can end up with the same substitute. Being able to state that distinction is enough here.

for a middle

Explain the mechanics end to end: normalise, transform with a shared secret, shape to the field, and record what produced the copy. Be ready to say why a per-job random assignment cannot give agreement across systems.

for a senior

Show what the design costs. The secret now carries the entire protection, a change to it invalidates every copy at once, and skipping normalisation produces failures that only affect some rows - the hardest kind to diagnose.

for a principal

Own the trade between a computed substitution and a stored one, and set the rule for the whole organisation: one shared rule set, one secret, versions stamped on every copy, and a policy for what a change to the secret obliges everyone to reproduce.

## Assigning a substitute versus deriving one There are two ways a job can decide what a real value becomes, and only one of them survives a system being processed on a different day. | Approach | How agreement is achieved | What happens six weeks later | |---|---|---| | **Assign per job** - draw a plausible value, remember it for the rest of this run | Only inside one run, from the job's in-memory record | The next job draws different values; the two copies no longer match | | **Derive from the value** - compute the substitute from the real value itself | By construction, everywhere, with no shared state | The same real value produces the same substitute, unchanged | Assignment is the intuitive design and it is the usual cause of copies that disagree. It works beautifully within one database and fails the moment a second system is copied. Derivation is the design that makes consistency a property rather than a coincidence: nothing has to be coordinated between the jobs, nothing has to be shipped between teams, and no record has to be kept alive between runs. *(Keeping a durable table of value-to-substitute pairs is a third option with quite different consequences - it makes the substitution reversible for whoever holds the table, which is a separate discipline with its own custody problem. This leaf is about agreement, and derivation achieves it without introducing that store at all.)* ## The four steps of a derived substitute 1. **Normalise the input.** Trim surrounding whitespace, fold case, strip punctuation and separators, and settle on one text-encoding form. `AC-0042`, `ac-0042 ` and `AC 0042` must all become one canonical string before anything else happens. 2. **Transform with a keyed one-way function.** Feed the canonical string plus one shared secret into a one-way transformation. Without the secret the result is reproducible by anyone; with it, only holders of the secret can reproduce it. 3. **Shape the result to the field.** The raw output is not a valid account number, phone number or member reference. Project it into the field's required length, alphabet and check-digit rule, so the value still parses and still passes validation downstream. 4. **Record which rules and which secret produced the copy**, so nobody later joins a copy made under one secret to a copy made under another. Steps 1 and 3 are where most implementations go wrong. Skipping normalisation means the copies agree only for values that were already spelled identically in both sources - and two systems almost never store a phone number or a reference in exactly the same shape. Skipping the shaping step produces substitutes that break parsers and validation, which sends engineers chasing defects that only exist in the copy. ## What the shared secret becomes Once substitutes are derived rather than assigned, the shared secret carries the whole protection. The reasoning is worth stating explicitly, because it surprises people: - The transformation rule is not a secret. It lives in code, in a configuration file, and in the heads of everyone who has read the job. - If there is no secret input, anyone holding a copy can take a list of plausible real values - every account number in a range, every name in a directory - run them through the published rule, and match the outputs against the copy. The substitution is then undone by anyone patient enough, without ever seeing the original data. - A secret input defeats that, because the attempt cannot be reproduced without it. So the secret is now as sensitive as the data the copies were built to protect, and holding both is equivalent to holding the real values. That is the trade the derivation makes: it replaces a stored mapping with a computed one, and concentrates the risk into a single value. **Where that value is stored, who may read it and how it is rotated is a security-management concern with its own answers**; what belongs here is the consequence for consistency. ## Changing the secret changes every copy Because every substitute is a function of the secret, changing it changes all of them at once. A copy produced before the change and a copy produced after it will not join, even though both jobs ran the same rules on the same source. Three practical consequences: - Treat a change of the secret as a **reproduction of every copy in the set**, not as an operation on one system. - **Stamp each copy with the version of the secret and the rule set** that produced it, and refuse to use two copies together whose stamps disagree. - Expect any long-lived derived artefact - a checked-in expectation file, an archived recording, a report snapshot - to be invalidated by the change, because the values inside it were derived under the old secret. The pay-off for accepting those consequences is large: any number of teams, on any number of systems, on any schedule, produce copies that agree, having exchanged nothing but a rule set and one secret.

  • What does the secret input buy over a published rule with no secret at all?
    Without a secret, the rule is reproducible by anyone. A holder of the copy can run every plausible real value through it and match outputs back to rows, recovering the original data without ever seeing it. The secret makes that attempt impossible to reproduce, which is why it becomes as sensitive as the data itself.
  • Why normalise the real value before transforming it, rather than after?
    The transformation is sensitive to every character. Two systems storing one reference as `AC-0042` and `ac 0042` would derive two unrelated substitutes, so the copies would silently fail to join for exactly those rows. Normalising first collapses the spellings into one canonical input, which is the only way the two derivations can land on the same result.
  • What happens to copies already in use when the shared secret is changed?
    Every substitute changes, so an old copy and a new one cannot be joined. Treat a change as a reproduction of the whole set rather than an operation on one system, stamp each copy with the version that produced it, and refuse to combine copies whose stamps differ. Long-lived expectation files and recordings are invalidated too.

Two cooks in different cities follow the same written recipe with the same secret spice blend, and produce the same dish without ever speaking to each other.

saying these in an interview costs you the question

  • Draws a random substitute per job and remembers it locally
  • Derives from the raw value without normalising it first
  • Treats the shared secret as ordinary, readable configuration
  • Thinks a published transformation alone makes a copy safe
  • Expects two independent jobs to agree by coincidence
  • Forgets to shape the output back to the field's format