Why do two systems masked by separate jobs stop agreeing on masked values over successive refreshes?
answer
- Nothing forces two jobs to stay aligned
- What changed on one side only
- Formatting differences before the transformation
- A missing row is not a mismatch
- Versions stamped on each produced copy
basics
~20 sNothing forces two independently owned jobs to stay aligned. A changed secret, an upgraded rule, a different normalisation or a local field patch on one side makes the same real value produce two substitutes, and the copies quietly stop joining.
solid answer
~50 sTwo jobs agree on the day someone sets them up identically, and diverge afterwards because they are maintained separately. The usual causes are a secret changed on one side only, a substitution rule upgraded in one job, different normalisation before the transformation, a different field width or check-digit rule, and a local patch made to unblock a rejected load. None of them announces itself: the copies still load, and the join simply matches fewer rows each cycle. **Pin four things and the divergence cannot happen silently**: one shared rule set consumed by every job rather than copied, one normalisation applied before transforming, one secret with a version, and a stamp on every copy recording the rule and secret versions and the point the source was taken from. Then verify: after each refresh, compare substitutes for identifiers present in both copies and reject the refresh below a full match.
code
pseudocode · 13 lines# every produced copy carries what made it
stamp(orders_copy) = { rules: 7, secret: 3, source_taken_at: "2026-04-01T00:00Z" }
stamp(billing_copy) = { rules: 7, secret: 3, source_taken_at: "2026-04-01T00:00Z" }
after each refresh:
if stamp(orders_copy).rules != stamp(billing_copy).rules
or stamp(orders_copy).secret != stamp(billing_copy).secret:
reject("copies produced under different definitions")
sample = 5000 identifiers present in BOTH copies # ignores timing gaps
match_rate = fraction of sample whose substitutes are equal
if match_rate < 1.0:
reject("substitutes diverged: " + failing_fields(sample))go deeper
Know that copies of two systems are produced by separate jobs, and that if those jobs are configured differently the same real value can end up with two different replacements.
Explain concrete causes and their signatures: a changed secret breaks everything at once, a normalisation difference breaks only some rows, and a local field override breaks one field. Be able to say which symptom points at which cause.
Demonstrate the diagnosis and the guard. Separate missing rows from mismatched values first, then show the pinned rule set, versioned secret and per-copy stamp, and put a match-rate check in the refresh so a bad copy never reaches a test suite.
Own who may change the substitution rules and what a change obliges. Argue for a single versioned definition consumed by every producer, a stamp that makes incompatible copies refuse to be combined, and the cost of the alternative in wasted investigation.
## Why separately produced copies drift apart Agreement between two copies is not a state a team reaches once; it is an invariant that has to be maintained against ordinary change. Two jobs, owned by two teams, touched by two backlogs, will diverge unless something forces them not to. The causes, roughly in order of how often they bite: 1. **The shared secret changed on one side.** Someone replaces it during a security exercise or an environment rebuild, and only the systems that were refreshed afterwards carry the new substitutes. Every copy now belongs to one of two eras. 2. **The substitution rule was upgraded in one job.** A wider field, a different alphabet, an added check digit, a swapped transformation - the change is entirely reasonable in isolation and silently breaks agreement with every copy that has not adopted it. 3. **Normalisation differs.** One job trims and folds case before transforming and the other does not, so the copies agree only for rows whose values were already spelled identically in both sources. This is the nastiest variant because it breaks *some* rows, which reads as flaky data rather than a broken configuration. 4. **Field shaping differs.** One job cuts the derived value to twelve characters and the other to sixteen. Both look correct; no substitute matches. 5. **A local override.** A team patches one field to get past a rejected load or an unhappy parser, means to remove it, and does not. Local overrides are the single most common cause of a set of copies that used to agree. 6. **A new system joined without adopting the shared rules.** Its first copy is produced by a job someone wrote from scratch. ## Separate a missing row from a mismatched value Before investigating, split the symptom in two, because the two have unrelated causes: | Observation | Likely meaning | |---|---| | An identifier exists in copy A and not in copy B | The copies were taken from the source at different points; the row simply was not there yet | | An identifier exists in both, with different substitutes | The rules, the normalisation or the secret diverged | Only the second is a consistency defect. Comparing substitutes **for identifiers present in both copies** removes the first from the picture entirely, and a team that skips this step spends its time chasing timing differences it should have expected. ## What has to be pinned - **One shared rule set, consumed rather than copied.** Every job reads the same versioned definition of what happens to each field. A job that carries its own edited copy of the rules has already drifted; it just has not been noticed yet. - **One normalisation, applied before the transformation**, defined in that same shared rule set - trimming, case folding, separator handling and text-encoding form. - **One secret, versioned.** A change is a whole-set operation: every copy is reproduced, or none is. - **A stamp on every copy** recording the rule-set version, the secret version and the point in time the source was taken from. Refuse to use two copies together whose stamps disagree; that single rule converts a silent mismatch into an immediate, readable refusal. - **No local overrides.** Where a team genuinely needs a different treatment for a field, it goes into the shared rule set as a new version that every job picks up, and every copy is reproduced. ## Verify after every refresh, not when a test fails The check is small and belongs to the production of the copies: 1. Take a sample of identifiers present in both copies - a few thousand is plenty. 2. Compare their substitutes and compute the match rate. 3. Require a full match. Anything less rejects the refresh, and the report names the field and the sample rows that disagreed. 4. Compare the stamps as well, so a mismatch of versions is caught even when the sample happens to agree. Rejecting the refresh is the important part. The alternative - letting the copies through and learning about the divergence from a failing cross-system test - moves the discovery to the worst possible place: an engineer who did not produce the copies, working on a feature that is not the problem, with no reason to suspect the data. A team that has been through that once usually adopts the match-rate check the same week. ## The organisational shape The technical fix is easy; the durable fix is deciding who owns the rules. A single owner publishes a versioned rule set and the check; teams contribute field rules to it and may not fork it. The moment two teams can each change how a shared identifier is replaced, the copies are one busy sprint away from disagreeing, and the failure will be discovered by whoever is unlucky rather than by whoever changed something.
- Two copies disagree, but only for rows created since the last refresh. What does that tell you?Probably nothing about the substitution rules. Rows one copy has never seen are a timing difference between when the two sources were taken, not a divergence in how values are replaced. Restrict the comparison to identifiers present in both copies; if the match rate is then perfect, the rules still agree and the gap is scheduling.
- Who should own the substitution rules when six teams each produce their own copy?One owner publishes a versioned rule set and the verification check; teams contribute new field rules into it and consume it rather than forking it. Local edits are precisely how two copies stop agreeing. Pair that with a stamp on every copy and a refusal to combine copies whose versions differ, so an unadopted change fails loudly instead of quietly.
Two clocks set from one reference on the first morning still disagree by the summer, unless both keep resynchronising to the same reference rather than to each other.
saying these in an interview costs you the question
- Assumes two jobs stay aligned because they started aligned
- Blames the application when a cross-system join thins out
- Changes the shared secret on one system only
- Keeps a local field override to unblock a load
- Checks agreement only after someone reports a failure
- Confuses a row one copy never received with a mismatch