A near-match matcher can never be made transitive without fusing distinct people: how should the platform define what 'same customer' means?
answer
- two relations, not one
- identity must stay an equivalence
- compare a derived key
- advisory matches feed review
- false merge costs more than false split
basics
~20 sKeep two relations. Identity is defined so that it is an equivalence by construction — equality of a normalized key — and the near-match rule stays advisory, feeding review and search but never partitioning the data or naming anything other systems store.
solid answer
~50 sThe mathematics forces the choice rather than merely informing it: anything used as identity is used as a key, and a key built on a relation that is not an equivalence has no stable meaning. So define identity as equality of a derived key — that is the kernel of a function, which is reflexive, symmetric and transitive for free — and accept that it under-merges. Then keep the near-match rule as a separate, deliberately non-transitive **candidate** relation that feeds a review queue, search and suggestions, and is never allowed to partition the data or mint identifiers. The alternative, taking the equivalence closure of the matcher's pairs automatically, buys recall at the price of chained merges, and it should only be adopted with a stated precision budget, because the two error directions are not equally cheap: a false merge mixes two people's records, while a false split leaves visible duplicates.
go deeper
Take away the rule of thumb: the relation you key on has to be an equivalence, so identity is defined by comparing a derived key rather than by a similarity score.
Be able to explain why comparing a normalized key is transitive for free, and why a threshold rule is not, so the two cannot play the same role in a system.
Argue the trade-off with evidence: what closure does to the class-size distribution, what the normalizer gives up in recall, and how a review queue recovers it without touching the key space.
Make and defend the platform-wide call: which relation is allowed to define identity, what the advisory relation may never be used for, how identifiers are minted and remapped, and which error direction the automatic side is biased toward.
## Why the matcher cannot be the identity relation Identity is not just an opinion the pipeline holds; it is the thing other systems key on. Invoices, entitlements, counts, joins and exported references all assume that 'the customer' is a well-defined object. That assumption is precisely the statement that the underlying relation is an equivalence: the classes exist, they are disjoint, and every row is in exactly one. A near-match rule is reflexive and symmetric but not transitive, so it has no classes, and any grouping derived from it is a property of the traversal. Keying a platform on that means keying it on an execution detail. That is the reason this is a design decision rather than a tuning exercise: no threshold makes a distance rule transitive, so the question is not 'how do we fix the matcher' but 'which relation is allowed to define identity'. ## Three exits, and what each costs 1. **Identity as the kernel of a normalization function.** Define 'same customer' as 'equal normalized key'. Because it compares a derived value with equality, it is an equivalence by construction and nothing else has to be checked. The cost is recall: genuine duplicates that normalize differently stay separate, and the quality of the grouping becomes the quality of the normalizer. 2. **Identity as the equivalence closure of the matcher's pairs.** Take every reported pair, close it, and treat the resulting classes as customers. This maximises recall and buys it with chaining: merges propagate along chains, one spurious pair fuses two whole classes, and the closure can never split what it joined. Adopting it responsibly means limiting which pairs are ever proposed, monitoring the class-size distribution, and keeping a correction path at the source of the pairs. 3. **Two relations, with different privileges.** Identity is an equivalence by construction as in the first exit; the near-match rule is retained as an explicitly non-transitive advisory relation — 'possibly the same as' — that surfaces candidates for review, powers search and suggestions, and is never used to partition or to name anything. Human or higher-confidence confirmation promotes a candidate by changing the identity key, not by bypassing it. The third exit is usually the right default at scale, because it is the only one that lets the advisory relation be as loose as it needs to be without that looseness reaching the key space. ## The two error directions are not symmetric | | False merge (two people fused) | False split (one person duplicated) | |---|---|---| | What the user sees | another person's data under their record | their history in two places | | Recoverable? | poorly — references to the merged identifier have already spread | yes — merging later is the normal path | | Detection | often only when someone complains | visible in duplicate-rate metrics | | Severity | a correctness and confidentiality incident | an operational annoyance | Because the costs are asymmetric, the automatic side of the system should be biased toward splitting, and the recall that bias gives up should be recovered through the advisory relation and review rather than by loosening the identity rule. Stating that asymmetry explicitly is most of the decision. ## What to promise the rest of the platform - **The identity relation is an equivalence.** Other systems may treat 'the customer' as an object, join on it, and count it. - **The advisory relation is not, and is documented as not being one.** No consumer may group by it, count it, or store an identifier derived from it. - **Identifiers that leave the platform are minted and mapped**, not derived from whichever member currently sorts first, so that a later merge changes the mapping rather than orphaning every stored reference. - **Merging is an event, not a silent recomputation.** When two identity classes are genuinely joined, the systems holding the old identifiers need to hear about it, which is a far easier contract to honour when merges are rare and deliberate than when they are a by-product of a threshold. The underlying principle is small and worth saying plainly in a design review: a relation you intend to key on must be an equivalence, and if the rule you have is not one, you either derive a rule that is or you stop treating its output as identity. Everything else in this decision follows from that.
- Why is 'same normalized key' transitive without any extra work?Because it is the kernel of a function: the rule is 'the derived values are equal', and equality of values is reflexive, symmetric and transitive. If `f(a) = f(b)` and `f(b) = f(c)` then `f(a) = f(c)` follows from equality alone, so any relation defined by comparing a derived value is automatically an equivalence.
- Why is a false merge budgeted more tightly than a false split in an automated pipeline?A false merge mixes two people's records, and the merged identifier propagates to every system that copied it, so the damage spreads and is awkward to unwind. A false split leaves duplicates: visible, measurable, and fixed by merging later. Asymmetric costs justify an automatic side biased toward splitting, with recall recovered through review.
- What breaks if a consumer groups and counts by the advisory 'possible duplicate' relation?Its counts stop being a property of the data. Without transitivity the relation has no classes, so any grouping it appears to produce is an artefact of the traversal, and two consumers computing the same metric can disagree while both are running correct code. That is why the advisory relation must be documented as non-partitioning.
saying these in an interview costs you the question
- Lets the fuzzy matcher define the primary identity relation
- Treats a false merge and a false split as equally cheap
- Plans to fix chaining by tuning the threshold indefinitely
- Exposes a group identifier whose meaning changes between runs
- Assumes any matching rule can be made transitive with more rules
- Adopts closure for recall without a stated precision budget