A merged customer group must expose one stable identifier: what property must the choice of canonical representative have?
answer
- one name per class
- chosen by members, not order
- same class, same canonical value
- membership change can move it
- derived value versus minted identifier
basics
~20 sThe representative must be determined by the class's members alone — never by arrival order, worker or wall-clock time. Formally it is a function whose value is equal for two rows exactly when they are in the same class.
solid answer
~50 sA **canonical representative** is one designated element, or one derived canonical value, standing for a whole class. The requirement is that the choice be a function of the class: any two rows in the same class must yield the same representative, and rows in different classes must not. Put the other way round, a canonicalizing function `f` with '`f(a) = f(b)` exactly when `a` and `b` are in the same class' gives you both a comparison rule and an identifier at once. Choosing 'whichever row we saw first' satisfies neither — it depends on the traversal, so the identifier moves between runs. The remaining subtlety is that even a correct derived representative is only stable while the class is: adding one member with a smaller key moves it, which is why a class whose identifier escapes to other systems is usually given a minted identifier and a stored mapping instead.
go deeper
Remember the test: two rows in the same group must produce the same representative, so the choice can never depend on which row was read first.
Explain the function view — comparing a derived canonical value for equality is automatically an equivalence — and use it to check whether a candidate rule can ever disagree between two members.
Weigh a derived representative against a minted identifier: recomputable with no state but it moves with membership, versus stable but needing a stored mapping and the work of keeping it in step.
Own the contract the identifier makes to everything downstream, including what happens to references when two classes merge, and require field promotion onto a merged record to distinguish agreed values from hidden conflicts.
## What a representative is for Once the rows are partitioned, downstream systems need one name for each class: something to store on an invoice, to join on, to show in a user interface. Picking one member of the class to play that role is the **canonical representative**; deriving a value that stands for the class — a normalized form, a composed key — is **canonicalization**. Both answer the same need, which is to talk about a class without listing it. The requirement is a single sentence: **the representative must depend only on the class**. Two rows that belong together must produce the same representative, and that must remain true whichever row you start from and whichever machine computes it. ## The function view makes the requirement checkable A clean way to state it: a canonicalizer is a function `f` over rows such that `f(a) = f(b)` holds exactly when `a` and `b` are in the same class. Two consequences fall out immediately. - **Any such function defines the relation.** Comparing canonical values for equality is itself an equivalence relation — reflexive, symmetric and transitive — because equality of values is one. So if you can canonicalize, you have also specified the grouping rule, for free and with transitivity guaranteed. - **The check becomes mechanical.** For a candidate choice rule, ask whether two members of the same class can ever disagree. 'The first row we encountered' disagrees as soon as the traversal changes. 'The row whose identifier is smallest under a fixed total order' does not: every member of the class computes the same minimum from the same set. ## Choices that fail, and why - **Arrival order.** The first row seen depends on shard boundaries and worker counts. Identical input, different representative. - **A worker-local sequence number or a process-local counter.** The same class computed on a different machine gets a different name. - **The row with the most recent processing timestamp.** The timestamp records when the pipeline touched the row, not anything about the customer, so it changes on every run. - **A freshly generated identifier per run.** Stable within a run, meaningless across runs, and downstream references rot silently. What these share is a dependence on something outside the class's own content. ## Derived versus minted identifiers Even a correct, purely derived representative moves when the class changes. If the rule is 'smallest key in the class' and a merge brings in a member with a smaller key, the representative changes, and any system that stored the old one now points at nothing. | | Derived from members | Minted for the class | |---|---|---| | Recomputable from data alone | yes | no — requires stored state | | Stable when membership changes | no | yes | | Needs a member-to-class mapping | no | yes | | Suits references held outside the pipeline | poorly | well | The usual resolution is to use both: a derived canonical value for comparing and deduplicating inside the pipeline, and a minted identifier with a stored mapping for anything that escapes to other systems. The mapping is exactly the extra state the stability guarantee costs, and it should be an explicit design decision rather than an accident. ## Well-definedness on the class The representative is not only a name; it is often the row whose field values get promoted onto the merged record. That is where a second, quieter requirement lives: a value attached to a class is only meaningful if every member agrees on it. A birth date, a country, a tax status taken from the representative silently becomes 'the class's value' even when other members carry a different one. The discipline is to classify each field before promoting it: 1. **Class-level and agreed** — every member carries the same value, so taking the representative's copy is safe and any member would do. 2. **Class-level but conflicting** — the members disagree, which is evidence about the merge itself; promoting one value hides that evidence, so the conflict belongs in the output as a conflict. 3. **Row-level by nature** — a source system's own identifier, a capture timestamp — which belongs to the member and should be kept per member rather than promoted at all. A merged record built without that classification looks authoritative and quietly encodes whichever row happened to win the representative contest.
- When is a derived representative preferable to a minted class identifier?When the grouping must be recomputable from the data alone, with no stored state to keep in step, and the consumers are inside the pipeline where a moving name costs nothing. As soon as an identifier is handed to a system you do not control, mint it and store a member-to-class mapping, because a derived name moves whenever membership changes.
- What makes a field attached to a merged class well defined?Every member of the class yields the same value for it. If two members disagree, the field is a property of the row rather than of the class, and taking the representative's copy hides a conflict that is itself evidence about whether the merge was right. Such fields should be surfaced as conflicts or kept per member.
saying these in an interview costs you the question
- Uses whichever row arrived first as the representative
- Names the class with a worker-local sequence number
- Assumes a derived representative never moves when membership changes
- Promotes one member's field value without checking the others
- Mints a fresh identifier on every run and calls it stable