How do you decide whether a store migration's identifier remapping must be an invertible bijection rather than a deliberate many-to-one merge?
answer
- who still holds the old identifiers
- reversibility is a retention commitment
- a merge spends injectivity permanently
- store the preimage set instead of an inverse
- you can merge later, never un-merge
basics
~20 sDecide by who still holds the old identifiers and what must be answerable later. A merge is a one-way loss of injectivity: the inverse stops existing, and only preimages you deliberately store can answer lineage questions afterwards.
solid answer
~50 sThis is a contract decision dressed as a data question. A **bijection** keeps the old identifier recoverable, so rollback is a re-run in the other direction, lineage questions have answers, and anything outside your boundary that still quotes an old identifier can be served. A deliberate **many-to-one merge** buys deduplication and a smaller, truer model, and its cost is exact: injectivity is gone, so no inverse exists — only the **preimage**, the set of old identifiers behind a row, and that has to be stored on purpose because it cannot be recomputed later. The questions that actually decide it are who outside your control still holds old identifiers, whether rollback means undoing the mapping or restoring from a backup, whether the merge is semantically true or merely convenient, and how long the mapping must remain answerable. The middle path is common: merge in the model, but keep the preimage as a durable attribute so lineage survives the loss of the inverse.
go deeper
Understand the fork first: keeping one new row per old record preserves the ability to go back, while combining records loses it. Which is acceptable is a decision, not a detail.
Say which property each option has and what follows: a merge is a non-injective map, so no inverse exists, and only a stored set of old identifiers can answer lineage questions later.
Show the operational consequences: rollback that depends on the inverse, back-translation for identifiers held outside the system, and enforcing uniqueness where the identifiers actually land.
Own the trade-off end to end: a retention commitment with a stated horizon, who is permitted to mint identifiers in the new space, and the asymmetry that you may merge later but can never un-merge.
## What each choice actually is, mathematically Strip away the migration and two options remain. | | Bijection onto the new identifiers | Deliberate many-to-one merge | |---|---|---| | Injective | Yes — each old record keeps its own row | No — several old records share a row | | Inverse | Exists and is unique | Does not exist; only preimages do | | Rollback | Re-run the mapping in the other direction | Requires a backup or a stored preimage | | External references | Old identifiers can be translated on demand | Translation is ambiguous unless preimages are kept | | Model quality | Carries the old store's duplicates forward | Produces the truer, smaller model | | Ongoing cost | A derivable or stored mapping, kept as long as anyone quotes old identifiers | A preimage set per row, kept for as long as lineage is asked for | The asymmetry that makes this a real decision is that **the merge is not reversible later**. You can always merge a bijection afterwards; you cannot un-merge once the preimages are gone. Reversibility is therefore the option with the higher running cost and the lower regret. ## The questions that decide it 1. **Who holds old identifiers outside your boundary?** Links already sent out, identifiers quoted in other teams' stores, references held by integrating systems, anything printed on a document a person keeps. Each of these is a party that can present an old identifier long after the migration and expect an answer. This is the single strongest force toward a bijection, and it is also the one most often discovered late. 2. **What does rollback mean here?** If the plan is to reverse the migration by running the mapping backwards, you have just required a bijection. If the plan is to restore the old store from a backup and replay, a merge is affordable — provided the replay's own correctness does not depend on the inverse. 3. **Is the merge semantically true?** Merging two records because they genuinely describe the same entity is a modelling improvement. Merging them because the new identifier scheme cannot tell them apart is data loss with a nicer name, and the two are easy to confuse in a review because the row counts look identical. 4. **For how long must this be answerable?** A translation promise is a retention promise. A stored mapping table or a preimage set is an asset that must be backed up, migrated again in future, and kept accessible for as long as the promise stands. 5. **Is the inverse derivable or must it be stored?** An encoding that embeds the old identifier gives back-translation for free and cannot drift; an arbitrary remapping needs a table, whose own uniqueness constraint is what actually enforces the injectivity you claimed. ## The middle path, and why it is usually right Most real migrations should merge in the model and keep the lineage anyway: the new row holds the **set** of old identifiers it absorbed. Note what that object is. It is a relation rather than a function inverse, so it answers "which old records became this one" and it is fine for it to have more than one answer. It gives you audit, incident forensics and translation of externally held identifiers, without pretending the mapping can be undone. The converse trap is worth naming too. A migration that keeps every duplicate in order to stay bijective has chosen fidelity to the old store over fidelity to reality, and it leaves the deduplication to be done later by whoever inherits it — at which point exactly this decision arrives again, with less context and more dependants. ## How to present the decision - **Name the property you are choosing**, not just the behaviour: "this mapping is injective and we keep a derivable inverse" is checkable; "we keep track of the old identifiers" is not. - **State the retention commitment in the same sentence as the merge.** A merge with no preimage retention is a decision to make certain questions permanently unanswerable, and that should be made explicitly rather than discovered during an incident. - **Write down who may mint identifiers** in the new space during and after the cutover. A second producer changes which mathematical property actually holds, regardless of what the migration's own code does. - **Make the reversible option's cost visible**, since it is a running cost against a one-off saving; that comparison, not the mathematics, is what the decision usually turns on. The mathematics does not choose for you. What it does is make the cost exact: a merge spends injectivity, injectivity is what an inverse is made of, and nothing you build afterwards can buy it back.
- The merge is already agreed. What single thing most preserves your options afterwards?Store the preimage: the set of old identifiers that each new row absorbed, as a durable attribute or side table. It is a relation rather than a function inverse, so a row may list several, which is exactly right. That preserves lineage, audit and translation for externally held identifiers, and it is the one artifact that cannot be reconstructed once the old store is retired. Decide its retention period at the same time, since the promise is only as good as the data behind it.
- Why is choosing the bijection the lower-regret option even when it costs more?Because the operations are asymmetric. Merging a bijective mapping later is always possible: you still hold both identifiers and can collapse them under a rule chosen with more information. Un-merging is not, once the preimages are gone, since no inverse exists to recover what shared a row. So the reversible choice trades a recurring cost for the ability to make the other decision later, which is usually worth more than the storage it consumes.
- What makes a merge a modelling improvement rather than data loss?Whether the records genuinely describe the same entity. If they do, the merge removes duplication the old store carried and the new model is truer. If they are merged because the new identifier scheme cannot distinguish them, two real entities have been conflated and the row counts will not reveal it. The test is to state the merge rule as a claim about the domain and see whether anyone with domain knowledge will sign it.
- How does allowing a second producer in the new identifier space change the decision?It changes which property actually holds, independently of the migration's code. Identifiers minted elsewhere have no old counterpart, so the mapping stops being surjective onto the live set and any back-translation must be explicitly partial. It also means the interim or final identifier space is now a published contract with more than one writer, so uniqueness must be enforced where the identifiers land rather than assumed by whichever component happens to produce most of them.
saying these in an interview costs you the question
- Calls a merge reversible because the old store still exists today
- Treats deduplication as free without naming the lineage cost
- Assumes nobody outside the system holds old identifiers
- Plans to reconstruct preimages after retiring the old store
- Confuses keeping a mapping table with the mapping being injective
- Decides on storage cost alone, ignoring the retention promise