A migration remaps every old record identifier to a new one: what does injectivity of that mapping guarantee, and what breaks without it?
answer
- no two old records may share a row
- distinct in, distinct out
- one-to-one, not onto
- image size equals domain size
- a collision is a silent entity merge
basics
~20 sInjective means no two distinct old identifiers are sent to the same new identifier. Lose it and two separate records land on one row in the target: one write overwrites the other, references converge, and no rollback can tell them apart.
solid answer
~50 sA mapping is **injective** (one-to-one) when `f(a) = f(b)` forces `a = b` — distinct old identifiers always produce distinct new ones. For a finite set of records that is the same as saying the image has exactly as many elements as the domain, so a single lost element is a single collision. Without injectivity two old records occupy one new row: either a uniqueness constraint rejects the second write, which is the loud and preferable failure, or the later write silently overwrites the earlier one and the migration reports success while two entities have been merged. Everything that pointed at the two old records now points at one, counts stop reconciling, and the old identifier can no longer be recovered from the new one, so the migration is no longer reversible even in principle. Injectivity says nothing about the target space being fully used — that is surjectivity.
code
pseudocode · 13 linesseen = empty map // new identifier -> the old one that claimed it
collisions = 0
for each oldId in oldStore:
newId = remap(oldId)
if seen contains newId:
report collision(oldId, seen[newId], newId)
collisions = collisions + 1
else:
seen[newId] = oldId
// collisions == 0 means: injective over the identifiers scanned,
// which is evidence about this data, not a proof over the domaingo deeper
Recall the one-line definition: distinct inputs must give distinct outputs, so no two old records may end up sharing a new identifier. Then name one consequence of breaking it, such as one record overwriting another.
Explain the mechanics: domain, codomain and image; why the image of a finite domain having fewer elements is exactly a collision; and why a uniqueness constraint in the target turns a silent merge into a failed write.
Show you would derive the property rather than sample for it: an encoding that can be decoded back to the old identifier, lossy transforms treated as collision generators, and a reconciliation that separates collisions from records that were never mapped.
Frame it as a contract: what the platform promises about identifier uniqueness across stores, who is allowed to mint identifiers during the cutover, and what the organisation gives up permanently the first time a merge is accepted as expected behaviour.
## The three sets a mapping fixes A mapping — a function — is not merely a piece of code that returns a value. It fixes three sets, and most arguments about a migration go wrong because only one of them was ever named. - **Domain** — the identifiers the mapping is defined on: every old record that exists, and every one that will be created while the migration runs. - **Codomain** — the space the results are allowed to live in: whatever shape of identifier the new store accepts, such as any 128-bit value or any string up to 64 characters. - **Image** — the identifiers the mapping actually produces. It is a subset of the codomain, and usually a vanishingly small one. Being a function at all already forbids one direction: **one old identifier may not be sent to two different new ones**. A remap that draws from a sequence generator on each call is not a function of its input at all — re-run it and the same record gets a different identifier — so none of the properties below apply to it. Injectivity forbids the other direction. ## Injective, stated in the migration's terms A mapping `f` is **injective** when `f(a) = f(b)` implies `a = b`; equivalently, distinct inputs always yield distinct outputs. In migration terms: **no two old record identifiers are ever remapped onto the same new identifier**. Over a finite domain there is an equivalent and far more checkable form — the image has exactly as many elements as the domain. One element missing from that count is exactly one collision. | Property | What it says about the remap | What it buys the migration | |---|---|---| | Injective | Distinct old identifiers get distinct new ones | Old records stay separate; the old identifier is recoverable from the new one | | Surjective | Every identifier the target space allows is produced | Almost never wanted here; it only means nothing is left spare | | Bijective | Both at once | Every new identifier names exactly one old record, so the migration is reversible | Note what the table does not claim. Injectivity is about **collisions**, not about coverage. A remap producing a thousand distinct identifiers inside a space of 2^128 is perfectly injective and nowhere near surjective, and that is the normal, healthy case. ## What a lost injection actually does Two old records colliding on one new identifier is not a rounding error; it is a silent join of two entities. Roughly in the order you notice: 1. **The second write wins, or is rejected.** If the target enforces uniqueness on the identifier, the second insert fails and the migration stops loudly. If it does not, the later row overwrites the earlier one and the run reports success. 2. **References converge.** Rows that referred to two different parents now refer to one. Ownership, permissions and billing links follow the merge wherever they lead. 3. **Counts stop agreeing.** The old store holds more rows than the new one and no per-row comparison explains the gap, because both old records now compare equal against the same new row. 4. **Rollback and audit die.** With a collision there is no way to say which old record a given new row came from, so the mapping cannot be undone even partially. ## Establish it; do not hope for it Injectivity is a claim about the whole domain, so a scan over today's data is evidence and never a proof. - Prefer a remap that is **injective by construction** — one that keeps the old identifier inside the new one, for instance by pairing it with a namespace that can be read back out. This holds only if the encoding is **uniquely decodable**: a plain concatenation of namespace and key is *not* injective when the delimiter can occur inside either part, because two different splits produce the same string. Length-prefixing, or a delimiter excluded from both parts, repairs it. - Treat anything **lossy** as a collision generator by design: truncating, lowercasing, stripping punctuation, or rounding a timestamp all send many inputs to one. - A **digest** of the old identifier is injective in practice but never by construction — the output space is finite and the input space is not, so what you are relying on is a probabilistic argument, not the mathematical property. - Add a **uniqueness constraint** in the target as a backstop, so a violated assumption fails at the first offending row rather than at the first customer report. ## What injectivity still does not give you - It does not mean every identifier the target can hold is used; that is surjectivity, and conflating the two is the most common slip on this material. - It does not make the mapping **total**: a remap that silently skips records is defined on too small a domain, a different defect with the same symptom of missing rows. - It says nothing about whether the record's **contents** survived; only the identifier is in scope. - It gives you a reverse lookup in principle, on the image only — you still have to store or derive that inverse before anything can use it.
- Is prefixing every old key with its namespace an injective remap?Only if the encoding is uniquely decodable. Plain concatenation is not: if the delimiter can appear inside the namespace or the key, two different splits collapse to the same string, so two distinct old keys can produce one new identifier. Length-prefix each part, or use a delimiter that is excluded from both, and the pair can be read back out — which makes the map injective by construction rather than by hope.
- A verification run scans every current record and finds no duplicate outputs. Has injectivity been proved?No. It has been checked on the records that existed at scan time. Injectivity is a property of the whole domain, including records created after the scan and any input shape the old store still permits. Treat the scan as a smoke test and get the guarantee elsewhere — from a construction that is injective by argument, plus a uniqueness constraint in the target that turns any surviving collision into a failed write rather than a silent merge.
- How is a remap that is not total different from one that is not injective, if both end with missing rows?Not injective means two old records arrived at one new row: the row exists and now conflates two entities. Not total means the mapping was undefined for some old records, so those rows never arrived at all. The first corrupts data that looks present; the second loses data that is visibly absent. They need different fixes — a different remap versus a widened domain — so it is worth separating them before reconciling counts.
A cloakroom where every coat gets its own ticket number: injective means no two coats are ever issued the same number, though most numbers in the book are never issued at all. The day two coats share a number, someone goes home in the wrong coat.
saying these in an interview costs you the question
- Uses injective and surjective as two names for the same property
- Thinks injective means the target identifier space must be fully used
- Believes a clean scan of today's records proves injectivity for all inputs
- Calls a digest of the old identifier injective by construction
- Treats a collision as harmless because the records looked similar anyway
- Assumes lowercasing or truncating an identifier preserves uniqueness