Why does comparing mapped objects by reference identify a row reliably inside one tracked set but not across two?
answer
- same object versus same row
- the map makes the two questions coincide
- coincidence ends at the set boundary
- compare by identifier plus type across sets
- unsaved key breaks hash-based collections
basics
~20 sInside one tracked set the identity map guarantees one instance per row, so reference comparison is a valid row-identity test. Across two sets one row has two instances and the test reports 'different' - compare by identifier at that boundary.
solid answer
~50 sReference comparison asks 'is this the same object?'. Inside one tracked set that question happens to have the same answer as 'is this the same row?', because the identity map materialises exactly one object per key — which is why reference comparison works, and why it works so consistently that code comes to depend on it. The moment operands can come from **different sets**, the coincidence ends: two instances exist for one row, reference comparison says they differ, and membership tests, deduplication and identity-keyed lookups all quietly give the wrong answer. The rule of thumb: compare by **identifier** wherever operands may come from different sets, and keep reference comparison for inside one operation. Watch the object whose key is not assigned yet: putting it in a hash-based collection before its key exists, and hashing on that key, corrupts the collection when the key appears.
go deeper
Learn the scope: inside one tracked set the same row is one object, so reference comparison works. Outside it, the same row can be two objects and the comparison quietly says they are different.
Explain why the identity map makes 'same object' and 'same row' coincide, and list the boundaries — a second set, a kept object, a rebuilt one — where they come apart again.
Recognise the silent signature: a membership test or a removal that does nothing, only on paths where objects cross boundaries. Apply identifier comparison at those boundaries rather than banning reference comparison outright.
Set the policy: what may cross a set boundary at all, and what identity means for objects that do. Prefer transferring identifiers or plain shapes so the question rarely arises, and keep a real business key where the domain offers one.
## Why reference comparison works at all here Comparing two references asks whether they point at the same object in memory. That is normally a much stronger question than "do these describe the same database row", and normally the wrong test. Inside a tracked set it happens to be exactly right, because the identity map enforces **one object per row key for the life of the set**: if two references describe the same row, the map guarantees they *are* the same reference. So reference comparison is not a hack inside one set — it is the cheapest correct identity test available, and it is why so much mapper-backed code contains it. ## Why it stops being right at the boundary The guarantee is scoped to one set, and nothing outside a set enforces anything. Two references to one row appear whenever: - two tracked sets are open at once, each having materialised the row; - an object was loaded by an earlier set, kept, and is now compared with one loaded by the current set; - an object came back from another layer — a queue, a cache of your own, a call boundary — and is compared with a freshly loaded one; - an object was rebuilt from a transfer shape rather than loaded at all. In every one of those cases the two objects describe one row and are different objects, so reference comparison answers "no" to a question whose true answer is "yes". ## What that breaks, concretely - **Membership tests** report absent for an element the collection describes. - **Deduplication** by identity keeps two entries for one row, so a batch writes twice or a total counts twice. - **Identity-keyed maps** — a lookup side table keyed by the object — miss, and the code takes the "not seen yet" branch. - **Removal** from a collection silently does nothing, because the element to remove is not the element held. These failures share a signature: nothing throws, the wrong branch is simply taken, and it only happens on paths where an object crossed a boundary. That is why they escape tests written entirely inside one set. These failures also tend to appear late, because the boundary that produces the second instance is often introduced long after the comparison was written: a job is split out, a caching layer is added, a call becomes asynchronous. The comparison did not change; the guarantee under it did. ## What to compare instead | Test | Correct when | Fails when | |---|---|---| | Reference comparison | Both operands loaded by the same tracked set | Operands come from different sets or were rebuilt | | Identifier comparison, type included | Both operands are saved rows | Either operand has no key assigned yet | | A natural business key, when the model has a real one | Any operands, including unsaved | The model has no genuinely unique, immutable business key | Identifier comparison is the workhorse: include the type, because the number 7 names a row in every table. Where the model genuinely has an immutable natural key, comparing on it is stronger still, because it also works before a key is assigned. ## The unsaved-object trap An object created in memory has no key yet. If it goes into a hash-based collection while its hash derives from that key, and the key is then assigned when the object is written, the object's hash changes while it sits in the collection — and the collection can no longer find it, including to remove it. There are two honest ways out: keep hashing stable and independent of the assigned key, or keep unsaved objects out of hash-based collections until they have one. The first is usually preferred because it survives every path an object can take; the second depends on nobody forgetting. When the key is assigned, and by what mechanism, is a separate topic; the consequence here is only that it can appear *after* the object exists. ## Interview framing Say the scope out loud: reference comparison is a correct identity test *inside one tracked set* and a broken one outside it, because the one-object rule is what made it correct in the first place. Then name the boundary cases that produce a second instance, describe one silent failure — a membership test that misses — and give the rule you actually apply: compare by identifier wherever operands might not share a set. That answer shows you understand the guarantee rather than having memorised a prohibition.
- Why include the type when comparing by identifier?Identifiers are unique only within their own table, so two rows in different tables routinely share the number 7. Comparing identifiers alone makes unrelated objects compare equal, which is worse than the problem being fixed. The identity map itself is keyed by type plus identifier for the same reason.
- What is the tell-tale signature of a cross-set identity bug in production?Nothing throws. A membership or removal test silently takes the wrong branch, or a duplicate appears, and only on paths where an object crossed a boundary — a background job, a cached object, a rebuilt one. Tests written entirely inside one set never reproduce it.
- If comparison by identifier is safer, why not use it everywhere and forget reference comparison?It is a reasonable default, but it needs care of its own: it is undefined for objects with no key yet, and it makes two unsaved objects compare either all-equal or all-different depending on how you write it. Inside one set, reference comparison remains correct and needs no such handling.
saying these in an interview costs you the question
- Says reference comparison is always wrong for mapped objects
- Says reference comparison is always safe because the mapper guarantees one instance
- Compares identifiers without considering the type
- Puts unsaved objects in hash-based collections keyed on the assigned identifier
- Expects a membership test to throw rather than silently miss