How should a mapped class define equality and hashing when its identifier stays empty until the row is written?
answer
- the key arrives mid-life
- a hash must never change
- lost inside its own container
- empty keys all look equal
- or assign the key at construction
basics
~20 sNever let the hash depend on a key that is filled in later: it changes when the row is written and strands the object inside any hash container it already joined. Keep the hash stable, and compare identifiers only when both exist.
solid answer
~40 sThe difficulty is that a store-assigned identifier arrives in the middle of an object's life. If hashing reads that field, an object placed in a hash-based container before the write lands in a bucket computed from an empty key and quietly becomes unfindable once the key appears — the container still holds it, but lookups and removals miss. Equality on the key has its own trap: every unwritten object compares equal to every other one. The workable shapes are a **constant hash** with equality falling back to reference identity while the key is absent; equality over a genuinely stable, non-null business value when the model has one; or removing the problem by computing the identifier in the application at construction, so the field never changes at all.
go deeper
Remember that an identifier supplied by the store does not exist while the object is new, so anything computed from it before the write is meaningless and will change later.
Explain the mechanics: a hash-based container picks a bucket once, from the value at insertion time, and never revisits that choice when a field changes afterwards.
Demonstrate a fix in a real model — a stable hash, identifier equality guarded by both sides having one, and a rule about which collections an unwritten object may join.
Treat it as a modelling decision rather than a coding trick: if identity has to hold across processes, caches and messages, an identifier computed up front removes a whole class of defects, at the price of giving up store-side numbering.
This is the classic consequence of a key that arrives late. An object is created, used, put into collections, passed around — and only at flush does the store hand back the value that identifies it. Any behaviour derived from the identifier therefore *changes* partway through the object's life, and hash-based containers do not tolerate that. ## The bucket is chosen once A hash-based container computes an element's hash when the element is added and uses it to pick a bucket. It does not revisit that decision. So: 1. The object is created with an empty key and hashes to, say, bucket A. 2. It is added to a set — perhaps the collection on the other end of a link. 3. The unit of work flushes; the insert runs; the layer writes the generated key onto the object. 4. The hash now points at bucket B. 5. A lookup, a containment test or a removal computes bucket B, searches it, and finds nothing. The object is still in the container, holding memory, reachable by iteration — but unreachable by lookup, and impossible to remove. Iterating and comparing still finds it, which is exactly why this bug survives casual testing. ## Equality on an empty key is its own trap If equality says "same class, same identifier", then two freshly created objects both compare equal, because both identifiers are empty. Put five of them in a set and four disappear. The same rule also makes an object equal to a *different* unwritten object and unequal to a reloaded copy of itself once the key exists — the verdict for a given pair changes as rows are written, which is unavoidable, but collapsing all unwritten objects into one is not. ## The three shapes that work | shape | equality | hash | when it fits | |---|---|---|---| | stable hash, guarded key equality | by identifier when both objects have one, otherwise reference identity | a constant, or derived from the class | store-assigned keys, the general default | | business-value equality | over a value present at construction and never edited | over the same value | a model with a genuinely stable natural value | | identifier assigned at construction | by identifier, always | by identifier, always | the application computes the key up front | The third shape is the reason this question sits with key timing rather than with equality in general: the whole problem is created by *when* the identifier appears, and a key that exists from birth removes it. ## Why a constant hash is acceptable It sounds wrong, and it usually is for a general-purpose value type. For mapped objects it is fine, because of what these containers actually hold: the children of one parent, the members of one small collection — tens of elements, not millions. Every element lands in one bucket and lookup degrades to a scan of that bucket, which at those sizes is unmeasurable. The rule that matters more is the one the container depends on: **an element's hash must not change while the element is in the container.** A stable, unselective hash keeps that promise; a selective hash that changes at flush does not. ## Practical rules - Do not put unwritten objects into hash-based containers that outlive the write, unless the hash is stable. - Make equality symmetric and reflexive at every stage: an unwritten object must equal itself. - Do not use "all fields" equality: two distinct rows can legitimately carry identical data, and merging them hides duplicates. - If you adopt a business value, be honest about whether it is really immutable. A value that can be corrected later fails in the same way as a late key, only more quietly. - Remember that a reloaded object and an instance you already hold should compare equal once both carry the same identifier — that is precisely what the identity map is trying to preserve. ## Where layers differ Data-access layers differ in how much of this they hide: some keep one instance per row within a unit of work so that reference identity is enough for most code, while others hand back fresh instances freely and lean entirely on the class's own equality. Code that will be read by people working in both worlds should not depend on the first behaviour. ## The short version The hash must be stable for the object's whole life; equality may use the identifier only when it exists, and must fall back to identity rather than declaring all unwritten objects the same. The cleanest escape is to stop the key from arriving late.
- Why does the object become unfindable rather than duplicated?Because the container chose a bucket once, when the element was added. Filling in the key changes the hash but not the placement, so a later lookup searches the bucket the new hash points at. The element is still held and still iterable — it is simply unreachable by lookup, and removal fails too.
- Is a constant hash really acceptable?For mapped objects, yes, at the sizes these collections actually reach. Every element shares one bucket, so lookup degrades to a scan — unmeasurable for the tens of children on an object, unacceptable for a container of millions. Correctness first: a stable dull hash beats a changing selective one.
- When is a business value safe to use for equality instead?When it is present at construction, never edited, and unique for the concept rather than merely unique in today's data. If it can be corrected later, equality built on it breaks exactly as a late-arriving key does, only less visibly.
saying these in an interview costs you the question
- Hashes the identifier and stores unwritten objects in a set
- Says two objects with empty identifiers are the same object
- Believes a container re-buckets elements when a field changes
- Uses reference identity only, so a reloaded row is a stranger
- Compares every mapped field and calls identical rows one object