What problem does the Identity Field pattern solve when mapping a database row to an in-memory object, and why can't you just rely on the object's own reference identity or equals() instead?
answer
- PK mirrored in memory
- answers 'which row am I?'
- drives INSERT vs UPDATE
- surrogate vs natural key
- transient object has no id yet
basics
~20 sIdentity Field stores the database primary key inside the object so the mapper always knows which row that object represents, because loading the same row twice can create two different object instances that need to be recognized as the same thing.
solid answer
~50 sIdentity Field adds a property to a mapped object holding the value of the underlying table's primary key. Its job is to answer 'which row does this object correspond to?' so the mapping layer can decide INSERT vs UPDATE, and so two separately loaded objects representing the same row are recognized as the same entity. Reference identity isn't reliable here: the same row can be fetched by two different queries, in two sessions, or after a serialization round trip (e.g. a web form), producing distinct object instances that are still 'the same' business entity. Two brand-new unsaved objects have no database identity yet either. The identity field is usually the table's primary key - autoincrement integer, sequence, or UUID - and is normally kept separate from, or handled carefully in, equals()/hashCode(), because a transient object's identity value can change once persisted.
go deeper
Should state, in plain terms, that the object keeps a copy of the row's primary key and that this is how the ORM knows which row to update.
Should explain the INSERT-vs-UPDATE decision the identity field drives, and know the equals()/hashCode() pitfall for transient objects.
Should reason clearly about surrogate vs natural key trade-offs and connect the identity field to the Identity Map within a session/unit of work.
Should discuss identity-field design choices at a systems level: distributed ID generation, offline-first UUID generation, and the migration cost of ever changing a natural key used as identity.
## What an Identity Field is An Identity Field is a property added to a mapped class - conventionally called `id` - whose value mirrors the primary key of the table row that object represents. Mechanically, the lifecycle looks like this: 1. An application creates a new domain object with no identity value set (null, zero, or another 'unsaved' sentinel). 2. When the object is saved, the mapper issues an `INSERT`; the database (or the application, for client-assigned keys such as UUIDs) produces a primary key value, and the mapper writes that value back into the object's identity field. 3. From that moment the object is 'persistent': future `SELECT`, `UPDATE`, and `DELETE` statements target the row identified by that field, and the mapper can tell apart 'this object needs an INSERT' from 'this object needs an UPDATE' purely by checking whether the identity field is set. ## Two notions of sameness The pattern exists because in-memory objects and database rows have two different notions of sameness that do not naturally line up. - **An object** has reference identity (is this the exact same object in memory) and, if you define one, value equality (do two objects have equal fields). - **A database row** has row identity, defined by its primary key, which is stable and independent of the row's current column values. Without an explicit identity field, there is no way for application code to ask 'are these two Customer objects, loaded by two different queries, actually the same customer row?' - they might be two distinct object instances with (momentarily) identical field values, or one might be stale. The identity field is the anchor that survives across queries, across sessions, and across process boundaries (for example, an object serialized to JSON for a web form and then deserialized on form submission still carries its id, so the server knows which row to update). ## Surrogate versus natural key The main design trade-off is surrogate versus natural key as the identity field. - **A surrogate key** (autoincrement integer, database sequence, or UUID) decouples identity from business data: identity never changes even if every business attribute of the row changes, and it stays a single, simple column regardless of how many attributes make the entity 'unique' in business terms. The cost is that it's meaningless outside the application, requires an extra column, and - for autoincrement/sequence keys - complicates distributed or offline object creation, since the ID often isn't known until the row is actually inserted (this is one reason systems that create IDs offline, like mobile apps with local-first writes, favor UUIDs despite their larger size and index-fragmentation cost). - **A natural key** (e.g. an email address or national ID) avoids the extra column and is meaningful on its own, but breaks the assumption that identity should never change: if the 'natural' value must later be edited, every foreign key referencing it has to cascade, and composite natural keys make foreign-key mapping and `equals()` implementations noticeably more complex. ## The two failure modes 1. **The most common production failure mode** is building `equals()`/`hashCode()` naively off the identity field. If a newly created (transient) object is placed into a `HashSet` or used as a `HashMap` key before it is saved, its hash code is computed from a null or default identity value; once the object is persisted and the identity field is populated, its hash code changes, and the object effectively 'disappears' from the collection it was already stored in, because the collection is still bucketed by the old hash. This is a well-known gotcha with JPA/Hibernate entities and is usually fixed by excluding the identity field from `equals()`/`hashCode()` until it is set, or by using a business key, or by overriding `equals()` to fall back to reference equality for transient instances. 2. **A second failure mode** is conflating identity with business meaning: reusing a mutable natural key as the identity field means an innocuous business change (a customer changes their email) turns into a data-integrity migration across every table that stores a foreign key to it. ## Where you have already seen it Hibernate/JPA's `@Id` annotation, Rails ActiveRecord's implicit `id` column, and Django's automatic `pk` field are all direct applications of this pattern - and in each case the identity field also powers the ORM's first-level identity map, so within one unit of work, two loads of the row with id=42 return the very same object reference, not just equal ones. The pattern isn't ORM-specific either: a REST API returning {"id": 42, ...} is exposing the identity field to clients precisely so a later PUT or PATCH request can say which resource it means to modify.
- How does an Identity Field interact with an ORM's first-level cache / Identity Map within one session or unit of work?The Identity Map keys its cache by the identity field's value, so when the mapper is asked to load a row it has already loaded in this session, it returns the exact same object reference instead of building a new one. This guarantees that within one unit of work, changes made through one reference are visible through every other reference to the same row, and prevents two conflicting in-memory copies of the same entity from silently diverging.
- What goes wrong if you put the identity field straight into equals()/hashCode() before the object has been saved?A transient object's identity field is null or a default sentinel, so its hash code is computed from that placeholder value. If the object is inserted into a hash-based collection before being saved, then saved (which changes the identity field and therefore the hash code), the collection's internal bucket no longer matches, and lookups for that object silently fail even though it's still physically present in the collection.
- How do surrogate keys like autoincrement or UUID differ as an identity field, and which would you pick for a distributed, offline-capable system?Autoincrement/sequence keys are compact and index-friendly but require a round trip to the database (or a central sequence) to be assigned, which is awkward when multiple nodes create rows offline and later sync. UUIDs can be generated locally with negligible collision risk, so they suit distributed or offline-first systems, at the cost of larger storage, worse index locality, and less human-readable identifiers.
Like a hospital wristband: two different photographs of the same patient (two object instances loaded separately) are recognized as the same person because they carry the same wristband number, not because the photos look identical.
saying these in an interview costs you the question
- Says two objects loaded from the same row are 'different' with no way to reconcile them
- Builds equals()/hashCode() purely off the raw id field without handling the transient/unsaved case
- Confuses Identity Field with the Identity Map / first-level cache pattern
- Assumes a natural key can never cause problems as an identity field
- Can't explain how the mapper decides INSERT vs UPDATE