skip to content

An unsaved Entity (say, a new `Order` object) doesn't yet have a database-assigned id — it's null until the first save. If equals() is implemented purely by comparing id, what breaks when two such transient, not-yet-persisted Order instances are compared or placed in a hash-based collection, and how do teams typically handle Entity equality before an id exists?

level: middleimportance: must knowfreq 65%

answer

  1. transient vs persistent entity
  2. null id before first save
  3. UUID assigned at construction avoids null-id gap
  4. hashCode shifts after id assignment breaks HashSet lookups

basics

~20 s

If equals() only checks the id and the id is null before saving, then two different unsaved orders both look 'equal' or the equality logic becomes unreliable, which can cause objects to get lost in Sets/Maps once the id is assigned. Teams usually fall back to comparing object references until an id exists, or generate the id in code (like a UUID) before saving so it's never null.

solid answer

~50 s

Comparing purely on a nullable id breaks for transient (not-yet-persisted) Entities: two distinct new Order objects both have id == null, so a naive equals() either mishandles nulls or reports both as indistinguishable, corrupting Sets/Maps that rely on distinct identity before the first flush. A subtler failure: if hashCode() also depends on id, an object inserted into a HashSet while transient gets rehashed once the database assigns a real id, so lookups after save silently fail even though the object is still present. Common fixes: (1) generate the identity client-side before persistence — a UUID or domain-specific ID value object created in the constructor — so id is never null and never changes; (2) fall back to Object identity (reference equality, i.e., don't override equals()) for transient instances, accepting that business equality is only meaningful post-persistence; (3) use a natural key temporarily if one exists and is guaranteed unique pre-save. Client-side UUID generation is the most robust and widely used approach because identity is available at construction time.

go deeper

for a junior

Should recognize that a not-yet-saved object might not have its id yet and that this could cause equality confusion; doesn't need to know ORM session/flush mechanics.

for a middle

Should be able to implement a correct equals()/hashCode() for an Entity that handles the transient (null-id) case safely and explain at least one concrete fix (client-generated id).

for a senior

Should know the Hibernate/JPA-specific version of this problem (hashCode shifting across a flush), and be able to choose and justify a strategy (UUID vs. reference-equality fallback vs. natural key) for a given system's constraints.

for a principal

Should evaluate the identity-generation strategy as an architectural decision with cross-cutting consequences — index locality, distributed-id-generation across services without central coordination, migration cost of switching key strategies — and set the convention for the team.

## Where the simple rule collides with persistence Entity equality sounds simple — 'compare by id, ignore everything else' — but it collides with a common wrinkle: many persistence technologies (relational databases with auto-increment primary keys, or ORMs like JPA/Hibernate with `@GeneratedValue`) don't assign identity until the object is actually written to the database. Before that first save/flush, the object is 'transient' and its id field is null. ## What goes wrong, step by step 1. If `equals()` is `return this.id != null && this.id.equals(other.id)` (a common 'safe' pattern), two different transient `Order` instances — freshly constructed, never saved, both with `id == null` — will both evaluate the guard as false against each other, so `equals()` returns false for both. 2. That specific pattern is safe in isolation, but it produces a subtler defect: because `equals()` and `hashCode()` effectively return 'not equal to anything', a `HashSet/HashMap` built while the entity is transient will silently 'lose' the entity after it's saved and its id changes from null to a real value, because `hashCode()` (which typically also depends on id) shifts, and the collection's bucket no longer matches. 3. The object is technically still there, but `contains()`, `remove()`, and `get()` stop finding it. ## Why the persistence layer works this way This exists because ORMs need a stable storage handle, and the cheapest way to get uniqueness at scale in a relational database is an auto-incrementing surrogate key assigned by the database at insert time. But the domain model conceptually needs identity from the moment the Entity is created in memory, because business logic (deduplication, equality checks in service code, adding an Entity to a Set before it's persisted) routinely runs before the first save. ## The three ways out, and what each costs The trade-off between resolutions is real. - **Falling back to Object/reference equality** for the transient case is simplest and avoids the hashCode-shifts-after-save bug entirely, since reference equality never changes — but `equals()` is meaningless as 'business equality' while the object lives in memory before persistence, which surprises developers who expect DDD Entities to always compare 'properly'. - **Generating identity client-side** — a UUID (v4 or v7), a Snowflake-style id, or a typed ID wrapping one of those — solves this cleanly: identity exists at construction time, is immutable for the object's entire lifetime including before-and-after persistence, and `equals()`/`hashCode()` work correctly from the moment the constructor runs. The cost: compact sequential integer keys are worse for some database index locality/page-fill characteristics and worse for humans reading foreign keys in ad hoc SQL, and you take on the responsibility of guaranteeing uniqueness yourself (UUIDs make collisions astronomically unlikely, so this is rarely a real cost). - **A weaker third option** — comparing on a natural/business key known pre-save — only works when such a key genuinely exists and is guaranteed unique before persistence. ## The failure signature in production In production, the failure signature of the naive approach shows up as 'objects disappearing from collections after save' or 'duplicate detection working in unit tests but not in integration tests.' A frequently-cited example: a shopping cart implemented as `Set<OrderLine>` where OrderLine is an Entity with a database-generated id. The cart correctly rejects duplicate line-additions while being built in memory (each OrderLine is a distinct transient object, distinct by reference), but after the cart is saved and reloaded within the same session, `hashCode()` (now based on the freshly-populated id) differs from what it would have been pre-save; if the same in-memory Set instance straddles that boundary, `contains()` checks against it can spuriously report 'not found' for an OrderLine that's actually present. The standard, well-documented fix in the Hibernate community is exactly what's described above: give every Entity a client-generated UUID (or a business key) assigned at construction, and never base `equals()`/`hashCode()` on a database-generated surrogate key alone.

  • Why is generating the id with a database sequence/auto-increment still popular despite this problem?
    Sequential integer/bigint keys are compact, fast to index, and produce naturally clustered inserts in many relational engines, and they're easy for humans to read in logs and ad hoc queries. Teams often accept the transient-equality complexity as a manageable cost, mitigating it by not overriding equals()/hashCode() until id is guaranteed non-null, or by never holding transient entities in hash-based collections.
  • Does using UUIDs for identity have any real downsides in high-throughput systems?
    Random (v4) UUIDs can hurt clustered-index insert performance and increase page fragmentation compared to sequential keys, because inserts land at random points in the index rather than appending at the end; time-ordered variants like UUIDv7 mitigate this by keeping a roughly monotonic prefix while still being generatable client-side.
  • How does this transient-identity problem interact with a Value Object embedded inside the Entity?
    It doesn't, because a properly modeled Value Object has no identity to be null in the first place — its equality is always structural and available immediately at construction, which is one more reason mis-modeling something as an Entity when it should be a Value Object tends to surface this class of bug unnecessarily.

Like a passport application before it's approved — the person exists and is themselves the whole time, but the passport number (the 'official id') only gets assigned once processing completes; if you tried to identify people solely by passport number, every applicant would look identical (no number yet) until their passport arrived.

saying these in an interview costs you the question

  • Doesn't realize an ORM-generated id can be null before first save
  • Uses the entity's auto-generated id in hashCode() without considering pre-save state
  • Assumes equals()/hashCode() 'just work' for JPA entities without any special handling
  • Can't name at least one strategy (client-side UUID, reference equality fallback) for the transient-identity gap

context