skip to content

You override equals() and hashCode() on a JPA entity using its database-generated primary key, add a brand-new instance to a HashSet, and then persist it. Why can the set stop finding that object, and how do you write equals/hashCode so it cannot happen?

level: middleimportance: must knowfreq 66%

answer

  1. Bucket chosen at add(), never recomputed
  2. Generated id: null -> 42 at flush = hash changes
  3. Objects.equals(null, null) == true -> all transients collapse
  4. Constant hashCode + "id != null && id.equals"
  5. Or assign UUID in constructor (UUIDv7 for index locality)

basics

~20 s

Before persist the generated id is null, so hashCode is computed from null; at flush the id is assigned and hashCode changes, but the object still sits in the bucket chosen by the old hash — contains() and remove() miss it. Fix: constant hashCode plus id-equality, or an identifier assigned before persist.

solid answer

~50 s

`HashSet` picks a bucket from `hashCode()` at insertion time and never re-buckets. A database-generated id is `null` when the object is transient, so the hash reflects "no id"; `persist()` plus flush writes the real id, the hash changes, and the object is now in the wrong bucket — `contains(x)` and `remove(x)` return false for the very instance that is inside the set. Equality is broken too: with a naive `Objects.equals(id, other.id)`, two different transient instances both have `null` ids and compare equal, so a `Set` collapses them into one element. Two fixes are accepted. **Constant hashCode**: return a fixed value (e.g. `getClass().hashCode()`), and in `equals` return true only when `this == other` or both ids are non-null and equal — never treat two nulls as equal. All instances of the type share a bucket, which is fine for aggregate-sized sets. **Assigned identifier**: generate a `UUID` (or use a natural key) in the constructor so the id is non-null and immutable from birth, and hash it normally.

code

java · 17 lines
java
@Entity
class Order {
    @Id @GeneratedValue private Long id;

    @Override public int hashCode() { return Objects.hash(id); }        // BROKEN
    @Override public boolean equals(Object o) {
        return o instanceof Order other && Objects.equals(id, other.id); // BROKEN
    }
}

Set<Order> set = new HashSet<>();
Order o = new Order();
set.add(o);
em.persist(o);
em.flush();          // id: null -> 42, hashCode changes
set.contains(o);     // false, although the set still contains o
set.size();          // 1 — it is in there, just in the wrong bucket

go deeper

for a junior

Recall that a generated id is null before persist and set afterwards, and that a HashSet gets confused when hashCode changes. Naming one safe fix is enough.

for a middle

Explain bucketing precisely, name both bugs (mutating hash and null==null equality), and write the constant-hashCode implementation correctly on demand.

for a senior

Connect it to real symptoms: duplicate children in a mapped Set, remove() that silently does nothing, orphans left after a cascade, and how you would migrate an existing entity to assigned UUIDs.

for a principal

Argue the modelling stance — assigned identifiers make entities well-behaved values everywhere in the system, and weigh that against UUID index fragmentation and time-ordered UUID variants.

## The mechanics of the failure `HashSet` is backed by a `HashMap`. On `add(e)` it calls `e.hashCode()`, maps the hash to a bucket index, and stores the entry there. It does **not** re-hash existing entries when their state changes — there is no way for it to observe such a change. Lookups (`contains`, `remove`, iteration-plus-equals) recompute the hash of the probe object and go straight to that bucket. Now take the classic "IDE-generated equals/hashCode over the @Id" entity with `@GeneratedValue`: 1. `new Order()` — `id` is `null`. `hashCode()` returns the hash of "null id", say bucket 3. 2. `set.add(order)` — stored in bucket 3. 3. `em.persist(order)` and a flush — Hibernate obtains the identifier from the sequence/identity column and writes it into the field. `hashCode()` now returns the hash of `42`, say bucket 17. 4. `set.contains(order)` — probes bucket 17, finds nothing, returns **false**. The object is still physically in the set and iteration will show it, but every hash-based operation misses it. `remove` fails silently; a subsequent `add` of the same object inserts a duplicate. This is not a Hibernate bug. It is a violation of the `hashCode` contract: *the hash of an object must not change while it is a key in a hash structure*. Entities with generated ids mutate exactly that field. ## The second, quieter bug: null-id equality A typical generated implementation is `Objects.equals(this.id, other.id)`. For two **different** brand-new objects both ids are `null`, so `Objects.equals(null, null)` is `true` — every transient instance of the type is equal to every other. Add three distinct new `OrderLine` objects to a `Set` before flushing and you keep one. If that set is a mapped `@OneToMany`, you have just silently dropped two rows. Equality must therefore be explicit: identity first, then *both* ids non-null and equal. ## Fix 1 — constant hashCode, id-based equals ```java @Override public int hashCode() { return getClass().hashCode(); } @Override public boolean equals(Object o) { if (this == o) return true; if (!(o instanceof Order other)) return false; return id != null && id.equals(other.getId()); } ``` The hash is constant for the whole lifetime, so the contract cannot be violated. `equals` is reference equality while transient (each new object is only equal to itself) and id equality once persisted — which is exactly the semantics you want. This is the pattern Hibernate's own documentation recommends when identifiers are generated. The cost is real but usually irrelevant: every instance of the type hashes to one bucket, so a `HashSet` degrades to linear scan. For a `Set<OrderLine>` inside an aggregate (tens of elements) this is nothing; for a set of 100k entities it matters, and you should not be holding 100k entities in a set anyway. One subtlety: use `getClass().hashCode()` carefully with proxies — a lazy proxy is a generated subclass, so its `getClass()` differs. Because Hibernate's proxy *delegates* `hashCode()` to the underlying entity implementation, this normally resolves correctly; if you want to be explicit, return a literal constant instead. ## Fix 2 — assign the identifier yourself ```java @Id private UUID id = UUID.randomUUID(); ``` Now the id is non-null and immutable from construction, so a normal `hashCode` over it is stable and `equals` needs no null handling. Entities behave sanely in sets, in caches, and after detachment, with no bucket pathology. The trade-offs are storage and index characteristics — random UUIDs fragment B-tree indexes and are wider than a `bigint`, which is why time-ordered variants (UUIDv7 / ULID) are commonly used instead. Hibernate 6 supports `@UuidGenerator(style = TIME)` for exactly this. ## Fix 3 — a true business key If the domain has an immutable, unique, non-null attribute (order number, ISBN, an ISO country code), hash and compare on that. It is stable by definition and reads naturally. The failure mode is picking something that turns out to be mutable or non-unique later — email addresses are the standard cautionary tale. ## What to say in an interview Name the mechanism (bucket chosen at insert, hash mutated at flush), name the second bug (two nulls compare equal), then give the two accepted fixes and say which you would pick and why. Mentioning that Hibernate's `PersistentSet` is a `Set` and therefore inherits the same trap when you map `@OneToMany(mappedBy=...) Set<Child>` shows you have seen it in a real codebase.

  • Doesn't a constant hashCode destroy HashSet performance?
    It collapses the set to a linked list, so lookups become O(n) for that type. In practice entity sets are aggregate-sized — the lines of an order, the roles of a user — where linear scan over a few dozen elements costs nothing. If you genuinely hold tens of thousands of entities in a set, the real fix is an assigned identifier, or not keeping them in a set at all.
  • Why must equals return false when both ids are null instead of treating them as equal?
    Because two distinct transient instances would otherwise be equal to each other, so a Set would keep only one of them and the rest would never be inserted into the database. Returning false for a null id makes a transient object equal only to itself, which matches the reference-equality semantics the persistence context provides once it is managed.
  • Does the same trap apply to a Set mapped as @OneToMany?
    Yes. Hibernate's PersistentSet wraps a HashSet, so children added before flush are bucketed by their pre-persist hash. Broken equals/hashCode there produces silently deduplicated children, failed remove() calls that leave orphans, and confusing behavior after merge. Using List with an order column sidesteps the hashing entirely but brings its own delete-and-reinsert costs.

A HashSet files an object in a drawer chosen by its hash, like filing a folder under the customer's phone number. Persisting changes the phone number but nobody re-files the folder — you go to the new drawer and the folder isn't there, even though it never left the cabinet.

saying these in an interview costs you the question

  • "Objects.hash(id) is fine, the id is set right after persist" — that is exactly the mutation that breaks the bucket
  • Treating two null ids as equal, collapsing distinct new entities in a Set
  • Claiming HashSet re-hashes its elements when their state changes
  • "Just use the ordered List instead" as the only answer, without knowing why the Set failed
  • Using mutable business fields to stabilise the hash instead of an immutable identifier

context