skip to content

If you put an object into a HashSet (or as a HashMap key) and then mutate a field that its hashCode depends on, what happens when you later try to look it up or remove it, and why?

level: middleimportance: must knowfreq 70%

answer

  1. Bucket chosen at insert time, never re-checked
  2. Mutate hashCode field → object stuck in old bucket
  3. contains/get/remove hash to new bucket → not found
  4. Stranded entry = unretrievable leak
  5. Fix: immutable keys / no mutable fields in hashCode

basics

~20 s

The object gets 'lost' in the set. Hash collections place keys into buckets by hashCode. Change a field used in hashCode and the new hashCode points to a different bucket, so lookups, contains, and remove no longer find it.

solid answer

~40 s

A HashSet/HashMap stores each key in a bucket chosen from its hashCode. When you insert, the collection records the object in the bucket matching the hashCode it had at insertion time. If you then mutate a field that participates in hashCode, the object's hashCode changes, but the collection does not re-bucket it; it stays in the old bucket. A later contains/get/remove computes the new hashCode, goes to the new (wrong) bucket, and fails to find the entry, returning false/null even though the object is physically still there. The entry becomes a stranded, unreachable 'leak' you can't remove by key. The fix is to never include mutable fields in hashCode, or to use immutable objects as keys. If you must mutate, remove the key first, mutate, then re-insert.

go deeper

for a junior

Knows that mutating an object that's a HashMap key or in a HashSet can make it 'disappear' from lookups, and that immutable keys avoid the problem.

for a middle

Explains the bucket mechanism: bucket is chosen from hashCode at insert time and never updated, so a changed hashCode sends lookups to the wrong bucket. Knows remove also fails and that it's a leak.

for a senior

Articulates the design rule (exclude mutable fields from hashCode, prefer immutable keys), the remove→mutate→re-insert workaround, and that iteration still reaches the entry while key access does not.

for a principal

Frames it as an invariant of hash-based structures, generalizes to TreeMap/compareTo and to distributed/persistent hashing, and sets a codebase policy (immutable keys, id-based identity, defensive design) to prevent the class of bug.

## The setup: what a hash collection actually does A **hash collection** (`HashMap`, `HashSet`, which is backed by a `HashMap`) stores entries in an array of **buckets**. To decide which bucket an object goes into, it calls the object's **`hashCode()`** method — a method every Java object has that returns an `int`. The collection takes that int, spreads/mixes it, and reduces it modulo the array length to get a bucket index. Two objects with the same hashCode land in the same bucket (a **collision**); within a bucket the collection uses **`equals()`** to tell entries apart. So a hash collection relies on a simple but strict assumption: **an object's hashCode does not change while it is stored.** The collection records *where* it put the object exactly once — at insertion time — and never re-checks. ## What goes wrong when you mutate a hashCode field Suppose `hashCode()` is computed from a field `name`. You insert the object; the collection computes `hashCode()` = (say) 42, reduces it to bucket 7, and stores the object in bucket 7. Now you call `obj.setName(...)`, changing a field that `hashCode()` reads. The object's `hashCode()` is now (say) 99, which would reduce to bucket 3. But the collection **did not move the object** — it is still sitting in bucket 7. Nothing notified the collection of the change; there is no mechanism for that. Now you do `set.contains(obj)` (or `map.get(obj)`, or `set.remove(obj)`): 1. The collection computes the object's *current* hashCode → 99 → bucket 3. 2. It looks in bucket 3 and finds nothing matching. 3. It returns `false` / `null` — even though the very same object is physically still in the collection (in bucket 7). The entry is now **stranded**: you can't find it, can't `get` it, can't `remove(key)` it. Iterating the whole collection still visits it (iteration walks every bucket), but key-based access is broken. In a long-lived map this is a genuine **memory leak**: the entry stays reachable from the map but is unretrievable, so it's never removed and never collected. ## Why even `remove` fails People expect at least `remove` to clean up. It doesn't: `remove` also hashes the key to find the bucket, so it goes to bucket 3 and finds nothing. The only ways to get rid of a stranded entry are to mutate the field *back* to its original value (restoring the old hashCode), iterate-and-`Iterator.remove`, or clear the whole collection. ## The rules that follow 1. **Don't include mutable fields in `hashCode()` (and `equals`).** Base identity on fields that never change after construction — ideally a stable id assigned at creation. 2. **Prefer immutable keys.** Strings, boxed primitives, `record`s with immutable components, enums — these are the safest hash keys because their hashCode can never change. 3. **If you truly must mutate a key**, follow the remove → mutate → re-insert dance so the collection re-buckets it. ## Note on equals The same hazard applies to `equals` if `equals` and `hashCode` use the same fields (as the contract requires). But hashCode is the field that controls *bucketing*, so it's the one that determines whether the entry can be found at all. `TreeMap`/`TreeSet` (sorted, comparison-based — not hash-based) have an analogous but separate hazard: mutating a field used by `compareTo`/`Comparator` corrupts the tree ordering.

  • Why doesn't remove(key) clean up the stranded entry?
    remove also computes the key's current hashCode to locate the bucket. Since the hashCode changed, it searches the new (wrong) bucket and finds nothing, so it removes nothing. The entry stays in the old bucket.
  • Can you still get the object out at all?
    Yes — iteration walks every bucket, so a for-each or Iterator still visits it, and Iterator.remove() can delete it. Only key-based access (get/contains/remove) is broken.

saying these in an interview costs you the question

  • Saying the object is physically removed or garbage-collected when mutated — it stays in the collection.
  • Claiming remove(key) will still work after the field changed.
  • Thinking the collection automatically re-buckets entries when a field changes — there is no such mechanism.

context