How can a mutated hash key cause a memory leak, and how would you detect and prevent this class of bug in a long-running service?
answer
- remove(key) no-ops → entry strongly referenced forever
- Reachable but logically dead = JVM leak shape
- Detect: rising post-GC heap floor + growing map size
- Heap dump in MAT → dominator tree shows the map
- Prevent: immutable/id keys, encapsulated re-key, bounded caches, hashCode lint
basics
~20 sIf a stored key's hashCode changes, you can no longer remove it by key, so the entry stays in the map forever — a slow leak as the map grows. Prevent it with immutable/id-based keys; detect it via heap growth and heap dumps showing maps full of unreachable-by-key entries.
solid answer
~50 sWhen code mutates a field used by a stored key's hashCode, the entry can no longer be located, so remove(key) silently does nothing and the entry stays strongly referenced by the map. In a long-running service that repeatedly inserts and 'removes' such keys, the failed removals accumulate; the map grows without bound — a classic Java leak where objects are reachable but logically dead. Detection: watch retained heap and map sizes over time (metrics, GC logs, OOM dumps); a heap dump analyzed in MAT shows a HashMap with a huge bucket count and entries that the application can't reach by key. Prevention is design-level: use immutable keys (records, Strings, enums) or id-based equals/hashCode so keys can't be mutated in place; encapsulate any necessary re-keying behind a method (remove→mutate→reinsert); add static-analysis/code-review rules flagging mutable fields in hashCode; and bound caches (size/TTL, or WeakHashMap) so even a mistake can't grow unbounded.
go deeper
Can state that a lost key means remove won't work so the entry stays around; may not connect it to a service-level leak.
Explains the reachable-but-dead leak shape and that remove silently fails, and names immutable keys as the fix.
Adds detection (heap trend, map-size metrics, heap dump in MAT) and a layered prevention plan, and reasons about WeakHashMap as a mitigation.
Designs the hazard out at the model level (immutable value types, id identity, unrepresentable mutation of keys), institutes guardrails (lints, bounded caches, heap-trend alerting), and frames it as systemic resilience, not a one-off fix.
## How a changed hashCode becomes a leak Recall the mechanic: a hash collection buckets a key by its `hashCode()` at insertion and never re-buckets it. If a stored key's hashCode-relevant field is mutated, the entry stays in its **old** bucket while every key-based operation now hashes to a **different** bucket. Crucially, **`remove(key)` also fails** — it hashes to the new bucket and removes nothing. So the entry remains, still **strongly referenced** by the map's internal array. In a short program that's a curiosity. In a **long-running service** it's a **memory leak**: imagine a request handler that, per request, inserts a session/connection/entity into a map, mutates one of its hashCode fields during processing, then calls `map.remove(it)` in a finally block to clean up. The remove silently no-ops every time. Over hours, millions of 'cleaned up' entries pile up. The objects are **reachable** (so the GC won't collect them) but **logically dead** (no code can ever use them by key). Retained heap climbs until `OutOfMemoryError: Java heap space`. This is the defining shape of a JVM leak: not 'forgot to free memory' (the GC handles freeing) but **unintended reachability** — a live data structure clinging to objects you believe you removed. ## Detection 1. **Trend the heap, not a snapshot.** Enable GC logging / JFR; watch **retained (live-after-GC) heap** trending upward over time. A leak shows a rising floor after each full GC, not just sawtooth churn. 2. **Instrument suspect maps.** Export the `.size()` of long-lived caches/registries as a metric. A monotonically growing map that 'should' stay bounded is the smoking gun. 3. **Heap dump analysis.** Trigger `-XX:+HeapDumpOnOutOfMemoryError` or `jmap` and open the dump in **Eclipse MAT** / VisualVM. Look at the **dominator tree**: a `HashMap`/`Node[]` retaining a huge object graph. The tell is a map whose entries the application has no remaining key-path to. 4. **Reproduce in a soak test.** Drive the insert/mutate/remove path under load and assert the map size returns to baseline; if it doesn't, you've localized it. ## Prevention (in priority order) 1. **Immutable or id-based keys.** Use `record`s, `String`s, enums, or base `equals`/`hashCode` on an immutable id. If a key can't be mutated in place, this leak is impossible by construction. This is the real fix. 2. **Encapsulate re-keying.** Where mutation of a keyed field is unavoidable, force all mutations through a method that does `remove → mutate → put` atomically; never expose a raw setter on a stored key. 3. **Lint / review gates.** Add a static-analysis rule (or PR-review checklist item) flagging non-final fields referenced in `hashCode`. Catch it before it ships. 4. **Bound the blast radius.** Make long-lived maps **defensive**: cap size with an eviction policy (LRU/TTL via a cache library), or use `WeakHashMap` / soft-keyed caches so even a logic mistake can't grow unbounded. Bounding doesn't fix the bug but turns a fatal leak into a tolerable one. 5. **Observability as a habit.** Treat every unbounded long-lived collection as a liability: give it a size metric and an alert. ## Why this matters at a principal level The individual bug is junior-to-mid to explain; the **systemic** angle is senior/principal: choosing an identity strategy for the whole domain model (immutable value types, id-based entities), making the dangerous pattern *unrepresentable* (no setters on keys), and putting guardrails (lints, bounded caches, heap-trend alerts) so a single careless commit can't take down the service hours later. The goal is not to remember the trick — it's to design so the trick never has to be remembered.
- Why won't the garbage collector reclaim the stranded entries?They remain strongly reachable through the map's internal node array. The GC only collects unreachable objects; 'unretrievable by key' is not the same as 'unreachable'. The map root keeps them alive.
- Would using WeakHashMap fix the underlying bug?Not the bug itself, but it caps the damage: a WeakHashMap key with no other strong reference becomes collectible, so stranded entries can eventually be cleared. The correct fix is still immutable/id-based keys; WeakHashMap is a blast-radius limiter.
saying these in an interview costs you the question
- Calling it a 'native memory leak' — it's ordinary heap, objects reachable from a live map.
- Expecting the GC to clean up unretrievable-but-reachable entries.
- Treating it as purely a coding bug rather than a design/observability problem in a long-running service.
- Believing bounding the cache (LRU/TTL) fixes correctness — it only limits the blast radius.