Why must you correctly override both equals() and hashCode() for objects stored in a HashSet, and what breaks if you don't?
answer
- hashCode picks the bucket, equals confirms within it
- Contract: equal objects => equal hash codes
- Override BOTH or get phantom duplicates / failed contains
- Mutating a hashCode field after insert 'loses' the element
- Use immutable fields / records / Objects.hash
basics
~10 sA HashSet uses hashCode() to find an element's bucket and equals() to confirm a match. If your objects don't override both consistently, the set can store duplicates or fail to find items you added.
solid answer
~50 sHashSet decides membership in two steps: it calls hashCode() to pick a bucket, then equals() to compare against the few elements there. The contract is: equal objects must have equal hash codes. If you override equals() but not hashCode(), two logically-equal objects can get different hash codes, land in different buckets, and the set never compares them — so it stores both as 'duplicates'. If you override hashCode() but not equals(), the default equals() uses reference identity, so distinct-but-equal objects are again treated as different. Either way de-duplication and contains() break. A subtler bug: mutating a field used in hashCode() after insertion changes the object's bucket, stripping the JVM's ability to find it — it becomes a lost element still consuming memory. The safe practice is to make hashCode/equals consistent and based on immutable fields, ideally using immutable value objects (or records, which generate both correctly).
code
java · 19 lines// BROKEN: equals overridden, hashCode forgotten
class Point {
final int x, y;
Point(int x, int y) { this.x = x; this.y = y; }
@Override public boolean equals(Object o) {
return o instanceof Point p && p.x == x && p.y == y;
}
// no hashCode() -> identity-based, inconsistent!
}
Set<Point> s = new HashSet<>();
s.add(new Point(1, 2));
System.out.println(s.contains(new Point(1, 2))); // false (!) different buckets
System.out.println(s.size()); // grows with 'duplicates'
// FIXED: a record generates a correct, consistent pair automatically
record PointR(int x, int y) {}
Set<PointR> ok = new HashSet<>();
ok.add(new PointR(1, 2));
System.out.println(ok.contains(new PointR(1, 2))); // truego deeper
Knows both methods must be overridden together for custom objects in a HashSet to behave correctly.
Explains the bucket-then-equals lookup and the equal-objects-imply-equal-hashcodes contract, and the duplicate/failed-contains symptoms.
Diagnoses all failure modes including the post-insertion mutation bug, recommends immutable fields/records, and reasons about collision behaviour and performance.
Sets codebase conventions (value objects, records, generated equals/hashCode), understands HashMap treeification of collision chains, and weighs identity vs. equality semantics in domain modelling.
## The two methods, defined Every Java object inherits two methods from `Object`: - **`hashCode()`** returns an `int` — a numeric summary of the object. The default implementation derives it from the object's memory identity, so two different instances almost always get different numbers. - **`equals(Object o)`** returns a boolean — whether two objects are 'the same' logically. The default implementation is `this == o`, i.e. **reference identity**: true only if they are literally the same object in memory. ## How HashSet uses them When you `add(x)` or `contains(x)`, HashSet (via its HashMap): 1. Calls `x.hashCode()` and maps it to a **bucket** (a slot in an internal array). 2. Looks **only inside that bucket** and uses `equals()` to compare `x` against the elements already there. 3. If an equal element is found, `x` is a duplicate (not added / `contains` returns true). Otherwise it is a new element. The critical insight: **`equals()` is only ever consulted within the bucket that `hashCode()` selected.** If two objects don't share a hash code, they never even get compared. ## The contract The `Object` documentation mandates: **if `a.equals(b)` is true, then `a.hashCode() == b.hashCode()` must also be true.** (The reverse need not hold — unequal objects *may* share a hash code; that's a collision and is fine.) ## What breaks when you violate it **Case 1 — override equals() but NOT hashCode():** Two logically-equal objects (say two `Point(1,2)` instances) now satisfy `equals`, but inherit the default identity-based `hashCode`, so they get *different* hash codes. They land in *different* buckets. HashSet never compares them, so it stores **both** — your set contains apparent duplicates, and `contains(new Point(1,2))` may return false even though an equal point is inside. **Case 2 — override hashCode() but NOT equals():** The two objects now share a bucket, but the default `equals` uses reference identity, so they're judged different. Again both are stored. De-duplication fails. **Case 3 — the consistency bug (mutation):** Suppose `hashCode()` uses a mutable field. You add the object (it goes to bucket A based on its current hash), then **change that field**. Its hash now points to bucket B, but it physically still sits in bucket A. Now `contains(theSameObject)` looks in bucket B, doesn't find it, and returns false — the element is **lost** while still occupying memory and inflating `size()`. This is why hashCode/equals should rest on **immutable** fields. ## The correct recipe - Override **both** together, using the **same set of fields**. - Base them on fields that don't change while the object is in the set (prefer immutability). - Make `equals` reflexive, symmetric, transitive, and consistent. - Let tooling generate them: IDEs, `java.util.Objects.hash(...)` / `Objects.equals(...)`, Lombok's `@EqualsAndHashCode`, or — best — a **`record`**, which auto-generates both correctly from its components. ## Quick mental model `hashCode` gets you to the right shelf fast; `equals` confirms it's the exact book. Both must agree on what 'the same' means, or you'll shelve duplicates and lose books.
- Is it legal for two unequal objects to have the same hashCode?Yes. That is a hash collision and is perfectly valid — the contract only requires equal objects to share a hash code, not the reverse. HashSet resolves collisions by using equals() within the bucket. Many collisions just degrade performance toward O(n) (or O(log n) once HashMap treeifies a long chain).
- Why are records a good fit for HashSet elements?Records are immutable and auto-generate equals() and hashCode() from their components, so the pair is always consistent and based on unchanging fields — eliminating both the missing-override and the mutation bugs.
hashCode is the aisle number in a library; equals is reading the title to confirm the exact book. If a book has the wrong aisle number, no one will ever look in the right aisle to find it — it's effectively lost on the shelves.
saying these in an interview costs you the question
- Overriding only one of the two methods
- Basing hashCode/equals on mutable fields and then mutating them while stored
- Believing equals() is enough for HashSet to de-duplicate
- Thinking identical hash codes mean the objects are equal
- Returning a constant from hashCode() 'to be safe' (degrades the set to O(n))