During code review, a diff gives a SKU record value-based equality but leaves its hashing identity-based — why will a hash-set dedupe still emit duplicate SKUs?
answer
- which half does the set consult first?
- identity hashing: one hash per instance
- equal records, different buckets, never compared
- tests that reuse one instance cannot see it
basics
~20 sThe dedupe set routes by hash before it ever checks equality. Identity-based hashing gives each record instance its own hash, so two value-equal SKU records land in different buckets, are never compared, and both survive. Value equality without matching hashing changes nothing.
solid answer
~50 sThe diff is half a change. A hash-set insert computes the element's hash, goes to that bucket, and compares equality only against what is already there. Identity hashing derives the hash from the instance itself, so two value-equal SKU records — the same product arriving from two supplier feeds — hash apart, are never compared, and are both stored: the dedupe emits duplicates while the new equality code sits unused. Worse, the existing tests almost certainly pass, because tests tend to reuse one instance, and identity hashing is perfectly consistent with itself. As a reviewer I would block the diff and require the matching value-based hash over the same fields, plus a test that inserts two independently constructed equal records and asserts the set holds one. Equality and hashing are one unit; a diff touching one must touch both.
go deeper
Be ready to recall that hash sets need equality and hashing to agree, and that defining only one half leaves the container silently broken for value semantics.
An interviewer expects the full insert trace: hash routes each equal record to its own bucket, equality is never consulted, both survive. Explain why identity hashing is self-consistent yet wrong here.
Demonstrate reviewer judgment: block one-sided diffs, articulate the mechanism in the review comment, demand the two-independent-instances set test, and explain why green tests gave false confidence.
Own the prevention policy: steer key types toward compiler-derived value forms where equality and hashing cannot drift, and make the pairing an enforced review or lint gate rather than tribal knowledge.
## Why half a key type is worth zero Most environments give every object a default equality and a default hash that agree with each other: equality means "the same instance", and the hash is derived from the instance's identity. That inherited pair is internally consistent — and consistently useless for value semantics. The diff under review changes only one half: SKU records now compare equal when their product code, size and color match. But the hash still answers per-instance. The equality-hash contract — keys that compare equal must produce equal hashes — is now violated by every pair of distinct-but-equal instances in the system. ## Tracing the dedupe The catalog job builds a hash set to deduplicate SKUs merged from several supplier feeds. Two feeds deliver the same product as two separate record instances: 1. Insert record A: hash(A) routes to bucket 7; bucket empty; store A. 2. Insert record B, value-equal to A: hash(B) is identity-derived, routes to bucket 23; bucket empty; store B. The set consulted equality zero times. Both records survive, the "deduplicated" catalog carries phantom duplicates, and downstream systems see the same product twice. No error, no warning — the set behaved correctly given inconsistent answers from its key type. ## Why the tests are green This defect has a signature reason for escaping test suites: **identity hashing agrees with identity, and tests love reusing instances.** A test that inserts the same reference twice, or that asserts `a equals b` on two records without ever putting both into a hash-based container, exercises nothing the diff broke. The bug only fires when two *independently constructed* equal instances meet inside a hash-routed structure — exactly the production condition (two feeds, two deserializations) and exactly what a convenient unit test avoids. Green tests here prove consistency of the old identity pair, not correctness of the new value semantics. The test to demand in the same diff: construct two equal records through two separate code paths, insert both into a hash set, assert size one. A stronger companion is a property test over random field values asserting that equal records report equal hashes. ## The reviewer's rule, and its mirror image The operative review rule is mechanical: **a diff that touches equality or hashing must touch both, derive both from the same fields, and ship the two-instance container test.** Some ecosystems make the pairing hard to get wrong — compiler-generated value classes and record types derive both together — which is why reviewers push key types toward those forms. The mirror-image diff is worth recognizing too: value-based hashing over identity-based equality. Curiously, that one *honors* the letter of the contract — equal-by-identity objects trivially hash equal — yet the dedupe still fails, because the set defers the final verdict to equality, and identity equality never declares two distinct instances equal. The lesson generalizes: the contract is necessary, not sufficient. Correct behavior needs both halves expressing the *same* notion of sameness, not merely a non-contradictory pair. ## What to write in the review A useful comment names the mechanism, not just the rule: "Hash sets route by hash before comparing; with identity hashing, two equal SKUs never meet, so this dedupe is a no-op for cross-feed duplicates. Please add the value-based hash over the same fields and a test inserting two independently built equal records." That framing survives the author's likely rebuttal — "but equality is defined now" — because it explains why the equality code is unreachable for the case that matters.
- Why did the existing unit tests pass despite the broken pairing?Because identity hashing is consistent with itself, and the tests never created the failing condition: two independently constructed equal instances inside one hash-routed container. Tests that reuse a single reference, or assert equality without a hash set, exercise nothing the diff broke. The bug needs distinct-but-equal instances to meet in a bucket-routed structure — a production condition, not a convenient test setup.
- What test would have caught this in the diff?Build two equal SKU records through separate construction paths, insert both into a hash set, and assert the set's size is one. Add a property test generating random field values and asserting that records comparing equal also report equal hashes — that assertion is the contract stated as executable code.
- Is the mirror-image diff — value-based hashing over identity-based equality — also broken?Behaviorally yes, though it technically honors the contract: identical instances hash equal trivially. But the set's final verdict is equality's, and identity equality never calls two distinct instances equal, so the dedupe still keeps both. It shows the contract is necessary but not sufficient — both halves must express the same notion of sameness.
saying these in an interview costs you the question
- Value-based equality alone makes a hash set deduplicate correctly
- Duplicates surviving a hash set mean the hash function collides too much
- Green unit tests prove a key type is safe for hash-based containers
- The set compares every pair of elements, so hashing only affects speed