skip to content

A service fetches a list of records and deduplicates it with new Set(records), but duplicates still get through. Explain why, and how you would deduplicate by content instead.

level: seniorimportance: should knowfreq 42%

answer

  1. fresh parse, fresh reference
  2. the container never inspects fields
  3. name the key yourself
  4. first-wins versus last-wins
  5. serialized keys are order-sensitive

basics

~20 s

A Set matches objects by reference, and every parsed record is a fresh object, so nothing collapses. Deduplicate on a derived primitive identity instead: keep a Set of ids, or build a Map keyed by that identity and take its values.

solid answer

~40 s

`new Set(records)` compares objects by reference, and every record produced by parsing a response is a distinct object, so two payloads describing the same entity are two entries and the Set removes nothing. The fix is to name the identity explicitly. If the records carry a stable key, filter against a `Set` of seen ids, or build `new Map(records.map(r => [r.id, r]))` and take `[...map.values()]` — the Map form also decides last-wins for you, while the seen-Set is first-wins. If there is no natural key you have to synthesize one, and a `JSON.stringify` key is fragile: property order changes the string, `undefined` and functions are dropped, `NaN` becomes `null`, and cycles throw. Better to build the key from the specific fields that define identity in your domain.

code

javascript · 22 lines
javascript
const records = [
  { id: 1, name: 'a' },
  { id: 2, name: 'b' },
  { id: 1, name: 'a2' },
];

console.log(new Set(records).size); // 3 — reference identity, nothing removed

// First-wins, order preserved
function dedupeBy(items, keyOf) {
  const seen = new Set();
  return items.filter(i => {
    const k = keyOf(i);
    if (seen.has(k)) return false;
    seen.add(k);
    return true;
  });
}
console.log(dedupeBy(records, r => r.id).map(r => r.name)); // ['a', 'b']

// Last-wins
console.log([...new Map(records.map(r => [r.id, r])).values()].map(r => r.name)); // ['a2', 'b']

go deeper

for a junior

Remember the core fact: a Set removes duplicate objects only when it is literally the same object twice, so deduplicating fetched records needs a key like an id.

for a middle

Write the dedupe two ways — a seen-Set filter and a Map keyed by id — and explain which one keeps the first copy and which keeps the last.

for a senior

Diagnose the failure from the symptom, choose first-wins or last-wins deliberately, name the JSON.stringify key pitfalls, and bound a long-lived seen-Set before it becomes a leak.

for a principal

Treat repeated records as an upstream contract problem: decide who owns the identity, whether the API should guarantee it, and what the dedupe window means for correctness across retries and replays.

## Why the Set does nothing Set membership for object values is reference identity. Two objects are "the same value" only when they are literally the same object in memory: ```js const a = JSON.parse('{"id":1}'); const b = JSON.parse('{"id":1}'); new Set([a, b]).size; // 2 ``` Every record that comes out of a parse, a mapper, a spread (`{...r}`) or a merge is a new object. So a `Set` of records deduplicates only in the one case that is never true for fetched data: the very same reference appearing twice in the array. The symptom in production is a Set-based dedupe that silently does nothing and a UI showing repeated rows. ## Step one: decide what identity means The real question is not which container to use but **what makes two records the same record**. Usually there is a domain answer — the primary key, the message id, the normalized email — and the dedupe should be written against it explicitly rather than inherited from reference semantics. ### First-wins with a seen-Set ```js function dedupeBy(items, keyOf) { const seen = new Set(); return items.filter(item => { const k = keyOf(item); if (seen.has(k)) return false; seen.add(k); return true; }); } dedupeBy(records, r => r.id); ``` This keeps the **first** record for each key and preserves the original order. It is the right shape when earlier data is more trustworthy, or when you are streaming and cannot look ahead. ### Last-wins with a Map ```js const byId = new Map(records.map(r => [r.id, r])); const unique = [...byId.values()]; ``` Re-setting an existing key overwrites the value while leaving the key's original position, so this keeps the **last** record for each key in first-seen order. That is usually what you want when later payloads are fresher — a paginated feed where a record was updated between pages. Choosing between first-wins and last-wins is a real decision with visible consequences, and an interviewer will ask which one your code implements. Be able to answer without re-reading it. ## When there is no natural key Sometimes the records genuinely have no id and identity is "all the fields". The tempting one-liner is a serialized key: ```js const seen = new Set(records.map(r => JSON.stringify(r))); ``` It works often enough to be dangerous. The failure modes are worth naming: - **Property order is significant.** `{a:1,b:2}` and `{b:2,a:1}` serialize differently and will not dedupe, even though nothing about the data differs. - **Values are lost.** Properties whose value is `undefined`, a function, or a symbol are omitted entirely, so two records that differ only there collapse into one. - **Some values change shape.** `NaN` and `Infinity` become `null`; a `Date` becomes an ISO string; a `Map` or `Set` becomes `{}`. - **Cycles throw.** A self-referential record raises a `TypeError` rather than producing a key. If you must serialize, canonicalize first — pick the identity fields explicitly and build the key from them in a fixed order: ```js const keyOf = r => `${r.tenant}|${r.type}|${String(r.externalId)}`; ``` An explicit template also documents, in one readable line, exactly what your system considers a duplicate. Guard against separator collisions if a field can contain the delimiter. ## Operational considerations A `seen` Set holds every key it has ever been shown for as long as it lives. That is harmless for a one-shot dedupe of a page of results, and it is a slow leak when the same Set spans a long-running stream or a server process handling an unbounded feed. In that shape you need a bound: reset the Set per batch or per window, or cap it and evict. Deciding the window is a design choice — it defines how far apart two occurrences can be and still count as duplicates. Finally, ask whether the dedupe belongs in the client at all. Duplicates arriving from an API are frequently a symptom upstream — a join fanning out rows, pagination overlapping, an at-least-once delivery guarantee. Client-side dedupe is the right patch while you fix the source, but a senior answer names the source as the real fix and treats the Set as the mitigation.

  • Your dedupe keeps the wrong copy when a record was updated between pages — what changed?
    You are running a first-wins strategy. A seen-Set filter keeps the earliest occurrence and discards every later one, so a corrected record arriving on page two is thrown away. Switch to keying a `Map` by the identity and setting each record as you go: the last write for a key wins while the key keeps its original position, so ordering is preserved and the freshest copy survives.
  • Why is JSON.stringify a risky way to build the identity key?
    Because the string depends on more than the data. Property insertion order changes the output, so equal records can produce different keys; `undefined`, function and symbol values are dropped, so different records can produce the same key; `NaN` becomes `null`; and a cyclic object throws. Build the key from the specific identity fields in a fixed order instead, which is both stable and self-documenting.
  • The dedupe runs over a long-lived stream and memory grows — what do you do?
    The seen-Set retains every key it has ever observed, so an unbounded stream means unbounded growth. Give it a window: reset per batch, per time slice, or cap the number of retained keys and evict the oldest. Choosing that window is a product decision — it defines how far apart two occurrences may be and still count as the same event.
  • Is client-side deduplication the right fix at all?
    It is a mitigation, not usually the cure. Repeated records normally mean something upstream — a join fanning out rows, overlapping pagination, or at-least-once delivery. Deduping in the client keeps the UI correct today, but the durable fix is to make the source emit each entity once or to give responses a stable identity and cursor. Say both: patch now, and name the upstream defect.

saying these in an interview costs you the question

  • Says Set compares object contents, so it should have worked
  • Reaches for JSON.stringify keys without naming its pitfalls
  • Cannot say whether the dedupe keeps the first or last copy
  • Ignores unbounded growth of a long-lived seen-Set
  • Treats client-side dedupe as the complete fix

context