skip to content

Set Semantics and Uniqueness

Set stores unique values under SameValueZero equality, which means NaN deduplicates but two structurally identical objects do not. The one-liner [...new Set(arr)] is a common interview warm-up, and the follow-up is always about what "unique" actually means here.

part ofJavaScriptoverview, primer and where to startread it →
on this pageshow

questions

4

In JavaScript, what does the expression [...new Set(myArray)] produce, and what are the limits of it as a deduplication idiom?

level: juniorimportance: must knowfreq 72%

answer

  1. constructor takes any iterable
  2. one copy per value
  3. spread reads it back out
  4. first-occurrence order, never sorted
  5. objects compared by reference

basics

~20 s

Spreading an array into a Set and back yields a new array with duplicates removed and first-occurrence order preserved, because a Set stores each value only once. It collapses only values the Set treats as equal; distinct objects with identical contents survive.

solid answer

~40 s

`new Set(myArray)` builds a Set from any iterable, and a Set keeps at most one copy of each value, so duplicates are dropped as it is filled. Spreading it back with `[...set]` (or `Array.from(set)`) gives a plain array again, and because a Set iterates in insertion order the survivors appear in first-occurrence order — `[...new Set([3,1,3,2])]` is `[3,1,2]`, not sorted. The idiom is exact for strings, numbers, booleans, `null`, `undefined` and symbols, and it even collapses `NaN`, which `indexOf` cannot find. Its limit is object elements: a Set compares objects by reference, so two separately-created objects with identical properties are two distinct entries and nothing is removed. It also returns a new array rather than mutating the original.

code

javascript · 10 lines
javascript
const arr = [3, 1, 3, 'a', 'a', NaN, NaN, 1];
console.log([...new Set(arr)]); // [3, 1, 'a', NaN]

// Objects are not deduped by content:
const users = [{ id: 1 }, { id: 1 }];
console.log([...new Set(users)].length); // 2

// The same reference twice is deduped:
const u = { id: 1 };
console.log([...new Set([u, u])].length); // 1

go deeper

for a junior

Be able to write [...new Set(arr)] from memory, say that it returns a new array with duplicates removed, and state plainly that the order is first-occurrence order, not sorted.

for a middle

Explain the mechanics: the constructor consumes any iterable and calls add, add ignores values already present, and spreading reads the Set's iterator back out in insertion order. Name the object-reference limit.

for a senior

Show the judgment about when the idiom is the wrong tool — deduping records by a business key, needing last-wins, or a long-lived seen-Set that grows without bound — and give the alternative you would ship instead.

for a principal

Frame duplicate-removal as a question of what identity means for the data: who defines the key, whether dedupe belongs in the client or at the source, and what it costs when the definition of "same record" later changes.

## The idiom, piece by piece `[...new Set(myArray)]` is three separate language features cooperating. 1. **`new Set(iterable)`** — the `Set` constructor accepts any iterable and calls `add` for every element it yields. `add` is a no-op when the value is already present, so the resulting Set holds each distinct value exactly once. 2. **A Set is itself iterable** — it has a `Symbol.iterator` that yields its values, so it can be spread. 3. **Array spread `[...]`** — drains that iterator into a fresh array. ```js const arr = [3, 1, 3, 2, 1]; const set = new Set(arr); // Set(3) { 3, 1, 2 } const out = [...set]; // [3, 1, 2] ``` `Array.from(new Set(arr))` is the identical operation with a different spelling, and it is the one to reach for when you also want a mapping function: `Array.from(new Set(arr), n => n * 2)`. ## Order is insertion order, not sorted order A `Set` iterates in the order values were first inserted. That makes the idiom **stable and first-wins**: the earliest occurrence of each value determines its position, and later duplicates are silently dropped. This surprises people who expect a sorted result — `[...new Set([3,1,2])]` is `[3,1,2]`. If you want sorted output you must sort afterwards, and if you want *last*-wins you cannot use this idiom directly, because a Set has no way to "re-add" a value at the end (`add` on an existing value leaves its position alone). ## Which values actually collapse A Set decides membership with SameValueZero, which behaves like `===` with one deliberate change: `NaN` is equal to itself. That makes the idiom stronger than the hand-rolled alternative: ```js [...new Set([NaN, NaN])]; // [NaN] — deduped [NaN].indexOf(NaN); // -1 — indexOf uses === [NaN].includes(NaN); // true — includes uses SameValueZero ``` `+0` and `-0` are also treated as the same value, so `[...new Set([0, -0])]` has length 1. Different types never collapse: `1` and `'1'` are two values, because no coercion happens anywhere in a Set. ## The hard limit: objects are compared by reference ```js const users = [{ id: 1 }, { id: 1 }]; [...new Set(users)].length; // 2 — nothing removed ``` Each object literal creates a distinct object, and the Set sees two distinct references. The idiom removes duplicates only when the *same* object appears more than once in the array: ```js const u = { id: 1 }; [...new Set([u, u])].length; // 1 ``` Deduplicating objects by content therefore needs a derived primitive key rather than the objects themselves. ## Other things worth knowing before you use it - **Any iterable works, not just arrays.** `new Set('aabbc')` iterates the string and gives `Set { 'a', 'b', 'c' }`, because strings are iterable. - **It is a copy.** The original array is untouched; the elements themselves are not cloned, so the new array holds the same object references. - **Sparse arrays are densified.** Array iteration yields `undefined` for holes, so `[...new Set([1, , 1])]` is `[1, undefined]`. - **A Set holds strong references** to whatever you put in it, so a long-lived Set used as a "seen" ledger keeps its entries alive; bound it or clear it if it grows without limit. - **`add` returns the Set itself**, which is why `set.add(1).add(2)` chains; `delete` returns a boolean saying whether something was removed; `size` is a property, not a method (`set.size`, never `set.size()`). ## When to reach for something else Use a plain filter when the notion of duplicate is not raw value equality — for example, "same `id`" or "same normalized email". Use a `Map` when you want to keep one representative object per key rather than a bare list of keys. Keep the Set idiom for the case it is genuinely perfect at: collapsing a list of primitives while preserving the order in which they first appeared.

  • Does the resulting array keep the original order, and which occurrence of a duplicate wins?
    A Set iterates in insertion order, so the result follows the array's original order and the **first** occurrence of each value keeps its position; later duplicates are dropped without moving anything. It is not sorted — `[...new Set([3,1,2])]` stays `[3,1,2]`. If you need last-wins, build a `Map` keyed by the value instead, since re-setting a key overwrites without reordering.
  • How would you deduplicate an array of numbers without Set, and what breaks?
    The classic is `arr.filter((n, i) => arr.indexOf(n) === i)`, which keeps an element only at its first index. It works for ordinary numbers and strings, but `indexOf` uses `===`, so `NaN` is never found and every `NaN` survives. Switching to `findIndex` with `Object.is` fixes `NaN` but then treats `-0` and `0` as different, which the Set idiom does not.
  • Is `Array.from(new Set(arr))` different from `[...new Set(arr)]`?
    Functionally no — both drain the Set's iterator into a new array in the same order. `Array.from` additionally accepts a mapping function as its second argument, so `Array.from(new Set(arr), s => s.trim())` dedupes and transforms in one pass. Spread is shorter; `Array.from` is the one to use when you need that mapper or an explicitly array-like source.

saying these in an interview costs you the question

  • Claims the result comes back sorted
  • Says it deduplicates objects with equal properties
  • Thinks it mutates the original array in place
  • Writes set.size() instead of the size property
  • Assumes NaN cannot be deduplicated by anything

context

open as a page

What equality rule does a JavaScript Set use to decide whether a value is already present, and what surprising results does it produce for NaN, -0, and objects?

level: middleimportance: must knowfreq 60%

basics

~20 s

A Set uses SameValueZero: like === except NaN counts as equal to itself, and +0 equals -0. So NaN deduplicates, 0 and -0 collapse into one entry stored as +0, and objects are matched by reference, never by their contents.

open as a page

A service fetches a list of records and deduplicates it with new Set(records), but duplicates still get through. Explain why, and how you would deduplicate by content instead.

level: seniorimportance: should knowfreq 42%

basics

~20 s

A Set matches objects by reference, and every parsed record is a fresh object, so nothing collapses. Deduplicate on a derived primitive identity instead: keep a Set of ids, or build a Map keyed by that identity and take its values.

open as a page

How do you compute the union, intersection, and difference of two JavaScript Sets, both with the built-in Set methods and by hand?

level: middleimportance: nice to knowfreq 34%

basics

~10 s

Modern engines provide Set.prototype.union, intersection, difference and symmetricDifference, which return a new Set without mutating either operand. Before those existed the idioms were new Set([...a, ...b]) for union and [...a].filter(x => b.has(x)) for intersection.

open as a page