A list of user-visible names sorted with names.sort() puts "Zebra" before "apple" and orders accented words oddly. Why does JavaScript's default string ordering do that, and what does Intl.Collator do differently?
answer
- code-unit order, not alphabetical order
- capitals before lowercase, accents after z
- alphabets disagree across locales
- one collator, reused as the comparator
- numeric option for embedded digits
basics
~20 sThe default comparison orders strings by UTF-16 code unit, so every uppercase letter precedes every lowercase one and accented letters land after "z". Intl.Collator compares by locale collation rules instead, which is what a human reader expects.
solid answer
~40 s`Array.prototype.sort` with no comparator converts elements to strings and compares them by UTF-16 code unit. That is a numeric ordering of characters, not an alphabetical one: `'Z'` is 90 and `'a'` is 97, so all capitals sort first, and `'é'` at U+00E9 sorts after every ASCII letter. Locale rules differ further — Swedish puts å, ä and ö *after* z, while German phonebook collation treats ä like "ae". `new Intl.Collator(locale)` encapsulates those rules, and its `compare` property is already a bound function you can hand straight to `sort`. Useful options: `sensitivity` (`'base'` ignores case and accents), `numeric: true` so "item2" precedes "item10", `ignorePunctuation`, and `usage: 'search'` for matching rather than ordering. Build one Collator and reuse it — that is markedly faster than calling `localeCompare` inside the comparator.
code
javascript · 8 linesconst names = ['Zebra', 'apple', 'résumé', 'resume'];
console.log([...names].sort());
// ['Zebra', 'apple', 'resume', 'résumé'] - UTF-16 code unit order
const collator = new Intl.Collator('en');
console.log([...names].sort(collator.compare));
// ['apple', 'resume', 'résumé', 'Zebra'] - locale collationgo deeper
Know that a bare sort() compares strings by character code, which is why capitals come first, and that locale-aware ordering needs Intl.Collator or localeCompare. Be able to write arr.sort(new Intl.Collator('en').compare).
Explain the mechanism: UTF-16 code-unit comparison versus locale collation data, with a concrete disagreement such as Swedish placing ä after z. Name sensitivity and numeric and say why one reused Collator beats per-comparison localeCompare.
Demonstrate the systems view: which layer owns the ordering when a database and a client both sort, why paginated results shuffle when two collations disagree, and why collation is a display notion that must not become a storage key.
Own the decision of where canonical ordering lives across services, how locale-specific ordering interacts with pagination and caching, and what guarantees you are willing to make when collation data changes underneath you with a runtime upgrade.
## Why the default order looks wrong With no comparator, sorting coerces each element to a string and orders by UTF-16 code unit — effectively by character number. ```js ['Zebra', 'apple', 'résumé', 'resume'].sort(); // ['Zebra', 'apple', 'resume', 'résumé'] ``` `'Z'` is U+005A (90), `'a'` is U+0061 (97), so every capital precedes every lowercase letter. `'é'` is U+00E9 (233), above all ASCII, so `résumé` lands after `resume` — and, in a longer list, after words starting with `z`. Nothing here is a bug; it is a defined ordering that simply is not the ordering a human expects to read. ## Collation is locale data "Alphabetical" is not one order. A few examples of genuine disagreement: - In **Swedish**, å, ä and ö are distinct letters that come *after* z. - In **German phonebook** collation (`de-DE-u-co-phonebk`), ä sorts as though it were "ae". - Case, accent and punctuation can each be significant or ignorable depending on what you are doing. `Intl.Collator` is the interface to that data. ```js const collator = new Intl.Collator('en'); ['Zebra', 'apple', 'résumé', 'resume'].sort(collator.compare); // ['apple', 'resume', 'résumé', 'Zebra'] ``` ## compare is a bound function `collator.compare` is an accessor that returns a function already bound to its Collator, which is why `sort(collator.compare)` works without a wrapper — the usual `this`-loss trap does not apply. It returns a negative number, zero, or a positive number, exactly the contract `sort` expects. ## The options that matter - **`sensitivity`** — `'base'` treats case and accent differences as equal, `'accent'` distinguishes accents but not case, `'case'` the reverse, and `'variant'` (the default for sorting) distinguishes everything. - **`numeric: true`** — compares embedded digit runs as numbers, so `item2` precedes `item10` instead of following it. - **`ignorePunctuation: true`** — drops punctuation from the comparison, useful for title lists. - **`caseFirst`** — `'upper'` or `'lower'`, for locales where you want a deterministic case order. - **`usage: 'search'`** — tunes the collator for matching rather than ordering; combined with `sensitivity: 'base'` it gives accent- and case-insensitive comparison for a filter box. ```js const loose = new Intl.Collator('en', { usage: 'search', sensitivity: 'base' }); loose.compare('resume', 'RÉSUMÉ') === 0; // true — a match for filtering ``` ## Collator versus localeCompare `'a'.localeCompare('b', locales, options)` is defined as constructing an `Intl.Collator` and comparing. For a single comparison it reads well. Inside a comparator it is a trap: sorting *n* elements performs on the order of *n* log *n* comparisons, and the naive form re-derives the collator each time. Building one `Intl.Collator` and reusing `compare` moves the setup cost outside the loop — the documented reason the constructor exists as a separate object. ## Sorting for display versus keying for storage A collation result is a locale-specific opinion, so it is not a stable key. Two consequences in real systems: 1. If a list is paginated by a database that sorts with its own collation, sorting the page again in JavaScript with a different collator produces visibly inconsistent order across pages. Pick one authority. 2. Never persist "sorted position" derived from a collator — locale data is updated with the runtime, and a different user's locale gives a different order anyway. For deduplication or identity, do not use a case- and accent-insensitive collator as an equality test on stored keys; that is a display-level notion of sameness, and using it as a uniqueness rule silently merges distinct records. ## Ties A collator can report two different strings as equal — that is the whole point of `sensitivity: 'base'`. Sorting is stable, so equal elements keep their input order, which is usually what you want but means the visible result depends on the order the data arrived in. If you need a deterministic tiebreak, compare a secondary field explicitly.
- Why is `collator.compare` safe to pass directly to sort, when passing a method usually loses `this`?`compare` is an accessor property that returns a function already bound to its Collator instance. The specification defines it that way precisely so it can be handed to a higher-order function without a wrapper or an explicit `bind`. Reading it repeatedly returns the same bound function for a given Collator, so caching the collator caches the comparator too.
- When would you use a collator with usage: 'search' rather than for sorting?For matching, not ordering — a filter box or a lookup where the user typing "resume" should match "résumé". Combined with `sensitivity: 'base'` the collator reports a compare result of 0 for strings that differ only by case or accent. Treat that as a display-level match; it is not an identity rule for stored keys.
- Your backend paginates a sorted list and the frontend re-sorts each page with Intl.Collator. What goes wrong?Two collations disagree, so page boundaries stop lining up: an item that the database placed on page 2 may sort above items on page 1 once the client reorders. The list looks shuffled to the user even though each page is internally sorted. Pick one authority for order — sort server-side and render in the received order, or fetch the full set and sort once on the client.
saying these in an interview costs you the question
- Thinking the default sort is alphabetical for any language
- Calling localeCompare inside the comparator for a large array
- Assuming one alphabet order works for every locale
- Using an accent-insensitive collator as a uniqueness check
- Believing uppercase-first ordering is a sort() bug