In JavaScript, why does ['zebra', 'Apple', 'ähre'].sort() produce ['Apple', 'zebra', 'ähre'], and how do you get human-expected alphabetical order?
answer
- not alphabetical, mechanical
- compared by UTF-16 code unit
- uppercase ASCII sits below lowercase
- accented letters live above z
- use localeCompare for collation
basics
~10 sA comparator-less sort compares strings by UTF-16 code unit, so every uppercase ASCII letter precedes every lowercase one and accented letters land after z. Use a locale-aware comparator instead: arr.sort((a, b) => a.localeCompare(b)).
solid answer
~40 sThe default string ordering is code-unit ordering, not alphabetical ordering. `Array.prototype.sort` with no comparator compares the string forms of the elements position by position using their UTF-16 code units. `'A'` is 0x41 and `'z'` is 0x7A, so every uppercase ASCII letter sorts before every lowercase one; `'ä'` is 0x00E4, above the entire ASCII range, so accented and non-Latin characters land after `'z'`. That gives `['Apple', 'zebra', 'ähre']`. For human-readable order, pass a comparator built on `String.prototype.localeCompare`, which performs locale-aware collation: `arr.sort((a, b) => a.localeCompare(b))` puts `ähre` with the `a`s and no longer segregates capitals. `localeCompare` returns a negative, zero, or positive number — its sign is what matters, the magnitude is unspecified — so it drops straight into a comparator, including as a term in a chained multi-key comparison.
code
javascript · 10 linesconst names = ['zebra', 'Apple', 'ähre', 'apple'];
console.log([...names].sort());
// [ 'Apple', 'apple', 'zebra', 'ähre' ] — code-unit order
console.log([...names].sort((a, b) => a.localeCompare(b)));
// accented and capitalised forms group with their base letter
const collator = new Intl.Collator('de');
console.log([...names].sort(collator.compare)); // reusable, faster for long listsgo deeper
Recognise the symptom — capitalised entries bunched at the top, accented names at the bottom — and know the one-line fix, sorting with a comparator that calls localeCompare.
Explain the mechanism: default comparison is UTF-16 code unit by code unit, so ASCII uppercase occupies a lower range than lowercase and anything non-ASCII is higher than both. Contrast that with collation treating case and accents as secondary differences.
Weigh the choice per use case: collation for anything a user reads, code-unit order for canonical machine output, and a hoisted Intl.Collator when the list is long enough for per-comparison collation cost to matter.
Own the consistency problem: if the server sorts in the database and the client re-sorts in the browser, the two collations will disagree. Decide which side is authoritative and specify the locale explicitly instead of inheriting the environment default.
## What the default actually compares When `sort` gets no comparator, each element is converted to a string and the strings are compared with the same rules the `<` operator uses. That comparison is purely mechanical: walk both strings, and at the first position where the code units differ, the smaller code unit wins. If one string runs out first, it is the smaller one. ```js console.log('A' < 'a'); // true — 0x41 < 0x61 console.log('Z' < 'a'); // true — 0x5A < 0x61 console.log('z' < 'ä'); // true — 0x7A < 0xE4 ``` So the result `['Apple', 'zebra', 'ähre']` follows directly: capital `A` first because uppercase ASCII occupies 0x41–0x5A, then lowercase ASCII 0x61–0x7A, then everything above ASCII. Two consequences bite in real applications: - **Case segregation.** A name list sorts as `Álvarez`? no — plainly: `Adams, Baker, adams, baker`, with every capitalised entry above every lowercase entry, which no user expects. - **Non-ASCII exile.** Accented Latin letters, Greek, Cyrillic, and CJK all sort after the whole ASCII alphabet, in code-point order rather than in any language's alphabet. ## The fix: localeCompare `String.prototype.localeCompare(that)` compares two strings using locale-aware collation rules and returns a negative number, zero, or a positive number. That is precisely the comparator contract, so it plugs in directly: ```js const names = ['zebra', 'Apple', 'ähre']; names.sort((a, b) => a.localeCompare(b)); // ['ähre', 'Apple', 'zebra'] — ä collates with a, case is a minor difference ``` Collation is a linguistic operation, not a numeric one. It treats case and accents as *secondary* differences: `a` and `A` are the same letter that differ only in case, and `ä` is a variant of `a` in many locales, so both sort next to their base letter rather than far away from it. That is the behaviour users mean by "alphabetical". Two details worth stating: - **Only the sign is defined.** The specification does not fix the magnitude of `localeCompare`'s result, so never test `=== -1`; test `< 0`. - **Results are locale-dependent by design.** Different languages genuinely order letters differently, so the same array can legitimately sort differently under different locales. Pass the locale explicitly when the ordering must be predictable rather than dependent on the environment's default. ## Cheap alternatives, and when they are wrong A common shortcut is to normalise case first: ```js names.sort((a, b) => { const x = a.toLowerCase(), y = b.toLowerCase(); return x < y ? -1 : x > y ? 1 : 0; }); ``` This fixes case segregation and nothing else: `ähre` still sorts after `zebra`, because lowercasing does not change the code point of `ä`. It is acceptable only for data you know is ASCII — machine-generated identifiers, slugs, enum names. For anything a human typed, it is the wrong tool. Conversely, code-unit order is the *right* tool when you need a stable, locale-independent, machine ordering: sorting keys for a canonical serialisation, ordering hashes, or building a deterministic index. In those cases the whole point is that the result must not vary with the user's language settings. ## Cost Collation is much more expensive than a code-unit comparison, and the comparator runs on the order of n log n times. For long lists, hoist the collator instead of calling `localeCompare` on each pair: `Intl.Collator` exposes a `compare` method that is bound and reusable, and passing it directly as the comparator avoids re-resolving locale data per comparison. ```js const collator = new Intl.Collator('de'); names.sort(collator.compare); ``` ## Chaining Because `localeCompare` returns a signed number, it composes with other comparison terms exactly like a subtraction does — for instance as the primary term of a multi-key comparator, with a numeric term breaking its ties.
- Can you compare localeCompare's return value against -1 and 1?No. Only the sign is specified — negative, zero, or positive — while the magnitude is left to the implementation. Test `< 0` and `> 0`, or just return the value straight from a comparator, which reads only the sign anyway. Comparing with `=== -1` is a portability bug waiting for a different engine or locale.
- When is the default code-unit ordering the right choice?When you want a deterministic, locale-independent machine order: canonical key ordering for serialisation or hashing, sorting identifiers and slugs, or any output that must be byte-for-byte identical regardless of the user's language settings. Collation is the wrong tool there precisely because it varies by locale.
- Is lowercasing both strings before comparing an adequate fix?Only for ASCII data. It removes the uppercase-before-lowercase split, but `ä` keeps its code point above `z`, so non-ASCII text still sorts after the whole Latin alphabet. It also mishandles languages where case mapping is not one-to-one. For human-entered text, use collation instead.
saying these in an interview costs you the question
- Calling the default sort alphabetical or dictionary order
- Claiming sort lowercases strings before comparing them
- Testing localeCompare's result against exactly -1 or 1
- Assuming toLowerCase fixes accented-character ordering
- Expecting one collation order to satisfy every language