How do you test locale-sensitive sorting and matching in a customer or account list?
answer
- Character codes are not alphabetical order
- Rules differ between locales
- Same data, two layers, two orders
- Identical-looking names can compare unequal
- Golden order confirmed by a native reviewer
basics
~20 sBuild a golden ordered fixture per locale containing accents, case variation, a locale-specific digraph and punctuation, and assert the full order. Then assert every layer uses the same comparator, and that composed and decomposed spellings compare equal.
solid answer
~40 sSorting text by raw character codes gives a stable but meaningless order, so real products compare through locale-sensitive collation, whose rules differ per locale: an accented letter may be a variant of its base letter or a separate letter of the alphabet, a two-character sequence may sort as one unit, case handling and ignorable punctuation vary, and comparison strength decides whether accents and case count at all. Test with a golden ordered list per locale whose expected order a native reviewer confirmed once. Then test consistency, which is where the expensive defects live: if the query layer, the merge step and the display layer use different comparators, cursor paging duplicates and drops rows. Also test equality, because a name stored composed and another decomposed look identical but compare unequal, which silently defeats duplicate detection.
code
pseudocode · 8 lines// inconsistent: the store orders one way, the merge re-sorts another
page = store.query(orderBy = "name", collation = storeDefault, after = cursor)
merged = sortInMemory(page + cached, comparator = localeComparator(userLocale))
// consistent: declare the order once and pass it everywhere
order = localeComparator(account.locale, strength = ACCENT_AND_CASE)
page = store.query(orderBy = "name", collation = order, after = cursor)
merged = sortInMemory(page + cached, comparator = order)go deeper
Know that sorting by character codes is not the same as alphabetical order, and that the correct order depends on the locale. Recognising accents and case as the usual first surprises is enough here.
Explain collation rules and comparison strength, and describe a golden ordered fixture that deliberately mixes accents, case, punctuation and digits. Be ready to say who confirms the expected order.
Demonstrate the consistency angle: the same comparator across query, merge and display, a paging invariant over a few thousand rows, normalization before equality checks, and one automated build running under an unfamiliar default locale.
Own it as an interface contract: one declared collation per surface, agreed comparison strength for search versus display, and normalization at the storage boundary so duplicate detection cannot be defeated by two spellings of the same name.
### Ordering is a locale rule, not arithmetic Sorting text feels like a solved problem until the data leaves one locale. Comparing strings by their raw character codes gives a stable, meaningless order: uppercase before lowercase, accented letters exiled after the unaccented alphabet, and anything outside the source alphabet grouped by its position in the character set rather than by how a reader would look it up. **Collation** is the locale-sensitive comparison that produces a human ordering instead. Its rules differ per locale in ways that are invisible until you look: - some locales treat an accented letter as a variant of the base letter, others as a distinct letter with its own position in the alphabet; - some treat a two-character sequence as a single letter that sorts as a unit; - some order case-sensitively, some ignore case entirely, and which of upper or lower comes first is itself locale-defined; - punctuation, spaces and hyphens may be ignorable at the primary comparison level, so two names that differ only by a hyphen tie; - digits inside text may or may not compare numerically. Collation also has *strength*: whether a comparison distinguishes base letters only, base letters plus accents, or base letters plus accents plus case. Search, sort, grouping and duplicate detection often want different strengths, and mixing them by accident is where the defects come from. ### The second half of the problem: equality Two strings can look identical, print identically and still compare unequal, because the same character can be stored either as one composed character or as a base character followed by a combining mark. Normalizing to one form before comparing or storing is the fix; skipping it produces a duplicate-detection hole. In a utility billing system that hole has a visible consequence: a customer signs up twice, the duplicate check compares the two spellings as unequal, and the month ends with a duplicated side effect — two accounts, two meters attached to one supply point and two bills. The uniqueness constraint held; it was comparing the wrong thing. Case-insensitive matching has the mirror problem: the default case mapping of one region maps a letter to a different result than another region's rules, so a login or a duplicate check that lowercases before comparing can decide two identical inputs differ, or that two different inputs match. ### The consistency defect The subtle and expensive failure is not "the order is wrong" but "the order is different in two places". A store query orders by its own configured collation; an application merge step re-sorts with a different comparator; a client re-sorts again for display. With a stable order, cursor-based paging works. With two disagreeing orders, rows appear on two consecutive pages or on none, and an export drawn page by page ends up with duplicates and holes that look like a data-integrity bug rather than an ordering bug. ### How to test it - **Golden ordered lists per locale.** Build a small fixture of names deliberately containing accents, case variation, a locale-specific digraph, a hyphen, a space, a leading article and a digit; assert the full expected order. The expected order is an oracle a native reviewer has to confirm once — do not invent it from intuition. - **Assert one comparator.** A test that the query layer, the merge step and the display layer produce the same order over the same fixture catches the consistency defect that no single-layer test can. - **Paging consistency.** Walk a fixture of a few thousand rows page by page, collect the identifiers, and assert the set is exactly the source set with no duplicates and no omissions. - **Normalization round-trip.** Store the composed and decomposed spellings of the same name and assert the duplicate check treats them as one. - **Case-mapping trap.** Include a name whose case conversion differs between regions, and run the match under a non-default process locale. - **Search versus sort strength.** Assert explicitly that search ignores accents if that is the intent, and that the displayed sort does not. ### What to say about defaults The honest senior answer is that the defaults are the trap: an ambient locale inherited from the machine makes every one of these tests pass on the developer's laptop and fail on a server configured differently. Pin the locale and the collation explicitly at every comparison site, and run at least one automated build with a deliberately unfamiliar default locale so an accidental dependency on the environment fails there rather than in production.
- How does an ordering difference between layers show up as an apparent data-integrity bug?Cursor-based paging assumes one total order. If the query layer orders by one collation and a later merge or display step re-sorts by another, a row can land after the cursor on one page and before it on the next, so it appears twice or not at all. An export drawn page by page then contains duplicates and holes, and the first hypothesis is always corrupted data rather than a comparator mismatch.
- Two customer names look identical on screen but the duplicate check treats them as different. What do you test?Normalization. The same character can be stored as one composed character or as a base character plus a combining mark, and byte comparison calls those unequal. Assert that both spellings of the same name round-trip to one canonical form before storage and comparison, and that the duplicate check catches them, since the consequence here is a second account and a second bill for one supply point.
- Should search and the displayed sort use the same comparison strength?Usually not, and the decision should be explicit. Search normally wants a weaker strength so a query without accents or with different case still matches, while a displayed list usually wants accents and case to be significant so the order is stable and predictable. Write an assertion for each intent separately, otherwise a later change to one silently alters the other.
Alphabetical order is a local custom, not arithmetic. Two libraries with identical books can shelve them in genuinely different orders and both be right.
saying these in an interview costs you the question
- Assumes character-code order is alphabetical order
- Tests sorting only with plain unaccented source text
- Ignores that each layer may sort differently
- Compares raw bytes for duplicate detection
- Relies on the machine's default locale in tests
- Invents the expected order without a native reviewer