What normalization rules must an anagram check agree on before any code is written?
answer
- Ask what counts as the same before coding
- Same letters, different spacing and capitalisation
- Case, separators, digits, scope of alphabet
- One rule, called by every path
- Changing it invalidates stored canonical keys
basics
~20 sState the canonicalization first: case folding, whether whitespace and punctuation are dropped, whether digits count, and which characters are in scope at all. Two phrases are anagrams only relative to that rule, so leaving it unstated leaves the behaviour undefined.
solid answer
~50 sAnagram equivalence is not an absolute property, it is a property relative to a canonicalization. The textbook pairing of a single word with a two-word phrase only matches if you fold case and drop whitespace; keep either and the answer flips. So before writing the comparison I pin down four things: what case folding applies, whether separators and punctuation are ignored or significant, whether digits and out-of-scope characters are dropped or rejected, and what unit of text counts as "a character". Then I put that rule in **one** function that both the ingest path and the comparison path call, because a rule implemented twice will diverge on the next edit. In a feature that stores canonical forms as keys, the rule is effectively a data format: changing it later invalidates every stored key, so it needs a version or a rebuild plan rather than a quiet edit.
go deeper
Be ready to notice that a word and a two-word phrase can only be anagrams once case is folded and spacing is ignored, and to ask which rule applies before comparing.
Explain the concrete checklist — case, separators, digits, alphabet scope, unit of text — and show how each choice changes the answer for the same pair of inputs.
Demonstrate production judgment: one shared canonicalization called by every path, tests pinned on representative inputs per supported class, and a deliberate drop-versus-reject decision for out-of-scope characters.
Own it as a data contract. Once canonical forms are persisted, changing the rule is a migration with versioned keys or a rebuild pass, and the cost of that migration belongs in the decision to change it.
## Equivalence is relative, and that is the whole question The interview version of this is the pairing everyone has seen: one word and a two-word phrase declared anagrams of each other. They are — *if* case is folded and whitespace is dropped. Under a rule that treats a space as an ordinary character, the two are not anagrams at all, because one contains a space and the other does not. Neither answer is wrong; they answer different questions. That is the point a senior candidate is expected to reach unprompted: "are these anagrams?" is under-specified until the canonicalization is named. In a live feature — say a word game catching shuffled duplicate answers — the canonicalization is not a formatting detail. It is the definition of the product behaviour. ## The checklist to pin down 1. **Case.** Are the inputs folded to a single case before comparison? Folding is not always a simple per-character mapping — some scripts have characters whose folded form has a different length — so "lowercase everything" is a decision to be stated, not an obvious default. 2. **Whitespace and punctuation.** Dropped, or significant? Dropping is what makes multi-word phrases comparable; keeping is what makes the check exact. Pick one and write it down. 3. **Digits and symbols.** Do they participate, or are inputs containing them rejected before the comparison runs? Rejection is often better than silent dropping, because silent dropping makes two visibly different inputs match. 4. **Scope of the alphabet.** Which characters are in bounds at all? This is where a fixed small tally quietly fails: it defines a scope by accident rather than by decision. 5. **The unit being compared.** What counts as "a character" for tallying? Sequences that look the same on screen can be composed from different underlying units. You do not need to solve that here — you need to state which unit your rule counts, so the failure is a known constraint rather than a mystery bug report. ## One rule, one place The most common production failure is not choosing wrongly, it is choosing twice. A feature that stores a canonical form when an answer is accepted, and computes another canonical form when a new answer arrives, has two implementations of the same rule. They agree on the day they are written. Then someone fixes punctuation handling on the ingest side, the comparison side keeps the old behaviour, and the feature starts missing duplicates that its own tests still pass on — because the tests exercise both paths through the same helper the developer happened to use. The fix is structural: the canonicalization is one function, both paths call it, and the tests assert on that function directly with representative inputs from every class the feature claims to support — mixed case, spacing variants, punctuation, digits, accented letters. When the audience widens, those tests fail loudly instead of the feature silently producing false matches. ## The rule is a data format Once canonical forms are persisted — as duplicate-detection keys, as grouping buckets, as an index — the normalization rule has become part of the stored data's contract. Change the rule and every existing key was computed under the old one: new keys stop matching old keys, and duplicate detection silently degrades rather than erroring. The options are the usual ones for a format change: version the key so old and new coexist during a migration, or run a rebuild pass over the stored forms. Either way it is a planned migration, not a one-line edit, and treating it as a one-line edit is exactly the failure this framing is meant to prevent. ## Rejecting versus dropping One judgment call deserves its own mention. When an input contains characters outside the agreed scope, you can drop them, or you can reject the input. Dropping keeps the feature permissive and makes two visibly different submissions compare equal — sometimes desirable, sometimes a bug that reaches users as "why did it say this was a duplicate?". Rejecting keeps the equivalence honest at the cost of turning away input. Choose deliberately, state the choice in the spec, and make the error message say which characters were out of scope. ## What interviewers listen for A candidate who asks the clarifying questions before writing anything; who names case, separators and scope specifically rather than saying "I'd normalize it"; who puts the rule in one place; and who recognises that a stored canonical form makes the rule a migration concern. A candidate who starts coding the comparison without asking has answered a question nobody asked.
- Where should the canonicalization live in the system?In one function that both the ingest path and the comparison path call, so a stored key and a query key are always produced identically. Duplicating the rule at each call site is how a feature passes its tests and misses duplicates in production: someone edits one copy, both remain individually correct, and only the pairing across paths is broken.
- What happens when you change the normalization rule after launch?Every canonical form already stored becomes stale, because it was computed under the old rule and no longer matches newly computed keys. Detection degrades silently rather than failing loudly. Treat it as a data-format change: version the key so both generations coexist during a migration, or run a rebuild pass over the stored forms before switching.
- Should out-of-scope characters be dropped or rejected?Decide deliberately and write it down. Dropping is permissive and makes two visibly different inputs compare equal, which surfaces later as a confusing false duplicate. Rejecting keeps the equivalence honest but turns input away, so the error must say which characters were out of scope. The wrong answer is dropping by accident because the tally simply had nowhere to put them.
Asking whether two parcels weigh the same is meaningless until you agree on the unit and on whether the packaging is included. The comparison is trivial; the agreement is the work.
saying these in an interview costs you the question
- Assumes lowercase letters only and never says so
- Strips separators on one path but not the other
- Calls normalization a formatting detail rather than a contract
- Changes the rule without rebuilding stored canonical keys
- Starts coding the comparison before defining equivalence