Your pretrained word vectors have no entry for misspellings or rare surnames — what fixes it?
answer
- the vocabulary is closed at training time
- one shared placeholder for everything unseen
- build the word out of its pieces
- sum the character n-gram vectors
- inflection-heavy languages gain the most
basics
~20 sSwitch to vectors that represent a word as the sum of its character n-gram vectors. An unseen surname or typo still shares n-grams with trained words, so a vector is composed for it rather than a shared placeholder.
solid answer
~50 sWhole-word vectors have a closed vocabulary: anything unseen at training time collapses to one shared unknown vector, or is dropped. The fix is subword composition, as in fastText — during training, each word is also represented by the set of character n-grams it contains (typically lengths 3 to 6, with word-boundary markers), each n-gram gets its own vector, and the word's vector is the sum of its n-gram vectors plus a whole-word vector when the word was seen. At inference an unseen token is composed from whatever n-grams it does contain, so `Kowalczyk` or `recieve` still lands somewhere sensible. The gain is largest for morphologically rich languages, where a single stem takes dozens of inflected forms that whole-word training treats as unrelated types and splits the evidence across. The cost is real: more parameters, slower lookups, and spurious similarity between orthographically similar but unrelated words.
go deeper
Know that whole-word vectors have a fixed vocabulary and that anything unseen at inference gets one shared unknown vector. Be able to name character n-grams as the way to build a vector for a word that was never trained on.
Explain the composition rule — enumerate the word's character n-grams over a length range, sum their learned vectors, add the whole-word vector when it exists — and say why inflected forms then share statistical evidence through their stem.
Show you would quantify the out-of-vocabulary rate on production traffic and check which tokens it hits before changing the representation, and be able to name the domains where orthographic similarity actively misleads.
Own the trade across the system: extra parameters and slower lookups against coverage on the entities that carry business value, and whether normalisation upstream is a cheaper, more predictable fix than a new representation the team must now maintain.
## The failure mode A whole-word vector model has a fixed vocabulary, usually every type above a frequency cutoff in the training corpus. At inference, a token outside that vocabulary has no row. The usual handling is to map it to a single shared unknown vector or to drop it entirely, and both are bad in the same way: every unseen word becomes the *same* point, so the model cannot distinguish one novel surname from another, or a typo from a foreign word. That matters more than a vocabulary-size argument suggests, because the out-of-vocabulary rate is not evenly distributed. It is concentrated exactly where product value often lives: proper nouns, product codes, user-generated misspellings, newly coined terms, and inflected forms of words whose base form *was* in the corpus. ## Composing from character n-grams The subword approach keeps the same training objective but changes what a word's vector *is*. Wrap the word in boundary markers — `<where>` — and enumerate its character n-grams for a range of lengths, conventionally 3 through 6: `<wh`, `whe`, `her`, `ere`, `re>`, `<whe`, `wher`, and so on. Each distinct n-gram gets its own learned vector. A word's representation is the sum of its n-gram vectors, plus a vector for the whole word itself when that word appeared in training. Two properties follow. First, evidence is *shared*: every occurrence of `walking`, `walked` and `walker` trains the `walk`-family n-grams, so a rare inflected form benefits from its frequent siblings. Second, the representation is *compositional at inference*: an unseen token has no whole-word vector, but its n-grams are almost all known, so summing them produces a specific, non-degenerate vector. `Kowalczyk` lands near other Slavic surnames; `recieve` lands near `receive` because they share most of their n-grams. The boundary markers matter here — they let the model distinguish a prefix or suffix from the same letters occurring mid-word. Because the number of distinct n-grams is large, implementations typically map n-grams into a fixed number of buckets by hashing, accepting some collisions to bound the parameter count. ## Where the gain is largest Morphologically rich languages — Finnish, Turkish, Hungarian, Czech, Arabic — are the strongest case. A single lemma may surface in dozens or hundreds of inflected forms. Whole-word training treats every form as an unrelated type, so the corpus evidence for one concept is shattered across many low-count rows, and each row is poorly estimated. Subword sharing reassembles that evidence through the shared stem n-grams. English, with its comparatively thin inflection, gains much less from morphology and mostly gains on typos and proper nouns. The gain is also largest when the training corpus is small: sharing parameters across surface forms is a form of statistical pooling, and pooling matters most when per-type counts are low. ## The costs and the failure cases *Parameters and speed.* You now store n-gram vectors as well as word vectors, and every lookup for an unseen word requires enumerating and summing its n-grams rather than one array index. Model files are noticeably larger. *Orthographic false friends.* Summing shared substrings means words that look alike get pulled together whether or not they are related. Short words are especially exposed, because a handful of n-grams determines the whole vector. If your domain has meaningful codes that differ in one character — part numbers, ticker symbols, dosage strings — subword composition can actively hurt by making genuinely distinct identifiers near-identical. *It does not fix polysemy.* An unseen word gets a vector, but it is still one static vector per surface form. Composition solves coverage, not sense. ## How to decide Measure before switching. Compute the out-of-vocabulary rate on *real* production text rather than on a held-out slice of the training corpus — production text carries the typos and proper nouns the corpus was cleaned of. If OOV is a fraction of a percent and concentrated on tokens you do not care about, whole-word vectors plus a normalisation step (lowercasing, simple spelling correction, entity typing) may be cheaper and more predictable. If OOV is several percent and lands on the entities driving the task, subword composition is the right structural fix, and you should evaluate it on the downstream metric rather than on vector-similarity spot checks.
- When would character n-gram composition make your representations worse rather than better?When surface similarity does not imply semantic similarity in your domain. Part numbers, ticker symbols, dosage strings and short codes differing by one character share nearly all their n-grams, so composition drags genuinely distinct identifiers together. Short words are the most exposed, since few n-grams determine the whole vector. In those domains, keep identifiers out of the vector path entirely.
- How would you measure whether the out-of-vocabulary problem is worth fixing at all?Compute the OOV token rate on real production traffic, not a held-out slice of the training corpus — production carries the typos and proper nouns the corpus was cleaned of. Then check where those tokens land: OOV on filler words is harmless, OOV on the entities that drive the task is not. Decide on downstream metrics, never on vector-similarity spot checks.
- Does composing words from n-grams also fix a word having several meanings?No. Composition changes how a vector is *built*, not how many there are — you still get exactly one static vector per surface form, so a word with two unrelated senses remains a single compromise point. Coverage and sense are separate problems, and only the coverage one is addressed here.
Whole-word vectors are a phrasebook: a sentence not printed in it is unavailable. Subword vectors are an alphabet plus common syllables — you can still spell out a name you have never met.
saying these in an interview costs you the question
- Suggests just enlarging the vocabulary until nothing is unseen
- Thinks the unknown-word vector is a reasonable fallback
- Claims subword vectors also solve multiple senses per word
- Assumes surface similarity always implies related meaning
- Measures out-of-vocabulary rate on the training corpus itself