An English-trained tokenizer shatters Finnish words into many pieces — why, and what fixes it?
answer
- the code is fitted to a corpus
- slots went where the frequency was
- several morphemes, each combination rare
- fertility rises, effective context falls
- after pretraining, ids are welded to weights
basics
~20 sThe merge table encodes whatever recurred in the tokenizer's training corpus. Trained mostly on English, it holds English word pieces, so Finnish morphology falls back to short fragments and bytes. Fixes are upstream: rebalance the tokenizer corpus and spend vocabulary slots on the target languages.
solid answer
~50 sNothing about Finnish is unencodable — the cause is vocabulary allocation. Subword merges are learned by frequency on a corpus, so if that corpus is overwhelmingly English, the slots go to English stems and suffixes. An agglutinative language stacks several morphemes into one long word, and those combinations were rare in the training text, so they get no slots and are spelled out in short fragments. The result is high fertility: many tokens per word, which means the same content consumes several times more context and compute, and the model sees each Finnish word as an unstable pile of pieces rather than a stable unit. The fix is at tokenizer training time — upsample the target languages, size the vocabulary to cover them, and measure tokens-per-word per language as an acceptance criterion. After pretraining the cheap fix is gone: you cannot swap the tokenizer, only extend the vocabulary with new tokens and continue pretraining so the new embedding rows are learned.
go deeper
Know that a tokenizer's pieces come from its training corpus, so a mostly-English tokenizer has mostly English pieces and other languages break into more tokens.
Explain agglutination against a frequency-fitted merge table: each stacked morpheme combination is individually rare, so it earns no slot and is spelled out in fragments.
Show the operational picture — measure tokens per word per language, know that the remedy is corpus rebalancing and vocabulary sizing before pretraining, and that afterwards it costs vocabulary extension plus continued pretraining.
Own the allocation decision as a product commitment: vocabulary slots are a fixed budget spent once across every language and code style you intend to serve, and revisiting it later is a training run with regression risk, not a config change.
## The mechanism, not the language It is tempting to say agglutinative languages are "hard to tokenize". They are not. A merge table is a compression code fitted to a corpus, and a code fitted to English is simply a bad code for Finnish — the same way a Huffman table built from English letter frequencies compresses Finnish poorly. Every merge that exists is one that paid for itself in the training text. If Finnish, Turkish, Hungarian or Tamil made up a fraction of a percent of that text, their recurring morphemes never accumulated enough count to beat English candidates, and the vocabulary reflects that. Agglutination amplifies the effect. A Finnish or Turkish word may carry case, number, possession and several derivational suffixes in one orthographic word. In English those would be separate whitespace-delimited words, each with its own chance to earn a token. Stacked into one word, each *combination* is individually rare, so none of them earns a slot even if each morpheme is common. ## What the symptom looks like The observable is **fertility**: tokens produced per word. A well-served language sits near one to one-and-a-half tokens per word; a badly served one can run to five or eight. Three things degrade together: - **Context.** The same document occupies several times more positions, so less of it fits in the window and long-document tasks fail earlier. - **Compute and latency.** More tokens in and more tokens out, for identical content. - **Quality.** The model must compose meaning from unstable fragments whose boundaries shift with inflection, so the same stem is represented differently in different forms. Downstream, the language's tokens are also individually rarer during pretraining, so they receive fewer gradient updates. This is why the same model can look competent in one language and clumsy in another even when both were present in the pretraining data. ## Fixing it before pretraining At tokenizer-training time the levers are direct and cheap: 1. **Rebalance the tokenizer corpus.** The tokenizer is fit on a sample, and that sample does not have to mirror the pretraining mix. Upsampling target languages buys them merges. A common approach is sampling languages with a temperature that flattens the distribution rather than reproducing raw web proportions. 2. **Size the vocabulary for the language count.** Serving many scripts well from 32K slots is arithmetically hard; larger vocabularies exist substantially to buy coverage across languages and code. This trades against embedding and output-layer cost. 3. **Keep byte-level fallback.** It guarantees nothing is lost while you are still under-serving a language, converting a correctness problem into a cost problem. 4. **Measure per language and gate on it.** Tokens-per-word or tokens-per-character by language, computed on held-out text, is a cheap acceptance test and the thing nobody runs until a user complains. ## Fixing it after pretraining Once the model is trained, the tokenizer is effectively frozen: ids index embedding rows, so a different tokenizer means the ids mean different things and the model is destroyed. Dropping in a replacement is not an option, and any candidate who proposes it has revealed they do not know how ids bind to weights. What is actually done is **vocabulary extension with continued pretraining**: add new tokens for the target language's frequent sequences, grow the embedding and output matrices, initialise the new rows sensibly — a common choice is the mean of the embeddings of the pieces the new token replaces, which starts it in a plausible region rather than at random — and then continue pretraining on a mixture heavy in the target language until the new rows are trained and the old behaviour has not regressed. This is a real training run with real cost and real regression risk, which is exactly why the allocation decision deserves attention before pretraining rather than after. ## The judgement an interviewer is testing Three things separate a strong answer. First, correctly locating the cause in corpus statistics and vocabulary allocation rather than in the language. Second, knowing that the consequences are simultaneously about cost, effective context and quality, not just cost. Third, knowing that the remedy is cheap before pretraining and expensive after, and being able to describe the extension-plus-continued-pretraining path without pretending it is free.
- Why can't you just replace the tokenizer on an already-pretrained model?Because token ids are not names, they are indices into the embedding and output matrices. A new tokenizer assigns different ids to different pieces, so every lookup returns the wrong vector and the model's learned behaviour is gone. The only coherent paths are keeping the tokenizer, or extending the vocabulary — adding rows and continuing pretraining so the new entries are actually learned.
- How would you initialise the embedding rows for newly added tokens?Not randomly, if you can avoid it. A standard heuristic is to initialise each new token's embedding as the mean of the embeddings of the existing pieces it replaces, so the new row starts in a semantically plausible region and continued pretraining refines rather than discovers it. Random initialisation works but wastes compute and risks disturbing the existing representation space early in the run.
- Does fixing fertility for one language cost anything for the others?Yes, unless you also grow the vocabulary. Slots are a fixed budget: upsampling one language in the tokenizer corpus means its merges displace others, so previously well-served languages fragment slightly more. Growing the vocabulary avoids the zero-sum trade but pays in embedding and output-layer parameters, which is the underlying reason multilingual models tend to carry larger vocabularies.
- Is high fertility a quality problem or only a cost problem?Both, and the quality half is the one people miss. Beyond consuming more context and compute, fragmenting a language means the same stem is represented by different piece sequences across its inflections, so structure has to be reassembled rather than looked up, and each of those pieces sees fewer gradient updates during pretraining. Cost is measurable immediately; the quality gap shows up as the model simply being weaker in that language.
saying these in an interview costs you the question
- Blames the language for being hard to tokenize
- Proposes swapping in a new tokenizer on a pretrained model
- Thinks byte-level fallback already solves the fragmentation
- Treats high fertility as purely a billing issue
- Assumes more pretraining data in the language fixes tokenizer allocation