skip to content

Why do LLMs tokenize text into subwords instead of whole words or characters?

level: juniorimportance: must knowfreq 72%

answer

  1. fixed id space, unbounded text
  2. one wall at each extreme
  3. dictionary misses the long tail
  4. characters cost sequence length
  5. common whole, rare split

basics

~20 s

Subword tokenization keeps a fixed vocabulary that can still encode any input. Word-level vocabularies fail on words they never saw; character-level ones make sequences several times longer. Subword pieces keep common words whole and split rare ones.

solid answer

~50 s

A tokenizer has to map unlimited text onto a fixed, finite set of integer ids, and the two obvious choices both break. A word-level vocabulary hits an unknown-word wall: misspellings, new product names, identifiers in source code and most inflected forms fall outside the list and collapse into a single unknown token, destroying information. Character-level tokenization has no unknown problem, but it makes sequences roughly four to five times longer, which costs attention compute and burns the context window. Subword algorithms — BPE, WordPiece, Unigram — sit in between: frequent words stay one token, rare words decompose into reusable pieces such as `un` + `forgiv` + `able`, and the base alphabet of characters or raw bytes guarantees that anything can be spelled out. The vocabulary is learned once on a corpus before pretraining and then frozen.

go deeper

for a junior

Be ready to state both failure modes in one breath: word-level vocabularies cannot represent unseen words, character-level ones make sequences far too long. Then say subword pieces keep frequent words whole and split rare ones.

for a middle

Explain the mechanism, not just the motive: a base alphabet plus learned multi-character pieces, chosen by frequency on a tokenizer training corpus, with the vocabulary size fixed before pretraining.

for a senior

Show that you treat the tokenizer as a frozen production dependency. Talk about what it means that token boundaries reflect the tokenizer's corpus, and what breaks when your traffic looks nothing like that corpus.

for a principal

Own the framing that tokenization is a design choice made once with long consequences — it fixes the embedding and output-layer size, the effective context you get per document, and how well the model serves languages and code styles nobody measured at the time.

## The problem a tokenizer solves A transformer does not consume text. It consumes a sequence of integers, each of which indexes a row in an embedding matrix. That matrix has a fixed number of rows, chosen before pretraining starts and unchangeable afterwards without surgery on the model. So tokenization is the question: how do you map an unbounded space of possible strings onto a bounded set of ids, reversibly and cheaply? ## Word-level tokenization and the unknown-word wall The intuitive answer is a dictionary of words. Split on whitespace and punctuation, assign each distinct word an id, keep the most frequent N. This fails in three ways at once. First, coverage: natural text has a long tail. Names, misspellings, hashtags, chemical formulae, filenames and camelCase identifiers appear constantly and were not in the dictionary. Everything outside the list becomes a single unknown token, so `Zyprexa` and `qwertyuiop` become literally the same input, and the model can neither read nor generate them. Second, morphology: `run`, `runs`, `running`, `runner` occupy four unrelated rows with no shared structure, so anything learned about one is not shared with the others. Languages that build words by stacking suffixes make this far worse. Third, size: covering enough of the tail requires hundreds of thousands of entries, and each entry costs a full embedding row. ## Character-level tokenization and the length wall The opposite extreme is one token per character. The vocabulary is tiny, nothing is ever unknown, and morphology is fully exposed. But the sequence gets much longer — English averages roughly four to five characters per common subword token, so the same document becomes several times as many positions. Attention cost grows with sequence length, the usable context shrinks by the same factor, and the model has to spend capacity relearning that `t`,`h`,`e` is a word before it can do anything semantic. Character models exist and work, but they are not the economical choice at frontier scale. ## The subword compromise Subword tokenization takes the useful half of each. The vocabulary is built from a base alphabet — every character in the corpus, or in byte-level variants all 256 possible byte values — plus a set of learned multi-character pieces. Learning is data-driven: the pieces that earn a vocabulary slot are the ones that recur often enough in the training corpus to pay for themselves. The result has three properties interviewers expect you to name: - **Frequent words stay whole.** `the`, `because`, `government` are single tokens, so common text is compact. - **Rare words decompose into reusable pieces.** An unseen word is spelled from pieces the model has seen elsewhere, so some of its structure survives — a prefix like `anti`, a suffix like `ation`. - **Nothing is truly out of vocabulary,** because the base alphabet is always available as a fallback. The named algorithms differ in how they choose which pieces get a slot — byte-pair encoding merges the most frequent adjacent pair repeatedly, WordPiece scores candidate merges by how much they improve corpus likelihood, and Unigram starts from a large candidate set and prunes it — but all three land in the same design space. ## What follows from this Two consequences are worth internalising early. One: the tokenizer is a separate artefact, trained on its own corpus before the model, and frozen for the model's lifetime. You cannot swap it after pretraining, because the ids are baked into the embedding matrix. Two: token boundaries are an accident of corpus statistics, not of meaning. A word that was frequent in the tokenizer's training data is one token; a semantically identical word that was rare is five. That asymmetry propagates into everything downstream — how much text fits in context, how well the model handles a given language or a given code style, and why it sometimes appears blind to the individual letters inside a word it has never had to spell out. ## Getting the framing right in an interview The strongest short answer names both failure modes, not one. Candidates who only say "so we can handle unknown words" are half right and miss why we do not simply use characters; candidates who only say "to keep sequences short" miss the coverage argument. Say both, then name the mechanism — a learned vocabulary of frequent pieces over a guaranteed base alphabet.

  • Is the tokenizer learned during pretraining, or before it?
    Before, and separately. The tokenizer is fit on its own corpus, producing a vocabulary and merge rules; only then does model pretraining start, with the embedding matrix sized to that vocabulary. Because token ids are bound to embedding rows, the tokenizer is effectively frozen for the model's life — changing it means resizing embeddings and continuing pretraining, not a drop-in swap.
  • If subword tokenizers can spell anything, why do some models still define an unknown token?
    Character-based subword vocabularies only guarantee coverage of characters seen during tokenizer training, so a script absent from that corpus still has no representation and needs an unknown token. Byte-level variants remove the need entirely by starting from all 256 byte values. Many vocabularies keep an unknown id defined anyway, as a reserved slot that is never emitted in practice.
  • What does subword splitting give the model that a word-level vocabulary cannot?
    Shared structure. If `nationalize` and `nationalization` share the pieces `national` and `iz`, representations learned for one transfer to the other, and a form never seen in training is still assembled from familiar parts. A word-level vocabulary treats every inflection as an unrelated symbol, so morphology has to be memorised case by case, and unseen forms are lost entirely.

saying these in an interview costs you the question

  • Says tokens are always whole words or always four characters
  • Claims the tokenizer is a dictionary of real English words
  • Argues character-level is strictly better because it has no unknowns
  • Thinks the model learns its tokenization during pretraining
  • Believes token boundaries follow meaning rather than corpus frequency

context