skip to content

How would you choose a model's vocabulary size — 32K, 128K, or 256K?

level: principalimportance: should knowfreq 34%

answer

  1. two cost sides, one budget
  2. embedding and softmax scale with V
  3. shorter sequences, fewer decode steps
  4. coverage across scripts drives it up
  5. leave slots you have not used yet

basics

~20 s

Vocabulary size trades sequence length against parameter and compute cost. Larger vocabularies encode the same text in fewer tokens but enlarge the embedding and output layers and leave rare tokens under-trained. Pick it from the languages, scripts and code you must serve, then leave reserved slots.

solid answer

~50 s

There is no universally right number; it is an allocation decision driven by what you must encode. A larger vocabulary shortens sequences — fewer tokens per document means less attention work, more content per context window and fewer decode steps — but every extra slot adds a row to the embedding matrix and to the output projection, each of size d_model, and the output softmax runs over the whole vocabulary at every step. Broad multilingual and code coverage pushes the number up, because slots are a fixed budget shared across every script you care about; a narrow English-plus-code model can be well served by far fewer. The counter-pressure is that very rare tokens get few gradient updates and can behave erratically. As of mid-2026, open-weight models commonly sit between roughly 32K and 256K, with multilingual models at the high end. Whatever the number, reserve a block of unused special-token slots so future control tokens do not require resizing the model.

code

python · 9 lines
python
d_model = 4096

for vocab in (32_000, 128_000, 256_000):
    per_matrix = vocab * d_model
    print(
        f"vocab={vocab:>7}  "
        f"tied={per_matrix/1e6:6.0f}M params  "
        f"untied={2*per_matrix/1e6:6.0f}M params"
    )

go deeper

for a junior

Know the direction of the trade: a bigger vocabulary means fewer tokens per document but a bigger embedding and output layer. That framing alone is a solid junior answer.

for a middle

Be able to compute the cost — vocabulary times model width per matrix, doubled if untied — and explain why the output softmax over the whole vocabulary makes this more than a memory question.

for a senior

Bring the operational pressures: script and code coverage as the dominant driver, the rare-token tail as a real failure mode, and vocabulary size as a shape decision for your actual serving workload.

for a principal

Own it as a one-shot commitment across every language, script and code style the product must serve, made before pretraining and expensive to revisit; insist on reserved special-token slots and on tokenizer-versus-pretraining corpus alignment as release gates.

## What actually changes with vocabulary size Vocabulary size V appears in exactly two places, and understanding both is the whole question. **The parameters.** The input embedding is a V x d_model matrix; the output projection that produces logits is another V x d_model (often tied to the input embedding to halve the cost). At d_model = 4096, going from 32K to 256K adds roughly 0.9 billion parameters per matrix. On a very large model that is a rounding error; on a small model it can be a substantial share of total parameters, spent on lookup tables rather than on layers that compute. Small models are therefore much more sensitive to this choice than large ones. **The sequence length.** A bigger vocabulary compresses text into fewer tokens. That shortens every sequence, which reduces attention work, fits more content into a fixed context window and — importantly for latency — reduces the number of sequential decode steps needed to produce a given amount of text. Decode steps are serial, so this is a real user-visible win, not just a FLOP count. The two effects pull in opposite directions, and the curve of downstream quality against V is broad and flat in the middle. That is why reasonable teams land on quite different numbers. ## The pressures that actually decide it **Language and script coverage.** This dominates. Slots are one budget spent across every script you intend to serve. A model targeting English and a few European languages can afford a modest vocabulary; one targeting dozens of languages including CJK and Indic scripts cannot, because each script needs enough merges to avoid encoding near byte-by-byte. Multilingual models carry large vocabularies for exactly this reason. **Code.** Source code has its own high-frequency sequences — indentation runs, common identifiers, operators, bracket pairs — and they compete with natural language for slots. A model expected to be strong on code needs some of the budget deliberately spent there. **Model size.** Judge the embedding and output cost as a share of total parameters. At a few hundred million parameters a 256K vocabulary is disproportionate; at hundreds of billions it is negligible. **Serving shape.** If your workload is dominated by long inputs, shorter sequences pay off through prefill compute and context headroom. If it is dominated by long generations, they pay off through fewer serial decode steps. ## The rare-token failure mode A large vocabulary has a long tail of tokens that barely occur in the pretraining corpus — sometimes because they were merged from a corpus slice that was later filtered out. Those tokens' embedding rows receive almost no gradient updates and remain close to their initialisation. Documented cases exist of such under-trained tokens producing bizarre behaviour when they appear in a prompt: derailed generations, refusals to repeat the token, or unrelated output. The mitigation is hygiene rather than cleverness — train the tokenizer on a corpus that matches the pretraining mix, and audit the vocabulary for tokens with near-zero training frequency before shipping. ## Reserved and special tokens Whatever V you choose, part of it is not learned from text at all. Special tokens — sequence boundaries, turn or role markers, tool and control markers, padding, and any sentinel your training recipe needs — are added as first-class vocabulary entries at tokenizer construction time. They are never produced by merges, and pre-tokenization must protect them, so that a user who types the literal special-token string in ordinary text does not have it encoded as the control token. Treating that as a security-relevant boundary is the right instinct: user text must not be able to forge control structure. Beyond the tokens you need today, deliberately reserve a block of unused slots. It costs a small number of embedding rows and buys the ability to introduce new control tokens later without resizing the embedding and output matrices — which on a shipped model is otherwise a genuinely disruptive change. Several open-weight model families ship exactly this: a run of reserved special-token entries with no defined meaning at release. It is cheap insurance and a good signal in an interview that you have thought past the first training run. ## How to answer this well Refuse to give a single number without asking what the model serves. State the two cost sides accurately, name coverage as the dominant driver, note the rare-token tail as a real hazard, and close on reserved slots. Anchoring on "whatever the last model I read about used" is the answer that marks a candidate as repeating rather than reasoning.

  • Why are small models more sensitive to vocabulary size than large ones?
    Because the embedding and output matrices scale with vocabulary but not with depth. On a model of a few hundred million parameters, a 256K vocabulary can consume a large fraction of the total budget on lookup tables that do no computation, starving the layers that do. On a very large model the same matrices are a small share, so the sequence-length win usually dominates.
  • What are under-trained or glitch tokens, and how do they arise?
    They are vocabulary entries that appear almost never in the pretraining corpus, typically because the tokenizer was fit on a data mix that differs from what the model was ultimately trained on. Their embedding rows stay near initialisation, so prompting with them produces erratic output. The defence is aligning the tokenizer corpus with the pretraining mix and auditing token frequencies before release.
  • Why reserve unused special-token slots at tokenizer training time?
    Because adding a token later means growing the embedding and output matrices of a shipped model — a disruptive change that invalidates checkpoints and serving assumptions. A block of pre-reserved entries costs a handful of rows and lets you introduce new control tokens afterwards by assigning meaning to a slot that already exists. Several open-weight families ship exactly such a reserved block.
  • Why must special tokens be protected during pre-tokenization?
    Otherwise a user who types the literal special-token string in ordinary text has it encoded as the control token, letting untrusted input forge structural boundaries the model is trained to trust. Correct handling treats special tokens as vocabulary entries that only the framework may insert, and encodes any matching user text as ordinary characters. It is a trust boundary, not a formatting detail.

saying these in an interview costs you the question

  • Names a single number without asking what the model must encode
  • Thinks a larger vocabulary is free because sequences get shorter
  • Ignores that the output softmax runs over the whole vocabulary each step
  • Assumes every vocabulary entry is learned from text
  • Believes new special tokens can be added later at no cost

context