skip to content

How do WordPiece and Unigram pick vocabulary differently from BPE's merges?

level: middleimportance: should knowfreq 40%

answer

  1. three different scoring rules
  2. bottom-up twice, top-down once
  3. count versus likelihood gain
  4. prune what costs least to lose
  5. segmentation search enables sampling

basics

~20 s

BPE merges the most frequent adjacent pair. WordPiece merges the pair that most improves corpus likelihood, roughly pair frequency divided by the frequencies of its two parts. Unigram goes the other way: it starts from a large candidate vocabulary and prunes the tokens that cost the least likelihood.

solid answer

~60 s

All three end up with a subword vocabulary, but they select it under different criteria. **BPE** is greedy on raw counts: merge the most frequent adjacent pair, repeat. **WordPiece** keeps the same bottom-up merge loop but scores candidates by likelihood gain instead of count — in effect pair frequency divided by the product of its parts' frequencies — so a pair whose halves are individually rare scores higher than an equally frequent pair made of two very common pieces. It marks word-internal continuations, conventionally with a `##` prefix. **Unigram** is top-down and probabilistic: seed a large candidate vocabulary, fit token probabilities, then repeatedly drop the candidates whose removal costs the least corpus likelihood until the target size is reached. At encoding time it does not replay merges — it searches for the highest-probability segmentation, which also lets it sample alternative segmentations for regularisation. SentencePiece is the implementation commonly used for BPE and Unigram, and it treats the input as a raw stream with whitespace encoded as a marker character rather than pre-split on spaces.

code

python · 8 lines
python
pair_freq = 40          # 'un' followed by '##able'
freq_un = 400
freq_able = 50

bpe_score = pair_freq                              # raw count
wordpiece_score = pair_freq / (freq_un * freq_able)  # likelihood gain

print(bpe_score, round(wordpiece_score, 6))

go deeper

for a junior

Know the headline: BPE merges by frequency, WordPiece merges by likelihood gain, Unigram prunes a big vocabulary down. Naming the direction of each is enough.

for a middle

Be able to state WordPiece's score as pair frequency over the product of its parts and explain why that prefers pieces that co-occur more than chance, plus how Unigram's EM-and-prune loop works.

for a senior

Explain the encoding-time differences — merge replay, longest match, Viterbi search — and why only a probabilistic vocabulary gives you segmentation sampling for regularisation.

for a principal

Own the selection criteria you would actually apply: ecosystem compatibility, multilingual fertility and coverage guarantees usually decide this, not a benchmark delta between the algorithms themselves.

## Three selection criteria, one output shape Every subword tokenizer produces the same kind of artefact: a fixed vocabulary of pieces over a base alphabet, plus a rule for segmenting new text. The interesting differences are in **how a piece earns its slot** and **how segmentation is decided at encoding time**. ## BPE: greedy on frequency BPE grows the vocabulary bottom-up, merging the most frequent adjacent pair at each step and recording an ordered merge table. Encoding replays those merges in rank order. The criterion is raw count, with no model of the corpus behind it. The weakness this exposes: a pair can be frequent simply because both of its halves are extremely common, even though gluing them together explains nothing new. Merging `th` + `e` is unsurprising and adds little, yet it wins on count alone. ## WordPiece: greedy on likelihood gain WordPiece keeps the bottom-up loop but changes the score. It treats the corpus as generated by a unigram model over the current vocabulary and asks which merge most increases the likelihood of the data. Working through the algebra, the score reduces to the pair's frequency divided by the product of its two parts' frequencies — a mutual-information-flavoured quantity rather than a raw count. The practical effect is that WordPiece prefers pairs that *co-occur more than chance would predict*. A pair of pieces that are individually rare but almost always appear together scores highly; a pair of ubiquitous pieces that happen to land next to each other often does not. Vocabularies come out slightly more morpheme-like on the margin. WordPiece marks continuation pieces explicitly — the convention is a `##` prefix on any piece that is not word-initial, so `playing` might segment as `play` + `##ing`. Encoding is a greedy longest-match scan within each word against the vocabulary, not a merge replay. ## Unigram: top-down pruning under a probability model Unigram inverts the direction. It starts with a deliberately oversized candidate vocabulary — often generated by frequent substrings or a permissive BPE run — and fits a unigram probability distribution over those candidates using expectation-maximisation, where the hidden variable is the segmentation of each word. Then it computes, for each candidate token, how much corpus log-likelihood would be lost if that token were removed, and prunes a fraction of the least useful ones. Refit, prune, repeat, until the vocabulary hits the target size. Because it ends with an explicit probability for every token, encoding is a search rather than a replay: Viterbi over the lattice of possible segmentations to find the most probable one. That has a distinctive consequence — the tokenizer can also *sample* from the segmentation distribution instead of always taking the argmax. Training a model on sampled segmentations of the same word is subword regularisation, which makes downstream models more robust to how text happens to be split. BPE has an analogous trick in dropping merges at random during encoding, but it is bolted on rather than intrinsic. ## SentencePiece: the orthogonal axis SentencePiece is frequently listed as a fourth algorithm; it is not. It is an implementation that can train BPE or Unigram, and its distinctive contribution is a different treatment of the input. Rather than requiring language-specific pre-tokenization on whitespace, it consumes the raw character stream and encodes spaces as an ordinary symbol — conventionally the marker `▁` (U+2581) attached to the token that follows a space. Two things fall out. First, tokenization becomes language-agnostic: Japanese, Chinese and Thai have no whitespace to split on, and SentencePiece does not need any. Second, detokenization is exact and trivial — concatenate the pieces and turn the markers back into spaces — with no heuristics about where to reinsert whitespace. ## Choosing between them In practice the choice is driven less by measured quality gaps, which are modest, than by ecosystem and requirements: - **BPE (usually byte-level)** dominates modern generative models: simple, fast, deterministic, and byte-level variants give guaranteed coverage. - **WordPiece** is associated with the encoder-model lineage and its `##` convention. - **Unigram** is preferred where multilingual balance and subword regularisation matter, and where the probability model itself is useful. An interviewer asking this question wants to hear that you know the criterion differs — count, likelihood gain, likelihood loss — and that encoding differs correspondingly — merge replay, longest match, probabilistic search. Claiming they are "basically the same thing with different names" is the answer that fails.

  • Why can Unigram sample different segmentations of the same word while BPE normally cannot?
    Unigram keeps an explicit probability for every token, so a word defines a lattice of candidate segmentations each with a computable probability. Encoding can take the argmax or sample from that distribution. BPE has no probabilities — only an ordered merge table whose replay is deterministic — so varying its output requires an added mechanism such as randomly skipping merges during encoding.
  • What does the ## marker in a WordPiece vocabulary actually encode?
    That the piece is a word-internal continuation rather than a word start. It makes `ing` as a suffix a different vocabulary entry from `ing` beginning a word, which preserves position information that would otherwise be lost, and it makes detokenization unambiguous: pieces carrying the marker are joined to the previous piece, others get a preceding space.
  • Is SentencePiece an alternative algorithm to BPE and Unigram?
    No — it is an implementation that can train either. Its distinctive property is input handling: it consumes the raw character stream, encodes whitespace as an ordinary symbol rather than splitting on it, and therefore needs no language-specific pre-tokenization. That is what makes it the common choice for corpora including Japanese, Chinese or Thai, where whitespace does not delimit words.
  • Do these algorithms produce measurably different downstream model quality?
    Differences exist but are usually small compared with vocabulary size and the language balance of the tokenizer's training corpus. Unigram tends to produce slightly more morphologically plausible pieces and supports subword regularisation; byte-level BPE wins on simplicity and guaranteed coverage. Ecosystem compatibility and multilingual fertility are the deciding factors far more often than a benchmark delta.

saying these in an interview costs you the question

  • Says WordPiece and BPE differ only in the ## marker
  • Describes Unigram as growing a vocabulary by merging pairs
  • Calls SentencePiece a fourth tokenization algorithm
  • Thinks all three encode by replaying an ordered merge table
  • Claims one algorithm is universally better regardless of corpus

context