skip to content

When do character 3-5-grams beat word unigrams as features for matching messy product titles?

level: middleimportance: nice to knowfreq 34%

answer

  1. ask whether the token is a reliable unit
  2. one typo destroys a whole word feature
  3. overlapping windows, not whitespace
  4. a bounded number of fragments change
  5. the cost is non-zeros per row

basics

~20 s

Character n-grams win when the tokens themselves are unreliable: misspelled or run-together brand names, inconsistent punctuation, model codes. A typo changes only a few of a word's character n-grams, whereas it destroys the word unigram entirely, so overlap survives.

solid answer

~50 s

Word unigrams assume the token is the unit of meaning, which breaks down on product titles. `Samsng`, `Samsung-Galaxy` and `samsungalaxy` are three unrelated columns as word unigrams, so two listings for the same item can share almost nothing. Slice the same strings into overlapping character 3-to-5-grams and they share most of their features: a one-letter typo kills at most a handful of n-grams out of dozens. Character n-grams also cope with concatenation, casing noise, punctuation and languages that do not delimit words with spaces, and they pick up morphological variants for free. The price is a much larger and denser feature space — every string of length L yields roughly L n-grams per size, so rows have many more non-zero entries, and individual features are far less interpretable. Word features stay preferable when the text is clean prose and you want readable coefficients.

go deeper

for a junior

Know the difference in one line: word features split on whitespace, character features slide a fixed-width window over the raw string. Be able to say why a typo hurts one far more than the other.

for a middle

Explain the tradeoff quantitatively: roughly how many fragments a string of length L produces, why the non-zero count per row jumps, and why the range 3 to 5 is a sensible default band.

for a senior

Demonstrate the diagnosis. Show how you would spot that tokenisation is the bottleneck — near-duplicate records sharing no features — and how you would validate the switch rather than assert it.

for a principal

Weigh interpretability against accuracy explicitly. Decide when a few points of matching quality justify a model no analyst can explain feature-by-feature, and say who in the organisation that decision belongs to.

## Two ways to slice a string The counting step needs units. **Word n-grams** split on whitespace and punctuation and treat each resulting token — or each run of `n` consecutive tokens — as a feature. **Character n-grams** ignore token boundaries and slide a window of `n` characters along the raw string. For `galaxy` with `n = 4` you get `gala`, `alax`, `laxy`. Taking a range such as 3 to 5 means every one of those window sizes contributes features, so a single word emits many overlapping fragments. ## Why messy titles break word features Marketplace product titles are user-authored and adversarially inconsistent. The same phone appears as `Samsung Galaxy S21`, `SAMSUNG GALAXY-S21`, `Samsng Galaxy S 21`, `samsunggalaxy s21 (unlocked)`. Under word unigrams: - `Samsung` and `Samsng` are two entirely distinct columns. Their similarity is exactly zero — the representation has no notion that they nearly match. - `samsunggalaxy` shares nothing with `Samsung` plus `Galaxy`. - `S21` versus `S 21` is one token or two. Every variation destroys a whole feature. Since the informative words in a title are the brand and model, and those are exactly the tokens people mistype, word unigrams fail on precisely the cases that matter. Under character 3-5-grams, `Samsung` yields fragments like `sam`, `ams`, `msu`, `sung`, `amsu`, `msun`, and so on. `Samsng` loses the ones spanning the missing `u` but keeps `sam`, `ams`, `sng`, `msn`-style neighbours — a substantial fraction of the original set. The two strings now have high overlap instead of none. Concatenation is handled for the same reason: `samsunggalaxy` contains every internal fragment of both words, plus a few junk ones straddling the join. ## What character n-grams buy you generally - **Robustness to typos and spacing.** A single edit perturbs a bounded number of fragments. - **Morphology for free.** `run`, `running`, `runner` share `run`-fragments without any stemming step. - **No dependence on a tokeniser.** Useful for languages that do not separate words with spaces, for code-like identifiers, SKUs and part numbers, and for text with heavy punctuation. - **Graceful handling of unknown words.** A brand never seen in training still shares fragments with known strings, whereas as a word unigram it would be entirely out of vocabulary. ## What they cost - **Feature count.** A string of length `L` produces about `L - n + 1` fragments for each `n`, so a 3-to-5 range emits roughly `3L` features per string. The vocabulary of distinct fragments is large, though bounded by the alphabet in a way the word vocabulary is not. - **Density.** This is the practical bite. Word unigrams give a title maybe six non-zero entries; character 3-5-grams give it a hundred or more. The matrix has far more stored values, so memory and training time rise even though it is still sparse. - **Interpretability.** A coefficient on the word `refurbished` is a sentence you can put in a report. A coefficient on the fragment `urb` is not. If someone will ask you to explain the model to a non-technical stakeholder, this matters. - **Noise.** Fragments that straddle two words (`g s2` from `galaxy s21`) are position-dependent accidents that can fit noise. ## Choosing, and the middle path The question to ask is: *is the token a reliable unit here?* For clean editorial prose — news articles, long reviews, support tickets written in sentences — the answer is yes and word features are the sensible default: fewer, denser in meaning, readable. For short, noisy, user-entered strings — product titles, addresses, company names, search queries — the answer is no, and character n-grams are frequently the single biggest accuracy win available. The two are not exclusive. A common approach is to build both representations and concatenate them, letting the model use word features where the tokenisation held up and character features where it did not. Whichever you pick, the range is a hyperparameter: 3 to 5 is a common starting band because 2-grams are too generic (nearly every string contains `an`) and 6-grams start to behave like whole words again, losing the typo robustness that was the point. Tune it the way you tune anything else — on held-out folds, with the vocabulary fitted inside the fold. ## Word n-grams as the other axis Moving from word unigrams to word bigrams is a different fix for a different problem: it recovers short-range order, so `not good` becomes its own feature rather than the pair `not` and `good`, which a bag-of-words cannot distinguish from `good, not`. That helps sentiment and negation, and costs a large vocabulary increase for features that are individually rare. It does nothing at all for misspellings — which is exactly why the two decisions are made separately.

  • What does moving from word unigrams to word bigrams buy you, and what does it cost?
    It recovers short-range order, so `not good` becomes its own feature instead of two independent words that a bag-of-words cannot tell from `good, not`. That helps negation and fixed phrases. The cost is a vocabulary that grows sharply while each new feature is individually rare, so most bigrams are seen a handful of times and are pruned or overfit. It does nothing for misspellings.
  • Why is 3-5 a common character n-gram range rather than 2 or 7?
    Two-character fragments are too generic — almost every string contains them, so they behave like the function words idf already discounts. Six or seven characters approach whole words and lose the typo robustness that motivated the choice. Three to five keeps fragments distinctive enough to carry signal while short enough that a single edit only perturbs a few of them.
  • Why does a character n-gram matrix get slower to train even though it is still sparse?
    Sparsity is about the fraction of zeros; cost is driven by the count of non-zeros. A short title yields perhaps six word unigrams but well over a hundred character fragments across a 3-to-5 range, so every row has an order of magnitude more stored values. Training time scales with those, not with the nominal column count.

Comparing whole words is like matching people by their full legal name: one typo and you get no match at all. Comparing character n-grams is like matching on overlapping fragments of the name, where a single wrong letter still leaves most fragments intact.

saying these in an interview costs you the question

  • Thinks character n-grams understand meaning or morphology deliberately
  • Claims character n-grams are always better than word features
  • Ignores the jump in non-zero entries per document
  • Confuses character n-grams with word n-grams
  • Expects readable per-feature explanations from character fragments

context