How do skip-gram and CBOW differ, and which wins on a small, rare-word-heavy corpus?
answer
- one predicts many, or many predict one
- which side of the window is the input
- averaging blurs an individual contribution
- count the updates a rare word receives
- small data favours the slower objective
basics
~20 sSkip-gram predicts each context word from the centre word; CBOW predicts the centre word from the averaged context. Skip-gram creates more separate updates per rare word, so it wins on small rare-word-heavy corpora, while CBOW trains faster.
solid answer
~50 sBoth slide a window over raw text and learn one dense vector per word type, but they point the prediction in opposite directions. **CBOW** averages the surrounding context vectors and predicts the centre word from that single averaged input. **Skip-gram** takes the centre word and predicts each surrounding word as a separate training example. That difference decides the economics: for one window of width 5, CBOW produces one update and skip-gram produces up to ten, so skip-gram extracts far more signal per occurrence — which is exactly what a rare technical term needs. CBOW's averaging also smooths away the idiosyncratic contexts that make a rare word distinctive, and lets frequent neighbours dominate the average. So on a 5-million-token domain corpus full of rare terms, I would pick skip-gram; on a billion-token general crawl where speed matters and most words are frequent, CBOW is a reasonable, much cheaper choice.
go deeper
Be able to state the direction of each objective without hesitating: CBOW guesses the middle word from its neighbours, skip-gram guesses the neighbours from the middle word. Also know that neither needs labels.
Explain why the direction changes rare-word quality — count the gradient updates a rare word gets under each, and say what averaging does to an individual context word's contribution. Know that the window width is sampled, not fixed.
Demonstrate that you pick the objective from corpus size and vocabulary shape rather than habit, and that you tune the window for the similarity you actually want — substitutable neighbours versus topical ones. Say what you would measure to confirm the choice.
Own the framing question: is training your own word vectors worth it at all versus reusing a stronger pretrained representation? Argue the cost, the domain-vocabulary gain, and the maintenance burden of a corpus you now have to keep refreshing.
## The shared setup Both objectives come from the same family of shallow predictive word-vector models, and both rest on the distributional hypothesis: words that occur in similar contexts should get similar vectors. Training data is created mechanically from raw text with no labels. A window slides over the corpus; at each position one token is the *centre* word and the tokens within the window on either side are its *context* words. The model itself is deliberately tiny — there is no hidden non-linearity. There are two parameter matrices, each with one row per vocabulary word: an *input* (centre) matrix and an *output* (context) matrix. A score for a (centre, context) pair is the dot product of the centre word's input vector and the context word's output vector. Training pushes the score of observed pairs up and the score of non-observed pairs down. When training finishes, the input matrix is normally kept as "the word vectors" and the output matrix is discarded (summing or concatenating the two is an occasionally used variant). ## The two directions **CBOW (continuous bag of words)** takes the context and predicts the centre. The input vectors of all context words in the window are averaged into one vector, and that average is scored against the output vectors to predict which word sits in the middle. "Bag of words" is literal: the average throws away order, so `the cat sat on` and `on sat cat the` give the same input. **Skip-gram** reverses this. It takes the centre word's input vector and predicts each context word independently. One window of half-width 5 with 10 neighbours becomes 10 separate (centre, context) training pairs, each with its own gradient step. ## Why the direction changes the outcome *Updates per occurrence.* A word that appears 40 times in the corpus receives roughly 40 centre-word updates under CBOW and roughly 40 x (window size) targeted updates under skip-gram. For rare words that difference is the whole game — skip-gram simply squeezes more supervision out of every appearance. *Averaging blurs.* Under CBOW, a rare word appearing as *context* is one of several vectors going into an average, and the average is dominated by whatever frequent words share the window. Its individual contribution to the gradient is diluted. Skip-gram never averages, so each pair is scored on its own. *Cost.* CBOW does one forward/backward pass per window versus skip-gram's several, so CBOW trains several times faster on the same corpus. On very large corpora that speed can buy more epochs or a larger vocabulary, and frequent words have ample data either way, so the quality gap narrows. The original reports are consistent with this: CBOW is faster and slightly better on syntactic regularities for frequent words; skip-gram is better for rare words and small data. ## The window is not just a width Two details of the window matter more than candidates expect. First, the window is typically *sampled*, not fixed: for each centre word an actual half-width is drawn uniformly between 1 and the configured maximum. Words adjacent to the centre therefore fall inside the window in every draw, while words at the far edge only make it occasionally. The effect is a soft distance weighting for free, plus a cheaper average cost. Second, window size changes *what kind of similarity* you get. Narrow windows (1-2) capture words that are grammatically substitutable — nearest neighbours look like the same part of speech and the same syntactic slot. Wide windows (10+) capture topical relatedness — nearest neighbours are words from the same subject area rather than words you could swap in. Neither is more correct; you pick to match the downstream use. A retrieval-style feature usually wants the wide, topical setting; a tagging feature usually wants the narrow one. A third pre-processing step, subsampling of very frequent words, discards common function words with a probability tied to their frequency. This both speeds training and effectively widens the window, since discarded tokens no longer occupy window slots. ## Choosing in practice Ask two questions. How much text do I have, and how much of the vocabulary I care about is rare? A domain corpus of a few million tokens where the interesting vocabulary is long-tail technical terms points squarely at skip-gram. A very large generic corpus where the goal is decent vectors quickly points at CBOW. If the rare vocabulary is not just rare but *absent* at inference time, neither objective helps by itself — that is a vocabulary-coverage problem and needs subword-composed vectors instead.
- Beyond how much context each pair sees, what does the window width change?It changes the kind of similarity you get. Narrow windows of one or two words yield neighbours that are grammatically substitutable — same part of speech, same slot. Wide windows of ten or more yield topically related words you would not swap in. Pick the width to match the downstream task rather than defaulting to five.
- How does GloVe's objective differ from skip-gram's?GloVe is count-based rather than predictive. It first builds a global word-context co-occurrence matrix, then fits vectors by weighted least squares so that a pair's dot product plus two bias terms approximates the log of its co-occurrence count. A weighting function caps the influence of very frequent pairs and downweights the enormous count-one tail. Skip-gram never materialises the matrix; it samples local windows instead.
- Why does training maintain two vectors per word instead of one?Scoring a pair as the dot product of one shared vector with itself would reward every word for predicting itself, which is degenerate. Separate input and output matrices keep the centre role and the context role distinct. After training the output matrix is normally thrown away and the input matrix is shipped as the word vectors.
CBOW is a fill-in-the-blank exercise: read the sentence with a hole and guess the missing word. Skip-gram is the reverse quiz: given one word, name the words likely to surround it.
saying these in an interview costs you the question
- Says the two differ only in training speed
- Thinks skip-gram predicts the centre word from context
- Assumes CBOW's averaging preserves word order
- Believes the window is always the full configured width
- Claims CBOW is always better because it is faster