Why does a static word vector give 'bank' one vector for both of its meanings?
answer
- the table is keyed by spelling
- no mechanism reads the sentence
- every occurrence pulls the same row
- a frequency-weighted compromise, not a split
- context similarity, not meaning similarity
basics
~10 sThe model stores one row per word type, keyed by spelling. River-bank and money-bank occurrences both pull on that same row, so the result is a single compromise vector sitting between two unrelated neighbourhoods.
solid answer
~50 sThese models learn a fixed table with one vector per word *type*, not per occurrence. Nothing in the objective can tell the two senses apart — both are the same spelling, so every window containing either sense updates the same row. The vector converges to a frequency-weighted compromise between the river neighbourhood and the finance neighbourhood, and typically lands near neither cleanly; if one sense dominates the corpus, the rarer sense is effectively swamped. The same mechanism produces a second surprise: because the vector is defined purely by which words surround it, antonyms like `hot` and `cold` come out very close, since they appear in nearly identical contexts. So similarity here means "interchangeable in context", not "same meaning". This is the structural limit of static vectors, and it is exactly the gap that later contextual representations — which emit a different vector for each occurrence of a word — were built to close.
go deeper
Be able to say plainly that the model stores one vector per word type, chosen by spelling, so both senses of a word share a row. Know that this is called polysemy and that it is a known limitation, not a bug in your training run.
Explain the mechanism: every occurrence updates the same parameters, so the row converges to a frequency-weighted compromise. Also be ready for the antonym surprise and why substitutability, not meaning, is what cosine measures here.
Show judgment about when this actually damages a downstream system — negation-sensitive features, minority-sense domains, corpus shift — and name the concrete mitigations plus their limits before reaching for a heavier representation.
Own the evaluation argument: a team quoting analogy or nearest-neighbour eyeballing as evidence of representation quality is measuring the wrong thing. Define what task-level evidence would justify keeping or replacing a static-vector feature.
## One row per spelling A static word-vector model is, at inference time, a table: a word type maps to a fixed vector. There are as many rows as there are word types in the vocabulary, and the row is selected by the surface form alone. Nothing about the surrounding sentence is consulted, because there is no mechanism to consult it — the vector was fixed when training ended. During training, every occurrence of the spelling `bank` — in `sat on the river bank`, in `the bank raised interest rates`, in `bank of fog` — produces gradients on the *same* row. The objective has no notion of sense; it sees a token and looks up an index. So the row is pulled toward whatever makes it predictive of *all* those contexts at once. ## What the compromise actually looks like The result is not a vector that cleanly represents either sense, and it is not a vector that splits into two clusters given enough epochs — there is only one row, so there is nothing to split. It settles at a frequency-weighted compromise: roughly, the direction that minimises total loss across both context distributions, which sits between the two neighbourhoods and is a poor member of either. Two practical consequences follow. *Sense dominance.* If 95% of the corpus's `bank` tokens are financial, the vector is essentially the financial sense with a slight pull toward the geographic one. Retrieval or classification features built on it will behave as if the rarer sense does not exist. Change corpora — a hydrology corpus instead of a news corpus — and the same word's vector moves substantially. *Polysemy is common, not exotic.* This is not a corner case affecting a handful of puns. High-frequency words are the most polysemous ones in most languages, so the words your model sees most often are exactly the ones whose vectors are most blurred. ## The related surprise: antonyms are neighbours The same distributional logic produces an effect that trips up candidates who describe these vectors as "capturing meaning". `hot` and `cold`, `increase` and `decrease`, `always` and `never` come out as close neighbours, because they occur in near-identical contexts — anywhere one is grammatical, the other usually is too. Cosine similarity in this space measures *contextual substitutability*, not semantic agreement. If your downstream task is sentiment or negation-sensitive, that is a genuine hazard rather than a curiosity. ## What analogy arithmetic does and does not prove The famous demonstration is `king - man + woman`, whose nearest neighbour is `queen`. It is a striking result, and it does show that some relational structure is encoded as roughly consistent offset directions in the space. But it proves less than it appears to. The standard evaluation protocol *excludes the three query words* from the candidate set before taking the nearest neighbour — and `king` itself is usually the true nearest vector to the query point. Without that exclusion the analogy frequently "fails". Reported analogy accuracy also varies sharply by relation type: geographic and morphological relations score far better than abstract semantic ones, and the aggregate number hides that. Treat an analogy score as one weak, coarse diagnostic of the space, never as evidence that the model has learned relational reasoning — and never as a substitute for evaluating on your actual downstream task. ## Living with it, or leaving it If you are stuck with static vectors, the mitigations are all partial: train on a domain corpus so the dominant sense is the one you care about; disambiguate upstream with a rule or a tag when a critical few words are the problem; or accept that features derived from polysemous high-frequency words will be noisy and lean on other features. The honest answer at interview, though, is that this limitation is structural. One vector per type is the defining property of the approach, and no amount of tuning removes it. It is the specific problem that motivated moving to representations that produce a different vector for each occurrence of a word, and being able to name that progression — and the reason for it — is what the question is really testing.
- Why do 'hot' and 'cold' end up as close neighbours in such a space?Because the vector is determined entirely by the distribution of surrounding words, and antonyms are almost perfectly interchangeable in context — anywhere `hot` is grammatical, `cold` usually is too. Cosine similarity here measures substitutability, not agreement in meaning. That is a real hazard for sentiment or negation-sensitive features built on these vectors.
- How much does the king minus man plus woman result actually prove?Less than it looks. The evaluation excludes the three query words before taking the nearest neighbour, and the query point's true nearest vector is usually `king` itself — so the exclusion does a lot of the work. Accuracy also varies hugely by relation type, with morphological and geographic relations far outscoring abstract ones. Treat it as a coarse diagnostic, not proof of relational reasoning.
- If most of your corpus uses the financial sense, what happens to the geographic one?It is effectively swamped. The row settles near the dominant sense with only a slight pull from the rare one, so downstream features behave as though the minority sense does not exist. Retrain the same word on a hydrology corpus and its vector moves substantially — a reminder that these vectors are properties of a corpus, not of a language.
It is like a dictionary that allows exactly one definition per headword: the editor, forced to merge the riverside and the financial entries into a single line, writes something that fits neither reader.
saying these in an interview costs you the question
- Says the vector splits into two sense clusters given enough training
- Claims cosine similarity means the words mean the same thing
- Thinks the model consults the surrounding sentence at inference
- Treats an analogy score as proof of learned reasoning
- Assumes word vectors are properties of the language, not the corpus