skip to content

Why does an English-trained embedding model degrade on Portuguese and Japanese text?

level: middleimportance: should knowfreq 48%

answer

  1. the space only maps what training covered
  2. vocabulary was fitted on one language
  3. non-Latin script fragments hardest
  4. translations near each other is a trained property
  5. no error fires, so evaluate per language

basics

~20 s

Its training data and tokenizer vocabulary were built for English, so other languages fragment into unfamiliar subword pieces and land in a poorly organized region of the space. Vectors still come back looking normal, so the degradation is invisible without a per-language evaluation.

solid answer

~50 s

An embedding model's semantic space is shaped entirely by what it was trained on. A model trained predominantly on English learned fine-grained structure for English and almost none for languages it barely saw, so Portuguese and Japanese reviews get vectors that are technically valid and semantically mushy — similar items are not reliably closer than unrelated ones. Two mechanisms compound. First, the tokenizer: a vocabulary fitted on English fragments other languages into many short, low-information pieces, worst for non-Latin scripts, which both wastes the length budget and gives the model unfamiliar units. Second, cross-lingual alignment: a monolingual model has no reason to place a Portuguese sentence near its English translation, so mixed-language corpora cannot be searched with one query. None of this errors — scores stay in range, results still rank — so language coverage belongs in model selection up front, verified with a small labelled retrieval set per language rather than assumed from an aggregate benchmark score.

go deeper

for a junior

Know that an embedding model only works well on languages it was trained on, and that using an English-focused model on other languages degrades results without producing any error.

for a middle

Explain both mechanisms: a tokenizer vocabulary fitted on English fragments other scripts into unfamiliar pieces, and training coverage determines how finely the space is organized for a language.

for a senior

Show that you would catch this with per-language labelled evaluation sets from the real corpus, and distinguish per-language search from genuine cross-lingual alignment when choosing a model.

for a principal

Own the portfolio decision: one multilingual model with a single index and cross-lingual retrieval, several per-language models with better local quality, or a translate-then-embed pipeline — each with different operational surface, capacity trade-offs and migration cost.

## The failure is quiet, which is the point Run a catalogue of Portuguese and Japanese product reviews through a strong English-trained embedding model and everything appears to work. Vectors come back at the right dimensionality. Similarity scores land in a familiar range. Search returns ten results per query, ranked. The only thing wrong is that the ranking is close to arbitrary, and no component of the system can tell you that. This is why language is a **first-order selection criterion**, decided before any leaderboard comparison — not something you discover in production. ## Mechanism one: training distribution A model's vector space is a learned map of the text it saw. Where it saw a lot, the map has fine structure: near-synonyms separate, domain senses separate, paraphrases collapse together. Where it saw little, the map is coarse — everything in that language crowds into an under-differentiated region, so distances between two Japanese reviews carry far less information than distances between two English ones. A useful mental check: the model cannot distinguish meanings it was never trained to distinguish. Absence of a language in training is not a mild handicap; it removes the very structure that similarity search reads. ## Mechanism two: the tokenizer Subword vocabularies are fitted to a corpus. Fit one on mostly-English text and it will contain whole common English words and useful English morphemes, but only fragments for other languages. - **Portuguese**, sharing the Latin script, fares better than you might fear at the character level but still fragments: accented forms, verb inflections and common function words split into pieces the model treats as unfamiliar units. - **Japanese** is worse on two counts. The script is outside the vocabulary's focus, so text may fall back to byte-level or very short pieces, and Japanese has no whitespace word boundaries, so the tokenizer's segmentation may not align with meaningful units at all. Two consequences follow. Quality drops because the model reasons over unfamiliar fragments. And the effective input shrinks: the same sentence costs several times more tokens, so more of it risks running past the model's length limit. ## Mechanism three: cross-lingual alignment This one is distinct and often missed. Even a model that handles each language *individually* will not necessarily place a sentence near its translation unless it was trained to. **Cross-lingual alignment** means exactly that: "the delivery was late" and its Portuguese equivalent occupy nearly the same point in one shared space. Models get that property by training on parallel or paired multilingual data with a contrastive objective. Without it, your mixed-language corpus is effectively several disjoint indexes wearing one coat: a Portuguese query retrieves Portuguese documents at best, never the English document that answers it. Decide up front which you need. "Each language searched within itself" is a weaker requirement than "any query retrieves any language", and only the second demands genuine alignment. ## What to do instead **Choose a model trained for your languages.** Multilingual encoders exist across the size range and are trained with explicit multilingual and often parallel data. Check the model's documented language list, not just its headline score. **Evaluate per language, not in aggregate.** A single averaged benchmark number hides exactly the failure you care about. Build a small labelled set — a few dozen query/relevant-document pairs per language, drawn from your own corpus — and measure retrieval quality in each. This is the only reliable evidence, and it is cheap. **Watch the tokenizer explicitly.** Compare tokens-per-character across your languages with the model's own tokenizer. A language costing three or four times more tokens than English is a warning about both quality and length budget. **Weigh one multilingual model against per-language models.** A multilingual model spends fixed capacity across many languages, so at equal size it is often slightly weaker on any single language than a dedicated one — a trade-off sometimes called the curse of multilinguality. Against that, one model means one index, one operational path, and cross-lingual retrieval for free. Per-language models mean better per-language quality but language detection at ingest and query time, several indexes, and no cross-lingual matching. Most teams take the multilingual model unless one language dominates the traffic. **Do not translate as a reflex.** Machine-translating everything into English before embedding is a real option and sometimes wins, but it adds a lossy step, a cost per document, and a dependency, and it does not help when the source text is short, colloquial or full of product-specific terms — exactly the character of user reviews. ## What an interviewer is testing That you know embedding quality is a property of training coverage rather than an intrinsic property of the model, that you can name both the tokenizer and the alignment mechanisms, and above all that you would catch this with a per-language evaluation instead of trusting that no errors means it works.

  • How would you prove the degradation rather than suspecting it?
    Build a small labelled retrieval set per language from your own corpus — a few dozen queries with known relevant documents — and measure recall at k separately for each. Aggregate benchmark scores hide per-language collapse. Compare a candidate multilingual model against the incumbent on the same set. Also compare tokens per character across languages, since a large ratio predicts both quality loss and truncation risk.
  • If each language is searched only within itself, do you still need cross-lingual alignment?
    No — that is the weaker requirement, and per-language indexes with per-language models can serve it well, at the cost of language detection and several indexes to operate. Alignment becomes necessary only when one query must retrieve documents in other languages, or when documents mix languages within a single chunk, which is common in reviews and support threads.
  • Does a multilingual model cost you quality on English?
    Usually a little, at equal model size — fixed capacity is shared across many languages, an effect often called the curse of multilinguality. Whether that matters is an empirical question for your corpus, and larger multilingual models narrow the gap. Measure both candidates on your English evaluation set before assuming the specialist wins.

saying these in an interview costs you the question

  • Assuming a top benchmark score implies good coverage in every language
  • Expecting an error or warning when a language is unsupported
  • Believing a monolingual model still places translations near each other
  • Ignoring that non-Latin script inflates token counts several times over
  • Judging multilingual quality from one aggregate average instead of per language

context