skip to content

Which Cohere rerank model fits a corpus mixing English, French and Japanese?

level: middleimportance: should knowfreq 32%

answer

  1. one model, one mixed candidate pool
  2. query language need not match the document's
  3. no translation hop required
  4. evaluate quality language by language
  5. recall must be multilingual too

basics

~20 s

Use a multilingual rerank model such as rerank-v3.5, which covers 100-plus languages and scores a query in one language against documents in another. You send one mixed candidate list to one call — no per-language routing and no translation step.

solid answer

~50 s

Cohere's current rerank line, `rerank-v3.5`, is multilingual: it covers on the order of a hundred languages and, importantly, is **cross-lingual** — a French query can score an English or Japanese passage on its meaning rather than on shared surface tokens. That collapses what would otherwise be an ugly pipeline: no language detection to route the query, no per-language index, no machine-translation hop before or after retrieval. You send the whole mixed candidate list in one request. The older `rerank-english-v3.0` / `rerank-multilingual-v3.0` split is the previous generation's shape and is not something to reach for on new work. Two caveats worth voicing: quality is genuinely uneven across languages, so evaluate per language on your own labelled queries rather than trusting one aggregate number; and multilingual tokenizers spend more tokens per character on non-Latin scripts, which eats into the per-document token budget faster than the English case would suggest.

go deeper

for a junior

Know that Cohere offers a multilingual rerank model and that you can send one mixed-language candidate list rather than splitting by language.

for a middle

Explain cross-lingual scoring — a query in one language matching a document in another — and why that removes translation and per-language routing from the pipeline.

for a senior

Insist on per-language evaluation rather than an aggregate score, flag that non-Latin scripts consume the per-document token budget faster, and check that the retrieval stage is multilingual too.

for a principal

Own the language strategy end to end: one index versus per-locale indexes, where translation belongs if anywhere, and how you set quality bars per market when model coverage is genuinely uneven.

## The problem multilingual rerank removes A corpus in several languages traditionally forced awkward choices at every stage. Detect the query language and route to a per-language index? Then a French user never sees the English document that answers them. Machine-translate everything into English at ingestion? Then you pay translation cost, store a lossy copy, and rank against text nobody wrote. Translate the query at request time? Then you add a hop and inherit the translator's errors. A multilingual, cross-lingual reranker removes the choice: it scores meaning across the language boundary, so a single mixed pool of candidates can be ranked by a single call. ## What Cohere provides As of mid-2026 the model to name is `rerank-v3.5` — a unified rerank model with broad multilingual coverage (on the order of a hundred languages) that handles cross-lingual query-document pairs directly. The earlier generation split coverage across `rerank-english-v3.0` and `rerank-multilingual-v3.0`; that split is historical context rather than a decision you should be making today. Because model ids in this space rot quickly, keep yours in configuration and treat "which model" as a deployment concern, not something inlined at a dozen call sites. ## What this changes in the pipeline With cross-lingual reranking available, the sane architecture is: 1. **One index**, containing passages in whatever language they were written in, embedded with a multilingual embedding model so recall is also cross-lingual. 2. **One retrieval call**, returning a mixed-language candidate pool. 3. **One rerank call** with the user's query as they typed it, over the mixed pool. 4. Optionally, a **language preference applied after ranking** — for example, when scores are close, prefer the passage in the user's own language, or translate the final answer rather than the corpus. Step 4 is the part people forget. Cross-lingual relevance is not the same as user-appropriate output: a Japanese-speaking user may be best served by the English document's content delivered in Japanese, which is a generation concern, not a ranking one. ## Caveats to raise unprompted **Uneven quality.** Multilingual models are trained on uneven data. Performance on high-resource languages (English, Spanish, French, Chinese) is typically much stronger than on low-resource ones. If your business depends on a specific language, build a labelled query set *in that language* and measure it, rather than reading a single aggregate benchmark number. **Token budgets and scripts.** Tokenizers do not split all scripts equally; Japanese, Korean, Chinese, Thai and many Indic scripts consume more tokens per visible character than English. Since the rerank call truncates each document at a per-document token cap, the same physical passage length reaches the cap sooner in those languages. Size your passages in tokens, not characters, and check the distribution per language. **Mixed-language passages.** Real documents interleave languages — an English technical term inside a French sentence, code inside prose. Cross-lingual models generally handle this, but it is another reason to evaluate on your own data rather than assuming. **Recall is the ceiling.** Reranking cannot recover a document the retrieval stage never returned. If your embeddings are English-only, a cross-lingual reranker fixes nothing: the French passages were already lost. Multilingual retrieval and multilingual reranking are a pair, and the embedding side has to be handled at the recall stage. ## The short answer for an interview Name a multilingual rerank model, say that cross-lingual scoring means one mixed pool and no translation hop, and then immediately qualify it: measure per language, watch the token budget on non-Latin scripts, and make sure retrieval is multilingual too, because rerank only reorders what recall found.

  • Your embeddings are English-only but the reranker is multilingual. What happens?
    Very little improves. Reranking can only reorder the candidates retrieval returned, so if the vector stage never surfaces the French or Japanese passages, no reranker can rescue them. Multilingual retrieval and multilingual reranking go together: fix the embedding model at the recall stage first, then the cross-lingual reranker has something to work with.
  • Ranking is cross-lingual, but your user only reads Japanese. How do you handle that?
    Keep relevance and presentation separate. Let the reranker pick the most relevant passages regardless of language, then apply a language preference as a tie-break when scores are close, and have the generation step answer in the user's language while grounding on whatever passage was best. Filtering out other languages before ranking throws away your best evidence.
  • How do you validate multilingual rerank quality before rolling it out?
    Build a labelled query set per language, not one pooled set, because multilingual models are strongest on high-resource languages and an aggregate score hides a weak one. Measure a ranking metric per language against the pre-rerank baseline, and include cross-lingual pairs — query in one language, gold passage in another — since that is the capability you are actually buying.

saying these in an interview costs you the question

  • Machine-translates the whole corpus before reranking
  • Routes each language to its own English rerank call
  • Assumes uniform quality across all supported languages
  • Thinks rerank can surface documents recall never returned
  • Sizes passages by characters, ignoring script token costs

context