skip to content

How do you judge whether Command's multilingual coverage fits a new market?

level: principalimportance: nice to knowfreq 30%

answer

  1. A list is a hypothesis, not an SLA
  2. Fluent is not the same as accurate
  3. Real inputs, never translated test data
  4. Native reviewers catch what metrics miss
  5. Scripts differ in tokens, so in price

basics

~20 s

Treat a vendor's language list as a starting hypothesis, not a guarantee. Command A is marketed at 23 languages and the earlier Command R line at ten key business languages, but only a task-specific eval with native reviewers tells you whether quality, safety behaviour and token cost hold in that market.

solid answer

~50 s

Cohere positions Command as multilingual by design: the Command R generation was tuned for ten key business languages — English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic and Chinese — with broader pretraining coverage beyond them, and Command A is marketed as supporting 23 languages. Those numbers describe training and benchmark coverage, not a per-task service level. Before committing to a market I build a labelled eval set in that language from real traffic, score the actual task rather than generic fluency, and put native speakers on a sample, because automated metrics miss register, formality and idiom errors that destroy trust. I also measure tokenisation cost, since non-Latin scripts inflate token counts and therefore price, latency and effective context. Then I compare alternatives — a bigger model, translate-then-prompt, a translation-specialised variant, or a regional model — and let the numbers decide.

go deeper

for a junior

Know that Command models are marketed as multilingual and that the supported-language list describes training and benchmark coverage rather than guaranteed quality on your particular task.

for a middle

Be able to design the check: a labelled eval set of real in-market inputs, scored on the task metric rather than fluency, with per-language results kept separate from the aggregate.

for a senior

Add the operational dimensions — native-speaker review, cross-lingual retrieval as its own test, tokenisation cost per script, and per-locale production monitoring that catches regressions after a model upgrade.

for a principal

Own the market decision. Compare model tier, translate-then-prompt, a specialised variant and a regional model on cost and quality, budget the human review capacity per language, and be willing to say a market is not ready to launch.

## What a language list actually claims When a vendor says a model supports N languages, the claim is roughly: these languages were well represented in training data, and the model was evaluated on benchmarks in them. For Cohere's Command line, the Command R generation was explicitly optimised for **ten key business languages** — English, French, Spanish, Italian, German, Brazilian Portuguese, Japanese, Korean, Arabic and Simplified Chinese — with additional languages present in pretraining at lower depth. **Command A** is marketed at **23 languages**, a broader claim on a newer flagship. What the claim is *not*: a per-task guarantee, a uniform quality level across the list, a promise about dialects and regional variants, or a statement about safety and refusal behaviour in that language. A model can be fluent in a language and still be materially worse at *your* task in it — grounded extraction from a specific document type, or following a strict output schema — because fluency and task competence are different capabilities. ## Build the eval before you build the feature The decision procedure that survives contact with reality: **1. Define the task, not the language.** "Does it speak Polish" is unanswerable. "Does it extract these seven fields from Polish invoices with ≥95% field accuracy" is a test you can run. **2. Get real inputs.** Translated English test data is the classic self-deception: it is cleaner, more formal, and structurally English, so it flatters the model. Use genuine documents and genuine user phrasings from the market. **3. Score the task metric.** Field accuracy, schema-validity rate, citation correctness, refusal rate — whatever the product depends on. Fluency scores are not the product. **4. Put native speakers on a sample.** This catches what metrics cannot: wrong formality register (addressing a customer with the familiar form in a language where that is rude), calques that read as machine translation, culturally wrong examples, mishandled honorifics. In a customer-facing product these are trust-destroying even when the information is correct. **5. Test the failure paths in-language.** Does the model refuse appropriately? Does it hallucinate more when the retrieved context is in a different language from the question? Cross-lingual retrieval — Spanish question, English source documents — is a distinct capability from monolingual work and needs its own test. ## Tokenisation is a cost and capacity question A subtlety senior candidates often miss: tokenisers are trained on data distributions that skew English, so the same semantic content in a non-Latin script — Arabic, Japanese, Korean, Thai, Hindi — frequently costs more tokens than its English equivalent. Since billing, rate limits and the context window are all measured in tokens, that is a direct hit on three things at once: cost per request rises, latency rises with the longer sequence, and the effective amount of retrieved context you can fit falls. Measure tokens per typical request in each target language rather than assuming parity, and feed the real number into the business case. A market that looks marginal at English token rates can be clearly unprofitable at 1.7x. ## The alternatives you are choosing between A principal-level answer compares options rather than accepting or rejecting one: - **Same model, better prompting.** Instructing in the target language, and giving in-language few-shot examples, often closes more of the gap than a model upgrade. - **Move up a tier.** The flagship's broader language coverage may fix it, at a cost you can compute against the market's revenue. - **Translate-then-prompt.** Translate into English, run the pipeline, translate back. Cheap and often surprisingly strong, but it loses nuance and adds two failure points — and it can leak sensitive text through a third service. - **A specialised variant.** Cohere ships a translation-focused Command A variant; a purpose-built model can beat a general one on that narrow job. - **A regional or open-weight model.** For some languages a locally trained model genuinely outperforms the global flagship, and self-hosting it may also settle a residency requirement. - **Don't launch yet.** A legitimate outcome. Shipping a product that is subtly wrong in a customer's language damages the brand more than a delayed launch. ## Own it as an ongoing commitment, not a launch gate Language support is not a one-time check. Model snapshots change and can regress in a specific language while improving on average, so the per-language eval must run on every model upgrade, not only at launch. Instrument production per locale — refusal rates, thumbs-down rates, escalation-to-human rates, schema-failure rates — because the earliest signal of a regression comes from users in that market, and it will not show up in an aggregate dashboard dominated by English traffic. Budget the human review capacity for each supported language up front; a market you cannot evaluate is a market you cannot responsibly support.

  • Why is translated English test data a bad basis for a multilingual eval?
    Because it is not what your users write. Machine-translated test sets are cleaner, more formal and structurally English, so the model performs better on them than on genuine in-market inputs full of local abbreviations, dialect, code-switching and domain jargon. You end up validating a distribution you will never serve. Source real documents and real user phrasings, even if you can only assemble a few hundred.
  • How does tokenisation change the economics of supporting a non-Latin-script language?
    Tokenisers skew toward the training distribution, so the same content in Arabic, Japanese or Hindi can consume noticeably more tokens than its English equivalent. Cost per request, latency and rate-limit consumption all scale with that, and the effective context left for retrieved documents shrinks. Measure tokens per representative request per language and put the real multiplier into the business case rather than assuming parity.
  • A model upgrade improves your average eval score but you support eight languages. What do you check?
    Per-language scores, not the aggregate. Averages hide regressions in low-volume locales, and snapshots genuinely can improve overall while getting worse in one language. Run the full per-language suite, compare against the pinned baseline, and treat any single-language regression as a blocker or a per-locale routing decision — keeping that market on the old snapshot until it is fixed.

saying these in an interview costs you the question

  • Taking a vendor's language count as a quality guarantee
  • Evaluating with machine-translated test data instead of real inputs
  • Judging output on fluency rather than task accuracy
  • Assuming token counts and costs are the same across scripts
  • Checking language quality once at launch and never on upgrades

context