skip to content

Which models make up Cohere's Command family, and why pin a dated model ID?

level: middleimportance: must knowfreq 55%

answer

  1. Flagship, prior pair, and a small one
  2. Context window doubled at the top
  3. IDs end in month and year
  4. Undated names are moving targets
  5. Upgrades belong in an eval, not a surprise

basics

~20 s

Command A is Cohere's current flagship chat model with a 256K-token context, above the earlier Command R+ and Command R at 128K and the small Command R7B. IDs carry a date suffix; undated aliases follow the newest snapshot, so production should pin the dated ID.

solid answer

~40 s

The Command line is Cohere's generative side, alongside Embed and Rerank. As of mid-2026 the flagship is **Command A** (`command-a-03-2025`), a ~111B-parameter model with a 256K-token context window designed to serve on as few as two high-end GPUs. Below it sit the previous generation, **Command R+** (`command-r-plus-08-2024`, ~104B) and **Command R** (`command-r-08-2024`, ~35B), both with 128K context, plus **Command R7B** (`command-r7b-12-2024`) for latency- and cost-sensitive work. Cohere has also shipped task-specialised Command A variants for reasoning, vision and translation. The naming convention matters operationally: the `MM-YYYY` suffix identifies an immutable snapshot, while undated aliases such as `command-r` resolve to whatever is newest. Pin the dated ID in production so a vendor release cannot change your outputs, and treat every upgrade as a change that goes through your eval set.

go deeper

for a junior

Know that Command is Cohere's chat line, that Command A is the current flagship above Command R and R+, and that model IDs end with a month-year snapshot suffix.

for a middle

Explain the tiering and the context-window differences, and say precisely what an undated alias does — resolves to the newest snapshot, so behaviour can change without any deploy on your side.

for a senior

Show the upgrade discipline: pinned IDs everywhere, an eval set that gates snapshot changes, canary or shadow traffic before rollout, and tracked deprecation dates so pinning never becomes a forced migration.

for a principal

Own model selection as an economic decision. Set the routing policy between small and flagship tiers, decide when a 256K context beats better retrieval, and keep the organisation's model inventory and eval harness funded so vendor releases are routine rather than incidents.

## The shape of the family Cohere splits its catalogue into three jobs: **Command** generates, **Embed** vectorises, **Rerank** orders. Command is the chat line, and it is tiered the way most vendor line-ups are — a flagship, a previous generation still in service, and a small fast model. **Command A** (`command-a-03-2025`) is the current flagship as of mid-2026. Two facts define it commercially. It carries a **256K-token context window**, double the previous generation's 128K, which matters for the long-document workloads Cohere targets. And it was engineered for deployment efficiency: roughly 111B parameters, served on as few as two high-end GPUs, which is the point when the customer is running it in their own environment rather than paying per token. Cohere has since shipped task-specialised Command A variants — a reasoning-oriented one, a vision-capable one, and a translation-focused one — which are separate model IDs rather than modes of the base model. **Command R+** (`command-r-plus-08-2024`, ~104B) and **Command R** (`command-r-08-2024`, ~35B) are the previous generation, both at 128K context. They remain widely deployed because a lot of production RAG was built on them and re-evaluating a pipeline is expensive. R was the workhorse; R+ was the higher-capability option for harder reasoning and more reliable tool use. **Command R7B** (`command-r7b-12-2024`) is the small end: cheap, fast, and viable where the task is narrow — classification, routing, short summarisation, first-pass extraction. Note what is *not* here: the older `command` and `command-light` generate-era models predate this line entirely and are not what an interviewer means by "the Command family" today. ## Choosing between them The honest answer is that size is the last dial, not the first. Sequence: fix the prompt and the retrieval quality, then measure the smallest model on your own eval set, then move up only where it fails. Real selection criteria: - **Task difficulty.** Multi-step reasoning and multi-tool agent loops degrade fastest on small models; extraction and routing barely notice. - **Context need.** If your prompts genuinely run past 128K tokens, the flagship's 256K window is not a nice-to-have, it is the requirement. But most systems that think they need it actually need better retrieval. - **Latency budget.** A small model's time-to-first-token is a product feature in an interactive surface. - **Cost per call multiplied by call volume.** A 10x price difference is irrelevant at 1,000 calls a day and decisive at 10 million. - **Self-hosting hardware.** If you are deploying the weights yourself, the GPU count the model needs is a hard constraint that outranks quality preference. A common and effective pattern is tiering: route the bulk of traffic to a small model and escalate to the flagship on a confidence signal or an explicit hard-case classifier. ## Why the date suffix exists Cohere's IDs encode a snapshot: `command-r-08-2024` is a specific set of weights that does not change. Undated aliases like `command-r` and `command-r-plus` are conveniences that resolve to the most recent snapshot of that line — and "most recent" moves when Cohere ships. That mutability is the whole point of the interview question. A pipeline on an alias can change behaviour with no deploy on your side: prompts that were tuned against one snapshot drift, few-shot formats that were reliably followed start being reinterpreted, JSON that always parsed occasionally does not, and your regression suite — if it asserts exact strings — either breaks loudly or, worse, passes while quality slides. Nothing in your change log explains it, because nothing in your repository changed. So the rule is: **pin the dated ID in every environment that matters**, and make a model upgrade a deliberate, reviewed change. The upgrade procedure is the same shape as a dependency bump — read the release notes, run your eval set against the new snapshot, compare on the metrics you actually care about (task accuracy, refusal rate, output-format validity, latency, cost per request), shadow or canary a slice of traffic, then roll forward with the old ID one config flip away. Aliases still have a place: exploratory notebooks, demos, and internal tools where drift is cheap and staying current is convenient. Just do not let one leak into a production config file. ## Deprecation is the other half Pinning trades one risk for another: snapshots are eventually retired. That makes model IDs an inventory problem — know which IDs your services use, subscribe to the vendor's deprecation notices, and keep the eval harness warm enough that an upgrade takes days rather than a quarter. A team that pins and then never revisits ends up doing a forced migration under a deadline, which is the failure mode pinning was supposed to prevent.

  • When would you deliberately use an undated alias?
    In throwaway contexts where drift costs nothing: notebooks, demos, internal experiments, and docs examples that should not go stale. The convenience of always getting the newest snapshot is real there. It stops being acceptable the moment output feeds a user-visible surface, a stored record, or a downstream parser — anything whose behaviour you would have to explain if it changed overnight.
  • Pinning protects you from drift, but what risk does it introduce?
    Silent staleness ending in a forced migration. Pinned snapshots get deprecated on the vendor's schedule, not yours, so a service that pinned two years ago and never re-evaluated faces an unplanned upgrade under a retirement deadline. The mitigation is to treat model IDs as tracked dependencies: inventory which services use which snapshot, watch deprecation notices, and keep the eval harness runnable so an upgrade is a day's work.
  • How do you decide whether a smaller Command model is good enough?
    Measure, do not reason about it. Build a labelled eval set from real traffic, run both models against it, and compare on the metrics that matter to the product — task accuracy, output-format validity, refusal rate, p95 latency and cost per request. If the small model is within tolerance, take it and route only the hard cases up. Anything else is guessing about a difference you can just test.

saying these in an interview costs you the question

  • Naming command and command-light as the current chat family
  • Believing undated aliases are frozen to a fixed snapshot
  • Assuming the flagship's 256K context applies to Command R too
  • Picking the largest model without measuring the small one
  • Pinning a snapshot and never tracking its deprecation

context