skip to content

When would you pick Mistral Small over Mistral Medium or Large?

level: juniorimportance: should knowfreq 58%

answer

  1. cheapest rung that passes your evals
  2. escalate by exception, not by default
  3. one rung has downloadable weights
  4. route traffic, don't standardise on one tier
  5. measure on your task, not a leaderboard

basics

~20 s

Pick Mistral Small for high-volume, well-defined work — classification, extraction, routing, short generation — where cost and latency dominate and it is accurate enough. Small is also the only tier of the three with freely self-hostable Apache-2.0 weights.

solid answer

~50 s

Mistral's general line-up is a three-rung ladder: **Small** (24B class, cheapest and fastest, Apache-2.0 open weights), **Medium** (mid-tier, API-only — no public weights), and **Large** (flagship reasoning, weights available only under a research licence). Cost and latency fall sharply as you go down the ladder, capability rises as you go up. The working method is to start at Small with a representative evaluation set, measure task accuracy, and escalate only where Small measurably fails. Most production traffic — routing, tagging, extraction, summarising short documents, tool-argument filling — sits comfortably at Small, and pushing it to Large multiplies cost for no measurable win. Two tier facts drive architecture decisions beyond quality: Small is the rung you can lawfully self-host and freeze, and Medium's API-only status means committing to it is committing to the vendor. On la Plateforme the IDs are `mistral-small-latest`, `mistral-medium-latest` and `mistral-large-latest`, with dated snapshots behind each alias.

go deeper

for a junior

Know the ladder — Small, Medium, Large — and be able to say that you start with the smallest model that solves the task and move up only when it demonstrably fails.

for a middle

Explain the cost and latency gradient, that Medium is API-only while Small has open weights, and that production should pin dated model snapshots rather than -latest aliases.

for a senior

Show the escalation method: a real evaluation set, prompt and schema fixes before tier changes, and routing so only the hard slice of traffic reaches an expensive tier.

for a principal

Own the lock-in dimension — picking the open Small rung preserves a self-hosting exit and data-residency options, while standardising on API-only tiers is a vendor commitment with cost and deprecation exposure.

## The ladder Mistral's general-purpose models form a deliberate three-rung ladder, and each rung differs on three axes at once: capability, price/latency, and how you are allowed to obtain it. **Mistral Small** is the entry rung — a 24B-class dense model, published under Apache 2.0 from Small 3 onward, with multimodal image input in later revisions and a long context window. It is the cheapest per token on la Plateforme, the fastest to first token, and the only rung of the three you can download and run on your own hardware for commercial use. **Mistral Medium** sits in the middle: stronger than Small on reasoning and instruction-following at a mid-tier price. Its defining operational fact is that it is **API-only** — Mistral has not released weights for it, so there is no self-hosting path at all. **Mistral Large** is the flagship for the hardest reasoning, longest instructions and most complex tool orchestration. Weights for flagship releases have been published under the research licence, meaning you can inspect and benchmark them but not deploy them commercially without a negotiated agreement; the supported commercial path is the API. Specialists sit off the ladder rather than on it: Codestral for code completion, Pixtral for vision, `mistral-embed` for embeddings. Reaching for a specialist is a different decision from climbing the ladder. ## How to actually choose The answer interviewers want is a method, not a rule of thumb. 1. **Build an evaluation set first** — 50-200 real inputs from your workload with known-good outputs or a grading rubric. Without it, every tier discussion is vibes. 2. **Start at Small.** It is the cheapest experiment and often sufficient. Measure task success, not general benchmark scores; a model that scores lower on a public leaderboard can win on your extraction task. 3. **Escalate only on measured failure**, and escalate one rung. Many teams jump straight from Small to Large and never discover Medium was enough. 4. **Try prompt and structure changes before a tier change.** Better instructions, few-shot examples, schema-constrained output or splitting one hard call into two easy ones frequently recovers the gap at a fraction of the cost. 5. **Route rather than standardise.** Production systems commonly run Small for the bulk of traffic and escalate specific request classes — or specific low-confidence outputs — to a higher tier. That keeps the average cost near Small's while keeping the tail quality near Large's. ## The cost shape Per-token prices differ by roughly an order of magnitude between the bottom and top of the ladder, and output tokens cost more than input tokens at every rung. Two consequences: first, on high-volume classification-style workloads the tier choice dominates your bill far more than any prompt optimisation; second, a chatty verbose model at a high tier is doubly expensive, so capping output length matters most exactly where tokens are dearest. Latency follows the same ordering. For interactive UX — autocomplete, inline suggestions, streaming chat first-token time — the Small rung is often chosen for responsiveness even where a larger model would answer slightly better. ## The licensing axis people forget Because only Small has permissive weights, the tier decision is also a lock-in decision. Choosing Small keeps a self-hosting escape hatch open: air-gapped or data-residency-constrained deployments, a version you can freeze indefinitely against vendor deprecation, and no per-token bill. Choosing Medium or Large is choosing an API relationship, whether directly with Mistral or through a cloud marketplace that resells its models. ## Model IDs and pinning On la Plateforme the tiers are addressed as `mistral-small-latest`, `mistral-medium-latest` and `mistral-large-latest`, each an alias that follows the newest snapshot of that tier, with dated snapshot IDs behind them. Open models have their own IDs such as `open-mistral-7b`, `open-mixtral-8x7b` and `open-mixtral-8x22b`. In production, pin the dated snapshot rather than the `-latest` alias: an alias that silently rolls forward can change your outputs, your token costs and your prompt's behaviour without a deploy on your side. Use `-latest` in development, pinned snapshots in production, and re-run your evaluation set before moving a pin. ## A weak answer versus a strong one A weak answer ranks the tiers by size and stops. A strong one says: measure on your own task, start low, escalate by exception, route mixed traffic, and be aware that only the bottom rung leaves the self-hosting door open.

  • Would you point production traffic at mistral-large-latest?
    No — pin a dated snapshot instead. A `-latest` alias rolls forward to a new model version without any deploy on your side, which can change output style, quality and token consumption mid-flight. Use the alias while developing, pin the dated snapshot in production, and move the pin deliberately after re-running your evaluation set against the new version.
  • Small fails on about five percent of your requests. What do you do before moving everything to a higher tier?
    Route rather than upgrade wholesale. Identify what distinguishes the failing five percent — longer inputs, harder reasoning, rarer formats — and escalate only that class, or escalate on a confidence or validation signal such as a failed schema parse. You keep near-Small average cost with near-flagship tail quality. Also retry prompt structure and output-schema constraints first; those often close the gap for free.
  • Why might a team choose Mistral Small even where Medium is measurably more accurate?
    Because Small is the only one of the two with Apache-2.0 weights. That buys self-hosting for data-residency or air-gapped requirements, a version you can freeze against vendor deprecation, fine-tuning you own outright, and a fixed infrastructure cost instead of per-token billing. Medium is API-only, so adopting it forecloses all of those options.

saying these in an interview costs you the question

  • Defaulting to the flagship tier for every workload
  • Assuming Medium weights can be downloaded and self-hosted
  • Choosing a tier from public leaderboards instead of your own evals
  • Pointing production at a -latest alias that silently rolls forward
  • Treating the tier ladder as the only lever, ignoring prompt and schema fixes

context