skip to content

Mistral

Mistral is the European open-weight-plus-commercial option: a compact model line-up you can either call over an OpenAI-shaped API or download and run yourself. Learn it for the mixture-of-experts questions and for the "what if we self-host?" branch of any vendor-choice discussion.

on this pageshow

explore

questions

19

What fields does a minimal Mistral chat completions request require?

level: juniorimportance: must knowfreq 72%

answer

  1. only two fields are truly required
  2. model plus an ordered message array
  3. system is a message, not a parameter
  4. choices[0].message.content, plus usage
  5. max_tokens optional, unlike Anthropic

basics

~10 s

Mistral's POST /v1/chat/completions needs only two body fields: model (for example mistral-large-latest) and messages, an ordered array of role/content objects. Authentication is a bearer API key; max_tokens, temperature and the rest are optional.

solid answer

~40 s

The endpoint is `POST https://api.mistral.ai/v1/chat/completions` on la Plateforme, authenticated with an `Authorization: Bearer $MISTRAL_API_KEY` header. The body requires `model` and `messages`; everything else has a server-side default. `messages` is an ordered list of objects with a `role` — `system`, `user`, `assistant` or `tool` — and `content`. A system instruction is just a message with `role: "system"` at the front of the array, not a separate top-level parameter as in Anthropic's Messages API. The reply comes back OpenAI-shaped: `choices[0].message.content` holds the text, `choices[0].finish_reason` says why generation stopped, and `usage` reports `prompt_tokens`, `completion_tokens` and `total_tokens`. The call is stateless — the server keeps no conversation, so each turn you resend the whole history including the previous assistant replies.

code

python · 15 lines
python
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

resp = client.chat.complete(
    model="mistral-large-latest",
    messages=[
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is a sparse mixture-of-experts model?"},
    ],
)

print(resp.choices[0].message.content)
print(resp.choices[0].finish_reason, resp.usage.total_tokens)

go deeper

for a junior

Be able to name the two required body fields, model and messages, and to read the reply out of choices[0].message.content. Know that the API key travels in an Authorization bearer header.

for a middle

Explain the four roles, why the call is stateless and what that costs you per turn, and how to branch on finish_reason before trusting the content. Know that max_tokens is optional here.

for a senior

Show that you log the returned model snapshot and usage per request for cost attribution and reproducibility, and that you pin dated model snapshots rather than -latest aliases where behaviour changes would break you.

for a principal

Own the history-management policy: how far back conversations are replayed, when turns get summarised or dropped, and how that trade of input-token spend against answer quality is measured rather than guessed.

## The endpoint and authentication Mistral's hosted platform (la Plateforme) exposes text generation at `POST https://api.mistral.ai/v1/chat/completions`. Authentication is a single header, `Authorization: Bearer $MISTRAL_API_KEY`; there is no per-request signing and no key-in-query-string form. Keys are created in the console and are account-scoped, so treat one as a production secret and keep it out of browser code. The request and response envelopes deliberately mirror the OpenAI chat-completions shape, which is why most OpenAI-oriented client libraries can be pointed at Mistral by changing the base URL. That similarity is a convenience, not a contract: the parameters that are genuinely Mistral's own (`safe_prompt`, the assistant-message `prefix` flag) have no OpenAI counterpart, and OpenAI extensions that Mistral has never implemented are simply not available. ## The two required fields `model` is a string naming the model to run, such as `mistral-large-latest` or `mistral-small-latest`. The `-latest` suffixed aliases float to the newest snapshot of that tier, while dated snapshot names pin behaviour; pin in production if reproducibility matters more than free upgrades. `messages` is an ordered JSON array. Each element carries a `role` and `content`. Order is meaningful — it *is* the conversation — and the array must end with something the model can continue from, normally a `user` message. ## The roles - `system` — standing instructions: persona, format rules, refusal policy. In Mistral's API this is an ordinary message placed first in `messages`. This is a real shape difference from Anthropic's Messages API, where `system` is a top-level request parameter, and it is a frequent porting bug. - `user` — the human turn, or whatever your application is feeding in. - `assistant` — a previous model reply that you are replaying as history. A trailing assistant message is also the hook for prefix continuation. - `tool` — the result of a function call being handed back to the model. ## Optional parameters worth knowing `max_tokens` caps the *completion* length and is optional; if you omit it the model generates until it stops naturally or hits the context ceiling. This differs from Anthropic's Messages API, where `max_tokens` is mandatory. `temperature` and `top_p` shape sampling, `stream` switches the response to server-sent events, `stop` takes stop sequences, `random_seed` requests reproducible sampling, `response_format` requests JSON output, and `safe_prompt` toggles Mistral's built-in guardrail instruction. ## The response object A non-streaming success returns: - `id` — the completion id, worth logging for support requests. - `object` — `"chat.completion"`. - `model` — the snapshot that actually served the request, which is how you discover what a `-latest` alias resolved to. - `choices` — an array; each entry has `index`, `message` (with `role: "assistant"` and `content`), and `finish_reason`. - `usage` — `prompt_tokens`, `completion_tokens`, `total_tokens`. This is what billing counts, and it is the number to log per request if you want cost attribution later. `finish_reason` is the field juniors most often ignore. `"stop"` means the model finished on its own. `"length"` means it was cut off at `max_tokens` or the context limit — the text is truncated, and if you were parsing JSON it is now unparseable. `"tool_calls"` means the model wants a function invoked instead of emitting prose. Branch on it before you touch `content`. ## Statelessness There is no server-side thread. Every call is independent, so a multi-turn chat means resending the entire history each time: system message, all prior user turns, all prior assistant replies, then the new user turn. Two consequences follow. First, input tokens grow with conversation length, so cost per turn climbs — trimming or summarising old turns is a real engineering decision, not premature optimisation. Second, if you forget to append the assistant's own reply back into the array, the model loses its memory of what it just said and starts contradicting itself. ## Common mistakes Reading `choices[0].text` (that is the shape of the old legacy completions endpoints, not chat), assuming `max_tokens` is required, putting the system prompt in a top-level field, and parsing `content` without first checking `finish_reason`. Also: the reply is a list because the schema allows multiple candidates, but you will normally be reading index 0.

  • Does Mistral remember the conversation between calls?
    No. The chat completions endpoint is stateless — nothing is stored server-side between requests. Each turn you resend the full `messages` array, including the assistant replies you received earlier. That is why input token count, and therefore cost per turn, grows as a conversation lengthens, and why dropping old turns or summarising them is a deliberate design choice.
  • Which finish_reason values should application code branch on?
    `"stop"` means the model completed normally. `"length"` means it hit `max_tokens` or the context ceiling and the text is truncated — never parse structured output in this case, retry with a higher cap or a shorter prompt. `"tool_calls"` means the model emitted a function call rather than prose, so `content` may be empty and you should read the tool call instead.
  • Where does a system instruction go here compared with Anthropic's Messages API?
    In Mistral's chat completions, the system instruction is an ordinary element of `messages` with `role: "system"`, normally first in the array. Anthropic's Messages API instead takes `system` as a top-level request parameter alongside `messages`. Porting code between the two without moving that field is a common source of silently ignored instructions.

saying these in an interview costs you the question

  • Says max_tokens is mandatory, as in Anthropic's Messages API
  • Assumes the API stores conversation history server-side
  • Puts the system prompt in a top-level system parameter
  • Reads generated text from choices[0].text instead of message.content
  • Parses the content without ever checking finish_reason

context

open as a page

How do you call Mistral's /v1/embeddings endpoint, and what does mistral-embed return?

level: juniorimportance: must knowfreq 58%

basics

~20 s

POST /v1/embeddings on api.mistral.ai with a bearer API key, model "mistral-embed" and an "input" array of strings. The response carries a data array of 1024-dimension float vectors, each tagged with the index of the string it came from, plus a token usage block.

open as a page

In Mistral's chat completions response, which fields carry a tool call?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Mistral sets choices[0].finish_reason to "tool_calls" and fills choices[0].message.tool_calls. Each entry has an id, a function.name, and function.arguments — a JSON-encoded string you must parse before running your code.

open as a page

How do you assemble a full reply from Mistral's streamed chat completion chunks?

level: middleimportance: must knowfreq 58%

basics

~20 s

Set stream true on Mistral's chat completions request and the reply arrives as chat.completion.chunk objects over server-sent events. Concatenate choices[0].delta.content across chunks in arrival order; the last chunk carries finish_reason, and the stream closes with data: [DONE].

open as a page

Porting tool_choice "required" from OpenAI to Mistral — which value do you use?

level: middleimportance: must knowfreq 58%

basics

~20 s

Use "any". Mistral's documented tool_choice values are auto (the default when tools are present), any (the model must call a tool), and none (tools stay declared but unused); "required" is OpenAI's spelling for the any mode.

open as a page

Which Mistral models are Apache-2.0 and which need a commercial licence?

level: middleimportance: must knowfreq 66%

basics

~20 s

Mistral runs two licence tracks. Apache-2.0 releases — Mistral 7B, Mixtral 8x7B and 8x22B, Mistral NeMo, Mistral Small, Pixtral 12B — are free to self-host commercially. Flagship and specialist weights ship research-only or non-production and need a paid licence.

open as a page

When would you pick Mistral Small over Mistral Medium or Large?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Pick Mistral Small for high-volume, well-defined work — classification, extraction, routing, short generation — where cost and latency dominate and it is accurate enough. Small is also the only tier of the three with freely self-hostable Apache-2.0 weights.

open as a page

In Mistral's chat completions, what does response_format json_object guarantee?

level: middleimportance: should knowfreq 50%

basics

~20 s

Only syntactic validity. Setting response_format to {"type": "json_object"} on a Mistral chat completions request constrains the output to parseable JSON, but not to your field names, types or required keys — and truncation at max_tokens can still leave it unparseable.

open as a page

What does Mistral's safe_prompt flag change about a chat completions call?

level: middleimportance: should knowfreq 42%

basics

~20 s

safe_prompt is a boolean on Mistral's chat completions request, off by default. When true, Mistral prepends a fixed safety instruction to the conversation, steering the model away from harmful, unethical or prejudiced output. It is guardrail prompting, not a classifier.

open as a page

In Mistral's FIM completions endpoint, what do prompt and suffix do?

level: middleimportance: should knowfreq 42%

basics

~20 s

Mistral's POST /v1/fim/completions takes prompt as the code before the cursor and suffix as the code after it; the model generates only the middle span that joins them. It is a Codestral-only endpoint, separate from chat completions, built for IDE-style inline completion.

open as a page

What does Mistral's /v1/ocr endpoint return for a multi-page PDF?

level: middleimportance: should knowfreq 32%

basics

~20 s

Mistral's document OCR route returns a pages array — one entry per page, each with the page's extracted content as Markdown, its index, page dimensions, and optionally the embedded figures as base64 images. Usage is reported and billed per page processed, not per token.

open as a page

Why does Mixtral 8x7B total 46.7B parameters instead of 56B?

level: middleimportance: should knowfreq 52%

basics

~20 s

The 8x7B name is a naming convention, not multiplication. Only the feed-forward blocks are replicated into 8 experts; attention and embedding weights are shared across them, so Mixtral 8x7B holds 46.7B total parameters and uses about 12.9B per token.

open as a page

How does Mistral's assistant-message prefix flag steer a completion's opening?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Adding prefix: true to a trailing assistant message in Mistral's chat completions request makes the model continue that text rather than start a fresh reply. It pins the opening tokens — an opening brace, a language, a persona, a resumed truncation.

open as a page

When would you run Mistral embeddings through the batch jobs API instead of /v1/embeddings?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Use batch jobs for large offline work with no latency requirement — a corpus backfill or re-embed. You upload a JSONL file of requests, create a job naming the target endpoint, and collect an output file later, at roughly half the synchronous price and without hammering per-minute rate limits.

open as a page

With Mistral tool_choice "any" on every turn, why does your agent never finish?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Because "any" forces a tool call on every request. The model can never return a plain text answer, so each round trip yields another call and the loop has no natural exit. Flip back to "auto" after the forced turn.

open as a page

Your Mistral tool loop 400s after tool_calls — what is wrong with the history?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Almost always a broken pairing in the messages array: the assistant message carrying tool_calls was not re-sent verbatim, a tool_call_id does not match one the server issued, or some of the returned calls got no role "tool" reply at all.

open as a page

Can you ship Mistral's Codestral weights inside a commercial product?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Not from the open weights alone. Codestral's downloadable weights carry the Mistral Non-Production License, which permits development, testing and research but excludes production and commercial use. Commercial paths are a paid Mistral licence, the hosted Codestral API, or an Apache-2.0 model instead.

open as a page

When is self-hosting Mistral's open weights better than calling la Plateforme?

level: principalimportance: should knowfreq 42%

basics

~20 s

Self-hosting wins when data must not leave your network, when sustained volume makes fixed GPU cost cheaper than per-token billing, or when you need a frozen version immune to vendor deprecation — and only for the Apache-2.0 tier, which caps the capability you can lawfully run.

open as a page

Why does Codestral offer both codestral.mistral.ai and the la Plateforme endpoint?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

They are two front doors with different keys and different billing. codestral.mistral.ai is the dedicated code endpoint aimed at IDE-style completion and takes a Codestral API key; api.mistral.ai is the general la Plateforme endpoint using your normal workspace key and standard per-token billing. Keys are not interchangeable.

open as a page