skip to content

Chat completions API

Calling Mistral for text: the chat completions endpoint on la Plateforme is close enough to OpenAI's shape that most SDKs drop straight in. What you have to learn are the few parameters that are genuinely Mistral's own.

on this pageshow

questions

5

What fields does a minimal Mistral chat completions request require?

level: juniorimportance: must knowfreq 72%

answer

  1. only two fields are truly required
  2. model plus an ordered message array
  3. system is a message, not a parameter
  4. choices[0].message.content, plus usage
  5. max_tokens optional, unlike Anthropic

basics

~10 s

Mistral's POST /v1/chat/completions needs only two body fields: model (for example mistral-large-latest) and messages, an ordered array of role/content objects. Authentication is a bearer API key; max_tokens, temperature and the rest are optional.

solid answer

~40 s

The endpoint is `POST https://api.mistral.ai/v1/chat/completions` on la Plateforme, authenticated with an `Authorization: Bearer $MISTRAL_API_KEY` header. The body requires `model` and `messages`; everything else has a server-side default. `messages` is an ordered list of objects with a `role` — `system`, `user`, `assistant` or `tool` — and `content`. A system instruction is just a message with `role: "system"` at the front of the array, not a separate top-level parameter as in Anthropic's Messages API. The reply comes back OpenAI-shaped: `choices[0].message.content` holds the text, `choices[0].finish_reason` says why generation stopped, and `usage` reports `prompt_tokens`, `completion_tokens` and `total_tokens`. The call is stateless — the server keeps no conversation, so each turn you resend the whole history including the previous assistant replies.

code

python · 15 lines
python
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

resp = client.chat.complete(
    model="mistral-large-latest",
    messages=[
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is a sparse mixture-of-experts model?"},
    ],
)

print(resp.choices[0].message.content)
print(resp.choices[0].finish_reason, resp.usage.total_tokens)

go deeper

for a junior

Be able to name the two required body fields, model and messages, and to read the reply out of choices[0].message.content. Know that the API key travels in an Authorization bearer header.

for a middle

Explain the four roles, why the call is stateless and what that costs you per turn, and how to branch on finish_reason before trusting the content. Know that max_tokens is optional here.

for a senior

Show that you log the returned model snapshot and usage per request for cost attribution and reproducibility, and that you pin dated model snapshots rather than -latest aliases where behaviour changes would break you.

for a principal

Own the history-management policy: how far back conversations are replayed, when turns get summarised or dropped, and how that trade of input-token spend against answer quality is measured rather than guessed.

## The endpoint and authentication Mistral's hosted platform (la Plateforme) exposes text generation at `POST https://api.mistral.ai/v1/chat/completions`. Authentication is a single header, `Authorization: Bearer $MISTRAL_API_KEY`; there is no per-request signing and no key-in-query-string form. Keys are created in the console and are account-scoped, so treat one as a production secret and keep it out of browser code. The request and response envelopes deliberately mirror the OpenAI chat-completions shape, which is why most OpenAI-oriented client libraries can be pointed at Mistral by changing the base URL. That similarity is a convenience, not a contract: the parameters that are genuinely Mistral's own (`safe_prompt`, the assistant-message `prefix` flag) have no OpenAI counterpart, and OpenAI extensions that Mistral has never implemented are simply not available. ## The two required fields `model` is a string naming the model to run, such as `mistral-large-latest` or `mistral-small-latest`. The `-latest` suffixed aliases float to the newest snapshot of that tier, while dated snapshot names pin behaviour; pin in production if reproducibility matters more than free upgrades. `messages` is an ordered JSON array. Each element carries a `role` and `content`. Order is meaningful — it *is* the conversation — and the array must end with something the model can continue from, normally a `user` message. ## The roles - `system` — standing instructions: persona, format rules, refusal policy. In Mistral's API this is an ordinary message placed first in `messages`. This is a real shape difference from Anthropic's Messages API, where `system` is a top-level request parameter, and it is a frequent porting bug. - `user` — the human turn, or whatever your application is feeding in. - `assistant` — a previous model reply that you are replaying as history. A trailing assistant message is also the hook for prefix continuation. - `tool` — the result of a function call being handed back to the model. ## Optional parameters worth knowing `max_tokens` caps the *completion* length and is optional; if you omit it the model generates until it stops naturally or hits the context ceiling. This differs from Anthropic's Messages API, where `max_tokens` is mandatory. `temperature` and `top_p` shape sampling, `stream` switches the response to server-sent events, `stop` takes stop sequences, `random_seed` requests reproducible sampling, `response_format` requests JSON output, and `safe_prompt` toggles Mistral's built-in guardrail instruction. ## The response object A non-streaming success returns: - `id` — the completion id, worth logging for support requests. - `object` — `"chat.completion"`. - `model` — the snapshot that actually served the request, which is how you discover what a `-latest` alias resolved to. - `choices` — an array; each entry has `index`, `message` (with `role: "assistant"` and `content`), and `finish_reason`. - `usage` — `prompt_tokens`, `completion_tokens`, `total_tokens`. This is what billing counts, and it is the number to log per request if you want cost attribution later. `finish_reason` is the field juniors most often ignore. `"stop"` means the model finished on its own. `"length"` means it was cut off at `max_tokens` or the context limit — the text is truncated, and if you were parsing JSON it is now unparseable. `"tool_calls"` means the model wants a function invoked instead of emitting prose. Branch on it before you touch `content`. ## Statelessness There is no server-side thread. Every call is independent, so a multi-turn chat means resending the entire history each time: system message, all prior user turns, all prior assistant replies, then the new user turn. Two consequences follow. First, input tokens grow with conversation length, so cost per turn climbs — trimming or summarising old turns is a real engineering decision, not premature optimisation. Second, if you forget to append the assistant's own reply back into the array, the model loses its memory of what it just said and starts contradicting itself. ## Common mistakes Reading `choices[0].text` (that is the shape of the old legacy completions endpoints, not chat), assuming `max_tokens` is required, putting the system prompt in a top-level field, and parsing `content` without first checking `finish_reason`. Also: the reply is a list because the schema allows multiple candidates, but you will normally be reading index 0.

  • Does Mistral remember the conversation between calls?
    No. The chat completions endpoint is stateless — nothing is stored server-side between requests. Each turn you resend the full `messages` array, including the assistant replies you received earlier. That is why input token count, and therefore cost per turn, grows as a conversation lengthens, and why dropping old turns or summarising them is a deliberate design choice.
  • Which finish_reason values should application code branch on?
    `"stop"` means the model completed normally. `"length"` means it hit `max_tokens` or the context ceiling and the text is truncated — never parse structured output in this case, retry with a higher cap or a shorter prompt. `"tool_calls"` means the model emitted a function call rather than prose, so `content` may be empty and you should read the tool call instead.
  • Where does a system instruction go here compared with Anthropic's Messages API?
    In Mistral's chat completions, the system instruction is an ordinary element of `messages` with `role: "system"`, normally first in the array. Anthropic's Messages API instead takes `system` as a top-level request parameter alongside `messages`. Porting code between the two without moving that field is a common source of silently ignored instructions.

saying these in an interview costs you the question

  • Says max_tokens is mandatory, as in Anthropic's Messages API
  • Assumes the API stores conversation history server-side
  • Puts the system prompt in a top-level system parameter
  • Reads generated text from choices[0].text instead of message.content
  • Parses the content without ever checking finish_reason

context

open as a page

How do you assemble a full reply from Mistral's streamed chat completion chunks?

level: middleimportance: must knowfreq 58%

basics

~20 s

Set stream true on Mistral's chat completions request and the reply arrives as chat.completion.chunk objects over server-sent events. Concatenate choices[0].delta.content across chunks in arrival order; the last chunk carries finish_reason, and the stream closes with data: [DONE].

open as a page

In Mistral's chat completions, what does response_format json_object guarantee?

level: middleimportance: should knowfreq 50%

basics

~20 s

Only syntactic validity. Setting response_format to {"type": "json_object"} on a Mistral chat completions request constrains the output to parseable JSON, but not to your field names, types or required keys — and truncation at max_tokens can still leave it unparseable.

open as a page

What does Mistral's safe_prompt flag change about a chat completions call?

level: middleimportance: should knowfreq 42%

basics

~20 s

safe_prompt is a boolean on Mistral's chat completions request, off by default. When true, Mistral prepends a fixed safety instruction to the conversation, steering the model away from harmful, unethical or prejudiced output. It is guardrail prompting, not a classifier.

open as a page

How does Mistral's assistant-message prefix flag steer a completion's opening?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Adding prefix: true to a trailing assistant message in Mistral's chat completions request makes the model continue that text rather than start a fresh reply. It pins the opening tokens — an opening brace, a language, a persona, a resumed truncation.

open as a page