What fields does a minimal Mistral chat completions request require?
answer
- only two fields are truly required
- model plus an ordered message array
- system is a message, not a parameter
- choices[0].message.content, plus usage
- max_tokens optional, unlike Anthropic
basics
~10 sMistral's POST /v1/chat/completions needs only two body fields: model (for example mistral-large-latest) and messages, an ordered array of role/content objects. Authentication is a bearer API key; max_tokens, temperature and the rest are optional.
solid answer
~40 sThe endpoint is `POST https://api.mistral.ai/v1/chat/completions` on la Plateforme, authenticated with an `Authorization: Bearer $MISTRAL_API_KEY` header. The body requires `model` and `messages`; everything else has a server-side default. `messages` is an ordered list of objects with a `role` — `system`, `user`, `assistant` or `tool` — and `content`. A system instruction is just a message with `role: "system"` at the front of the array, not a separate top-level parameter as in Anthropic's Messages API. The reply comes back OpenAI-shaped: `choices[0].message.content` holds the text, `choices[0].finish_reason` says why generation stopped, and `usage` reports `prompt_tokens`, `completion_tokens` and `total_tokens`. The call is stateless — the server keeps no conversation, so each turn you resend the whole history including the previous assistant replies.
code
python · 15 linesimport os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.chat.complete(
model="mistral-large-latest",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is a sparse mixture-of-experts model?"},
],
)
print(resp.choices[0].message.content)
print(resp.choices[0].finish_reason, resp.usage.total_tokens)go deeper
Be able to name the two required body fields, model and messages, and to read the reply out of choices[0].message.content. Know that the API key travels in an Authorization bearer header.
Explain the four roles, why the call is stateless and what that costs you per turn, and how to branch on finish_reason before trusting the content. Know that max_tokens is optional here.
Show that you log the returned model snapshot and usage per request for cost attribution and reproducibility, and that you pin dated model snapshots rather than -latest aliases where behaviour changes would break you.
Own the history-management policy: how far back conversations are replayed, when turns get summarised or dropped, and how that trade of input-token spend against answer quality is measured rather than guessed.
## The endpoint and authentication Mistral's hosted platform (la Plateforme) exposes text generation at `POST https://api.mistral.ai/v1/chat/completions`. Authentication is a single header, `Authorization: Bearer $MISTRAL_API_KEY`; there is no per-request signing and no key-in-query-string form. Keys are created in the console and are account-scoped, so treat one as a production secret and keep it out of browser code. The request and response envelopes deliberately mirror the OpenAI chat-completions shape, which is why most OpenAI-oriented client libraries can be pointed at Mistral by changing the base URL. That similarity is a convenience, not a contract: the parameters that are genuinely Mistral's own (`safe_prompt`, the assistant-message `prefix` flag) have no OpenAI counterpart, and OpenAI extensions that Mistral has never implemented are simply not available. ## The two required fields `model` is a string naming the model to run, such as `mistral-large-latest` or `mistral-small-latest`. The `-latest` suffixed aliases float to the newest snapshot of that tier, while dated snapshot names pin behaviour; pin in production if reproducibility matters more than free upgrades. `messages` is an ordered JSON array. Each element carries a `role` and `content`. Order is meaningful — it *is* the conversation — and the array must end with something the model can continue from, normally a `user` message. ## The roles - `system` — standing instructions: persona, format rules, refusal policy. In Mistral's API this is an ordinary message placed first in `messages`. This is a real shape difference from Anthropic's Messages API, where `system` is a top-level request parameter, and it is a frequent porting bug. - `user` — the human turn, or whatever your application is feeding in. - `assistant` — a previous model reply that you are replaying as history. A trailing assistant message is also the hook for prefix continuation. - `tool` — the result of a function call being handed back to the model. ## Optional parameters worth knowing `max_tokens` caps the *completion* length and is optional; if you omit it the model generates until it stops naturally or hits the context ceiling. This differs from Anthropic's Messages API, where `max_tokens` is mandatory. `temperature` and `top_p` shape sampling, `stream` switches the response to server-sent events, `stop` takes stop sequences, `random_seed` requests reproducible sampling, `response_format` requests JSON output, and `safe_prompt` toggles Mistral's built-in guardrail instruction. ## The response object A non-streaming success returns: - `id` — the completion id, worth logging for support requests. - `object` — `"chat.completion"`. - `model` — the snapshot that actually served the request, which is how you discover what a `-latest` alias resolved to. - `choices` — an array; each entry has `index`, `message` (with `role: "assistant"` and `content`), and `finish_reason`. - `usage` — `prompt_tokens`, `completion_tokens`, `total_tokens`. This is what billing counts, and it is the number to log per request if you want cost attribution later. `finish_reason` is the field juniors most often ignore. `"stop"` means the model finished on its own. `"length"` means it was cut off at `max_tokens` or the context limit — the text is truncated, and if you were parsing JSON it is now unparseable. `"tool_calls"` means the model wants a function invoked instead of emitting prose. Branch on it before you touch `content`. ## Statelessness There is no server-side thread. Every call is independent, so a multi-turn chat means resending the entire history each time: system message, all prior user turns, all prior assistant replies, then the new user turn. Two consequences follow. First, input tokens grow with conversation length, so cost per turn climbs — trimming or summarising old turns is a real engineering decision, not premature optimisation. Second, if you forget to append the assistant's own reply back into the array, the model loses its memory of what it just said and starts contradicting itself. ## Common mistakes Reading `choices[0].text` (that is the shape of the old legacy completions endpoints, not chat), assuming `max_tokens` is required, putting the system prompt in a top-level field, and parsing `content` without first checking `finish_reason`. Also: the reply is a list because the schema allows multiple candidates, but you will normally be reading index 0.
- Does Mistral remember the conversation between calls?No. The chat completions endpoint is stateless — nothing is stored server-side between requests. Each turn you resend the full `messages` array, including the assistant replies you received earlier. That is why input token count, and therefore cost per turn, grows as a conversation lengthens, and why dropping old turns or summarising them is a deliberate design choice.
- Which finish_reason values should application code branch on?`"stop"` means the model completed normally. `"length"` means it hit `max_tokens` or the context ceiling and the text is truncated — never parse structured output in this case, retry with a higher cap or a shorter prompt. `"tool_calls"` means the model emitted a function call rather than prose, so `content` may be empty and you should read the tool call instead.
- Where does a system instruction go here compared with Anthropic's Messages API?In Mistral's chat completions, the system instruction is an ordinary element of `messages` with `role: "system"`, normally first in the array. Anthropic's Messages API instead takes `system` as a top-level request parameter alongside `messages`. Porting code between the two without moving that field is a common source of silently ignored instructions.
saying these in an interview costs you the question
- Says max_tokens is mandatory, as in Anthropic's Messages API
- Assumes the API stores conversation history server-side
- Puts the system prompt in a top-level system parameter
- Reads generated text from choices[0].text instead of message.content
- Parses the content without ever checking finish_reason