In Mistral's chat completions, what does response_format json_object guarantee?
answer
- it constrains grammar, not meaning
- parseable is not the same as correct
- the prompt still owns the field names
- check why generation stopped first
- schema-bound mode is the stronger option
basics
~20 sOnly syntactic validity. Setting response_format to {"type": "json_object"} on a Mistral chat completions request constrains the output to parseable JSON, but not to your field names, types or required keys — and truncation at max_tokens can still leave it unparseable.
solid answer
~50 s`response_format: {"type": "json_object"}` puts the model in JSON mode: the reply is a JSON document rather than prose wrapped in a markdown fence. That is a syntax guarantee, not a schema guarantee — the model can still invent field names, nest differently, return a string where you wanted a number, or omit a key you consider required, and json_object mode will happily produce valid JSON that your parser accepts and your domain code then chokes on. Two practical rules follow. You must still describe the shape you want in the prompt, ideally with an example object, because the mode constrains grammar and your prompt constrains semantics. And you must check `finish_reason` before parsing: a `"length"` truncation cuts the document mid-object, so a parse failure there is a token-budget bug, not a model failure. When you need real field-level guarantees, use Mistral's schema-bound structured-output form and still validate on receipt.
code
python · 25 linesimport json, os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.chat.complete(
model="mistral-large-latest",
messages=[{
"role": "user",
"content": (
"Extract the invoice. Reply with keys "
"invoice_id (string), total (number). "
'Example: {"invoice_id": "A-1", "total": 12.5}'
),
}],
response_format={"type": "json_object"},
max_tokens=512,
)
choice = resp.choices[0]
if choice.finish_reason == "length":
raise RuntimeError("truncated: raise max_tokens")
data = json.loads(choice.message.content)
assert "invoice_id" in data and "total" in datago deeper
Know that setting response_format to json_object makes Mistral return a JSON document instead of prose with a markdown fence, and that you still parse it yourself.
Explain that the guarantee is syntactic only — key names, types and required fields are not enforced — and that you must describe the shape in the prompt and check finish_reason before parsing.
Show the production pattern: generous max_tokens, finish_reason branch, defensive parsing plus schema validation on receipt, and raw-output logging so field drift after a model change is detectable.
Own where the contract lives. Decide when a schema-bound structured-output mode is worth its rigidity versus json_object plus validation, and make model-upgrade regression testing of output shape a standing requirement rather than an incident response.
## What the parameter is `response_format` is a request-body object on Mistral's `POST /v1/chat/completions`. Setting it to `{"type": "json_object"}` switches on JSON mode. In the Python SDK it is passed as `response_format={"type": "json_object"}` alongside `model` and `messages`. Omit it and you get free-form text, which for a JSON-shaped instruction usually means an explanatory sentence, a ```json fence, the object, and a closing remark — three of which your parser has to strip. ## The guarantee, stated precisely JSON mode guarantees **the output parses as JSON**. That is the entire promise. It is a grammar constraint applied during decoding: at each step the tokens that could not continue a valid JSON document are excluded, so the model cannot emit a stray preamble, a markdown fence, or a trailing comma. What it does not constrain is everything semantic: - **Field names.** Ask for `user_id` and you may get `userId`, `id`, or `user`. - **Types.** `"42"` instead of `42`, `"true"` instead of `true`. - **Shape.** A bare object where you wanted an array of them, or an extra wrapper key like `{"result": {...}}`. - **Required keys.** A field you consider mandatory can simply be absent. - **Values.** JSON mode has nothing at all to say about whether the content is correct. A perfectly-formed object full of fabricated values is a complete success by this parameter's standard. The failure mode this creates is nastier than plain-text parsing failure, because it is silent. Unfenced prose fails loudly at `json.loads`. A valid object with `userId` instead of `user_id` sails through the parser and fails three layers down, in code that assumed the key existed. ## You still have to prompt for the shape Because the mode constrains grammar and your prompt constrains meaning, JSON mode does not free you from describing the output. In practice: state the exact keys, state the type of each, and include one small example object in the system or user message. The example is worth more than the description — it fixes casing, nesting and cardinality in one shot. A useful companion trick on Mistral specifically is prefix continuation: ending the messages array with an assistant message flagged as a prefix and containing `{` removes even the possibility of a preamble. It composes with JSON mode rather than replacing it. ## The truncation trap The single most common production incident with JSON mode is not a model failure at all. If generation hits `max_tokens`, the document is cut mid-object and the response is unparseable — and `finish_reason` will say `"length"`, not `"stop"`. So the correct receiving code is: read `finish_reason` first. If it is `"length"`, do not attempt to parse — retry with a larger cap, or reduce what you asked for. Only on `"stop"` does a parse failure indicate something interesting. Teams that skip this check spend a long time blaming the model for what is a budget setting, and it gets worse as prompts grow, because longer inputs push toward longer outputs. ## When you need real guarantees When field-level adherence actually matters, Mistral also supports a schema-bound structured-output form of `response_format`, where you supply a JSON Schema and the decoder is constrained to documents matching it. That upgrades the guarantee from "is JSON" to "matches this shape", which eliminates the key-name and type drift that json_object leaves open. It costs you some flexibility — schemas have constraints on what constructs are supported, and a schema that is too elaborate is harder for the model to fill sensibly than a flat one. Even then, validate on receipt. Schema conformance does not imply semantic correctness, and defensive parsing at the boundary is cheap. ## Practical checklist 1. Set `response_format` to json_object (or the schema-bound form when you need shape guarantees). 2. Describe the exact keys and types in the prompt, and include an example object. 3. Give `max_tokens` real headroom for the largest plausible document. 4. On response, check `finish_reason` before parsing. 5. Parse defensively, then validate against your own model — Pydantic, a JSON Schema validator, whatever your stack has. 6. Log validation failures with the raw output. Field-drift regressions after a model upgrade are invisible without them. ## The interview framing What is being probed is whether you understand that constrained decoding operates on grammar, not on meaning. A candidate who says "JSON mode means I get my object" has skipped the distinction. A candidate who says "it means it parses; I still validate the shape, and I check finish_reason before I even try" has the right model of the feature.
- Your JSON-mode call returned unparseable output. What do you check first?`finish_reason`. If it is `"length"`, generation was cut off at max_tokens and the document is truncated mid-object — a budget problem, not a model problem, fixed by raising the cap or asking for less. Only when it is `"stop"` is an unparseable body worth investigating further.
- Do you still need to describe the schema in the prompt when json_object mode is on?Yes. The mode constrains the output to valid JSON; nothing about it tells the model which keys you want. Without an explicit description — best paired with one small example object — you get valid JSON with drifting key casing, wrapper objects and inconsistent types across calls.
- How does prefix continuation interact with JSON mode on Mistral?They compose. JSON mode blocks non-JSON tokens during decoding, while ending the messages array with an assistant prefix message containing `{` pins the opening explicitly. Adding a `}` stop sequence bounds the other end. Belt and braces for a single flat object, at the cost of some assembly care on receipt.
saying these in an interview costs you the question
- Claims json_object mode enforces your field names and types
- Skips describing the desired shape in the prompt
- Parses the body without checking finish_reason first
- Blames the model when a truncated document fails to parse
- Treats valid JSON as evidence the values are correct