skip to content

How does Mistral's assistant-message prefix flag steer a completion's opening?

level: seniorimportance: should knowfreq 33%

answer

  1. the last message is the assistant's, not the user's
  2. you write the opening, the model continues
  3. kills the preamble before the JSON
  4. best paired with stop sequences
  5. never build it from user input

basics

~20 s

Adding prefix: true to a trailing assistant message in Mistral's chat completions request makes the model continue that text rather than start a fresh reply. It pins the opening tokens — an opening brace, a language, a persona, a resumed truncation.

solid answer

~50 s

Normally the last element of `messages` is a `user` turn and the model begins a reply from scratch. With Mistral's prefix continuation you instead end the array with `{"role": "assistant", "content": "...", "prefix": true}`, and the model treats that text as the already-emitted start of its own answer and continues from it. That gives you deterministic control over the opening, which is where most format failures happen: emit `{` to kill the "Sure, here's the JSON" preamble, emit a French opening clause to lock the response language, emit a persona's first words for role-play, or resend a reply that stopped with `finish_reason: "length"` so generation resumes mid-thought. Pair it with `stop` sequences to bound the continuation. The safety caveat is real: you are putting words in the model's mouth, which weakens its alignment behaviour, so never build a prefix from untrusted user input.

code

python · 16 lines
python
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

resp = client.chat.complete(
    model="mistral-small-latest",
    messages=[
        {"role": "user", "content": "Give me three fruits with colours as JSON."},
        {"role": "assistant", "content": "{", "prefix": True},
    ],
    stop=["}"],
    max_tokens=200,
)

print(resp.choices[0].message.content)

go deeper

for a junior

Know that Mistral lets you end the messages array with an assistant message flagged prefix true, and that the model then continues that text instead of starting a new reply.

for a middle

Explain why the opening tokens dominate the rest of the output, and name concrete uses: forcing an opening brace, locking the response language, pairing with stop sequences to bound both ends.

for a senior

Reach for it to resume a length-truncated completion without re-paying for tokens, and be explicit that the prefix is application-controlled text that untrusted input must never reach.

for a principal

Own the trade: format compliance rises while honest abstention falls, and the feature is vendor-specific. Decide whether it belongs behind a capability check in a multi-provider abstraction and what it costs you in portability.

## The mechanism Every chat model is, underneath, continuing a token sequence. A chat template renders your `messages` array into that sequence and then opens an assistant turn, and the model fills the rest. Prefix continuation exposes that seam: you supply the first part of the assistant turn yourself, and the model resumes from your text instead of from an empty turn. In Mistral's chat completions API you do this by making the final element of `messages` an assistant message carrying `"prefix": true`: - normal call: `[..., {"role": "user", "content": "List three fruits as JSON"}]` - prefix call: `[..., {"role": "user", "content": "List three fruits as JSON"}, {"role": "assistant", "content": "{", "prefix": true}]` The flag only makes sense on the *last* message, and only on an assistant one. Earlier assistant messages in the array are ordinary conversation history and need no flag. ## Why the first tokens are worth controlling Autoregressive generation is path-dependent: the first few tokens condition everything after them. A model that begins with "Sure! Here is the JSON you asked for:" has committed to prose, and the JSON that follows is now wrapped in chatter or a markdown fence. A model that begins with `{` has committed to an object. You did not persuade it with instructions; you removed the branch. That is the general shape of every good use of this feature. **Format locking.** Prefixing `{` or `[` is the cheapest way to strip preambles and fences. It composes with JSON-mode response formatting rather than competing with it. **Language locking.** Multilingual models drift toward English when the instruction is in English. A prefix of a few words in the target language pins the output language far more reliably than "respond in French" ever does. **Persona and role-play.** Starting the assistant turn with the character's voice keeps subsequent turns in that register, which matters in long conversations where a system prompt slowly loses influence. **Resuming a truncation.** When a reply comes back with `finish_reason: "length"`, you can send the partial text back as a prefix and let generation continue from exactly where it stopped, rather than regenerating from the beginning and paying for the same tokens twice. This is the most operationally valuable use and the one senior candidates are expected to reach for. **Skipping known boilerplate.** If every answer must open with a fixed header, emit the header as prefix rather than spending output tokens and hoping the model reproduces it verbatim. ## Bounding the continuation A prefix controls the start, not the end. Combine it with the `stop` parameter so the model halts at a known boundary — for example prefixing `{` and stopping at `}` when you want exactly one flat object, or stopping at a newline when you want one line. Between the two you have pinned both ends of the output shape without relying on the model to obey a formatting instruction. Also remember that the prefix text you send is part of the answer you assembled, not something the model generated. When you store the turn or show it to a user, make sure the prefix appears exactly once — check the returned `content` against the prefix you sent rather than blindly concatenating, or you will ship a duplicated opening brace and a parse error. ## The safety and injection caveat Putting words in the model's mouth is precisely the shape of a jailbreak. An assistant turn that begins "Certainly, here are the step-by-step instructions:" bypasses the refusal that the model would otherwise have produced first, because refusals are themselves a way of opening a turn. Alignment training conditions the model's *choice* of opening; a prefix takes that choice away. The hard rule that follows: **the prefix must never be derived from untrusted input.** It is application-controlled text, on the same trust level as your system prompt. If a user can influence it — through a template variable, a stored profile field, a retrieved document — you have handed them a channel that is strictly more powerful than an ordinary prompt injection, because it speaks with the assistant's own voice. A related judgement point: prefixing also suppresses signals you might want. If a model would have refused, or would have said "I don't have enough information", forcing an opening brace means you get a fabricated object instead of an honest refusal. Format compliance goes up; you should expect uncertainty signalling to go down, and you should test for that rather than assume the trade is free. ## Portability This is a Mistral-specific field. Not every OpenAI-shaped provider accepts it, and generic OpenAI client libraries have no schema slot for it, so it usually has to be passed as an extra body parameter. If your codebase abstracts over several providers, prefix continuation belongs behind a capability check rather than in the common path.

  • How would you use it to recover from a reply that ended with finish_reason 'length'?
    Resend the conversation with the truncated text as a trailing assistant message carrying prefix true, and raise max_tokens. Generation resumes from where it stopped instead of restarting, so you pay output tokens only for the remainder. Reassemble carefully so the already-received portion is not duplicated in what you store.
  • Why is prefix continuation a security-sensitive feature?
    It lets the caller author the opening of the model's own turn, which is exactly how refusal-bypass jailbreaks work — begin the assistant turn with compliance and the refusal branch never gets taken. Treat the prefix as application-controlled text at the same trust level as the system prompt, and never interpolate untrusted user or retrieved content into it.
  • What do you lose by forcing an opening brace on every structured call?
    Honest abstention. If the model would have said it lacks the information to answer, forcing `{` converts that refusal into a fabricated object. You gain parse-rate and lose an uncertainty signal, so measure hallucination rate alongside format compliance rather than assuming the trade is free.

saying these in an interview costs you the question

  • Thinks prefix goes on a user message rather than an assistant one
  • Sets prefix on an earlier history message instead of the last one
  • Believes the model rewrites or paraphrases the prefix text
  • Interpolates untrusted user input into the prefix
  • Says prefixing has no effect on refusal or safety behaviour

context