What does Mistral's safe_prompt flag change about a chat completions call?
answer
- a boolean, and it defaults to off
- injects text, does not score anything
- additive to your own system message
- a prompt, therefore dilutable
- watch the false-refusal rate
basics
~20 ssafe_prompt is a boolean on Mistral's chat completions request, off by default. When true, Mistral prepends a fixed safety instruction to the conversation, steering the model away from harmful, unethical or prejudiced output. It is guardrail prompting, not a classifier.
solid answer
~50 s`safe_prompt` is a request-body boolean on `POST /v1/chat/completions`, defaulting to false. Setting it true makes the platform inject a standard safety instruction ahead of your conversation — in substance, assist with care, respect and truth, and avoid harmful, unethical, prejudiced or negative content. Two consequences matter operationally. First, it is *additive*: it does not replace your own system message, so your instructions and the guardrail both apply and can conflict, which usually shows up as extra refusals or hedging on legitimate but sensitive domains such as medicine, security research or law. Second, it is a prompt, so it is soft — a long adversarial context can dilute it. It also costs a small number of input tokens on every request. Treat it as one cheap layer, and put a real moderation check on user-facing input and output rather than relying on it alone.
code
python · 15 linesimport os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.chat.complete(
model="mistral-small-latest",
messages=[
{"role": "system", "content": "You are a support agent for a bank."},
{"role": "user", "content": "How do I dispute a charge?"},
],
safe_prompt=True,
)
print(resp.choices[0].message.content)go deeper
Know that safe_prompt is a boolean on the request, defaults to false, and switches on a built-in safety instruction. Do not confuse it with a content filter that blocks requests.
Explain that it works by injecting instruction text into the context, that it stacks with rather than replaces your system message, and that it therefore costs a few input tokens per call.
Show that you would A/B it against your evaluation set and watch false refusals in sensitive-but-legitimate domains, and that you keep an independent output check rather than trusting an in-context guardrail.
Own the layering: which safety decisions are prompt nudges, which are classifier gates with auditable scores, and which are product policy. Be able to justify where the boolean flag sits in that stack and what it is not allowed to be relied on for.
## What the flag is `safe_prompt` is a boolean field in the body of a Mistral chat completions request, sitting alongside `model`, `messages`, `temperature` and the rest. It defaults to `false`. In the Python SDK it is passed as `safe_prompt=True` to `client.chat.complete(...)`; over raw HTTP it is `"safe_prompt": true` in the JSON body. It is one of the small handful of parameters that are genuinely Mistral's own rather than inherited from the OpenAI-shaped envelope. If you are running an OpenAI-oriented client against Mistral's base URL, this is exactly the kind of field you have to pass through as an extra body parameter, because the OpenAI schema has no equivalent. ## What it actually does When true, the platform prepends a fixed safety instruction to the conversation before the model sees it. The instruction is guardrail prose — the gist is to assist with care, respect and truth, respond usefully but securely, avoid harmful, unethical, prejudiced or negative content, and promote fairness and positivity. You do not author it and you cannot edit it; the flag is on or off. The crucial mental model: **this is prompt engineering performed on your behalf, not a separate safety system**. No classifier inspects the request, nothing is scored, and no request is blocked at the API boundary because of this flag. All that happens is that extra instruction text enters the context window, and the model's behaviour shifts because of it. ## Why that distinction matters in production Three practical consequences follow from "it is only a prompt". **It is additive, and it can fight your own system message.** Your `role: "system"` instruction is still there. If your product is a security-training assistant and your system prompt says "explain exploit techniques for defensive education", the injected guardrail pushes the other way. The visible symptom is not an error but a behaviour change: more refusals, more hedging, more "I can't help with that" on requests your product legitimately needs answered. Teams that flip the flag on globally and then wonder why quality regressed in one vertical have usually hit this. **It is soft.** Because the guardrail is in-context text, its influence competes with everything else in the window. A long conversation, an adversarial user turn, or a deliberately crafted assistant prefix can all dilute or override it. It raises the cost of eliciting bad output; it does not make it impossible. Never describe it in an interview as a guarantee. **It costs tokens.** The instruction is billed as input on every request that carries the flag. On a short, high-volume endpoint that overhead is a measurable percentage of prompt cost; on long-context calls it is noise. It is a small number either way, but it is not free, and it is worth knowing which side of that line your traffic sits on. ## Where it fits in a real safety design A production layering usually looks like: input screening on untrusted user text, a system prompt that states the product's own policy explicitly, optionally `safe_prompt` as a cheap extra nudge, output screening before anything is rendered to a user, and logging of refusals so you can tell over-refusal from genuine blocks. Mistral additionally offers a dedicated moderation model you can call as a separate classification step — that is the right tool when you need an actual allow/deny decision with a score you can threshold and audit, because a boolean prompt flag gives you no signal at all about what it did. The judgement question an interviewer is really probing is whether you can distinguish three different things that all get called "safety": alignment baked into the model weights, in-context instruction (which is what `safe_prompt` is), and an external classifier gating the request. They fail differently, they are observable to different degrees, and only the third one gives you an auditable decision. ## Evaluating whether to turn it on Do not flip it on by belief. Run your evaluation set both ways and compare two numbers: harmful-completion rate and false-refusal rate on legitimate in-domain requests. For a general consumer assistant the trade is usually favourable. For a specialist tool operating in a sensitive-but-legitimate domain it often is not, and a well-written domain-specific system prompt plus a real output classifier beats the generic guardrail. Either way you should be able to show the numbers rather than assert the policy.
- If safe_prompt is enabled, can you still get a policy-violating completion?Yes. It only inserts instruction text into the context; nothing inspects or blocks the request. A long conversation, adversarial phrasing, or a forced assistant prefix can all outweigh it. It raises the effort required, it does not guarantee anything, so user-facing systems still need output screening.
- How would you measure whether turning it on is a net win for your product?Run your evaluation set with the flag off and on, and compare two rates: harmful or policy-violating completions, and false refusals on legitimate in-domain prompts. A generic guardrail often improves the first and worsens the second, and in specialist domains such as security or clinical content the second effect can dominate.
- Does it replace or override the system message you send?Neither — it is additive. Your system message stays in the conversation and the guardrail instruction is prepended alongside it. Both influence the model simultaneously, which is why conflicting instructions surface as hedging and inconsistent refusals rather than as an API error.
saying these in an interview costs you the question
- Claims safe_prompt runs a moderation classifier over the request
- Says it is enabled by default on la Plateforme
- Believes it replaces or overrides your own system message
- Treats it as a hard guarantee against policy-violating output
- Ignores that it adds input tokens to every request