skip to content

In Mistral's FIM completions endpoint, what do prompt and suffix do?

level: middleimportance: should knowfreq 42%

answer

  1. Prefix and postfix, not a chat turn
  2. Cursor has code on both sides
  3. Only the middle comes back
  4. Codestral models only
  5. Chat-shaped response envelope

basics

~20 s

Mistral's POST /v1/fim/completions takes prompt as the code before the cursor and suffix as the code after it; the model generates only the middle span that joins them. It is a Codestral-only endpoint, separate from chat completions, built for IDE-style inline completion.

solid answer

~40 s

`POST /v1/fim/completions` is Mistral's fill-in-the-middle route, and it exists because inline code completion is not a conversation. You send `prompt` — everything to the left of the cursor — and `suffix` — everything to the right — and the model returns only the span that belongs between them, already conditioned on both sides. Contrast that with a chat call, where the model sees only what precedes and cannot know it must stop cleanly before an existing closing brace. The endpoint accepts `model` (a Codestral model such as `codestral-latest`; ordinary chat models are not valid here), plus generation controls like `max_tokens`, `temperature` and `stop`. The response is chat-completion shaped: a `choices` array whose `message.content` holds the generated middle, with a `finish_reason`. `suffix` is optional — omit it and you get plain left-to-right continuation.

code

bash · 10 lines
bash
curl https://api.mistral.ai/v1/fim/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -d '{
    "model": "codestral-latest",
    "prompt": "def add(a, b):\n",
    "suffix": "\n\nprint(add(1, 2))\n",
    "max_tokens": 32,
    "temperature": 0
  }'

go deeper

for a junior

Recall that prompt is the code before the cursor and suffix the code after it, and that the model fills in only the gap between them. Name the route as /v1/fim/completions.

for a middle

Explain why chat completions cannot do this — no visibility of what follows the cursor — and describe the response envelope with choices[0].message.content and finish_reason. Know that only Codestral models serve it.

for a senior

Talk about making it usable in an editor: windowed context instead of whole files, debouncing and cancelling stale requests, small max_tokens, streaming for first-token latency, and trimming completions that overlap the suffix.

for a principal

Own the economics and the fallback story: prefix and suffix are both billed input, so context windowing is a cost lever at fleet scale, and you need a defined degradation path when the completion service is slow or unavailable so the IDE stays usable.

## Why a separate endpoint exists An IDE completion is a different problem from a chat turn. When the cursor sits inside an existing function, the model must produce text that is coherent with the code *after* the cursor too — it has to stop before the closing brace that is already there, respect the return type declared below, and not re-emit code that follows. A chat completion only ever sees a prefix, so it cannot do this reliably. Fill-in-the-middle (FIM) is the training-and-serving pattern that solves it, and Mistral exposes it as its own route: `POST https://api.mistral.ai/v1/fim/completions` ## The two fields that define it - **`prompt`** — the *prefix*: everything before the cursor. Despite the name, it is not an instruction; it is literal source code. - **`suffix`** — the *postfix*: everything after the cursor. Optional. When present, the model conditions on it and generates only the bridging span. The model's output is the **middle** alone. It does not repeat the prefix and does not repeat the suffix, so your editor inserts the returned text verbatim at the cursor. This is the single most common misunderstanding: people expect the whole file back and then write string-diffing code to recover the insertion, which is unnecessary. When you omit `suffix`, the call degrades gracefully to ordinary left-to-right continuation — useful for "finish this line" at end-of-file. ## Model restriction FIM is a capability of Mistral's Codestral code models; the endpoint is not a generic route that any model can serve. Passing a general chat model id will not give you fill-in-the-middle behaviour. Prefer the family alias (`codestral-latest`) over a dated point version so your integration does not pin itself to a snapshot that will be superseded. ## Other request parameters Beyond `model`, `prompt` and `suffix`, the request carries the usual generation controls: - `max_tokens` — cap the length of the inserted span. Inline completion wants this small; a 2000-token completion is not something a developer will accept at the cursor, and it costs latency you cannot hide. - `temperature` — low values are the norm for code completion; determinism matters more than variety when the suggestion appears while the user is typing. - `stop` — stop sequences, useful for cutting a completion at a newline or a closing delimiter when you only want a single statement. - `random_seed` where you need reproducibility in tests. Streaming is available for this endpoint too, which is what an editor plugin actually wants: render the first tokens as ghost text immediately rather than blocking on the full span. ## Response shape The response mirrors the chat-completions envelope rather than inventing a new one: ``` { "id": "...", "model": "codestral-latest", "choices": [ {"index": 0, "message": {"role": "assistant", "content": " return a + b\n"}, "finish_reason": "stop"} ], "usage": {"prompt_tokens": 42, "completion_tokens": 6, "total_tokens": 48} } ``` Read `choices[0].message.content` and insert it. `finish_reason` tells you whether the model stopped naturally, hit `max_tokens` (`length`), or hit a stop sequence — a `length` finish on an inline completion usually means your cap is too tight or the model is rambling and you should discard the suggestion rather than insert a truncated statement. ## Operational notes for editor integrations Latency dominates the user experience. Three levers actually help: keep `max_tokens` small, trim the prefix and suffix to a window around the cursor rather than shipping the whole file on every keystroke, and debounce requests so you are not firing one per character. Cancel in-flight requests when the cursor moves — an answer for a stale cursor position is worse than no answer, because inserting it corrupts the buffer. Billing is per token like any other call: both the prefix and the suffix count as input tokens, so a large `suffix` is not free. That is the real cost pressure behind windowing the context rather than sending the entire file. Finally, validate before inserting. A completion that duplicates the first line of your `suffix` should be trimmed client-side; models occasionally overlap the boundary, and a cheap prefix/suffix overlap check saves the user from broken code.

  • What does the model return if you omit suffix entirely?
    It falls back to plain left-to-right continuation: with only `prompt` supplied there is nothing on the right to join up to, so the model just continues the code. That is the right call for end-of-file or end-of-line completion, but for a cursor inside an existing block you lose the whole point of FIM — the model no longer knows what it must stop before.
  • Your editor plugin sends the whole file as prefix and suffix on every keystroke. What goes wrong?
    Cost and latency. Both prefix and suffix are billed as input tokens and both must be processed before the first output token appears, so per-keystroke full-file calls are slow and expensive. Window the context around the cursor, debounce keystrokes, cap `max_tokens`, stream the response, and cancel in-flight requests when the cursor moves.
  • How do you tell whether an inline completion was cut off rather than finished?
    Check `finish_reason` on the choice. `stop` means the model ended naturally or hit one of your stop sequences; `length` means it hit `max_tokens` and the span is truncated mid-statement. Truncated code should generally be discarded rather than inserted, or the cap raised for that request class.

saying these in an interview costs you the question

  • Thinking the endpoint returns the whole rewritten file
  • Believing any chat model serves FIM completions
  • Treating prompt as an instruction rather than code before the cursor
  • Sending the entire file on every keystroke
  • Assuming suffix is required for the call to work

context