How do you call Mistral's /v1/embeddings endpoint, and what does mistral-embed return?
answer
- One POST, an array of strings
- Fixed width, no dimension knob
- 1024 floats per input
- index is the only join key
- usage.prompt_tokens is what you pay
basics
~20 sPOST /v1/embeddings on api.mistral.ai with a bearer API key, model "mistral-embed" and an "input" array of strings. The response carries a data array of 1024-dimension float vectors, each tagged with the index of the string it came from, plus a token usage block.
solid answer
~40 sIt is a single POST to `https://api.mistral.ai/v1/embeddings`, authenticated with `Authorization: Bearer $MISTRAL_API_KEY`. The body needs two fields: `model` (`mistral-embed`) and `input`, which takes an **array** of strings, so you batch many texts into one round trip instead of one call per text. The response is an object with `data`, an array of entries each holding `object: "embedding"`, the `embedding` float array, and an `index` — that index is the only thing tying a vector back to your input, and entries come back in request order. `mistral-embed` produces a fixed 1024-dimension vector; there is no dimensionality parameter on that model. A `usage` block reports `prompt_tokens`/`total_tokens`, which is what you are billed on, and each input has a maximum length — an over-long string errors rather than being silently truncated.
code
bash · 7 linescurl https://api.mistral.ai/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-d '{
"model": "mistral-embed",
"input": ["Embed this sentence.", "And this one too."]
}'go deeper
Be able to name the route, the two body fields (model and input), and the fact that input takes an array. Say plainly that mistral-embed returns 1024 floats per string.
Explain the response shape — data entries with embedding and index, plus a usage block — and why index is the join key. Know that dimensionality is fixed and that billing counts input tokens only.
Show the ingestion loop: chunk to the sequence limit, batch to the token-per-minute allowance, back off on 429, and keep retries idempotent against your own chunk ids so a partial batch failure does not duplicate work.
Own the consequences of a fixed 1024-dimension vector: storage and index sizing for the whole corpus, and the fact that changing embedding model later means a full re-embed. Decide up front whether that lock-in is acceptable or whether a configurable-width model is worth it.
## The endpoint Mistral's embedding surface is one HTTP route on la Plateforme: `POST https://api.mistral.ai/v1/embeddings` Auth is a bearer token: `Authorization: Bearer $MISTRAL_API_KEY`, the same workspace key used for chat completions. Content type is `application/json`. ## The request body Two fields matter: - `model` — the embedding model id. The general-purpose one is `mistral-embed`. - `input` — an **array of strings**. Sending an array is the normal case, not an optimisation: one HTTP call can carry many texts, which matters enormously when you are embedding a corpus, because per-request overhead and rate-limit accounting are both per call. That is it. Unlike the chat endpoint there is no temperature, no streaming, no tools — embedding is deterministic given the model, so none of the sampling parameters exist here. ## The response body ``` { "id": "...", "object": "list", "model": "mistral-embed", "data": [ {"object": "embedding", "embedding": [0.018, -0.024, ...], "index": 0}, {"object": "embedding", "embedding": [...], "index": 1} ], "usage": {"prompt_tokens": 15, "total_tokens": 15} } ``` The key operational detail is `index`. There is no echo of your original text and no id you supplied — the only correspondence between input and output is positional, carried by `index`. So when you fan a batch out, keep your own parallel list (or a dict keyed by index) and rebuild the pairing on the way back. Code that assumes `data` is already aligned to the input list is *usually* right, but reading `index` explicitly is the contract-correct way. ## Dimensionality `mistral-embed` emits a fixed **1024-dimension** float vector. The request has no dimension knob for this model: you cannot ask for 256 or 512 dimensions and you cannot ask for a quantised output type. That is a fact worth knowing before you size storage or pick an index configuration — the vector width is a property of the model, not a request parameter. As of mid-2026 Mistral also ships a code-specialised embedding model, Codestral Embed, whose API does expose `output_dimension` and `output_dtype` so you can trade recall for footprint. If your interviewer asks "can I shrink Mistral vectors?", the accurate answer is "not with `mistral-embed`; that is a capability of the code embedding model, and otherwise you shrink them yourself downstream." ## Usage, billing and limits `usage.prompt_tokens` is the token count of everything you sent, and embeddings are billed per input token — there are no output tokens to pay for. There is a maximum sequence length per input string; an input beyond it is rejected with a 4xx error rather than truncated, so chunking is your responsibility before the call, not the API's. A single request also has a practical ceiling on how many strings and how many total tokens it will accept, which is why bulk pipelines send batches of a few dozen to a few hundred chunks rather than the whole corpus in one body. Rate limits on la Plateforme are workspace-level and expressed in both requests and tokens per minute; a `429` on the embeddings route is handled like any other — exponential backoff with jitter, and reduce batch size if it is the token limit you are hitting rather than the request limit. ## Practical shape of an ingestion loop 1. Chunk documents to sit inside the model's sequence limit. 2. Group chunks into batches sized to your token-per-minute allowance. 3. POST each batch; read `data[i].index` to re-attach vectors to chunk ids. 4. Retry `429`/`5xx` with backoff, and make retries idempotent by keying on your own chunk id. 5. If the corpus is large and not latency-sensitive, push the same requests through the asynchronous batch jobs API instead, which runs them offline at a discount. One consistency rule outlives the API detail: query vectors and document vectors must come from the same model. Mixing vectors produced by `mistral-embed` with vectors from any other model in one index produces silently meaningless similarity scores.
- You send 200 strings in one call and get 200 vectors back — how do you know which vector belongs to which string?By the `index` field on each entry in `data`, which is the position of the input in the array you sent. The API echoes neither your text nor any id of your own, so you keep a parallel list (or map index to your chunk id) and rejoin on the way back. Entries arrive in request order in practice, but reading `index` is the contract-correct join.
- Can you ask mistral-embed for shorter vectors to save storage?Not on that model — `mistral-embed` returns 1024 dimensions with no dimensionality parameter. If you need a smaller footprint you either reduce dimensions yourself downstream, quantise in your vector store, or use Mistral's code embedding model, which as of mid-2026 exposes `output_dimension` and `output_dtype` for exactly that trade.
- What happens if one string in your input array is longer than the model's maximum sequence length?The request fails with a client error rather than the input being quietly truncated, and because the whole array is one request, the entire batch fails with it. So chunk before you call, validate lengths client-side, and consider isolating suspicious inputs so one oversized document cannot poison a batch of hundreds.
saying these in an interview costs you the question
- Thinking you must send one string per request
- Believing a dimensions parameter shrinks mistral-embed output
- Assuming the response echoes the original input text
- Expecting temperature or streaming on the embeddings endpoint
- Mixing vectors from different models in one index