skip to content

Your Mistral tool loop 400s after tool_calls — what is wrong with the history?

level: seniorimportance: should knowfreq 44%

answer

  1. the server rejects it before generating
  2. pairing, not prompting
  3. one block, not loose messages
  4. echo the assistant turn verbatim
  5. every id needs an answer

basics

~20 s

Almost always a broken pairing in the messages array: the assistant message carrying tool_calls was not re-sent verbatim, a tool_call_id does not match one the server issued, or some of the returned calls got no role "tool" reply at all.

solid answer

~40 s

Mistral validates the structure of the conversation you resend, not just the last message. Every `role: "tool"` message must sit after an assistant message that actually contains a `tool_calls` entry with the same `tool_call_id`, and every call in that assistant message needs a matching tool reply before the next assistant turn. The four ways teams break it: rebuilding the assistant turn from `content` only and dropping `tool_calls`; minting your own ids instead of echoing the server's; answering `tool_calls[0]` and ignoring the rest; and context trimming that evicts the assistant tool-call turn while keeping its tool results (or the reverse). Fix it by appending the assistant object exactly as returned, then one tool message per call, and by treating that assistant turn plus its results as one indivisible unit whenever you trim history.

go deeper

for a junior

Know that the tool result must reference the tool_call_id the response gave you, and that the assistant message containing tool_calls has to stay in the conversation you resend.

for a middle

Explain the grammar the API enforces — assistant tool_calls turn, then one role "tool" message per id — and name the reconstruct-the-assistant-message bug as the usual cause.

for a senior

Diagnose from the outgoing messages log, cover the multi-call and tool-crashed paths, and describe the trimming interaction where a sliding window slices a tool exchange in half.

for a principal

Push for a history type that cannot express an invalid exchange — atomic blocks, a single append helper, tests over the parallel-call path — so the invariant is enforced by structure rather than by every engineer remembering it.

## The failure mode The loop works on the first request, the model returns `finish_reason: "tool_calls"`, you run the function, you send the follow-up — and the API rejects the request with a 400 before any tokens are generated. This is a *request validation* failure, not a model failure, which is why no amount of prompt tweaking helps and why the same code shape fails deterministically every time. Mistral checks that the `messages` array you resend describes a coherent tool exchange. ## The invariant you must satisfy A conversation containing tool use has a rigid grammar: 1. An assistant message whose `tool_calls` list holds one or more calls, each with a server-issued `id`. 2. Immediately following it, one message per call with `role: "tool"`, `tool_call_id` equal to that call's `id`, and a string `content`. 3. Only then the next assistant turn (or the next user turn). Every `tool_call_id` must resolve to a call that exists earlier in the array, and every call should be answered. Break either direction and the request is malformed. ## The four ways real code breaks it **Rebuilding the assistant turn instead of echoing it.** The most common cause. Code that normalizes messages into an internal `{role, content}` shape silently discards `tool_calls`, because on a tool turn `content` is usually empty. The history then contains an empty assistant message followed by tool results that point at ids nothing declares. Fix: store and resend the assistant object as the API returned it, including `tool_calls`, and make your internal message type able to carry it. **Minting your own ids.** Some developers generate a UUID per invocation for their own tracing and put it in `tool_call_id`. The id is not yours to choose — it is the handle the server issued in the response, and any other value fails to resolve. Keep your trace id in a separate field of your own logging, not in the wire message. **Answering only the first call.** When the model returns two or three calls, `tool_calls[0]` handling leaves the remaining calls unanswered. Iterate the list, execute each (in parallel if your tools are independent), and append a tool message for each id. **Context trimming that cuts across the unit.** A sliding window that keeps the last N messages will eventually slice between an assistant tool-call turn and its results, leaving orphans on one side. The assistant tool-call message plus all of its tool results form one atomic block: evict them together or keep them together. The same applies to summarisation — a summariser that replaces old turns with prose must not swallow half a tool exchange. ## Failures that are not history-shaped Before concluding it is pairing, rule out two neighbours. A tool `content` that is not a string (a dict or a number passed straight through from your function) is a type-validation failure — serialize with `json.dumps` first. And a call whose arguments you could not parse still needs an answer: return a tool message describing the parse failure rather than skipping the id, otherwise you convert a recoverable model mistake into a hard 400. ## Making it structurally impossible The durable fix is not defensive checks scattered through the loop; it is a history object that cannot represent an invalid state. Practical shape: - Append the raw assistant message from the response — never a reconstruction. - Have one function that takes the assistant message plus a map of `id -> result string` and returns the full list of messages to append, raising if any id is missing from the map or any key is not a real call id. - Give the trimmer knowledge of blocks rather than messages, so its unit of eviction is "assistant turn plus its tool results". - Assert the invariant in a unit test with a fixture containing two parallel calls; the multi-call path is the one that never gets exercised by hand. ## Diagnosing it fast in production Log the outgoing `messages` array with contents truncated but roles and ids intact. Nearly every instance of this bug is visible in three seconds from that log: an assistant turn with no `tool_calls`, a tool message whose id appears nowhere else, or a call id with no answering message. Because the request is rejected before generation, it costs no tokens — so a spike in these is cheap in money and expensive in latency and user-facing errors, and it should page on error rate, not on spend. ## What a senior answer includes Name the invariant, name at least two concrete ways code violates it, and then move to prevention: echo rather than rebuild, treat the tool exchange as an atomic block in trimming, and cover the multi-call path in tests. Candidates who only say "you must include the tool_call_id" have seen the docs; candidates who mention the trimming interaction have run an agent long enough for the window to fill.

  • Your context trimmer drops the oldest messages. How do you keep it from causing this?
    Make its unit of eviction the tool exchange, not the message. Group the assistant message that carries `tool_calls` with all of its `role: "tool"` results into one block and evict or retain the block as a whole. If a summariser compresses old turns, it must replace whole blocks with prose too — never half of one, which leaves either orphan results or unanswered call ids.
  • A tool crashed and you have no output. What goes in the messages array?
    Still a `role: "tool"` message with that call's `tool_call_id` — content describing the failure, for example that the service timed out and the model should retry or choose another approach. Skipping the message leaves the call unanswered and turns a recoverable runtime error into a malformed request. The model handles an error string far better than your loop handles a 400.
  • How do you distinguish this 400 from one caused by a bad tool schema?
    Look at what changed and when. A schema problem fails on the very first request that carries the `tools` array, before any tool call exists; a pairing problem only appears on the follow-up request after the model returned `tool_calls`. Logging the outgoing message roles and ids separates them immediately, since a pairing failure shows a visible orphan.
  • Does the failed request cost you tokens?
    No — validation happens before generation, so a rejected request bills nothing. That makes the bug invisible on a spend dashboard and visible only as an error rate and as user-facing latency, which is why the alert for it belongs on request failures rather than on usage metrics.

saying these in an interview costs you the question

  • Rebuilds the assistant turn from content and drops tool_calls
  • Generates its own tool_call_id values
  • Answers only the first of several tool calls
  • Trims history by message count across a tool exchange
  • Skips the tool message entirely when the function threw

context