With Mistral tool_choice "any" on every turn, why does your agent never finish?
answer
- the exit condition can never fire
- the parameter is per request, not per chat
- forcing has no expiry
- auto on the second hop
- none to close the loop
basics
~20 sBecause "any" forces a tool call on every request. The model can never return a plain text answer, so each round trip yields another call and the loop has no natural exit. Flip back to "auto" after the forced turn.
solid answer
~40 s`tool_choice: "any"` is a hard constraint on the response: Mistral must emit a tool call, so `finish_reason` comes back as `"tool_calls"` every time and the prose answer that normally ends the loop is unreachable. Since `tool_choice` is a per-request parameter and not conversation state, the fix is to stop hard-coding it in your client: set `"any"` only on the request where a call is genuinely mandatory — the first hop, or a structured-extraction step — and rebuild subsequent requests with `"auto"` so the model can decide it has enough information. If you need a guaranteed prose turn to close the loop, send the final request with `"none"`. Independently of all this, cap iterations, because forcing is not the only way a loop fails to terminate.
go deeper
Know that "any" forces a tool call, so a response under it is never a final text answer, and that the setting is sent per request rather than fixed for the conversation.
Explain why the loop's exit test cannot fire under forcing and show the fix: force only the hop that needs it, then send "auto", and close with "none" when you need a guaranteed prose turn.
Bring the operational view — hop-count instrumentation, an iteration cap independent of tool_choice, and the superlinear cost of resending a growing history on every hop.
Own the policy: where forcing is legitimate at all, what the escape-hatch tool looks like so a compelled model has a correct choice, and how per-request budgets and caps are enforced centrally rather than per feature team.
## The symptom An agent that always calls a tool, forever. Latency climbs, the token bill climbs with the growing history, and the user never sees an answer — often with the model re-calling the same tool with slightly reworded arguments because your history keeps telling it that it must call something and it has already learned everything the tool can tell it. ## The mechanism `tool_choice` in Mistral's chat completions request takes `"auto"`, `"any"`, `"none"`, or an object naming a specific function. `"any"` means the model must produce a tool call. It is not a hint, and it does not lapse once one call has been made — it applies to whichever request carries it. So if your client wrapper sets `"any"` on every call to the API, the exit condition your loop is watching for (`finish_reason: "stop"`, or an assistant message with no `tool_calls`) is structurally unreachable. You have written a loop whose termination test can never pass. This is easy to introduce because forcing genuinely is the right setting for the *first* request in many designs: a research agent that must always search before answering, a router that must always pick a handler, an extractor whose only acceptable output is a filled schema. The mistake is scoping the setting to the client or the conversation instead of to the request. ## The fix, in order of preference **Scope the parameter per request.** Build the request fresh each hop. Hop 1 sets `"any"`; every later hop sets `"auto"`. Nothing about the conversation state needs to change — the model reads the tool results already in the history and, freed from the constraint, answers when it has enough. **Close with `"none"` when you need a guaranteed final answer.** If your product cannot tolerate an nth tool call — say you are on a latency budget or you have hit your own iteration cap — send one last request with `"none"`. The tools stay declared, calling is barred, and the model must summarise what it has. This turns "ran out of iterations" into a coherent answer instead of an error page. **Cap iterations anyway.** Forcing is the deterministic cause of non-termination, but not the only one: a tool that returns unhelpful output can make an `"auto"` model retry indefinitely too. A hard cap (typically a handful of hops) with a `"none"` closing turn is the belt-and-braces version, and it also bounds worst-case cost per user request. ## Detecting it in production Instrument hop count per user request and alert on the tail. A distribution where the modal conversation uses one or two tool hops but a slice sits pinned at the cap is the signature. Correlate with `tool_choice` in your request logs; if the pinned slice is all `"any"`, you have found it in one query. Because each hop resends the entire growing history, cost per conversation grows superlinearly with hops, so this bug is expensive in a way a per-request cost dashboard hides. ## The related quality problem Even a single forced turn has a cost worth naming: a model compelled to call something when no declared tool fits will pick the nearest and invent arguments rather than decline. If your tool surface does not cover the whole input distribution, either keep `"auto"` and write tool descriptions good enough that abstention is natural, or declare an explicit no-applicable-tool function so that forcing still has a correct answer available. Combining that escape hatch with per-request scoping gives you determinism where you need it without manufacturing confident nonsense on the inputs you did not anticipate. ## What a strong answer sounds like State the mechanism in one line — forcing makes the terminating response unreachable — then show that you know `tool_choice` is request-scoped, that `"none"` is the closing move, and that an iteration cap belongs there regardless. Mentioning the cost shape (history resent every hop) and the forced-call quality risk marks someone who has operated one of these loops rather than read about one.
- You keep "auto" throughout and the loop still runs away. What now?Cap hops and close with a `"none"` request so the model must summarise what it has. Runaway under `"auto"` usually means the tool results are unhelpful — empty result sets, errors the model reads as retryable, or output too large to be usable — so also look at what the tool returns. The cap bounds cost while you fix the underlying tool.
- Why is a runaway tool loop disproportionately expensive?Every hop resends the entire conversation, and the conversation grows with each assistant tool-call turn and each tool result. Input tokens per hop therefore rise as the loop runs, so total cost grows faster than linearly in hop count. A loop that stalls at ten hops costs far more than ten times a one-hop conversation, which is why the cap is a spend control as much as a correctness one.
- When is forcing on the first hop genuinely the right call?When prose is never a valid output for that step: structured extraction into a known schema, a router that must select a handler, or a policy that every answer be grounded in a retrieval call. In those cases forcing removes a whole class of failure. Scope it to that one request, and where the model might have nothing to call, declare an explicit escape-hatch tool.
saying these in an interview costs you the question
- Thinks "any" applies once and then lapses
- Sets tool_choice globally in the client wrapper
- Blames the model or the prompt for the loop
- Skips an iteration cap because auto usually terminates
- Believes forcing a call improves answer quality generally