skip to content

When should you force an LLM to call a tool instead of letting it decide?

level: juniorimportance: should knowfreq 45%

answer

  1. a per-turn control, not a global setting
  2. the ungrounded answer is the quiet failure
  3. required call with nothing to call
  4. let the last turn produce prose
  5. forcing does not fix arguments

basics

~20 s

Force a call on any turn where the answer must come from a system of record. Left free to choose, a model will often answer an order-status question from context or invention; requiring the lookup tool on that turn removes the option and grounds the reply in real data.

solid answer

~50 s

Tool choice is a per-turn control your code owns, not a fixed setting: on each request you can leave the decision to the model, require that it call *some* tool, require one *named* tool, or forbid tools entirely. The reliability use is grounding. Ask a ticketing support agent "has my refund gone through?" with tools merely available, and a common failure is a fluent, confident, entirely fictional status — the customer cannot tell. Forcing the order-lookup tool on the first turn of any order conversation makes the reply conditional on a real record. Force narrowly, though: if you require a call when none fits, the model still emits one and invents arguments to satisfy the requirement. And forbid tools on the closing turn, so the loop can actually produce prose and terminate. Forcing improves grounding; it does not improve argument quality, so keep validating.

go deeper

for a junior

Know that you can require the model to call a tool on a given turn, and say why: without it, the model may answer a data question from invention rather than from the record.

for a middle

Explain that tool choice is set per request and describe the sequence — force the grounding call, release control so the model can reason, disable tools so the last turn produces prose.

for a senior

Show the cost side: a forced call on a turn where none fits produces a spurious call with invented arguments, so forcing should be conditional on signals your harness already has, and paired with boundary validation.

for a principal

Frame it as where determinism belongs. A step you force identically every time is a workflow step that should live in code, not in the model; reserve model discretion for the choices that genuinely vary, and hold a measured budget for spurious versus ungrounded answers.

## Tool choice is a per-turn decision The useful framing is that on every request you decide how much freedom the model has over tool use. The four positions are: let the model decide, require that it call something, require one specific tool, and disable tools. Providers expose this control under different names, but the important property is the same everywhere — it is set per request, so a well-built harness varies it across the turns of a single conversation rather than picking one setting at startup. ## The failure that forcing fixes Models do not reliably recognise the boundary of their own knowledge. Give an event-ticketing support agent a set of tools and ask "has my refund gone through?", and one of the most common production failures is that no tool is called at all. The model has seen an enormous number of refund conversations and can generate a plausible, specific, well-formatted status — a date, an amount, a processor name — that is entirely fabricated. Nothing about the output signals that it was invented, which is precisely what makes it dangerous: a hallucinated *tool call* usually fails loudly at the boundary, while a hallucinated *answer* ships straight to the customer. Forcing the lookup tool on the first turn of any order-scoped conversation removes the option. The model cannot answer until a real record is in the context. This is the single highest-value use of the control, and it is why grounding-critical turns should not be left to the model's discretion. ## Forcing has a cost If you require a call on a turn where no call is appropriate — the customer said "thanks, that's all" — the model still has to emit one. It will pick the least implausible tool and invent arguments to fill the required fields, because the alternative (emitting nothing) has been taken away. Forcing therefore trades one failure mode for another: fewer ungrounded answers, more spurious calls. The way out is to force narrowly and conditionally. Your harness usually knows when grounding is required — the customer supplied an order number, the conversation is scoped to a booking, the previous turn ended in a question about status. Force on those turns and hand control back immediately afterwards, so the model can decide for itself whether a refund is warranted. ## Disabling tools matters just as much The mirror image is the closing turn. Once the data is in context, or once you have hit a budget, you want prose for the user rather than another call. Forbidding tools on that request guarantees the model produces a final answer, which is the difference between a loop that terminates and one that keeps circling because "one more lookup" always looks locally reasonable. The same setting is useful for a summarisation or explanation pass over results already gathered. ## If you always force the same tool, do not use the model A useful design smell: if turn one of every conversation forces the same specific tool with arguments that are already determined by your routing logic, that step is not a decision at all. Call the tool in code, put the result in the prompt, and start the model on turn two. You save a round trip, remove a class of argument errors entirely, and make the step deterministic and testable. Forcing a specific tool is best reserved for the case where the model still chooses the arguments even though it does not choose whether to call. ## What forcing does not fix Forcing controls *whether* a call happens and sometimes *which* tool. It says nothing about whether the arguments are right. A forced lookup can still be issued against an id the model invented, and a forced call under pressure is if anything more likely to carry made-up arguments. Grounding and argument validation are separate defences and you need both: the forced call gets a real record into the conversation, and the boundary check makes sure the call itself referred to a record that exists. ## Practice as of mid-2026 All major providers expose per-request tool-choice control, and the pattern above — force to ground, release to reason, disable to conclude — is standard in agent harnesses. What is still contested is how aggressively to force in long multi-turn agents, where over-forcing produces measurable spurious-call rates and under-forcing produces confident invented answers; teams tune this against their own trace data rather than from a general rule.

  • What goes wrong if you require a tool call on every turn of a conversation?
    Two things. The model emits calls on turns where the right move was to answer, inventing arguments to satisfy the requirement, so your spurious-call and invalid-argument rates both rise. And the loop struggles to terminate, because the model is never allowed to produce the final natural-language answer. Force on the turns that need grounding, then release control, and disable tools on the closing turn.
  • Why is a hallucinated answer more dangerous than a hallucinated tool call?
    A hallucinated tool call usually hits your validation gate — an unknown id or an out-of-enum value fails loudly and never reaches the side effect. A hallucinated answer bypasses every check you own and goes straight to the user, formatted as confidently as a true one. That asymmetry is the argument for forcing grounding calls on any turn whose answer depends on your data.
  • When is forcing a specific named tool a sign you should not be using the model for that step at all?
    When the tool and its arguments are both determined by your own routing logic. If turn one always forces the same lookup with an id you already hold, call it in code, put the result in the prompt, and start the model afterwards. You remove a round trip, a class of argument errors, and a source of nondeterminism, and the step becomes testable like any other function.

saying these in an interview costs you the question

  • Treats tool choice as one global setting for the whole app
  • Assumes a model will call a tool whenever it needs data
  • Forces a tool call on every turn of the conversation
  • Thinks forcing a call also guarantees correct arguments
  • Never disables tools, then wonders why the loop will not end

context