In an LLM agent, what is a hallucinated tool call versus a wrong-tool pick?
answer
- does the named tool exist at all
- one fails routing, one fails judgement
- registry check versus description rewrite
- one is countable automatically, one needs a reference
- adding overlapping tools raises the second
basics
~20 sA hallucinated tool call names a tool that does not exist in the definitions the agent was given. A wrong-tool pick is a valid call to a real tool that cannot answer the request. Different causes, different fixes.
solid answer
~50 sA hallucinated tool call is one the harness cannot even route: the agent asks for `submit_medwatch_3500a` when the only registered tool is `report_adverse_event`. The name came from the model's world knowledge, not from the tool list, so the failure happens before any argument matters. A wrong-tool pick executes perfectly — the tool is real, the arguments validate, the result comes back — and is simply the wrong instrument for the question, usually because two tool descriptions overlap or one is vaguely worded. The fixes diverge accordingly. Hallucination is a harness problem: validate the requested name against the registry and return "no such tool, the available ones are…" so the next turn can correct. Wrong-tool selection is a tool-surface problem: sharpen descriptions, state when *not* to use each tool, and delete near-duplicates. Rewriting descriptions will not stop an invented name, and registry validation will not stop a bad choice.
go deeper
Be able to state the test out loud: does the requested tool exist in the definitions or not? Give one example of each case and name the fix that goes with it.
Explain why the two sit at different layers — one is a routing failure the harness catches in code, the other a comprehension failure caused by ambiguous tool descriptions — and why each fix is useless against the other.
Talk about detectability: hallucinated names are counted automatically and should approach zero, while wrong-tool rate needs reference labels and creeps back whenever someone adds an overlapping tool to the registry.
Own tool-surface governance as a standing concern — who may add a tool, what review checks for overlap with existing ones, and why an unmanaged registry converts a solved failure mode back into a recurring one.
## Two failures that look alike in a bug report Both arrive as "the agent called the wrong thing", and both are commonly reported as "the model hallucinated". They are different defects with different owners, and conflating them is why teams sometimes spend a week rewriting tool descriptions to fix a problem that a five-line registry check would have ended. ## Hallucinated tool call The agent emits a call whose name — or, in the milder version, whose parameter — does not appear in the tool definitions supplied for that turn. A pharmacovigilance triage agent that produces `submit_medwatch_3500a` is drawing on real domain vocabulary it absorbed in pretraining; the name is plausible, professionally shaped, and completely unroutable, because the registry only contains `report_adverse_event`. The mechanism is worth stating plainly: a language model generating a tool call is producing tokens, and unless the decoder is constrained to the registered names, nothing structurally prevents it from producing a name that merely sounds right. Pressure rises when the tool list is long, when names are near-synonyms of each other, or when the domain has a strong external vocabulary the model already knows. As of mid-2026, providers commonly constrain or validate tool-name generation against the supplied definitions, so pure name invention is much rarer on frontier models than argument-level errors — but it has not disappeared, especially in setups that parse tool calls out of free-form text rather than using a structured tool-calling interface. The fix is mechanical and belongs to the harness. Never dispatch on an unvalidated name. Look it up; if it is absent, return a structured error that names the tools that do exist, so the agent's next turn has the information it needed. Secondary hygiene: distinctive names that are not near-synonyms, and a loaded tool set no larger than the task needs. ## Wrong-tool selection Here everything mechanical works. The tool exists, the arguments validate, the call executes, a result comes back — and it is the wrong result, because the agent chose an instrument that cannot answer the question. Searching a product label when the answer lives in the published literature is not a crash; it is a bad decision that produces a confident, well-formed, useless observation. That is what makes it harder to catch: the trace looks healthy and only the outcome is wrong. The cause is almost always the tool surface rather than the model. Two tools whose descriptions could each plausibly cover the request. A description that says what the tool *is* but not when to reach for it. A tool whose name promises more than it delivers. An agent reading twelve one-line descriptions has no way to resolve an ambiguity the author left in. The fix is editorial. Treat each description as prompt text, because that is exactly what it is: say what the tool does, what it returns, and explicitly when *not* to use it in favour of its neighbour. Add a short worked example to the ones that are routinely confused. Namespace related tools so their relationship is visible. And when two tools genuinely overlap, the cheapest fix is deleting one. ## Why the distinction pays The two failures live at different layers, so the diagnostic signals differ too. A hallucinated call is detectable automatically and unambiguously: the name is not in the registry, full stop, and you can count these exactly. A wrong-tool pick needs judgement — a human or a reference trajectory saying what the right call would have been — because syntactically it is indistinguishable from a correct call. That difference in detectability has a practical consequence. Hallucinated calls should trend to near-zero once validation exists, and if they do not, something is wrong with the integration itself. Wrong-tool rates never reach zero; they are managed downward by curating the tool surface, and they creep back up every time someone adds a tool without checking whether it overlaps an existing one. ## The trap to avoid The common error is the blanket label. Calling both "hallucination" pushes both toward the same instinctive fix — a bigger model, a sterner prompt — when one wants a registry lookup in code and the other wants an hour of careful writing. Ask first whether the requested tool exists. That single question splits the two cases, and the answer decides who fixes it.
- Which of the two can you measure automatically, and why does that matter?Hallucinated calls: the name either is or is not in the registry, so the harness counts them exactly and for free. Wrong-tool picks need a reference — a human label or an expected trajectory — because the call is syntactically valid. That asymmetry sets expectations: hallucination rate should trend to near-zero once validation exists, while wrong-tool rate is managed downward by curating descriptions and never fully disappears.
- An agent invents a parameter name that isn't in the tool's schema. Which category is that?The same family as a hallucinated tool call — fabricating something the definitions never declared — and the same fix applies: validate against the schema and return an error naming the fields that actually exist. It shades toward the invalid-argument category, which matters only for tagging. Practically, both are caught by the same validation boundary before dispatch, so treat them as one row unless your incident data says otherwise.
saying these in an interview costs you the question
- Labeling every bad tool call a hallucination
- Fixing an invented tool name by rewriting tool descriptions
- Assuming registry validation also prevents poor tool choices
- Believing a stronger model removes the need to validate names
- Adding overlapping tools without checking existing descriptions