What are the main failure modes of a tool-using LLM agent, and what fixes each?
answer
- one name per failure, one fix per name
- invented tool versus real-but-wrong tool
- arguments fail the schema, not the model
- no-progress and goal drift differ
- a success claim is not evidence
basics
~20 sTool-using agents fail in recognizable categories: invented tool names, invalid arguments, the wrong real tool, non-progressing loops, drift from the original goal, and false claims of success. Each has a different fix, so classifying the failure is the first debugging step.
solid answer
~60 sI keep a short taxonomy with one fix per row. **Hallucinated tool call** — the agent names a tool that isn't in the registry; fix by validating the call before execution and returning "no such tool, here are the real ones". **Invalid arguments** — right tool, bad payload (natural language in a strict date field, missing required key); fix with schema validation plus an error that names the offending field. **Wrong-tool selection** — a valid call to a tool that can't answer the request; fix in the tool surface, by disambiguating descriptions and removing near-duplicates. **Non-progressing loop** — the same action or an A-B oscillation with no new information; fix by detecting no progress and changing the observation, not by allowing more attempts. **Goal drift** — constraints stated at step 1 are gone by step 60; fix by re-anchoring the goal each step. **False success** — it says "done" when the tool errored; fix with an external post-condition check. Tag every failed run with exactly one primary mode; the distribution tells you where to invest.
go deeper
Be able to list the categories by name and give one concrete example of each — an invented tool name, a bad date argument, a repeated identical action. Naming them plainly is most of what is being tested here.
Explain what distinguishes each mode mechanically and name the fix that owns it, including why schema validation fixes argument errors while better descriptions fix tool selection. Interviewers expect the mapping, not just the list.
Show how you use the taxonomy operationally: tagging failed runs with one primary mode, reading the distribution, and resisting the urge to rewrite the prompt when the data points at the tool surface or a missing verifier.
Own the argument that most agent failures are system-design failures rather than model failures, and defend where each mitigation should live — schema, harness, verifier, or approval gate — so fixes are not all pushed into one giant prompt that nobody can maintain.
## Why a taxonomy at all "The agent doesn't work" is not a bug report. Agents are long chains of model decisions and side effects, so a single failed run can have a dozen plausible causes, and without shared vocabulary every incident review restarts from zero. A taxonomy converts a vague complaint into a category, and each category points at one owner and one fix. That is the whole value: it is a triage artifact, not a theory of intelligence. ## The categories **1. Hallucinated tool call.** The agent emits a call to a tool that does not exist in the definitions it was given, or invents a parameter the schema never declared. A pharmacovigilance triage agent emitting `submit_medwatch_3500a` when the registry only holds `report_adverse_event` is the canonical shape: the invented name is plausible domain vocabulary absorbed in pretraining. Fix: validate every call against the registry before execution and return a structured error naming the tools that actually exist. Distinctive, non-overlapping names help; so does keeping the loaded tool set small. **2. Invalid arguments.** Correct tool, unusable payload — a wrong type, a missing required field, a fabricated identifier, or unparsed natural language in a strict field (`onset_date: "last Tuesday"` against an ISO-8601 schema). Fix: validate at the boundary and return an error that says which field failed and what shape it wants, so the next attempt has new information. The pathological variant is the identical retry — the agent re-sends the same malformed value four times because the error text was a generic "validation failed". **3. Wrong-tool selection.** A well-formed call to a real tool that cannot answer the request. This is a comprehension failure, not a formatting one, and it usually traces to overlapping or vague tool descriptions. Fix lives in the tool surface: disambiguate descriptions, add a worked example of when each applies, namespace related tools, and delete near-duplicates. **4. Non-progressing loops.** The agent repeats an identical action, or ping-pongs between two tools on the same input — searching a drug label, then the literature, then the label again with the identical query — or, in a GUI setting, reopens the same modal dialog forever. The trace grows; the state does not. Detection is a repeated-state or no-new-information check; the real repair is giving the agent a different observation (a summarized result, an explicit "this returned nothing, try a different query"), because more attempts on the same input produce the same output. **5. Goal drift and instruction forgetting.** Constraints stated in the opening instruction stop being honoured deep into a long run, as thousands of observation tokens dilute them. Fix: re-anchor — restate the goal and hard constraints in the most recent turn, or keep them in a durable artifact the agent re-reads each step. **6. Premature termination and false success.** The agent stops and declares the task complete when the work did not land — the write tool returned an error, or the verification step never ran. This is the most dangerous row because nothing looks broken. Fix: an external post-condition check that never trusts the final message. **7. Unsafe or irreversible action.** The agent takes a destructive step it should have escalated. Fix: an approval gate in front of the irreversible tools, not a prompt asking it to be careful. ## Turning it into an artifact Write it as a table: mode name, observable signature in the trace, likely cause, the single owning mitigation, and where that mitigation lives — prompt, schema, harness, or verifier. "Where it lives" is what makes it actionable, because it names the team that fixes it. Then tag every failed run with exactly one *primary* mode. Failures compound — a malformed argument causes a retry loop that exhausts the budget — so the rule is to tag the first mode in the causal chain, or the counting becomes meaningless. Once runs are tagged, the distribution is the roadmap: forty percent argument errors is a schema and error-message problem, not a model problem, and no amount of prompt rewriting will move it. ## What the taxonomy is not It is not a scoring rubric — it says what broke, not how good the run was. It is not a substitute for verification: naming "false success" does not detect it. And it is not universal. The generic rows above are a starting skeleton; a mature system adds domain-specific rows ("queried the wrong tenant's data", "submitted before the human review step") that its own incident history produced.
- A single failed run shows a malformed argument, then a retry loop, then budget exhaustion. Which mode do you tag it?The first mode in the causal chain — the malformed argument. The loop and the exhaustion are downstream symptoms, and tagging them would inflate the loop bucket while hiding the schema problem that actually caused the run to fail. If you tag every symptom, the distribution stops pointing at anything. Record the chain in the incident notes, but count one primary mode per run.
- How does the failure distribution change what you work on next?It reallocates effort away from prompt rewriting. A bucket dominated by invalid arguments is a schema and error-message problem: tighten types, use enums, make validation errors name the field. A bucket dominated by wrong-tool selection is a tool-description problem. A bucket dominated by false success means the harness has no post-condition check. Only when failures spread evenly across categories with no structural cause is model capability the honest answer.
- Is the generic taxonomy enough, or should each system have its own?The generic rows are a skeleton — they catch the mechanical failures every tool-using agent has. Domain rows come from your own incidents: querying the wrong tenant, skipping a mandatory review step, filing under the wrong regulatory category. Those are the ones that hurt in production and no external taxonomy will list them for you. Expect the table to grow for the first few months and then stabilize.
saying these in an interview costs you the question
- Calling every agent failure 'hallucination' regardless of cause
- Fixing a stuck loop by raising the iteration cap
- Treating a wrong-tool pick and an invented tool as one bug
- Assuming a larger model eliminates argument-validation failures
- Tagging one run with every symptom it displayed