skip to content

Agent Evaluation

How you tell whether an agent actually works: task-completion benchmarks, scoring the whole trajectory rather than just the final answer, tool-call accuracy, and the cost and latency it burned getting there. Interviewers ask because agents fail in ways a single output metric hides — hallucinated tool calls, stuck loops, right answer by luck.

part ofAI agentsoverview, primer and where to startread it →
on this pageshow

explore

questions

14

What are the main failure modes of a tool-using LLM agent, and what fixes each?

level: middleimportance: must knowfreq 70%

answer

  1. one name per failure, one fix per name
  2. invented tool versus real-but-wrong tool
  3. arguments fail the schema, not the model
  4. no-progress and goal drift differ
  5. a success claim is not evidence

basics

~20 s

Tool-using agents fail in recognizable categories: invented tool names, invalid arguments, the wrong real tool, non-progressing loops, drift from the original goal, and false claims of success. Each has a different fix, so classifying the failure is the first debugging step.

solid answer

~60 s

I keep a short taxonomy with one fix per row. **Hallucinated tool call** — the agent names a tool that isn't in the registry; fix by validating the call before execution and returning "no such tool, here are the real ones". **Invalid arguments** — right tool, bad payload (natural language in a strict date field, missing required key); fix with schema validation plus an error that names the offending field. **Wrong-tool selection** — a valid call to a tool that can't answer the request; fix in the tool surface, by disambiguating descriptions and removing near-duplicates. **Non-progressing loop** — the same action or an A-B oscillation with no new information; fix by detecting no progress and changing the observation, not by allowing more attempts. **Goal drift** — constraints stated at step 1 are gone by step 60; fix by re-anchoring the goal each step. **False success** — it says "done" when the tool errored; fix with an external post-condition check. Tag every failed run with exactly one primary mode; the distribution tells you where to invest.

go deeper

for a junior

Be able to list the categories by name and give one concrete example of each — an invented tool name, a bad date argument, a repeated identical action. Naming them plainly is most of what is being tested here.

for a middle

Explain what distinguishes each mode mechanically and name the fix that owns it, including why schema validation fixes argument errors while better descriptions fix tool selection. Interviewers expect the mapping, not just the list.

for a senior

Show how you use the taxonomy operationally: tagging failed runs with one primary mode, reading the distribution, and resisting the urge to rewrite the prompt when the data points at the tool surface or a missing verifier.

for a principal

Own the argument that most agent failures are system-design failures rather than model failures, and defend where each mitigation should live — schema, harness, verifier, or approval gate — so fixes are not all pushed into one giant prompt that nobody can maintain.

## Why a taxonomy at all "The agent doesn't work" is not a bug report. Agents are long chains of model decisions and side effects, so a single failed run can have a dozen plausible causes, and without shared vocabulary every incident review restarts from zero. A taxonomy converts a vague complaint into a category, and each category points at one owner and one fix. That is the whole value: it is a triage artifact, not a theory of intelligence. ## The categories **1. Hallucinated tool call.** The agent emits a call to a tool that does not exist in the definitions it was given, or invents a parameter the schema never declared. A pharmacovigilance triage agent emitting `submit_medwatch_3500a` when the registry only holds `report_adverse_event` is the canonical shape: the invented name is plausible domain vocabulary absorbed in pretraining. Fix: validate every call against the registry before execution and return a structured error naming the tools that actually exist. Distinctive, non-overlapping names help; so does keeping the loaded tool set small. **2. Invalid arguments.** Correct tool, unusable payload — a wrong type, a missing required field, a fabricated identifier, or unparsed natural language in a strict field (`onset_date: "last Tuesday"` against an ISO-8601 schema). Fix: validate at the boundary and return an error that says which field failed and what shape it wants, so the next attempt has new information. The pathological variant is the identical retry — the agent re-sends the same malformed value four times because the error text was a generic "validation failed". **3. Wrong-tool selection.** A well-formed call to a real tool that cannot answer the request. This is a comprehension failure, not a formatting one, and it usually traces to overlapping or vague tool descriptions. Fix lives in the tool surface: disambiguate descriptions, add a worked example of when each applies, namespace related tools, and delete near-duplicates. **4. Non-progressing loops.** The agent repeats an identical action, or ping-pongs between two tools on the same input — searching a drug label, then the literature, then the label again with the identical query — or, in a GUI setting, reopens the same modal dialog forever. The trace grows; the state does not. Detection is a repeated-state or no-new-information check; the real repair is giving the agent a different observation (a summarized result, an explicit "this returned nothing, try a different query"), because more attempts on the same input produce the same output. **5. Goal drift and instruction forgetting.** Constraints stated in the opening instruction stop being honoured deep into a long run, as thousands of observation tokens dilute them. Fix: re-anchor — restate the goal and hard constraints in the most recent turn, or keep them in a durable artifact the agent re-reads each step. **6. Premature termination and false success.** The agent stops and declares the task complete when the work did not land — the write tool returned an error, or the verification step never ran. This is the most dangerous row because nothing looks broken. Fix: an external post-condition check that never trusts the final message. **7. Unsafe or irreversible action.** The agent takes a destructive step it should have escalated. Fix: an approval gate in front of the irreversible tools, not a prompt asking it to be careful. ## Turning it into an artifact Write it as a table: mode name, observable signature in the trace, likely cause, the single owning mitigation, and where that mitigation lives — prompt, schema, harness, or verifier. "Where it lives" is what makes it actionable, because it names the team that fixes it. Then tag every failed run with exactly one *primary* mode. Failures compound — a malformed argument causes a retry loop that exhausts the budget — so the rule is to tag the first mode in the causal chain, or the counting becomes meaningless. Once runs are tagged, the distribution is the roadmap: forty percent argument errors is a schema and error-message problem, not a model problem, and no amount of prompt rewriting will move it. ## What the taxonomy is not It is not a scoring rubric — it says what broke, not how good the run was. It is not a substitute for verification: naming "false success" does not detect it. And it is not universal. The generic rows above are a starting skeleton; a mature system adds domain-specific rows ("queried the wrong tenant's data", "submitted before the human review step") that its own incident history produced.

  • A single failed run shows a malformed argument, then a retry loop, then budget exhaustion. Which mode do you tag it?
    The first mode in the causal chain — the malformed argument. The loop and the exhaustion are downstream symptoms, and tagging them would inflate the loop bucket while hiding the schema problem that actually caused the run to fail. If you tag every symptom, the distribution stops pointing at anything. Record the chain in the incident notes, but count one primary mode per run.
  • How does the failure distribution change what you work on next?
    It reallocates effort away from prompt rewriting. A bucket dominated by invalid arguments is a schema and error-message problem: tighten types, use enums, make validation errors name the field. A bucket dominated by wrong-tool selection is a tool-description problem. A bucket dominated by false success means the harness has no post-condition check. Only when failures spread evenly across categories with no structural cause is model capability the honest answer.
  • Is the generic taxonomy enough, or should each system have its own?
    The generic rows are a skeleton — they catch the mechanical failures every tool-using agent has. Domain rows come from your own incidents: querying the wrong tenant, skipping a mandatory review step, filing under the wrong regulatory category. Those are the ones that hurt in production and no external taxonomy will list them for you. Expect the table to grow for the first few months and then stabilize.

saying these in an interview costs you the question

  • Calling every agent failure 'hallucination' regardless of cause
  • Fixing a stuck loop by raising the iteration cap
  • Treating a wrong-tool pick and an invented tool as one bug
  • Assuming a larger model eliminates argument-validation failures
  • Tagging one run with every symptom it displayed

context

open as a page

In agent evals, what does pass^k measure that a single run per task hides?

level: middleimportance: must knowfreq 68%

basics

~20 s

pass^k is the share of tasks an agent solves on all k repeated attempts, so it measures reliability rather than best-case ability. A single run hides variance: roughly 90% per-attempt success clears all 8 attempts on only about 43% of tasks.

open as a page

Why score an AI agent's trajectory and not just its final answer?

level: middleimportance: must knowfreq 72%

basics

~20 s

Final-answer scoring cannot tell a clean four-step solution from a forty-step flail that stumbled onto the same output, and it credits answers reached by luck. Trajectory metrics score the path: tool choices, ordering, and wasted steps.

open as a page

How do you score tool selection and argument correctness in an agent trajectory?

level: middleimportance: must knowfreq 58%

basics

~20 s

Treat the reference trajectory's tool calls as ground truth. Precision is the share of the agent's calls that appear in the reference; recall is the share of reference calls the agent made. Then, on matched calls only, score arguments separately.

open as a page

An agent claims success after its tool returned HTTP 502 — what failure mode is this?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A false success claim: the agent narrates completion over a failed side effect. It is dangerous because nothing crashed, so the run looks green while the work never landed. Only an independent post-condition check catches it.

open as a page

In an agent eval harness, how do you reset state between repeated runs?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Every task owns a seeded initial state that the harness restores before each attempt — a container rebuilt from a per-task image, a database restored from a snapshot or rolled back in a transaction, a scratch filesystem recreated. Each concurrent rollout gets its own isolated instance.

open as a page

When does an agent eval task need a judge model instead of a programmatic verifier?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Use a programmatic verifier whenever success has a checkable end state — asserted database rows, a passing test suite, a valid schema. Reach for a judge model only for the residue that has no machine-checkable form, such as whether a drafted customer message was accurate and appropriately toned.

open as a page

In an LLM agent, what is a hallucinated tool call versus a wrong-tool pick?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A hallucinated tool call names a tool that does not exist in the definitions the agent was given. A wrong-tool pick is a valid call to a real tool that cannot answer the request. Different causes, different fixes.

open as a page

When should an agent eval harness stub, record, or call third-party tools live?

level: middleimportance: should knowfreq 48%

basics

~20 s

Stub cheap, self-owned behaviour where you control the contract; replay recorded real responses when fidelity matters and determinism is required; call live only in a small, separate suite whose job is detecting that the real API has drifted away from your stubs and recordings.

open as a page

After 60 steps an agent ignores its original constraints — what failure is that?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Goal drift, also called instruction forgetting: constraints stated at the start lose influence as later observations dominate the context. The agent still works competently, just toward a subtly different objective than the one it was given.

open as a page

When should trajectory comparison be exact-match, any-order, or subsequence?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Match as strictly as the task genuinely constrains, and no stricter. Exact-match only where every step and its position are mandated; any-order for independent steps; subsequence when relative order matters but extra steps in between are acceptable.

open as a page

When does an LLM judge over the full trace beat reference-trajectory matching?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Use a judge when the quality you care about cannot be expressed as string comparison — was the plan sensible, did the agent recover well, did it stop at the right time — or when the task has too many valid paths to enumerate references for.

open as a page

How do you split an agent eval suite across CI tiers under a fixed cost budget?

level: principalimportance: should knowfreq 38%

basics

~20 s

Spend the cheap tier on breadth and the expensive tier on depth. A small stubbed subset at one rollout per task runs on every change to catch breakage; the full suite at several rollouts, with judges and live-ish tools, runs nightly inside a hard spend and wall-clock cap.

open as a page

How heavily should trajectory metrics gate an agent release without freezing it onto one path?

level: principalimportance: should knowfreq 30%

basics

~10 s

Let outcome correctness and hard safety constraints block, and keep path-conformance metrics advisory or budgeted. Gating on similarity to a reference path selects for imitating yesterday's solution and scores genuine improvements as regressions.

open as a page