Why score an AI agent's trajectory and not just its final answer?
answer
- the answer is not the whole story
- one bit versus a whole path
- right result, forty wrong steps
- luck does not survive a rerun
- precision, arguments, ordering, waste
basics
~20 sFinal-answer scoring cannot tell a clean four-step solution from a forty-step flail that stumbled onto the same output, and it credits answers reached by luck. Trajectory metrics score the path: tool choices, ordering, and wasted steps.
solid answer
~50 sOutcome metrics answer one question — is the final answer or end state correct? Process (trajectory) metrics ask how the agent got there: did it choose the right tools, call them with the right arguments, in a workable order, and without waste. The gap matters because agents are stochastic. An agent that reaches the right answer by brute-forcing every tool is fragile: change the phrasing, the corpus or the model version and the luck evaporates. Outcome-only scoring also hides cost — a support agent that issues `list_tickets` three times in one episode passes the outcome check while burning triple the tokens and latency. And a pass/fail bit gives you nothing to debug; a trajectory tells you *which* step first went wrong. In practice: keep outcome correctness as the thing you ship on, and use process metrics to explain results, predict reliability and enforce budgets.
go deeper
Know the vocabulary: outcome metrics look at the final answer, trajectory metrics look at the steps taken to get there. Be able to give one example of a run that passes the first and fails the second.
Be ready to name concrete process metrics — tool-selection correctness, argument correctness, step count, redundant-action rate — and explain why a right answer via a forty-step path is a real defect rather than a cosmetic one.
Show you have used trajectories to debug production agents: pointing at the first deviating step, spotting repeated calls that signal the agent is not seeing its own results, and catching cost regressions before outcome scores move.
Own the framing that outcome correctness blocks releases while process metrics inform and budget them. Be able to argue why over-weighting trajectory conformance freezes the system onto yesterday's path and what you would measure instead.
## Two families of metric Every agent evaluation reduces to two questions asked at different granularity. **Outcome metrics** look only at the terminal state: did the agent return the expected answer, is the ticket now closed, does the file now contain the right content? They are cheap, unambiguous and are what a user actually cares about. **Process metrics**, also called trajectory metrics, score the sequence of steps that produced that state — the ordered list of tool calls, their arguments, the observations returned, and (where available) the model's intermediate reasoning. A trajectory is just the episode's execution record. Scoring it means comparing that record against expectations: a reference trajectory an expert authored, a set of constraints ("the ticket must be created before the log is attached"), a budget ("no more than eight tool calls"), or a rubric applied by a judge model. ## What outcome-only scoring hides **Right answer, wrong reasons.** With a stochastic policy, some fraction of correct answers are accidents — the agent guessed, or it happened to have the answer memorized and never used the tool that was supposed to retrieve it. Those passes do not generalize. Rerun the same task with a paraphrased prompt and they disappear. **Cost and latency blowups.** An agent that reaches a four-step answer in forty steps passes every outcome check and is ten times more expensive. Because token cost and wall-clock time are roughly linear in step count, path length is the single strongest predictor of the bill. **Side effects.** Outcome checks usually inspect one target object. An agent that reached the right end state by also mutating three unrelated records looks identical to a clean one unless you read the trajectory. **No debugging signal.** A failed outcome is one bit. A failed trajectory tells you the first step that deviated, which is where the fix lives — a bad tool description, a missing argument, a retrieval miss. ## The process metrics that matter - **Tool-selection correctness** — did the agent call the tools the task requires, and only those? Usually expressed as precision and recall over tool names against a reference. - **Argument correctness** — for the calls that matched, were the parameters right? Scored separately, because the right tool with wrong arguments is a different defect from the wrong tool. - **Ordering conformance** — were sequencing constraints respected, using exact, any-order or subsequence comparison depending on how much order the task actually fixes. - **Step count and efficiency ratio** — number of steps taken, often normalized as steps taken divided by steps in the reference. A ratio near 1.0 is ideal; 10.0 is a flail. - **Redundant-action rate** — the share of steps that repeat an earlier call with identical arguments and return no new information. This is the cleanest single detector of a spinning agent, and it is directly actionable: the fix is usually to surface prior results better in context. - **First-deviation index** — where in the episode the path first left the acceptable set. Aggregated across a suite, it shows whether failures are early planning errors or late execution errors. ## Why process predicts reliability Reliability is about the *distribution* of runs, not one run. A short, canonical path has fewer places to go wrong, so it survives repetition and small input perturbations. A long path multiplies per-step error: even at 97% per-step correctness, forty steps is a coin flip. Trajectory metrics therefore give you an early-warning signal — path length and redundancy creep upward before outcome scores fall, because the agent's slack absorbs the degradation until it cannot. ## Where process metrics mislead Most real tasks have more than one correct path. Scoring against a single reference trajectory penalizes an agent that found a better route — a newer model that solves in three steps what your reference does in seven will look like a regression under a strict comparison. Step count can also be gamed: an agent that crams work into one enormous call scores well on efficiency while being harder to debug and recover. And process scores are not what the user experiences; nobody ships an elegant trajectory that returned the wrong answer. ## Putting it together The standard arrangement is to compute both from the same run: outcome correctness is the blocking criterion, and process metrics act as diagnostics plus budget checks (step-count and redundant-action thresholds). Store the full trace for every eval run whether it passed or failed — the trajectory of a *passing* run is where you find the luck, and that is exactly the failure you would otherwise ship.
- How would you actually detect that a passing run got the right answer by luck?Look for the causal step in the trace: if the task requires retrieving a value and the trajectory contains no successful retrieval before the answer, the model produced it from parameters or guessed. Two supporting signals are reruns of the same task under paraphrase, and ablating the tool the answer supposedly came from — a truly grounded agent fails cleanly without it, while a lucky one keeps answering.
- If step count is a metric, what stops an agent from gaming it?Nothing, if it is scored alone. An agent can collapse work into one oversized call, or skip verification steps that the reference includes, and score better on efficiency while being less reliable. Efficiency is only meaningful as a tiebreaker among trajectories that already pass the outcome check and any required-step constraints, never as an independent objective.
- What is redundant-action rate and why is it a better loop signal than step count?It is the share of steps that repeat an earlier call with the same arguments and yield no new information. Step count is confounded by task difficulty — a hard task legitimately takes more steps — while repetition is difficulty-independent. A rising redundant-action rate almost always means the agent is not seeing or trusting its own prior results, which points at a context or result-formatting fix.
saying these in an interview costs you the question
- Says the final answer is all that matters if it is correct
- Assumes a correct answer proves the tool calls were correct
- Treats fewer steps as strictly better regardless of the task
- Confuses collecting traces with scoring them
- Uses process metrics as the blocking gate and ignores outcome