skip to content

In ADK's adk eval, what do tool_trajectory_avg_score and response_match_score measure?

level: middleimportance: must knowfreq 62%

answer

  1. one exact, one fuzzy
  2. tool calls compared name and args
  3. final text scored on unigram overlap
  4. defaults 1.0 and 0.8
  5. criteria live in a folder-scoped config

basics

~20 s

tool_trajectory_avg_score compares the tool calls the agent actually made — names and arguments, in order — against the ones recorded in the eval case, averaged over invocations. response_match_score compares the final response text to the expected text using ROUGE-1 similarity.

solid answer

~40 s

They are the two scorers `adk eval` applies to every recorded case, and they answer different questions. `tool_trajectory_avg_score` is an exact comparison: for each invocation, the actual sequence of tool calls with their arguments must match the recorded `tool_uses`, scoring 1.0 or 0.0, and the run reports the average across invocations. Its default pass threshold is 1.0, so one extra or reordered call fails the case. `response_match_score` scores the agent's final text against the recorded `final_response` with ROUGE-1 overlap, defaulting to a 0.8 threshold, because natural language will never match verbatim. You override either bar in a `test_config.json` `criteria` block placed beside the eval files. The pairing is the point: trajectory catches a right answer reached by the wrong path, text match catches the path being right while the answer degraded.

code

json · 6 lines
json
{
  "criteria": {
    "tool_trajectory_avg_score": 1.0,
    "response_match_score": 0.7
  }
}

go deeper

for a junior

Recall that ADK scores two things per case — whether the expected tool calls happened, and how closely the final text matches the recorded answer — and that both have default pass thresholds.

for a middle

Explain the mechanics: exact name-and-argument matching averaged over invocations at a 1.0 default, ROUGE-1 unigram overlap at 0.8, and criteria set in a test_config.json beside the files.

for a senior

Show you read the two metrics together to diagnose — trajectory-only failures mean a route change, response-only failures mean language or data drift — and that you re-record rather than quietly lowering bars.

for a principal

Own the limits: these are cheap deterministic tripwires, not quality judgments. Argue for what belongs at this tier versus a costlier judged tier, and for folder-scoped criteria so one loose bar cannot leak across the suite.

## What adk eval actually computes When `adk eval` replays a recorded case, it runs the agent against each recorded user turn and collects two things: the sequence of tool calls the agent emitted, and the final text it produced. Those are compared against what the eval case recorded, and turned into two named metrics. ## tool_trajectory_avg_score For each invocation, ADK lines up the actual tool calls against the `tool_uses` recorded under `intermediate_data`. The comparison is exact — tool name and argument values — and order-sensitive. An invocation scores 1.0 when the actual calls match the expected ones and 0.0 otherwise; the reported metric is the average over all invocations in the case, hence *avg*. The default pass threshold is **1.0**. That is a deliberate choice: for a deterministic tool path, anything less than a perfect match means the agent did something you did not sanction — called an extra tool, skipped one, or passed a different argument. It also makes the metric brittle by construction. An agent that calls `get_weather(city="London")` where you recorded `get_weather(city="london")` fails. An agent that calls a search tool twice because the first result was thin fails, even though the answer improved. When you knowingly accept variation, you lower the threshold, but a threshold below 1.0 on a two-call trajectory is a very coarse dial — 0.5 means "half the invocations may be wrong", not "a bit of wobble is fine". ## response_match_score The final response is scored with ROUGE-1 similarity against the recorded `final_response`: unigram overlap between the produced text and the expected text. The default threshold is **0.8**. Text scoring has to be fuzzy — no LLM emits the same sentence twice — but ROUGE-1 is a lexical measure, not a semantic one. Two consequences you should be able to name in an interview: - A correct answer phrased differently can score low. "London is cloudy, 15 degrees" against a recorded "It is 15C and cloudy in London" loses overlap on function words and formatting. - A wrong answer phrased similarly can score high. Swap the temperature and most unigrams still match. ROUGE-1 does not know which token carried the fact. So `response_match_score` is a **regression tripwire**, not a correctness oracle: it tells you the answer moved, and you look. Treating it as proof of correctness is the classic misuse. ## Setting criteria Neither threshold lives in the eval case. You put a `test_config.json` next to the eval files with a `criteria` object naming the metrics and their minimums; it applies to the eval files in that folder, and `adk eval --config_file_path` can point at one explicitly. Folder scoping is useful in practice: keep a folder of deterministic tool tests at trajectory 1.0, and a folder of open-ended conversational cases at a lower response bar, without touching a single case. When no config is supplied, the defaults above apply — trajectory 1.0, response match 0.8. ## Reading a failure The useful discipline is to read the two together: - **Trajectory fails, response passes** — the agent found the answer another way. Either the tools changed, or the model drifted to a different route. Often the correct response is to re-record, but first ask whether the new route costs more calls or touches something it should not. - **Trajectory passes, response fails** — the plumbing is intact and the language moved. Usually a prompt or model change; sometimes live tool data changed underneath a recorded case. - **Both fail** — treat it as a real behavioural regression, not a scoring artifact. `--print_detailed_results` gives the per-case breakdown you need to make that call rather than staring at a pass/fail summary. ## Where the metrics stop These two scorers are cheap, deterministic and offline, which is exactly why they belong in a fast loop. They cannot judge whether an answer was *good*, whether the agent was efficient, or whether a refusal was appropriate. Those judgments need a different instrument, and ADK's metric set is not the place to fake them. The productive framing in an interview: the ADK metrics are the regression net, and anything that needs judgment is a separate, more expensive tier you run less often.

  • Why is the default tool trajectory threshold 1.0 while the response threshold is 0.8?
    Tool calls are deterministic and machine-checkable: names and arguments either match what you sanctioned or they do not, so anything under a perfect match means unsanctioned behaviour. Natural language never reproduces verbatim, so a text metric has to allow slack; 0.8 ROUGE-1 overlap is the default compromise between catching real drift and firing on rewording.
  • An agent's response_match_score is 0.9 but the answer states the wrong temperature. Why did it pass?
    ROUGE-1 counts unigram overlap, not meaning. Change one numeric token in an otherwise identical sentence and the vast majority of unigrams still match, so the score barely moves. That is why the metric is a regression tripwire rather than a correctness check — factual assertions need a checker that looks at the fact, not the wording.
  • Your trajectory metric fails because the agent now calls a search tool twice instead of once. Lower the threshold or re-record?
    Decide from the cause, not the noise. If the second call is a genuine improvement you want to keep, re-record the case so the expected trajectory reflects the intended behaviour. Lowering the threshold silences the metric for every case in that folder, including ones where a stray extra call is exactly the failure you wanted to catch.

saying these in an interview costs you the question

  • Treating response_match_score as a semantic correctness check
  • Thinking trajectory scoring tolerates extra or reordered tool calls by default
  • Believing thresholds are configured inside the eval case file
  • Lowering thresholds to make a suite green instead of diagnosing the change
  • Assuming a passing eval run means the answer was good

context