skip to content

How do you score tool selection and argument correctness in an agent trajectory?

level: middleimportance: must knowfreq 58%

answer

  1. two failable parts per call
  2. ground truth is the reference call set
  3. substitution costs you twice
  4. score parameters only on matched calls
  5. IDs exact, free text fuzzy

basics

~20 s

Treat the reference trajectory's tool calls as ground truth. Precision is the share of the agent's calls that appear in the reference; recall is the share of reference calls the agent made. Then, on matched calls only, score arguments separately.

solid answer

~50 s

Score the two dimensions independently. **Tool selection** is a set-comparison problem: with the reference trajectory's calls as ground truth, precision = correct calls / calls the agent made, and recall = correct calls / calls the reference made. Low precision means the agent invokes tools it should not; low recall means it skips work. A helpdesk agent that calls `reset_password` where the reference calls `unlock_account` takes a double hit — a false positive for the tool it did call and a false negative for the one it did not. **Argument correctness** is then scored only over the calls whose tool name matched, because the right tool with wrong parameters is a different defect with a different fix (schema and description work, not routing work). Report it as exact-match on the argument object, or field-level accuracy when some fields are free text and need fuzzy or judge-based comparison.

code

python · 20 lines
python
from collections import Counter

def tool_precision_recall(predicted, reference):
    pred = Counter(c["name"] for c in predicted)
    ref = Counter(c["name"] for c in reference)
    tp = sum((pred & ref).values())
    precision = tp / max(sum(pred.values()), 1)
    recall = tp / max(sum(ref.values()), 1)
    return precision, recall

def argument_accuracy(predicted, reference):
    matched = [(p, r) for p, r in zip(predicted, reference) if p["name"] == r["name"]]
    if not matched:
        return None
    return sum(p["args"] == r["args"] for p, r in matched) / len(matched)

pred = [{"name": "reset_password", "args": {"user": "u1"}}]
ref = [{"name": "unlock_account", "args": {"user": "u1"}}]
print(tool_precision_recall(pred, ref))  # (0.0, 0.0)
print(argument_accuracy(pred, ref))      # None - no matched call to score

go deeper

for a junior

Be able to say that scoring a tool call means checking two things — the tool the agent picked and the parameters it passed — and that these are counted separately.

for a middle

Be ready to define precision and recall over tool calls with the reference as ground truth, explain what each error type means in practice, and show that a substituted tool counts as both a false positive and a false negative.

for a senior

Demonstrate that you read these per tool and per argument field, weight side-effecting false positives harder, and handle the messy parts: pairing repeated calls, tolerance for numeric and date parameters, and fuzzy scoring for free-text arguments.

for a principal

Own the limits of conformance scoring. Argue when precision and recall against a single reference stop being the right instrument, and what constraint- or rubric-based scheme you would adopt for tasks with several equally valid tool sets.

## Why two metrics, not one A tool call has two independently failable parts: *which* tool, and *what you pass it*. Collapsing them into a single "tool-call accuracy" number destroys the information you need to act. Wrong tool means the agent's routing is broken — the tool descriptions overlap, or the task framing does not distinguish them. Right tool with wrong arguments means the schema is ambiguous, a required field lacks an enum, or the agent lacks the context to fill a parameter. The remedies are different, so the metrics must be separate. ## Tool selection as precision and recall Take the reference trajectory's tool calls as the ground-truth set and the agent's calls as the predicted set. Then: - **True positives** — calls whose tool name appears in both. - **False positives** — calls the agent made that the reference did not (extra, wrong, or hallucinated tools). - **False negatives** — calls the reference made that the agent did not (skipped work). Precision = TP / (TP + FP); recall = TP / (TP + FN); F1 combines them. Use *multiset* comparison, not plain sets: if the reference calls a search tool twice and the agent calls it five times, three of those are false positives, and a plain set intersection would hide that. The two error types have very different production meanings. Low recall is under-action: the agent answers from memory instead of retrieving, or gives up early. Low precision is over-action, which is the more dangerous one when tools have side effects — an agent that calls a refund tool it should not have called scores as one false positive and one very unhappy customer. For that reason many suites weight false positives on mutating tools far more heavily than on read-only ones, or track a separate hard-fail flag for any unauthorized side-effecting call. ## Substitution is a double error The most common real failure is substitution, not omission. A helpdesk agent that calls `reset_password` where the reference calls `unlock_account` has done something plausible and wrong: the user is unlocked as a side effect but their password now no longer works. In the metric this registers twice — a false positive for `reset_password`, a false negative for `unlock_account` — which is the correct accounting, because both the wrong action and the missing action are real defects. It also means substitution errors depress precision and recall symmetrically, which is a useful fingerprint when you scan aggregate scores: a suite where both numbers dropped together usually has a routing problem between two similar tools, while recall alone falling usually means the agent stopped early. ## Argument correctness Score arguments only over the calls whose tool matched — arguments to a wrong tool are meaningless. Options, roughly in order of strictness: - **Exact match** on the whole argument object. Simple, brittle, and the right choice when the parameters are IDs and enums. - **Field-level accuracy** — score each parameter, then average. This distinguishes "one optional field wrong" from "every field wrong" and points directly at the offending parameter when you aggregate per field. - **Typed comparison** — exact for IDs and enums, numeric tolerance for amounts, normalized comparison for dates and casing, and semantic or judge-based comparison for free-text fields like a search query or a message body, where many phrasings are equally right. Always separate *invalid* arguments from merely *wrong* ones. An argument that violates the schema is caught by the runtime and is a validation failure; an argument that is well-formed but names the wrong account is a correctness failure that only the reference comparison catches. Reporting them together makes the schema-tightening work invisible. ## Matching calls before you can score arguments An overlooked mechanical detail: to score arguments you must first *pair* each agent call with a reference call. When a tool appears once, pairing is trivial. When it appears several times with different arguments, naive positional zipping mis-pairs and reports argument errors that are really ordering differences. Pair by best argument similarity within the same tool name, or make the reference explicit about which occurrence is which. ## Reading the numbers Interpret these per tool, not just in aggregate. A single overlapping pair of tools can drag a suite-wide precision score down while every other tool is perfect, and only the per-tool breakdown makes that visible. And remember the metric's ceiling: precision and recall against a reference measure conformance to *one* expert path. They are the right instrument for conformance and the wrong one for tasks where several tool sets are equally valid — there you fall back to constraint checks ("any retrieval tool must be called before answering") or a judge over the trace.

  • Why should false positives on mutating tools be weighted more heavily than on read-only tools?
    Because an unnecessary read costs tokens while an unnecessary write costs real-world state — a refund issued, an account disabled, a message sent. Weighting them equally lets an agent trade a dangerous extra write against a harmless extra read and still score well. Many suites go further and treat an unauthorized side-effecting call as a hard fail for the episode rather than a fractional precision penalty.
  • How do you score an argument that is free text, like a search query?
    Exact match is wrong there — many phrasings are equally good. Score it by whether it retrieves the required information (an outcome-flavoured check on the observation), by semantic similarity against the reference query above a validated threshold, or with a judge model given a narrow rubric. Keep it as a separate field-level score so that one fuzzy parameter does not turn the whole call's argument accuracy into a fuzzy number.
  • Your suite shows precision 0.62 and recall 0.94. What does that pattern suggest?
    The agent is doing all the required work plus extra: it rarely skips a needed call but frequently adds unnecessary ones. Typical causes are overlapping tool descriptions that make several tools look applicable, speculative parallel calls, or retries after results the agent did not recognize as sufficient. Look at the per-tool breakdown first — a single ambiguous pair usually accounts for most of the false positives.

saying these in an interview costs you the question

  • Reports one blended tool-call accuracy number and cannot decompose it
  • Scores arguments on calls where the tool name did not match
  • Uses set intersection so repeated calls never count as errors
  • Treats an extra read-only call the same as an extra refund call
  • Calls a well-formed but wrong argument a schema violation

context