skip to content

How would you measure tool-selection accuracy as teams keep adding tools to a shared agent?

level: principalimportance: should knowfreq 38%

answer

  1. one turn, one catalog, one decision
  2. some turns should call nothing
  3. confusion by pair, not one average
  4. ablate the catalog to get a curve
  5. new tools should have to earn entry

basics

~20 s

Score selection on its own: a labelled set of single user turns, a fixed catalog, one tool call each, compared against the gold tool — including turns whose correct answer is to call nothing. Re-run it at several catalog sizes and gate new tools on it.

solid answer

~50 s

Treat selection as a measurable unit separate from end-task success, because an agent can recover from a bad pick and can also fail for reasons unrelated to it. Build a labelled set of single user turns, each annotated with the tool a human says is correct — and include a substantial slice of turns where the right behaviour is to call nothing at all, since a growing catalog mostly increases false-positive calls. Score correct-tool rate, per-tool confusion, and abstention accuracy separately; BFCL-style suites treat irrelevance detection as its own category for exactly this reason. Then run the same set at deliberately different catalog sizes — 20, 100, the full catalog — to get your own decay curve rather than a published one. That curve is the governance instrument: a new tool ships only if it does not drop selection accuracy on the existing suite beyond an agreed threshold, which turns "add another tool" from a free action into one with a measured price.

go deeper

for a junior

Know that you can test tool choice directly: give the model a request, see which tool it calls, compare against the answer a human labelled. Be able to say why some test cases should expect no tool call.

for a middle

Describe the scoring breakdown — correct-tool rate, abstention accuracy, argument correctness conditioned on the right tool — and why selection is measured separately from whether the task eventually succeeded.

for a senior

Show you would ablate the catalog to derive your own decay curve, read per-pair confusion rather than the average, pin model versions across comparisons, and weight errors by blast radius rather than counting them equally.

for a principal

Own the incentive problem: adding a tool is a local decision with a global cost, so make the cost visible and gate on it. Set the admission threshold, require labelled examples with every new tool, mandate retirement of unused ones, and decide how the suite is refreshed as traffic drifts.

## Why measure selection in isolation The tempting metric is end-task success, and it is the wrong instrument for this question. An agent that picks the wrong tool often recovers, seeing the error and trying again, so end-task success under-reports selection failures while quietly paying for them in latency, tokens and side effects. Conversely an agent can pick perfectly and still fail on arguments, on a broken backend, or on synthesis. Isolating selection gives you an attributable signal: this number moved because the catalog changed. This is a deliberately narrow, single-turn measurement — one request, one catalog, one decision — and it is a different instrument from grading a whole episode's trajectory, which answers a different question about how the agent worked through a task. ## Building the labelled set Start from real requests. Sample production turns, cluster them, and have a human annotate the tool that should have been called. Three properties matter. - **Coverage across the catalog**, not just the popular tools. The tools that get misselected are usually the long tail nobody wrote test cases for. - **Deliberate near-neighbour density.** Include the pairs you already suspect collide, over-represented relative to traffic, because those are what regress. - **A real abstention slice.** Twenty to thirty percent of the set should be turns where the correct behaviour is to call no tool: chit-chat, requests outside the agent's remit, requests needing clarification first. As catalogs grow, the dominant new error is not picking the wrong tool for a valid request — it is inventing a tool call for a request that needed none. A suite without abstention cases scores that failure as a perfect zero. ## What to score - **Correct-tool rate** on turns that should call a tool. - **Abstention accuracy** on turns that should not — false-positive call rate is the number to watch as the catalog grows. - **Per-pair confusion**, as a matrix. Aggregate accuracy hides the actual defect: a two-point overall gain can be a large win on one pair and a regression on another. The matrix tells you which two tools to merge or namespace. - **Argument correctness, separately.** Conditioned on the right tool being chosen, were the arguments right? Keeping these apart prevents a schema problem from being misdiagnosed as a catalog problem. ## The decay curve is the deliverable Run the same labelled set against deliberately truncated catalogs: a 20-tool subset, 100, the full set. That ablation gives you *your* curve — where accuracy starts sliding for your model, your domain, your naming conventions. Published figures are directionally useful and never transferable, because the shape depends entirely on how much your tools overlap. The curve tells you whether you are still on the flat part (add tools freely) or past the knee (every addition costs accuracy and you need search, disclosure or scoping first). ## Turning it into governance The reason this matters at scale is organizational. Adding a tool is a local decision made by one team, and its cost is global and invisible: everyone else's selection accuracy drops slightly. Without measurement, the incentive is entirely one-directional and catalogs grow until the agent is unreliable. With measurement, you can set an admission policy: - A proposed tool must come with labelled examples of requests it should win. - The full suite is re-run with the tool added. If overall selection accuracy or any specific pair drops past an agreed threshold, the tool does not ship as-is — it gets renamed into a namespace, merged with the tool it collides with, or scoped to a toolset. - Retirement is part of the policy. Tools with no wins in the suite and no traffic get removed; a catalog that only grows is a catalog that only degrades. ## Keeping the instrument honest The set decays. Traffic shifts, tools change, and a suite that was representative a quarter ago is not now. Refresh it from recent production traces on a schedule. Pin the model version when comparing runs, because a provider model update moves selection accuracy on its own and will otherwise be attributed to your catalog change. And keep the suite cheap enough to run on every catalog change — single-turn selection scoring needs one model call per case with no tool execution, which is exactly the kind of eval you can afford to run constantly. ## The honest limitation Single-turn selection accuracy is a proxy. It does not capture whether the agent recovers well, whether the wrong pick was harmless or destructive, or whether a sequence of individually reasonable picks adds up to the wrong approach. Weight errors by blast radius — misselecting between two read-only lookups is not the same as firing an irreversible write — and read the number alongside outcome measures rather than instead of them.

  • Why include turns where the correct answer is to call no tool at all?
    Because the dominant new failure as a catalog grows is a false-positive call — the model finds something plausible for a request that needed none. A suite made only of tool-worthy requests cannot see that failure and will score a badly over-eager agent as perfect. Reserve a fifth to a third of the set for chit-chat, out-of-remit and clarification-first turns, and report abstention accuracy separately.
  • Why not just track end-task success and skip a separate selection metric?
    Because the two decouple in both directions. Agents often recover from a wrong pick, so end-task success hides selection errors while still paying their latency, token and side-effect cost; and tasks fail for argument, backend or synthesis reasons that have nothing to do with the choice. Selection scored in isolation is attributable — it moves when the catalog moves — which is what a governance gate needs.
  • How do you stop a provider model update from being misread as a catalog regression?
    Pin the model version for comparison runs and re-baseline explicitly when you upgrade. Selection accuracy shifts on model changes alone, sometimes by several points, so an unpinned suite will attribute that to whichever tool happened to ship that week. Keep the last baseline's raw results so an upgrade can be measured as its own experiment.
  • What should the admission policy do when a new tool genuinely collides with an existing one?
    Not block it outright — resolve the collision. Options in order: merge with the existing tool if they do the same work, namespace both so the domain is explicit, or scope them into different toolsets so they never co-occur. Re-run the suite after the fix and require the specific confused pair to separate, not just the average to recover.

saying these in an interview costs you the question

  • Measures only end-task success and infers selection quality from it
  • Omits cases where the correct behaviour is to call nothing
  • Reports a single aggregate accuracy with no per-pair breakdown
  • Compares runs across different model versions without pinning
  • Treats catalog growth as free because no metric ever objects

context