Why does an agent's tool-selection accuracy fall as its catalog grows from 20 to 400 tools?
answer
- flat first, then falls
- paid on every turn
- buried in the middle of the list
- several defensible candidates per intent
- overlap between tools, not wording
basics
~20 sEvery tool definition is loaded into the prompt at once, so hundreds of them cost tens of thousands of tokens and force the model to discriminate between many near-identical options in a single pass. Crowding plus semantic overlap pushes selection accuracy down.
solid answer
~50 sSelection accuracy stays roughly flat while a catalog is small and then falls off as it grows. An enterprise assistant spanning HR, IT, facilities and payroll can sit around 94% correct-tool accuracy with 20 tools and land near 61% once 412 are loaded. Two forces drive that. First, **crowding**: every definition is in the prompt before the user's request is even read, so the relevant one is buried in the middle of a very long list where attention is weakest, and the whole thing is paid for on every turn. Second, and usually the larger effect, **semantic overlap**: at 400 tools you inevitably have several plausible candidates for "request time off" or "look up an employee", and the model has to break a tie the catalog never disambiguated. The fix is structural — load fewer definitions per turn and give the model a way to find the rest — not a better sentence on each tool.
go deeper
Know that every tool definition sits in the prompt, so a big catalog is expensive and the model has more chances to pick the wrong one. Be able to say that hundreds of tools degrade selection accuracy rather than just costing money.
Explain both mechanisms: token crowding with mid-prompt attention loss, and semantic overlap between tools from different teams. Be ready to argue why a bigger context window fixes neither, and to give a rough sense of the flat-then-falling curve.
Show you would quantify the curve for your own catalog before choosing a remedy, separate crowding from overlap with an ablation, and treat wrong selections as side-effecting incidents rather than free retries.
Own the growth policy: who may add tools, what evidence a new tool must bring, and how you keep a shared catalog from eroding as a dozen teams contribute to it. Frame the tradeoff as catalog breadth against per-turn discrimination, with measurement as the arbiter.
## What is being measured Tool-selection accuracy, in the narrow sense used here, is a single-turn measure: given one user request and one catalog of tool definitions, does the model call the tool a human labeller marked as correct? It deliberately excludes whether the whole task eventually succeeded, because an agent can recover from a bad first pick, and can also fail a task for reasons that have nothing to do with which tool it chose. Isolating selection is what makes the decay visible at all. ## The shape of the curve The decay is not linear and not smooth. Small catalogs — roughly a dozen or two well-separated tools — sit near the model's ceiling, and adding a few more costs almost nothing. Somewhere past that, accuracy starts sliding, and it slides fastest when the newly added tools are near-neighbours of existing ones rather than genuinely new capabilities. A realistic enterprise pattern: 20 tools at about 94% correct selection, 400+ tools at about 61%. The exact numbers depend on the model, the catalog and the request mix, which is precisely why you measure the curve for your own catalog rather than trusting a published figure. ## Force one: crowding the context Tool definitions are just tokens in the prompt. A definition with a name, a paragraph of description and a JSON Schema of parameters commonly costs several hundred to a thousand tokens. Four hundred of those can be tens of thousands of tokens spent before the model has read a word of the user's request, on every single turn of the conversation. Three consequences follow. Cost and latency rise on every call, whether or not any tool is used. Prompt-cache economics get worse as the static block grows. And the relevant definition ends up somewhere in the middle of a very long, very homogeneous list — the position where long-context recall is empirically weakest, the failure the field calls context rot or being lost in the middle. ## Force two: semantic overlap Crowding is the obvious effect; overlap is usually the bigger one. Independent teams contribute tools, and at scale their vocabularies collide. HR ships a leave-request tool; payroll ships a wage-advance tool that its own docs describe in terms of "time off" and "pay period"; facilities ships a desk-booking tool whose description mentions "absence". None of them is wrong in isolation. Together they present the model with three defensible answers to "I'm off next Tuesday, sort it out", and nothing in the catalog says which one owns that intent. This is why polishing one tool's wording rarely moves the needle: the error lives in the *relationship between* definitions, not inside any one of them. ## Why a bigger context window is not the answer The reflex answer — "windows are huge now, just load everything" — misses both forces. A larger window does nothing about overlap; the three time-off candidates are equally ambiguous whether the prompt is 30k or 300k tokens. And it makes crowding worse rather than better, because the tokens are still paid for on every turn and long-context attention degradation is exactly what large windows expose you to. Capacity to fit the catalog is not the same as capacity to discriminate within it. ## What the decay costs you A wrong selection is not a free retry. The tool runs, so a wrong pick can have side effects — a payroll advance requested when the user wanted leave booked. It burns turns and tokens as the agent notices and backs out. And it degrades trust in ways users report as "the assistant is unreliable" long before any dashboard shows it, because the end-task sometimes still completes. ## The structural responses Everything that actually works reduces how many definitions are in front of the model at decision time, or makes the survivors distinguishable. Server-side tool search over deferred definitions keeps the catalog out of the prompt and injects only the few that match. Progressive disclosure loads a one-line metadata tier per tool and the full usage instructions only after activation. Namespacing and toolsets separate the colliding pairs by name and by which ones are eligible at all for a given role or request class. And an admission policy — a new tool has to earn its place against a labelled selection suite — stops the catalog from eroding quietly as teams keep adding to it.
- Would a much larger context window make this problem go away?No. A bigger window lets the definitions fit, but it does nothing about semantic overlap — three plausible time-off tools stay equally ambiguous. It also worsens the crowding side: the tokens are still spent on every turn, and long homogeneous lists are exactly where mid-prompt recall degrades. Fitting the catalog is not the same as discriminating within it.
- How would you tell crowding apart from overlap as the dominant cause in your own catalog?Ablate. Run the labelled selection set against a small catalog containing only the suspected colliding pairs, then against the full catalog with the collisions removed. If errors concentrate on a handful of tool pairs in a confusion matrix, it is overlap; if they are diffuse and worsen mainly as the prompt lengthens, it is crowding. Most real catalogs show both, with overlap dominant.
- Does the decay hit argument correctness the same way it hits tool choice?Not equally. Once the right tool is selected, its schema is in context and argument filling is largely a local problem, so argument accuracy degrades far more slowly with catalog size. That is why the two are scored separately: a catalog-scaling problem shows up as wrong-tool errors, while argument errors usually point at the schema instead.
saying these in an interview costs you the question
- Says a larger context window makes the problem disappear
- Treats prompt tokens as the only cause, ignoring semantic overlap
- Claims accuracy falls linearly and predictably with tool count
- Proposes rewriting each description as the complete fix
- Dismisses wrong selections as harmless because the agent can retry