skip to content

How many tools should one OpenAI request expose, and what breaks with too many?

level: principalimportance: nice to knowfreq 32%

answer

  1. Two costs, only one is tokens
  2. Definitions ride along on every turn
  3. Overlap, not count, is the real problem
  4. Merge, then subset, then split
  5. A stable block is a cheap block

basics

~20 s

Tool definitions are serialized into the prompt and billed as input tokens on every request, and selection accuracy falls as similar tools multiply. Keep the per-request set small and distinct; route or split across sub-agents rather than growing one surface.

solid answer

~50 s

Two costs scale with the tool surface. The first is mechanical: the whole `tools` array — names, descriptions and JSON Schemas — is serialized into every request, including every iteration of an agent loop, so a large surface is a fixed tax on each turn and on the transcript's context budget. The second is behavioural and worse: as tools overlap in purpose, the model picks the wrong one, invents arguments to fit an approximate match, or oscillates between two near-duplicates. Mitigations, in order of payoff: merge near-duplicate tools into one with a parameter; write descriptions that say when *not* to use the tool; select a task-relevant subset per request, often via embedding retrieval over tool descriptions; and split genuinely distinct domains into sub-agents each with a narrow surface. Keep the block stable and prefix-positioned so caching can absorb what remains.

go deeper

for a junior

Know that every tool you attach is text in the prompt, so it costs tokens on each request, and that a shorter, clearer list helps the model choose correctly.

for a middle

Explain both costs — linear token overhead per turn and non-linear selection error from overlapping tools — and the first fixes: merge near-duplicates, and write descriptions that say when not to use a tool.

for a senior

Show the operational loop: per-tool usage telemetry, an eval set that scores selection accuracy, retrieval-based subsetting with an always-on core, and awareness that a varying tool block breaks prefix caching.

for a principal

Treat the tool surface as a designed API with an unusual consumer — argue the decomposition into sub-agents or a router on domain and permission boundaries, and defend the cost and reliability tradeoff with measurements rather than instinct.

## Two different costs Ask "how many tools" and there are two answers, and only one of them is about tokens. **The token cost is linear and predictable.** Every element of `tools` — name, description, and the full JSON Schema for parameters — is serialized into the prompt on every request that carries it. Forty tools with thoughtful descriptions and nested schemas can easily be several thousand tokens. In an agent loop that runs eight iterations, you pay that eight times, on top of a transcript that is itself growing. It also eats context window you would rather spend on retrieved documents and tool results. **The accuracy cost is non-linear and is what actually breaks products.** Tool selection is a discrimination task performed from descriptions alone. When two tools are semantically adjacent — `search_docs` and `search_kb`, `get_user` and `fetch_customer` — the model has no ground truth to separate them, so it guesses, and the guess is stable in the wrong direction for whole classes of query. Symptoms: the wrong tool called with plausible arguments; arguments invented to bridge a near-fit; alternation between two similar tools across iterations; and increased no-tool answers because nothing looked clearly right. ## Reduce before you route The first move is not architecture, it is editing. Merge near-duplicates into one tool with a discriminating parameter — one `search(source: enum)` beats four source-specific search tools, because the enum makes the choice explicit and schema-constrained rather than implicit and prose-mediated. Delete tools nobody has called in a month; usage telemetry per tool name is cheap to collect and immediately actionable. Then rewrite descriptions as selection guidance rather than documentation. The highest-value sentence in a tool description is often negative: "Use for current inventory levels only; for historical trends use `sales_report`." That single line resolves a whole confusion class. Names matter for the same reason: distinct verbs and nouns discriminate better than a shared prefix with a suffix that varies. ## Then route When the catalogue is genuinely large — dozens of unrelated capabilities — the answer is to stop sending it whole. Two patterns dominate: **Per-request subsetting.** Embed each tool's description once, embed the incoming request, and attach only the top-k relevant tools. This keeps the surface small for the model while the catalogue stays large for the system. The risk is recall: if retrieval misses the right tool it becomes invisible, so pin a small always-on core (clarify, hand off, finish) and monitor for tasks that fail because the needed capability was never offered. **Sub-agents.** Split by domain — a billing agent, a search agent, a scheduling agent — each with a handful of tools, fronted by a router whose own tool surface is one delegate function per domain. Each model call then makes an easy choice from a short list. The cost is an extra hop of latency and the coordination burden of passing context between agents, so it pays off at the point where a flat surface has demonstrably stopped discriminating, not before. ## The caching angle Whatever surface remains, keep it stable and at the front of the prompt. Automatic prefix caching only helps when the leading bytes are identical across requests, so a tool block that is serialized in a nondeterministic order — iterating a Python dict built from a registry, say — defeats it silently. Sort deterministically, keep the tool block ahead of anything volatile like timestamps or user-specific preamble, and the fixed tax on each loop iteration is substantially discounted. Note that dynamic per-request subsetting works against this: each distinct subset is a distinct prefix. On high-volume traffic that tension is worth measuring — sometimes a slightly larger, always-identical surface is cheaper in practice than a minimal one that changes every request. ## Measure, do not guess None of the above is decidable from first principles for your domain. Build a small eval set of representative requests labelled with the tool that *should* be called, and score selection accuracy as you change the surface. That converts "is forty tools too many?" from an opinion into a number, and it catches the specific pairs the model confuses — which is the actionable finding, since fixing two descriptions usually beats re-architecting into sub-agents. ## The judgement to voice in an interview A large tool surface is a symptom of pushing an organizational decomposition problem into the prompt. The tools a model sees are an API you designed for a rather literal-minded consumer that reads only the docs and never the source. The right instinct is the one you would apply to any public API: fewer, better-named, well-documented operations with orthogonal purposes — and separate services when the domains are genuinely separate.

  • How would you decide between per-request tool subsetting and splitting into sub-agents?
    Subsetting keeps one agent and one transcript, so it is far simpler operationally — reach for it when the tools are numerous but the task is one coherent job. Sub-agents pay off when the domains have genuinely different context needs, different system prompts, or different permissions, since each agent then carries only what it needs. Cost of sub-agents is an extra hop and the work of passing context across the boundary, so do not adopt them before a flat surface has measurably stopped discriminating.
  • Does dynamically selecting tools per request hurt prompt caching?
    Yes, and the tension is worth measuring rather than assuming. Prefix caching depends on the leading bytes being identical, so every distinct tool subset is a distinct prefix and a distinct cache entry. At high volume a slightly larger but always-identical tool block can be cheaper overall than a minimal one that varies per request. Whichever you choose, serialize the block deterministically — a nondeterministic ordering defeats caching silently, with no error to tell you.
  • How do you measure whether your tool surface is actually too large?
    Build a labelled eval set: representative user requests paired with the tool that should be called, then score selection accuracy as you vary the surface. The valuable output is not the headline number but the confusion pairs — which two tools the model swaps. Those usually point at overlapping descriptions or names, and rewriting a couple of descriptions typically recovers more accuracy than any architectural change.

saying these in an interview costs you the question

  • Assuming tool definitions are free because they are not messages
  • Adding a new tool per endpoint instead of parameterising one
  • Believing the model reads your code, not just the description
  • Treating tool count as the problem when overlap is the real cause
  • Re-architecting into sub-agents before fixing ambiguous descriptions

context