Your agent's tool schemas cost 2,100 tokens per request — how do you cut that?
answer
- paid on every request, not once
- profile before you cut
- optional tail is usually the bloat
- flatten what mirrors your domain model
- smaller often scores better — verify it
basics
~20 sTool definitions are re-sent on every request, so their cost is per-turn, not one-off. Cut by removing parameters the handler can supply, flattening deep nesting, compressing per-field prose to decision-relevant text, and splitting overloaded tools — then verify accuracy did not drop.
solid answer
~50 sFirst, name the cost correctly: definitions are serialized into every request, so 2,100 tokens is 2,100 tokens per turn, competing for attention with the conversation itself. Then attack it in order of yield. Most bloat is thirty optional parameters that the handler could default or derive — delete them from the model's interface. Next, flatten nesting that mirrors your internal domain objects rather than any real grouping. Then rewrite descriptions so every sentence changes a decision, cutting implementation detail the model cannot act on. Finally, split tools that grew a mode switch and a matching set of conditionally-relevant fields. I have seen a 2,100-token schema fall to about 300 and *improve* tool-selection accuracy, because the smaller surface was also the clearer one. But that is an empirical claim, so I re-run a fixed set of representative turns after the edit rather than assuming smaller is better.
go deeper
Know that tool definitions are part of every request rather than registered once, so their size is a recurring cost in both tokens and context space.
Explain where the bloat lives — long optional tails, deep nesting, prose that documents instead of instructs — and the concrete cuts for each.
Show the empirical loop: profile the serialized definitions, cut in order of yield, then re-run a fixed set of turns including near-misses to confirm selection accuracy held. Be able to explain why a much smaller schema often scores better.
Own the budget: how much of the context window tool definitions may occupy, the review standard that keeps them there as teams add tools, and the point at which the problem stops being schema size and becomes catalog size.
## Why the number is per-turn Tool definitions are not registered once. They are part of the request, serialized into context ahead of the conversation on every call the agent makes. A twenty-turn task with 2,100 tokens of definitions pays that twenty times, and an agent loop with tool results flowing back pays it on every internal step too. Prefix caching softens the bill where the definition block is stable — it sits at the front and rarely changes — but caching reduces price, not context occupancy: those tokens still fill the window and still compete for the model's attention with the material that actually matters. That second cost is the one senior candidates raise. Long definition blocks push real conversation toward the middle of a long context, where models attend least reliably, and every extra optional field is one more thing the model has an opportunity to fill in. ## Where the tokens actually are Profile before cutting. Serialize the definitions exactly as sent, count tokens per tool and per field, and sort. The distribution is almost always lopsided, and the usual culprits are: - **A long tail of optional parameters.** Thirty optionals on one tool is a common shape and almost always a leaked internal API surface. - **Deep nesting.** Each level adds braces, a `properties` key, and its own `required` array before any content. - **Per-field prose that documents rather than instructs.** "The unique identifier of the record in the claims table" tells the model nothing it can act on. - **Verbatim duplication.** Six tools carrying the same three-sentence paragraph about the retry policy. - **Enum sprawl.** A fifty-literal vocabulary that changes twice a year. ## The cuts, in order of yield **1. Remove parameters the model should not be choosing.** Correlation ids, tenant ids, timestamps, channel flags, pagination cursors your loop manages — none of these need to be model-visible. Every field you move to the handler removes tokens *and* a fabrication surface. This is usually the biggest single win. **2. Flatten structure that mirrors your domain model.** Schemas that reproduce an internal object graph are paying for a hierarchy nobody needs. Keep nesting where the grouping is real and reused; collapse the rest into a handful of top-level fields. **3. Rewrite descriptions as instruction.** Apply the test: delete any sentence that would not change what the model does. Implementation notes, service names and table names almost never survive it. What stays is capability, preconditions, provenance and scope. **4. Split overloaded tools.** A tool with a mode parameter and two disjoint sets of conditionally-relevant fields is two tools wearing a coat. Splitting shortens each definition and makes selection easier, because each name and description now describes one thing. **5. Deduplicate shared policy.** Guidance that applies to every tool belongs once in the system prompt, not copied into each description. ## The counter-intuitive result A 2,100-token definition set with thirty optional fields can be beaten on accuracy — not just on cost — by a 300-token flattened version. This surprises people, but the mechanism is straightforward. The bloated schema was ambiguous: the model had to decide which of thirty optionals were relevant, and the answer was rarely stated. The flat one presents four decisions with clear answers. Clarity and compactness are correlated here far more often than they conflict. That correlation is not a law, though. Cutting too far removes the preconditions and scope statements that stop misfires, and the failure appears as a selection regression rather than a schema error, which makes it easy to misattribute. ## How to know you did not break it Hold a fixed set of representative turns — including the near-misses where the tool must *not* fire — and compare which tools got called with which arguments before and after the edit. Watch cost and accuracy together; a change that halves tokens and costs several points of correct selection is usually a bad trade, since a wrong call costs a full extra round trip plus a recovery turn. Change one thing at a time: tools interact, and a description you trimmed on tool A can change how often tool B fires. ## The limits of schema economy There is a floor. Once definitions are tight, remaining pressure comes from tool *count*, and that is a different problem with different instruments — loading strategies for large catalogs rather than shorter text. The distinction matters in an interview: schema economy is what you do first because it is cheap, safe and often improves quality; catalog-scale techniques are what you reach for when a lean set of definitions is still too many.
- Prompt caching covers the definition block. Does the token cost still matter?Yes. Caching lowers the price of re-sending a stable prefix, but the tokens still occupy the context window and still compete for attention with the conversation. It also has a sharp edge: editing any tool definition invalidates the cached prefix, so a bloated block that changes weekly gets less benefit than its size suggests. Treat caching as a discount on cost, never as a reason to stop pruning.
- You cut schemas from 2,100 to 300 tokens and selection accuracy dropped. What do you look at?Almost always the guidance you cut with the fields: preconditions and the explicit statements of when not to call. Compare the before and after text for scope sentences that disappeared, and restore them — they are cheap. Also check enums you replaced with free strings, since those were doing classification work, and confirm you did not trim a field the model genuinely needed to supply.
- When is splitting one tool into two the wrong move?When the two variants share most of their arguments and the choice between them depends on information the model does not yet have. Then you have converted an argument decision into a selection decision and made it earlier and harder. Splitting pays when the modes are genuinely disjoint — different fields, different preconditions — and hurts when it just duplicates a schema behind two names.
saying these in an interview costs you the question
- Believes tool definitions are sent once per session
- Cuts descriptions without re-measuring selection accuracy
- Assumes prompt caching makes schema size irrelevant
- Optimises tokens while leaving thirty optional fields
- Treats smaller schemas as automatically better