skip to content

How does server-side tool search over deferred definitions keep a 300-tool catalog out of context?

level: seniorimportance: must knowfreq 58%

answer

  1. definitions held back, not shipped
  2. the model asks before it sees
  3. lexical match, no vector index
  4. one tool must stay loaded
  5. a miss looks like a missing feature

basics

~20 s

Definitions are marked deferred and held server-side instead of being injected into the prompt. The model calls a small always-loaded search tool with keywords, the server keyword-matches the deferred definitions, and only the few hits are injected before the model calls one normally.

solid answer

~50 s

Rather than shipping 300 definitions in every request, you mark them deferred so the provider holds them and the prompt starts with only a tool-search tool plus the handful of tools needed on literally every turn. When a request arrives the model searches with keywords, the server matches over the deferred catalog, and the matching definitions are injected into context; the model then calls one of them through the normal tool path. The matching is deliberately keyword-based — regex and BM25 style — not embedding similarity: there is no index to build or refresh when a team adds a tool, exact identifier and domain-term matches are what tool metadata actually contains, and the behaviour is deterministic and debuggable. At least one tool must remain loaded or the model has no entry point at all, so the search tool itself is never deferred. The cost is an extra round trip; the win is a small, stable prompt prefix and far less to discriminate between.

go deeper

for a junior

Know the shape: instead of putting all 300 definitions in the prompt, the agent searches for the tool it needs and only the matches get loaded. Be able to say why that saves tokens.

for a middle

Explain the full round trip — deferred registration, keyword search, injection of the hits, ordinary call — and why the search tool itself can never be deferred.

for a senior

Argue the keyword-over-embeddings choice on index lifecycle, exact identifiers and debuggability, and name the zero-hit failure that presents to users as a missing feature. Know the threshold below which the machinery is not worth it.

for a principal

Own the catalog as a shared asset: who fixes metadata when searches miss, how injection is capped so crowding does not return, and how the always-loaded hot set is negotiated between teams competing for prompt real estate.

## The problem it solves The default contract of function calling is that every tool the model may use is described in the prompt of every request. That is fine at ten tools and untenable at three hundred: tens of thousands of tokens on every turn, a long homogeneous list where the relevant entry is buried mid-prompt, and a dozen plausible candidates for any given intent. Deferred loading breaks that contract. The catalog still exists, but the definitions live server-side until something asks for them. ## How the flow actually runs 1. You register the full catalog but mark the bulk of it deferred, so those definitions are not serialized into the prompt. 2. The prompt carries a small always-loaded set: the tool-search tool, plus any tool used on essentially every turn. 3. The model reads the user request and, finding nothing that fits in its loaded set, calls search with query terms — "leave", "time off", "pto", or a literal tool name it half-remembers. 4. The server matches those terms against the deferred definitions and returns, or directly injects, the top few full definitions. 5. The model now sees three or four candidates instead of three hundred, and calls one through the ordinary tool path. Nothing about the call itself is special — a deferred tool that has been loaded is just a tool. Definitions that have been pulled in stay available for the rest of the conversation, so the search cost is paid once per capability, not once per call. ## Why keyword matching rather than embeddings The instinct of anyone who has built retrieval is to embed the tool descriptions and do vector search. In practice the field settled on keyword matching — regex and BM25 flavours — for four reasons. - **No index lifecycle.** Every tool a team adds, renames or edits would require re-embedding and re-indexing, and a stale index silently hides new capabilities. Keyword matching reads the catalog as it stands. - **The queries are keyword-shaped.** A model searching for a tool emits domain nouns and near-identifiers, not paragraphs. Lexical match is a good fit for short, jargon-heavy queries; this is the regime where BM25 has always been competitive. - **Exact identifiers matter.** Tool names, namespaces and parameter names are literal strings. Embeddings blur exactly the distinctions — `create_report` versus `create_request` — that you most need preserved. - **Debuggability.** When a search misses, you can reproduce the match deterministically and see which term failed. An embedding miss is an opaque similarity score, and it drifts when the embedding model changes underneath you. That is a judgement about tool catalogs specifically, not a general claim that lexical search beats semantic search. Tool catalogs are small, curated, jargon-dense and edited constantly — the conditions that favour lexical matching. ## What must stay loaded The entry point cannot itself be deferred: if nothing is loaded, the model has no way to ask for anything. Beyond the search tool, keep a small hot set — the two or three tools used on almost every turn — because paying a search round trip for them is pure latency. Everything else defers. A useful sizing heuristic is that below roughly two or three dozen tools the machinery costs more than it saves, and you are better off loading the lot. ## Failure modes worth naming The characteristic failure is a **search miss that reads as an absent capability**. The model searches "book annual leave", finds nothing because the catalog metadata says "absence request", and tells the user the assistant cannot do that — a false negative that looks like a product gap rather than a retrieval bug. Mitigations: instrument searches that return zero hits and treat them as a metadata backlog; instruct the agent to retry with broader or alternative terms before concluding nothing exists; and offer a way to enumerate a namespace as a fallback. The second failure is **over-broad queries** that pull twenty definitions back in and recreate the crowding you were avoiding. Cap how many definitions a single search may inject. ## Costs and interactions The direct cost is one extra model turn plus the search itself before real work starts. Against that: the static prompt prefix becomes small and stable, which is good for both cost and prompt caching, and per-turn token spend drops sharply for conversations that touch only a few capabilities. Latency-sensitive paths that always use the same tool should keep it loaded rather than searching for it every session. A related pattern for very large catalogs is to let the agent discover definitions from files in its execution sandbox and read only the ones it needs. That is the same principle — hold references, fetch on demand — applied through the filesystem instead of a search tool.

  • What breaks when the search returns nothing for a request the catalog does actually cover?
    The agent concludes the capability does not exist and tells the user so — a retrieval bug that presents as a product gap. Guard it three ways: log every zero-hit search as a metadata gap to fix, instruct the agent to retry with broader or synonymous terms before giving up, and provide a fallback that lists a namespace's tools so it can browse rather than guess.
  • Why not embed the tool descriptions and do vector search over them instead?
    Because you would own an index that must be rebuilt whenever any team edits a tool, and a stale index hides new capabilities silently. Tool queries are short and jargon-heavy, which suits lexical matching, and exact identifier distinctions like create_report versus create_request are exactly what embeddings blur. Keyword matching is also reproducible when you debug a miss.
  • At what catalog size does this stop being worth it?
    Below roughly two or three dozen well-separated tools the extra round trip and the miss risk cost more than the tokens saved — load everything. The pattern earns its keep when the catalog is large enough that most turns touch a small fraction of it, which is the usual shape once several teams contribute.
  • How does deferred loading interact with prompt caching?
    Favourably, in general: the static prefix shrinks to a small, stable always-loaded set, so more of the conversation is cacheable and the cached block does not churn every time a team edits an unrelated tool. Definitions pulled in mid-conversation land after that prefix, so they extend the suffix rather than invalidating the cached head.

saying these in an interview costs you the question

  • Assumes vector search is automatically better than keyword matching
  • Forgets the model needs at least one loaded tool as an entry point
  • Claims deferred tools can be called without ever being loaded
  • Suggests truncating every description instead of deferring definitions
  • Ignores that a search miss looks like a missing capability to the user

context