An on-call agent keeps calling search_logs when get_metrics would answer. How do you fix it?
answer
- the definitions are the prompt
- say when not to use it
- name the alternative in the text
- synonyms cannot be disambiguated
- enums make the wrong call unsayable
basics
~20 sOverlapping tool definitions are usually the cause: the model reads names and descriptions as prompt text and cannot tell which surface owns the question. Rewrite each description to state what it is for and when to prefer the other, and merge genuine near-duplicates.
solid answer
~50 sTreat the tool definitions as prompt engineering, because at selection time that is exactly what they are — the model has the names, descriptions and argument schemas, and nothing else. If `get_metrics(service, window)` and `search_logs(query)` both read as "find out what is happening with a service", picking between them is close to a coin flip. The fixes, in order of payoff: give each description a positive scope *and* a negative one ("use for numeric time-series of a named metric; do not use for free-text search — that is search_logs"), so the discriminating rule is written down rather than inferred; constrain arguments with enums so the wrong tool is harder to express; add one few-shot exemplar covering exactly the confusable case; and if two tools genuinely overlap, delete one or merge them behind a single tool with a discriminating argument. Then re-run a labelled set of representative queries to confirm the change actually moved selection.
go deeper
Know that the model chooses a tool from its name, description and arguments alone, so a vague description is the usual reason it picks the wrong one.
Explain how to rewrite a description so it discriminates — scope, an explicit anti-scope naming the other tool, and trigger phrasings — and how enum arguments make the wrong call harder to express.
Demonstrate the diagnosis first: separate genuine selection failure from a vague task or an under-serving tool by reading the thought before the action, then fix the definitions and verify on a labelled set, watching for over-correction in the reverse direction.
Own the catalogue as a designed surface: a policy that near-synonymous tools get merged rather than described, that every definition carries an anti-scope, and that catalogue changes ship with the same regression evidence as a code change.
## The diagnosis comes first Before changing anything, establish that this is a *selection* problem. Pull the traces and look at what the model wrote in the thought immediately before the action. Three different diseases look identical from the outside. If the thought says "I need the error rate for billing-api" and the action is `search_logs`, the model knew what it wanted and picked the wrong surface — a genuine selection failure, and the tool definitions are the lever. If the thought is vague — "let me investigate the service" — the model has not decided what it is looking for, and the fix is upstream in how the task is framed, not in the catalogue. If the model tried `get_metrics` first, got an empty or confusing result, and fell back to `search_logs`, that is not a selection failure at all; the metrics tool is under-serving and the agent is compensating. Assume the first case for the rest of this. ## Descriptions are prompt text At the moment of choosing, the model sees the tool name, the description, and the argument schema. That is the entire basis for the decision. A description like "Search the logs." tells it what the tool does and nothing about when to reach for it, which is the only question actually being asked. A description that discriminates has three parts. **Scope**: what class of question this answers — "numeric time-series for a named metric on a named service over a time window". **Anti-scope**: what it is explicitly not for, naming the alternative — "not for free-text search across log lines; use search_logs for that". **A trigger phrasing or two**: the shape of request that should route here — "error rate, latency percentile, request volume". The anti-scope clause is the one most teams omit and the one that moves the number most, because ambiguity is by definition a comparison between two tools and only a cross-reference resolves it. Write these in the imperative and keep them short. Every token spent here is spent on every turn, and a rambling description competes with the task for attention. ## Make the wrong call harder to express Schemas do selection work too. If `get_metrics` takes `metric` as a free string, the model can invent one and the tool half-answers. If `metric` is an enum of the metrics you actually export, the model either picks a real one or discovers it cannot express the request — which pushes it toward the other tool for the right reason. Required arguments have the same effect: a tool that demands a `service` and a `window` signals that it answers a narrow, structured question, while a tool taking one free-text `query` signals the opposite. Argument shape communicates intent before a single word of description is read. ## Near-duplicates: delete rather than describe The worst version of this problem is three tools that are the same tool. `find_user`, `lookup_user`, `user_by_email` are synonyms in English, and no amount of description writing makes a model reliably choose among synonyms — you are asking it to guess an arbitrary convention. Merge them into one `get_user` with a discriminating argument (an identifier plus an enum of identifier kinds, or optional mutually exclusive fields) and the ambiguity is gone by construction rather than by persuasion. The general principle: prefer one tool with a well-typed argument over several tools that differ only in how you reach the same data. Every additional near-synonym in the catalogue is a fresh opportunity to be wrong, and it costs prompt tokens forever. ## Exemplars for the confusable case A ReAct prompt usually carries a few worked examples. Spend one of them on the boundary you are losing: a question that sounds like free-text investigation but is actually a metric lookup, with the thought that names the discriminating feature ("this asks for a rate over time, so metrics, not logs") and the correct action. Exemplars are expensive in tokens, so target the demonstrated confusion rather than adding generic ones. ## Verify, do not assume Every change here is a prompt change, and prompt changes regress silently. Keep a labelled set of representative on-call questions with the tool you expect for each, run it before and after, and look at whether the confusion moved or merely relocated — a description that over-corrects will start routing genuine log searches to metrics. Watch the reverse direction as explicitly as the one you set out to fix. ## What interviewers are checking That you reach for the tool definitions rather than for a bigger model or a longer system prompt; that you know a description is read as instruction, not documentation; that you are willing to delete a tool; and that you validate the change instead of declaring victory from one good trace.
- The two tools genuinely overlap for a class of questions. What then?Pick a canonical owner and say so in both descriptions — "for questions about error rate, prefer get_metrics; use search_logs only when no metric exists for the signal". Where the overlap is total rather than partial, that is evidence they should be one tool with a mode argument. Leaving a real tie unresolved just relocates the coin flip; the model needs a rule it can apply, not a hint that both are acceptable.
- Three tools are named find_user, lookup_user and user_by_email. What is the fix?Merge them. The names are synonyms in English, so the model is being asked to recall an arbitrary convention rather than to reason, and description text cannot repair that. One `get_user` taking an identifier plus an enum of identifier kinds — email, id, username — expresses all three intents with no ambiguity, costs fewer prompt tokens, and removes two chances to be wrong on every single turn.
- Would a bigger or newer model fix this on its own?Partly, and unreliably. Stronger models do disambiguate underspecified catalogues better, which is why the problem sometimes seems to disappear on an upgrade. But you have then bought a fix you cannot inspect, cannot version, and lose on the next model change — and the ambiguity still costs tokens and still misfires on the hard cases. Fix the definitions; the model upgrade then helps on top of a sound catalogue rather than compensating for a broken one.
- How would you tell this apart from the metrics tool simply returning poor results?Read the trace order. If the agent calls get_metrics first, gets something thin or confusing, and only then falls back to search_logs, the selection was right and the tool is under-serving — fix the tool's output, not its description. Genuine selection failure looks different: the thought names a metric-shaped need and the very first action is search_logs, with no prior attempt.
saying these in an interview costs you the question
- Blaming the model instead of the tool descriptions
- Writing descriptions that say what the tool does but never when to use it
- Keeping three synonymous tools and hoping better wording separates them
- Adding a rule to the system prompt rather than to the tool definition
- Declaring the fix successful from one improved trace with no labelled check