Why is a tool's description field a prompt rather than documentation?
answer
- it causes a decision, not a lookup
- the missing half is "when not to"
- preconditions before arguments
- models over-call plausible tools
- forty useful words beat three vague ones
basics
~20 sThe description is the model's only instruction about a capability, so it must be written to drive a decision: what the tool does, what must be true before calling it, where argument values come from, and which nearby requests it must not be used for.
solid answer
~50 sDocumentation explains a thing to someone who already decided to use it. A tool description has to *cause* the decision, correctly, with no other source of truth available. So it is written as instruction: one line on what the capability does, the preconditions that must hold, a sentence on where argument values legitimately come from, and — the part people omit — explicit negative guidance about the adjacent requests that must not trigger it. "Files a claim" is a caption, not an instruction. Expanded to forty words it can say that the incident must already have occurred, that a policy number is required, and that status checks and coverage quotes are out of scope. That rewrite alone usually fixes more misfires than any change to the agent loop, because it is the only channel through which the model can learn scope.
go deeper
Be able to take a bare description like "Files a claim" and expand it into something that states what the tool does and what it must not be used for.
Explain the four ingredients — capability, preconditions, argument provenance, negative guidance — and why models over-call tools whose boundary is never stated in text.
Show that you validate rewrites against a fixed set of user turns including near-misses, and that you know descriptions interact, so tightening one tool can suppress another.
Own the house template and the review standard for descriptions, and decide what belongs in a tool definition versus cross-cutting system-prompt policy so guidance does not drift out of sync with the tools it governs.
## The description does a job, not a favour A tool description is prompt text. It sits in context on every turn, and the model uses it to answer one question: given what the user just asked, should I call this, and with what? Nothing else about the tool is available to it. So the writing standard is not "is this an accurate account of the function" but "does this text reliably produce the right decision, including the decision *not* to call". That reframing changes what goes in. Reference documentation optimises for completeness and for a reader who can ask follow-up questions. A description optimises for a reader who cannot ask anything, will read it exactly once per turn, and is simultaneously reading a dozen other descriptions. ## The four things a good description carries **1. Capability.** What the tool actually does, in the caller's vocabulary, in one sentence. Prefer the domain's words over your codebase's words: "file a first notice of loss" beats "invoke the FNOL intake pipeline". **2. Preconditions.** What must already be true. For a claims-filing tool: the incident has occurred, the policy is active, the caller has supplied a policy number. Preconditions are how the model learns to gather missing information from the user before calling rather than guessing at it. **3. Argument provenance.** For each non-obvious field, where the value legitimately comes from — the user's own words, a prior tool result, a fixed vocabulary. This is the single most effective defence against invented arguments, and it belongs partly in the per-field descriptions inside the schema. **4. Negative guidance.** The requests that look like a match but are not. This is the field's highest-value content and the most commonly missing. Models are cooperative; a plausible-looking tool in scope will get called unless the text says otherwise. ## A worked rewrite Start with the description as most teams first write it: > Files a claim. Three words, syntactically fine, operationally useless. The model cannot tell whether "I want to know what happened with my claim from March" matches, nor whether a policy number is needed, nor which of six claim categories a burst pipe belongs to. Now the forty-word version: > File a first notice of loss against an active auto or property policy. Requires a policy number and an incident that has already occurred. Do not use to check the status of an existing claim, to quote or compare coverage, or to change policy details. Same capability, and now three decisions are determined by the text: when it applies, what must be collected first, and which adjacent intents route elsewhere. Note the shape — capability, preconditions, then the explicit "do not use" clause. That template travels well. ## Why models over-call, and why negative guidance works An instruction-tuned model biases toward being helpful with the affordances in front of it. If a request is roughly in the neighbourhood of a tool's stated purpose, the model will usually reach for it, because from its position calling something plausible looks better than declining. A description that only states the positive case leaves the entire boundary undefined and the model fills it in optimistically. Stating the boundary explicitly is not redundancy — it is the only place the boundary exists. The reverse failure exists too. Descriptions loaded with warnings and hedges can suppress a tool that should fire, leaving the model to answer from parametric memory when it should have looked something up. The calibration is empirical: change the text, re-run a fixed set of representative user turns, and look at what got called. ## Register and length Write in the imperative, in the second-order voice of an instruction to an assistant. Keep it dense: every sentence should change a decision. Fifty to eighty words is a typical landing zone for a non-trivial tool — long enough for preconditions and exclusions, short enough that it does not dominate the context shared with every other tool's definition. Avoid implementation detail (which service it hits, which table it writes) unless it changes the model's behaviour; the model cannot act on it and it costs tokens on every turn. A useful discipline: after writing, delete every sentence that would not alter what the model does. What survives is the description. ## Where descriptions come from Many frameworks generate the description from a docstring. That is convenient and usually wrong on first pass, because docstrings are written for maintainers — they explain how the function works and omit preconditions and scope, which are exactly the model-relevant parts. Treat generated text as a draft, then rewrite it as instruction and read the serialized definition to confirm what the model receives.
- Can you push that guidance into the system prompt instead of the tool description?Partly, and teams do — but it scales badly. System-prompt rules about specific tools drift out of sync when a tool changes, and they separate the instruction from the thing it governs. Scope and preconditions belong with the definition so they travel with it; the system prompt is better used for cross-cutting policy such as always confirming before irreversible actions.
- How do you know a description rewrite actually helped rather than moved the failure?Hold a fixed set of representative user turns — including the near-miss ones the tool must decline — and compare which tools fire before and after the edit. Descriptions interact, so a rewrite that fixes over-calling on one tool can suppress a neighbour. Judge the change on the whole set, not on the single prompt that prompted the edit.
- Should per-parameter descriptions repeat what the top-level description says?No — they should carry what the top-level cannot: the format of a value, its unit, its provenance, and what to do when the user has not supplied it. A field description reading "the policy number" adds nothing; one reading "policy number exactly as printed on the document; ask the user if not stated, never infer it" prevents a class of invented arguments.
saying these in an interview costs you the question
- Writes descriptions as one-line captions of the function
- Omits any statement of when not to call the tool
- Explains implementation internals the model cannot act on
- Assumes an auto-generated docstring is good enough
- Treats the description as documentation for engineers