skip to content

How does progressive disclosure shrink the context cost of a 400-tool agent catalog?

level: middleimportance: should knowfreq 46%

answer

  1. one line now, the manual later
  2. choosing costs less than using
  3. tens of tokens versus hundreds
  4. full docs arrive on activation
  5. the first tier must still discriminate

basics

~20 s

Load one cheap line per tool — name plus a one-sentence purpose — and keep the full parameter documentation, examples and usage rules out of context until the agent activates that tool. Most turns then pay for metadata only, not for the whole catalog.

solid answer

~50 s

Progressive disclosure splits a tool's definition into tiers that are paid for at different times. Tier one is metadata: a name and a one-line purpose, on the order of tens of tokens, loaded for every tool. Tier two is the full usage instructions — parameters, preconditions, worked examples, error semantics — often several hundred to a thousand tokens, injected only once the agent has selected that tool. Tier three is bulk material the tool references, such as scripts or reference documents in the agent's sandbox, read only when actually executed. At 400 tools the arithmetic is decisive: a 40-token metadata line each is about 16k tokens, while 900-token full definitions each would be around 360k. The same idea packages agent capabilities in general — the Agent Skills format loads a name and description first and the body of SKILL.md only on activation. The tradeoff is that tier one must be discriminative enough to select on, since that is all the model sees at decision time.

code

markdown · 15 lines
markdown
---
name: hr.request_leave
description: Submit a paid-time-off request for the signed-in employee.
---

## Parameters
- start_date (required, ISO-8601 date)
- end_date (required, ISO-8601 date, must be >= start_date)
- leave_type (required, one of: annual, sick, unpaid)

## Preconditions
The employee must have an approved manager on record.

## Errors
Insufficient balance is returned as a business error, not an exception.

go deeper

for a junior

Know that only a short name-and-purpose line is loaded for each tool, and the detailed parameter docs arrive after the agent picks that tool. Be able to state the token saving in rough terms.

for a middle

Lay out the tiers and the arithmetic, and explain that the always-loaded tier is what selection is actually made on — so it has to carry the distinguishing information itself.

for a senior

Discuss the failure modes you would instrument: interchangeable metadata lines, activation thrash, and calls made on metadata alone. Know where deferred tiers are stored and why a filesystem gives you ownership and versioning for free.

for a principal

Treat the metadata line as a scarce shared budget across teams and set the rules for it — what a one-liner must contain, who arbitrates when two teams' lines collide, and how the tier split interacts with search and namespacing across a catalog many groups contribute to.

## The principle A tool definition is not one indivisible blob. Some of it is needed to *choose* the tool; the rest is needed only to *use* it. Progressive disclosure separates those and defers the second part, so the routine cost of a large catalog collapses to the cost of choosing. ## The tiers **Tier 1 — metadata, always loaded.** The tool's name and a single sentence of purpose, plus perhaps its namespace. Budget it at tens of tokens. This tier answers exactly one question: is this the tool for the request in front of me? **Tier 2 — full instructions, loaded on activation.** Parameter list and types, required versus optional fields, preconditions and permissions, what the result looks like, common errors, and any usage rules the caller must respect. This is the expensive tier, commonly several hundred to a thousand tokens, and the agent only needs it for tools it is actually about to call. **Tier 3 — bulk resources, read on execution.** Reference documents, long examples, scripts the tool ships with. These live outside the prompt entirely — on a filesystem the agent can read from — and are pulled in only when the work requires them. ## The arithmetic For 400 tools with 40-token metadata lines, tier one costs roughly 16k tokens. Loading full 900-token definitions for all of them would be around 360k. A typical conversation touches a handful of tools, so tier two adds a few thousand tokens rather than hundreds of thousands. That is the entire argument, and it is why the pattern generalizes beyond tools: the Agent Skills format applies the same three-tier split to packaged capabilities, loading a name and description first and the full instruction body only when the skill is activated. ## How the deferred tiers are stored The common implementation is a filesystem the agent can read. Definitions live as files — say a directory per domain, with one markdown file per tool — and the agent reads the file for a tool once it has picked it. That gives you version control, per-team ownership of a directory, and the ability to update instructions without touching prompt-assembly code. It also composes with search: a keyword match returns a path, and the agent reads the path. ## Where it interacts with selection accuracy Progressive disclosure is a *context-cost* mechanism first and a selection mechanism second, and it can cut either way. It removes crowding — the model reads 400 short lines instead of 400 long blocks — but it also removes evidence: whatever distinguishes two similar tools must now fit in tier one, or the model is choosing between two indistinguishable one-liners. That is the discipline the pattern imposes. Tier one has to be discriminative, which usually means it names the domain, the object acted on, and the boundary against the nearest neighbour, in one line. ## Failure modes **Under-informative metadata.** Two tools whose tier-one lines both read "handle employee requests" cannot be told apart at decision time, no matter how good their tier-two docs are. **Activation thrash.** An agent that loads tier two for a tool, decides it is wrong, loads another, and repeats, spends more tokens than the naive approach would have. Cap activations per turn and treat a high thrash rate as a signal that tier one is too thin. **Instructions the agent never reads.** If tier two lives in a file the agent is not reliably prompted to read before calling, you get calls made on metadata alone, with predictable argument errors. Make the read a required step, not an optional convenience. ## When it is not worth it With a few dozen well-separated tools, tier splitting adds machinery, an extra read step, and a new failure mode for no meaningful saving. It pays when the catalog is large enough that most turns touch a small fraction of it — which is also the point at which per-team directory ownership starts to matter organizationally. ## How it combines with the other scaling levers Progressive disclosure, deferred definitions with server-side search, and namespacing are complementary rather than alternative. Search decides *which* tools are even candidates; namespaces make the candidates distinguishable; disclosure decides *how much* of each candidate is paid for and when. A large production catalog typically uses all three.

  • What has to be true of the metadata tier for this to work at all?
    It must be discriminative on its own, because it is the only thing the model sees when choosing. One line should convey the domain, the object acted on, and the boundary against the nearest similar tool. If two tools' metadata lines are interchangeable, no amount of detail in the deferred tier will fix the selection error — the decision has already been made.
  • How is this different from simply writing shorter tool descriptions?
    Shortening throws information away permanently; disclosure moves it rather than deleting it. The parameter semantics, preconditions and error rules still exist and still reach the model — just after selection, when they are actually needed. Shortening degrades usage accuracy to buy context; disclosure buys context without that trade, at the cost of an extra read step.
  • What does activation thrash look like and how do you catch it?
    The agent loads one tool's full instructions, decides it is wrong, loads another, and repeats — spending more tokens than loading everything would have. Instrument activations per turn and the ratio of activations to actual calls. A ratio well above one says the metadata tier is too thin to choose on, so fix tier one rather than capping activations alone.

saying these in an interview costs you the question

  • Thinks progressive disclosure means permanently shortening descriptions
  • Puts no distinguishing information in the always-loaded tier
  • Assumes the agent will read deferred instructions without being required to
  • Applies the pattern to a dozen tools where it only adds machinery
  • Confuses deferring the documentation with deferring the tool's availability

context