Your agent lists a case-file directory, then opens every document anyway — why?
answer
- the cheap tier carries no signal
- opening everything is rational under blind descriptors
- worse than preloading — you paid the listing too
- enrich offline, once per document
- watch expansion ratio and use rate
basics
~20 sAlmost always because the listing does not discriminate. Entries like doc_final_v3.pdf give the model nothing to choose on, so it expands everything — paying a listing round trip on top of loading the whole corpus, which is worse than preloading.
solid answer
~50 sThis is the characteristic failure of metadata-first loading: the cheap tier exists but carries no signal. If the listing is filenames alone, and the filenames are `scan_002.pdf`, `doc_final_v3.pdf`, `notes.docx`, the model cannot rank them against the task, and its safest strategy is to open all of them. You now pay strictly more than preloading — the listing hop, plus every payload. **The fix is at the listing, not the prompt.** Enrich each entry with what a selection decision actually needs: a one-line summary or first-page snippet, a document type, a date, and a size so the model can weigh expansion cost. Where the source cannot provide that, generate it once offline and store it with the item. Secondary fixes help but do not substitute: instruct the agent to expand at most *n* candidates and justify each, and expose a search or filter tool so it can narrow before listing.
go deeper
Understand that a list of names only helps if the names tell the model something. If every entry looks the same, the model has no basis to choose and will open them all.
Explain why over-expansion makes the tiered design worse than preloading — you pay the listing round trip and then load everything anyway — and name enrichment of the listing as the fix.
Diagnose from traces: rank the causes, choose between offline enrichment, filtering and bounded expansion, and define the expansion-ratio and use-rate metrics that prove the fix landed.
Treat catalog quality as platform work with an owner and a budget, not a prompt tweak, and decide what descriptive metadata every corpus must publish before agents are allowed to build on it.
## The symptom You build the tiered design correctly — a listing tool, then an open tool — and the traces show the agent listing 40 case files and then reading 38 of them. Token use is higher than the preloaded baseline you were trying to beat, latency is far worse because the reads are serial, and answer quality has not improved. ## The cause, in order of likelihood **1. Non-discriminative metadata.** The dominant cause. Selection is only cheap if the descriptors distinguish the candidates. Bare filenames from a real-world share drive rarely do: `doc_final_v3.pdf`, `scan_002.pdf`, `Copy of untitled.docx`. Faced with descriptors that carry no task-relevant signal, opening everything is the model's rational move — it is the only strategy that cannot miss the answer. **2. No cost signal.** If the listing does not say how large an item is, the model has no reason to be frugal. Sizes make expansion feel expensive and push it toward summaries or targeted queries. **3. No expansion policy.** Nothing in the instructions bounds how many items may be opened, so the model treats thoroughness as free. An explicit budget — open at most three, and state why each was chosen — changes behaviour markedly. **4. Recall anxiety from the task framing.** Prompts that say 'do not miss anything' or 'be exhaustive' convert a selection task into a coverage task. The model is doing what it was told. **5. Weak feedback from earlier fetches.** If opened documents come back as unstructured walls of text with no clear indication of relevance, the model cannot tell whether it has found what it needs, so it keeps opening. ## Fixing it at the source The durable fix is to make the listing carry decision-grade information. For a case-file corpus, an entry should look less like `doc_final_v3.pdf, 2.1 MB` and more like: identifier, document type (deposition, invoice, correspondence), date, party or matter, size, and a one-sentence summary. That is perhaps thirty tokens per entry — forty entries is 1,200 tokens, trivially affordable — and it is usually enough for the model to pick two. When the source system has no such metadata, generate it. Run a one-time enrichment pass that summarizes each document and stores the summary as a sidecar, then serve that in the listing. This is an offline cost paid once per document, against an online cost paid on every task. Refresh it when the document changes. ## Fixes that help but do not substitute - **Bounded expansion.** Instruct an explicit cap and require a stated reason per expansion. This limits the damage but, with genuinely uninformative metadata, means the agent picks blindly within the cap and may simply fail. - **Filter before listing.** Expose a query or filter over the listing so the agent narrows by date, type or keyword first. This shrinks the candidate set, though it does not make individual entries any more distinguishable. - **Return an excerpt rather than the full document.** A first-page or matched-region excerpt on open, with an explicit option to request the rest, keeps the mistake cheap when the model does over-expand. - **Deduplicate re-fetches.** Ensure the tool notes when an item is already in context rather than returning the payload a second time. ## Knowing whether you fixed it The metric is the **expansion ratio**: expansions per task divided by candidates offered, tracked before and after the change. Pair it with a **use rate** — of the documents expanded, how many the final answer actually cites. High expansion with low use is the signature of this failure. A healthy tiered system shows an expansion ratio well under one in ten and a use rate above half. One caution: fixing metadata quality can look like a prompt problem and get 'fixed' by adding sterner instructions. Instructions constrain a model that lacks information; they do not supply it. If the listing genuinely does not distinguish the candidates, no amount of instruction will make the selection correct — it will only make the failures quieter.
- The source system has no descriptions at all. What now?Generate them offline. Run a one-time enrichment pass that produces a one-sentence summary, a type label and a date for each item, store it as a sidecar, and serve that in the listing. It is a per-document cost paid once against a per-task cost paid forever. Refresh on change, and accept that a generated summary is approximate — it only has to be good enough to support a selection decision, not to replace the document.
- Would stricter instructions — 'open at most three files' — fix this on their own?No. A cap bounds the damage but does not supply the information the model lacks; it will pick three blindly and may miss the right one, turning a visible over-expansion into a quiet wrong answer. Caps are worth having, but as a guardrail on top of discriminative metadata, not as a substitute for it.
- What metric tells you the fix worked?Expansion ratio — items opened per task against items offered — tracked alongside a use rate measuring how many opened items the answer actually referenced. The failure signature is a high expansion ratio with a low use rate. After enrichment you want the ratio well under one in ten and the use rate above half, with answer quality holding or improving.
saying these in an interview costs you the question
- Blames the model instead of the listing's information content
- Thinks stricter prompt wording can replace missing metadata
- Believes any listing is automatically cheaper than preloading
- Omits size and date, leaving no cost or recency signal
- Never measures how many opened documents were actually used