skip to content

How does progressive disclosure let an agent work over 300 warehouse tables?

level: middleimportance: must knowfreq 56%

answer

  1. cheapest representation that supports the next decision
  2. catalog first, schemas second
  3. expand only the chosen branch
  4. metadata must discriminate, or everything expands
  5. watch the expansion ratio

basics

~20 s

Progressive disclosure gives the model identifiers and thin metadata first — 300 table names with row counts — and full DDL only for the two or three tables a query actually joins. Selection happens over cheap descriptors; expensive content is expanded on request.

solid answer

~40 s

Full DDL for 300 tables is tens of thousands of tokens the model mostly cannot use. Progressive disclosure splits the load into tiers. **Tier one** is the catalog: table names, row counts, maybe a one-line purpose. That is small enough to preload and rich enough for the model to reason about which tables a question touches. **Tier two** is the expansion: a describe-table tool that returns columns, types and keys for the handful it picked. **Tier three**, if needed, is sample rows or column statistics for the one table it is unsure about. Each tier is only entered when the previous one leaves a real ambiguity. The design constraint is that tier-one entries must be *discriminative* — if the metadata cannot separate the candidates, the agent expands everything and the saving evaporates.

go deeper

for a junior

Know the shape: the model first sees a list of names or IDs, then asks for the details of the few it picked. Be able to give one concrete example, like a file listing before opening a file.

for a middle

Explain the tiers and what belongs in each, and state the constraint that makes it work — the first tier has to carry enough signal for the model to choose without expanding everything.

for a senior

Bring the failure modes and the numbers. Name the expansion ratio, non-discriminative metadata, tier-one bloat and stale catalogs, and say how you would size each tier against a real budget.

for a principal

Own the interface contract. Decide what metadata the platform is required to publish for every corpus so agents can select cheaply, and treat catalog quality as an engineering deliverable rather than a prompt detail.

## The idea Progressive disclosure is a layering discipline for context: expose the cheapest representation that supports the next decision, and expand only along the branch the model actually takes. It is the practical shape most just-in-time loading takes, because pure JIT with no up-front information leaves the model guessing what exists at all. The canonical case is a warehouse-analytics agent. A mid-size warehouse has a few hundred tables; full DDL — every column, type, constraint and comment — runs to tens of thousands of tokens. A user asks for weekly returns by region. The agent needs maybe three tables. Preloading all 300 schemas spends the whole budget to use one percent of it. ## The tiers **Tier one — the catalog.** Table names, row counts, and ideally a short human-written purpose line. Three hundred entries at roughly ten tokens each is a few thousand tokens: small enough to sit in the prompt permanently. Its only job is to let the model say *which* tables plausibly matter. **Tier two — the structure.** For each selected table, a tool returns columns with types, primary and foreign keys. Three tables is a few hundred tokens. This is what the model needs to write a correct join. **Tier three — the specifics.** Sample values, cardinalities, or an enum's actual members, requested for one column when the model cannot tell what `status = 'R'` means. The same structure appears everywhere: a case-file agent lists filenames, dates and sizes before opening a document; a codebase agent sees a file tree before reading a file; a support agent sees ticket subjects before pulling transcripts. ## Why the tiering, rather than just fetching Without tier one the agent has no map. It either guesses identifiers — which produces failed fetches and wasted round trips — or issues broad searches whose results it cannot evaluate. Tier one is what makes selection cheap: a compact, stable, preloadable index over an expensive corpus. This is why the practical answer is almost never 'preload everything' or 'preload nothing', but 'preload the index, fetch the bodies'. ## The design constraint that decides whether it works **Tier-one entries must discriminate.** The whole scheme rests on the model being able to choose between candidates from metadata alone. A listing of `doc1.pdf, doc_final_v3.pdf, doc_final_v3_REAL.pdf` carries no signal, so the model opens all of them and you have paid for a listing round trip on top of loading everything anyway — worse than preloading. Discriminative metadata usually means one or more of: a human-meaningful name, a one-line description, a type or category, a date, and a size so the model can weigh the cost of expanding. Size matters more than people expect. If a listing tells the model an item is 40k tokens, a well-instructed agent will look for a cheaper way in — a summary, a targeted query — rather than swallowing it. ## Sizing the tiers There is a real budget question at tier one. A catalog of 300 tables is fine; a catalog of 300,000 files is not. When tier one itself does not fit, you add a tier *above* it — directories before files, schemas before tables, categories before items — or you replace the flat listing with a query interface that returns a bounded, ranked page of candidates. The rule is that every tier must fit comfortably and must let the model take exactly one decision. ## Failure modes to name - **Non-discriminative metadata**: covered above; the expansion fans out to everything. - **Tier-one bloat**: the 'thin' catalog quietly grows a description paragraph per entry and becomes the thing you were avoiding. - **Expansion without pruning**: the agent expands ten candidates, keeps all ten payloads in the window, and the session degrades exactly as if you had preloaded. - **Stale catalogs**: the listing is cached from an hour ago and references items that no longer exist, producing failed fetches the model then tries to work around. ## Measuring it The metric that tells you the tiering is earning its keep is the **expansion ratio**: how many tier-two fetches per task, against how many candidates tier one offered. A ratio close to one means the metadata is doing its job. A ratio that drifts toward the candidate count means the model cannot discriminate and you are paying for the indirection without getting the saving. Pair it with a use-rate — of the items expanded, how many were actually referenced in the answer.

  • What belongs in the tier-one listing beyond a name?
    Whatever lets the model choose without opening the item: a one-line purpose, a type or category, a date for recency, and a size so it can weigh the cost of expanding. Size is underrated — an agent told an item is 40k tokens will look for a summary or a targeted query instead of swallowing it whole. Anything that does not change a selection decision is bloat and belongs in tier two.
  • What if the tier-one catalog itself is too large to preload?
    Add a tier above it or replace it with a query. Directories before files, schemas before tables, categories before items — each level must fit comfortably and support exactly one decision. Alternatively expose a search or filter tool that returns a bounded page of candidates rather than the whole catalog. The structure is recursive; the constraint is that no single level ever blows the budget.
  • How would you know the tiering is not paying off?
    Track the expansion ratio — tier-two fetches per task against the number of candidates offered. If the agent routinely expands most of the catalog, the metadata is not discriminative and you are paying an extra round trip to end up with everything loaded anyway. Pair it with a use-rate: of the items expanded, how many were actually cited in the final answer.

saying these in an interview costs you the question

  • Thinks progressive disclosure just means smaller chunks
  • Puts full descriptions in the tier-one listing
  • Assumes bare filenames are enough metadata to choose from
  • Never prunes expanded payloads out of the window
  • Confuses tiered disclosure with ranking retrieval results

context