Why do agents load context through tools on demand instead of preloading it?
answer
- identifiers now, content later
- attention is a budget
- the needed subset is discovered, not known
- round trips are the price
basics
~20 sPreloading spends the attention budget on material the model mostly will not use, and quality degrades as windows fill. Just-in-time loading carries lightweight identifiers — paths, IDs, queries — and pulls full content through a tool only when a step actually needs it.
solid answer
~50 sPreloading means stuffing everything the agent *might* need into the prompt before the first turn. It is simple and it costs one round trip, but almost every token is dead weight: the model pays attention over material it will never use, and long-context evaluations consistently show usable quality falling well short of the advertised window. **Just-in-time loading inverts that.** The agent starts with cheap handles — file paths, record IDs, table names, a search query — and calls a tool to fetch the actual content only when a step requires it. The window then holds mostly relevant text, which is the thing that drives answer quality. The price is real: each fetch is another model round trip, so latency grows with the number of hops, and the agent can fetch the wrong thing. As of mid-2026, just-in-time is the default for open-ended agent work, with preloading reserved for small, stable, almost-always-needed material.
go deeper
Be able to say plainly that the agent starts with references — paths, IDs, names — and calls a tool to read the full content only when it needs it, instead of pasting everything in up front.
Explain the mechanics: what goes in the starting prompt, what the tool returns, and why filling the window with mostly-irrelevant text lowers answer quality even when it technically fits.
Show the judgment call. Name hit rate, subset size and latency budget as the deciding factors, and describe the hybrid — a preloaded catalog plus JIT bodies — rather than picking a side.
Own the measurement and the architecture. Decide what the team instruments (fetches per task, use-rate of fetched content), and how the choice interacts with caching economics, evaluation reproducibility and per-request cost targets.
## The two shapes of getting context in An LLM only sees what is in its context window. There are exactly two ways to get something in there: put it in before the model runs (**preloading**), or give the model a tool it can call to pull it in mid-run (**just-in-time**, or JIT, loading). Everything else — retrieval quality, caching, compaction — sits on top of that choice. Preloading is the older habit, inherited from single-shot prompting: assemble every document, schema, or API page the task might touch, concatenate it, send it. JIT loading instead puts *references* in the prompt — a directory listing, a set of record IDs, a table catalog, a search endpoint — and lets the model decide which of them to expand into full content. ## Why preloading stops working Three pressures push against preloading: **Attention is a budget, not free space.** Every token in the window competes for the model's attention. A useful framing is that the goal is the *smallest high-signal token set*, not the largest one that fits. When 95% of the prompt is irrelevant, the relevant 5% is harder to find and easier to contradict. **Effective context is smaller than advertised.** Frontier models ship million-token windows, but long-context benchmark families agree that usable quality holds over roughly 40–70% of the advertised length, and that it falls off a cliff rather than degrading smoothly. This degradation is commonly called *context rot*. A prompt that technically fits can still produce worse answers than a much shorter one. **You often cannot know in advance.** For a real agent task — investigate this incident, answer this question against the warehouse — the set of documents that matter is discovered *during* the work. Preloading forces you to guess the superset, and the superset is usually enormous. ## What JIT loading actually looks like The agent's starting prompt contains identifiers and enough metadata to choose between them. A file agent gets a directory listing. A database agent gets a table catalog with names and row counts. A support agent gets ticket IDs with subjects and dates. Then a tool — read this path, describe this table, open this ticket — returns the full payload, which is appended to the conversation as a tool result. Two properties matter. First, the model does the selection, so selection benefits from everything the model knows about the task, not just from a similarity score computed before the task started. Second, the content that lands in the window has already been *justified* by a decision — it is there because the model asked for it. ## What it costs JIT loading is not free, and a good interview answer names the costs: - **Round trips.** Each fetch is another inference call. A three-hop lookup is three times the time-to-answer of a preloaded prompt, and the hops are serial when later choices depend on earlier results. - **Wrong fetches.** The model can pick the wrong file, or open five when one would do. Every wrong fetch costs both a round trip and window space. - **Non-determinism.** Which context an answer used now varies per run, which complicates reproducing a bad answer and evaluating the system. - **Cache pressure.** A prompt that changes shape mid-session interacts badly with prefix caching unless the loaded content is appended behind a stable prefix. ## Choosing between them The honest rule is about *hit rate and size*. Preload material that is small, stable, and needed on nearly every turn — a short style guide, the two schemas the agent always joins, a company glossary. Load just in time when the candidate set is large, the needed subset is small, and which subset is needed depends on the task. Hybrids are normal and usually correct: preload a compact index or catalog so the model can choose without a round trip, and JIT-load the bodies. That is the same structure as progressive disclosure — identifiers and metadata up front, content on request. ## How you would know it is working Instrument it. Track how many fetches a task takes, how many of the fetched items were actually cited or used, and the ratio of loaded tokens to task-relevant tokens. A rising fetch count with a flat success rate means the agent is groping; a low use-rate on fetched content means the metadata it selects from is too weak to discriminate. Both are fixable, and neither is visible if you only look at the final answer.
- If the window is a million tokens, why not just preload everything anyway?Because effective context is smaller than advertised context. Long-context evaluations put usable quality at roughly 40–70% of the stated window, with cliff-shaped rather than gradual decay, and the failure is silent — the model answers confidently from the wrong part of a bloated prompt. Fitting is not the same as being used well. Cost and latency scale with prompt size too, so you pay more per turn for a worse answer.
- What would push you back toward preloading a fixed set of documents?A small, stable set with a high hit rate, under a tight latency budget. If a support agent needs the same 8k-token policy document on 95% of turns, fetching it costs an extra round trip on nearly every request and saves almost nothing. Preloading it also makes it part of a cacheable prefix, so the marginal cost drops further. Determinism helps too: fixed context is easier to evaluate and reproduce.
- How do you keep an agent from fetching the same document repeatedly across a long session?Make the already-loaded content visible and addressable — keep it in the conversation, or maintain a short manifest of what has been loaded this session with its identifier. Agents re-fetch mostly because they cannot tell what they already have. Deduplicating at the tool boundary (returning a short 'already loaded above' marker instead of the payload) both stops the loop and avoids paying window space twice.
saying these in an interview costs you the question
- Thinks a bigger window makes selection unnecessary
- Claims preloading is always cheaper because it is one call
- Ignores the extra latency of each fetch round trip
- Treats JIT loading as the same thing as retrieval ranking
- Assumes the model reliably picks the right file to open