skip to content

Preload 40k tokens of API reference, or take three round trips to fetch one endpoint?

level: seniorimportance: should knowfreq 50%

answer

  1. count hops, not dollars
  2. serial hops block; parallel hops are nearly free
  3. hit rate decides what gets preloaded
  4. pay tokens for the map, round trips for the territory

basics

~20 s

Count round trips, not dollars. Preloading costs one hop and a permanently degraded window; on-demand fetching costs a hop per hop and a cleaner window. Decide on hit rate, how serial the hops are, and the task's latency budget — then measure both.

solid answer

~50 s

Frame it as a **latency-versus-signal** trade measured in round trips. Preloading the whole reference is one inference call, fully deterministic, and cache-friendly — but 40k tokens of mostly irrelevant endpoints sit in every turn, crowding the few paragraphs that matter. Fetching on demand keeps the window clean, at the cost of one extra model round trip per hop, and those hops are typically *serial*: the agent must see the index before it knows which endpoint to open. Three serial hops on a two-second model is six seconds added to time-to-answer. The deciding variables are hit rate (if 90% of tasks need the same endpoint, preload it), how many hops the discovery actually takes, whether they can be parallelized, and whether a human is waiting. In practice you preload a compact endpoint index and fetch the bodies — one extra hop, most of the token saving.

go deeper

for a junior

Know that fetching context on demand adds an extra call to the model, and that each extra call costs time even when it saves tokens.

for a middle

Be able to lay out both sides — one call with a bloated window versus several calls with a clean one — and name hit rate and latency budget as the deciding factors.

for a senior

Reason in hops. Distinguish serial from parallel fetches, propose the preloaded-index hybrid, and say exactly what you would instrument before committing to either design.

for a principal

Own the budget. Set the latency SLO and the quality bar the design is optimizing against, decide where speculative prefetch and caching pay for their complexity, and accept the reproducibility cost that variable context imposes on evaluation.

## Why round trips are the right unit The instinct is to compare token cost in dollars, and that comparison usually says 'fetch on demand' loudly — 40k tokens per turn adds up. But token price is rarely what decides this in production. Two other things do: **answer quality**, which favours the small clean window, and **wall-clock latency**, which favours the fat static prompt. Latency is measured in model round trips, and a round trip is expensive in a way tokens are not: it is a full inference call, including think time. So count hops. A preloaded prompt answers in one call. A JIT agent that must list, then read the index, then open the endpoint page, then answer, is four calls. If each call is two to four seconds, the JIT path adds six to twelve seconds before the first useful token reaches the user. ## The variables that actually decide it **Hit rate.** How often is the same material needed? If nearly every task touches the same three endpoints, preloading them costs 3k tokens and saves a hop on every request. If the needed endpoint is unpredictable across 200 candidates, preloading all 200 to serve one is a terrible bargain. **Hop depth and seriality.** JIT costs are not one extra call; they are however many *dependent* calls discovery takes. Serial hops are the killer, because each one blocks on the previous result. If the agent can fire several fetches in one turn — it already knows three identifiers — the cost collapses to roughly one hop regardless of how many items are read. **Who is waiting.** An interactive assistant with a human watching has a latency budget measured in seconds; a batch pipeline processing overnight has none. The same architecture is wrong in one setting and right in the other. **Window pressure.** If the task is long-horizon, the 40k preloaded tokens are not a one-time cost — they persist through every turn, competing with the accumulating work product and pulling the session toward the degradation zone sooner. **Determinism.** Preloaded context is identical every run, which makes failures reproducible and evals stable. JIT context varies, so the same question can succeed and fail on different runs for reasons invisible in the final output. ## The hybrid that usually wins For an API reference specifically: preload a compact index — endpoint names, one-line purposes, maybe method and path shape — at a couple of thousand tokens, and fetch the full page for the one or two endpoints the task uses. That buys a single extra round trip instead of three, because the index removes the discovery hops, and it keeps roughly 95% of the token saving. The pattern generalizes: **pay tokens for the map, pay round trips for the territory.** A second lever is batching the fetch. If the index makes it obvious that two endpoints are needed, the agent should request both in the same turn rather than serializing. Whether the harness supports that is an implementation detail, but designing the tool surface so multiple identifiers can be requested at once is a deliberate choice that directly cuts hop count. ## Latency mitigations that do not change the architecture - **Speculative prefetch**: while the model is reasoning, the harness fetches the two most likely items so the tool result is already warm. Cheap in tokens only if you discard what is unused. - **Streaming**: it does not reduce hops, but it changes the perceived cost — the user sees progress during the middle hops rather than a blank wait. - **Warm caching of the preloaded part**: if the index is stable, it can sit behind a cacheable prefix, so the tokens cost near-nothing after the first turn and the argument shifts further toward preloading the index. ## What to measure before deciding Do not argue this from first principles when you can instrument it. Log, per task: number of fetches, which items were fetched, which were actually used in the answer, total tokens, and end-to-end latency split into model time and tool time. Two numbers settle most arguments — the **hit rate** of each preload candidate, and the **hop distribution** of the JIT path. If the hit rate is high and the hop count is above two, preload. If the hit rate is low, keep the fetch and work on cutting hops with a better index. ## The answer an interviewer wants Not 'JIT is better'. They want you to say it depends on hit rate, hop depth and who is waiting; to propose the index-plus-fetch hybrid; to name the measurement that decides it; and to be honest that preloading still wins for small, stable, high-hit-rate material under a tight latency budget.

  • How does parallelism change the arithmetic?
    Dramatically. The cost of JIT loading is the number of *dependent* calls, not the number of items. If the agent already knows three identifiers and can request all three in one turn, reading three documents costs roughly one hop instead of three. That makes the tool surface a latency decision: accepting a list of identifiers rather than one, and letting the model issue several calls in a turn, directly cuts time-to-answer.
  • Does prefix caching change which side wins?
    It shifts the balance toward preloading the stable part. If a 40k reference sits in an unchanging prefix, after the first turn its marginal token cost is small, so the dollar argument against preloading weakens. What it does not fix is the quality argument — cached tokens still occupy the window and still compete for attention. So caching makes preloading a *compact index* very attractive, and preloading the whole reference only somewhat less bad.
  • How would you cut a three-hop discovery path down to one?
    Remove the discovery, not the fetch. Preload or cache a compact index so the agent never needs a hop to learn what exists, and make identifiers predictable enough that it can address an item directly. Then batch: let one turn request every item the index made obvious. Speculative prefetch by the harness can absorb the remaining hop, at the cost of fetching some things nobody uses.

saying these in an interview costs you the question

  • Compares only token price and ignores wall-clock latency
  • Assumes JIT always adds exactly one round trip
  • Forgets that serial hops block on each other
  • Ignores that preloaded tokens persist across every turn
  • Claims caching removes the quality cost of a bloated prompt

context