skip to content

How would you set a default prefetching policy across a large product's surfaces, and validate that it was right?

level: principalimportance: should knowfreq 44%

answer

  1. a budget, not a switch
  2. classify surfaces before choosing defaults
  3. cost scales with rendered links
  4. pair latency with wasted-fetch ratio
  5. peak traffic multiplies speculation

basics

~20 s

Set defaults per class of surface rather than globally, budget the added bandwidth and origin load explicitly, and validate with paired numbers: navigation latency by whether the payload was warm, and the share of prefetches never used.

solid answer

~50 s

There is no single right default, so the policy should classify surfaces. Short, high-intent link sets — primary navigation, wizard steps, a detail page's siblings — can afford an early trigger. Long lists and feeds get intent-only prefetching or none, because the request count scales with the list. Personalised and expensive routes warm code but not data. Metered or slow connections opt out entirely. Then put numbers on it: added requests per page view, added bytes per session, the origin capacity the policy consumes at peak, and the share of prefetches never navigated to. Validate by rolling the change out in stages and watching both sides — navigation latency percentiles split by warm versus cold, against origin load and wasted-fetch ratio. Give teams one owned default and an explicit per-link escape hatch, and revisit when a surface changes shape.

go deeper

for a junior

Understand the shape of the tradeoff: prefetching makes clicks feel instant and costs data for every guess that turns out wrong.

for a middle

Be able to argue a default for a specific surface — how many links it renders, how likely each is to be taken, and whether the route is personalised.

for a senior

Bring the measurement: warm-versus-cold navigation latency in the field, the wasted-fetch ratio, and the origin load the policy adds at peak.

for a principal

Own it as a budget with defaults per surface class, a kill switch, and a review trigger when a surface changes shape — and say plainly that there is no universally right default.

Prefetching is usually discussed as a per-link switch. Across a large product it is really a **budget** — the product is spending its users' bandwidth and its own origin capacity to buy navigation latency — and budgets need an owner, defaults, limits, and a way to tell whether the spend is working. ## Start by classifying surfaces, not by picking a global default A default that is obviously right for a five-item navigation bar is obviously wrong for an infinite feed, because the cost scales with the number of rendered links while the benefit does not. | Surface class | Sensible default | Why | |---|---|---| | Primary navigation, few links | early trigger, code and data | small fixed cost, very high hit rate | | Wizard or flow with a known next step | eager on the next step only | one link, near-certain | | Detail page with sibling links | intent trigger | moderate count, real intent signal | | Long list or feed | intent only, or code only | request count scales with the list | | Personalised or expensive routes | code only | data is per-user, uncacheable, ages | | Anything with side effects | never | a speculative fetch performs the action | | Slow or metered connections | off | the user is paying directly for the guess | ## Put numbers on the spend Four numbers make the policy arguable rather than aesthetic: 1. **Added requests per page view** — how many speculative fetches a typical render of this surface generates. This is the one that surprises people on lists. 2. **Added bytes per session** — total speculative transfer, and the share of it discarded. 3. **Origin capacity consumed at peak** — speculative traffic multiplies with real traffic, so it lands hardest during a spike. Prefetches a CDN absorbs are cheap; prefetches that reach dynamic route handlers are not. 4. **Wasted-fetch ratio** — prefetched and never navigated to, over all prefetches. It is the honest measure of how good the guess is. ## Validate with paired measurements A speed number alone will always look good, because prefetching does make warm navigations faster. Pair it: - Navigation latency percentiles, split by whether the payload was already in hand. If warm and cold navigations are close, the prefetch is buying little. - The wasted-fetch ratio for the same surface over the same window. - Origin request rate and error rate attributable to speculative traffic. - A field view rather than a lab one: on a fast office connection almost nothing looks expensive. Roll it out in stages — one surface, then a percentage of traffic — and hold the origin metrics next to the latency ones. Be ready to revert a surface rather than the whole policy. ## Guardrails that keep it from becoming an incident - A concurrency cap and a low-priority schedule, so speculation cannot delay what the user is looking at. - Cancellation when a link scrolls away or the user navigates elsewhere. - A kill switch that can disable speculation product-wide without a deploy, because the failure mode is amplified load at peak. - A hard rule that side-effecting targets are never speculatively fetched, enforced in review rather than by convention. - Discarding speculative state on a context change such as sign-out or workspace switch. ## The organisational part The technical policy fails if every team re-decides it per link. What works is one owned default per surface class, expressed in whatever shared link component the product has, plus an explicit escape hatch that is easy to use and visible in review. Escape hatches used routinely are a signal the default is wrong for that surface, not that engineers are careless. And the policy should be revisited when a surface changes shape: a page that grew from twelve links to two hundred has silently changed its bill without anyone editing the prefetch setting. ## What a strong answer admits That this is a judgment call with no universally correct answer; that the right default depends on the audience's connections as much as on the architecture; that the benefit is bounded by how much of the click latency is actually network; and that a route which is slow because of server work will still feel slow when it is warm only if the speculation had enough lead time. Saying prefetch everything is a signal that the candidate has never watched their own origin absorb it.

  • Why is a global prefetch default risky in a large product?
    Because cost scales with the number of rendered links while benefit scales with hit rate, and those diverge wildly by surface. The same setting that costs five requests on a navigation bar costs hundreds on a feed. A global default also hides changes: a surface that grows from a dozen links to hundreds silently multiplies its bill with no code change to review.
  • How does speculative traffic behave during a traffic spike?
    It multiplies with real traffic, so the extra load arrives exactly when the origin has least headroom. If the prefetched routes are dynamic rather than cacheable, each speculative fetch is real server work. That is why a product-wide kill switch that does not need a deploy is worth more than a careful default.
  • What would tell you a prefetching policy is not worth its cost on a surface?
    Warm and cold navigation latency being close in field data, a high wasted-fetch ratio, or a measurable share of the origin's load coming from fetches never used. Any one of those means you are spending bandwidth and capacity for milliseconds — and on metered connections you are spending someone else's money to do it.

saying these in an interview costs you the question

  • Sets one global prefetch default and never revisits it per surface.
  • Justifies the policy with navigation speed alone, with no cost number.
  • Ignores that speculative load scales up during a traffic spike.
  • Measures only in the lab, where bandwidth is free and fast.
  • Has no way to disable speculation without a deploy.
  • Treats frequent per-link opt-outs as carelessness rather than a wrong default.