skip to content

A stale prerendered page is hit by a burst of requests at once. How many out-of-band renders run, and what do the requests get?

level: seniorimportance: should knowfreq 52%

answer

  1. the stampede arrives at peak popularity
  2. refresh as a per-path singleton
  3. the rest of the burst is not queued
  4. cost scales with paths due, not views

basics

~20 s

Typically one render per path rather than one per request: implementations keep a single refresh in flight for a path and answer the rest of the burst from the stored copy. Render and data-source load therefore tracks paths refreshed, not traffic.

solid answer

~50 s

The burst does not become a burst of renders. Implementations keep **one refresh in flight per path**: the first request that makes the path due starts a render, and while it runs, further requests for the same path do not start another — they are answered from the stored copy, exactly as they would be otherwise. Nothing is queued, so those requests are not slower. The important consequence is where load lands: with rendering inside the request, a thousand simultaneous requests mean a thousand renders and a thousand rounds of data access; here they mean one. The coalescing is scoped to the path, though, so a burst spread over forty stale paths can start up to forty renders, and separate server instances usually track in-flight work separately, so the same path can be refreshed once per instance unless they share that state.

go deeper

for a junior

The key idea is that a crowd arriving at once does not multiply the work. One refresh runs for the page, and everybody in the crowd is handed the copy that is already stored.

for a middle

Explain the rule and its consequence: one render in flight per path, the other requests served normally, and a render bill that follows how often pages become due rather than how often they are viewed.

for a senior

Show where it breaks — many paths becoming due together, in-flight state that is per process rather than shared, and renders that outlast the freshness window — and name the metrics that would have told you.

for a principal

Treat refresh concurrency as a capacity question. Decide the ceiling on concurrent renders and how expiries are spread, so a single content event cannot turn into a self-inflicted load test on the data source.

## What a burst would do without coalescing Suppose every request that found a stored copy due started its own refresh. A popular page whose freshness window has just elapsed receives hundreds of requests in the same second; each one begins a render; each render queries the same data source for the same data and produces identical HTML. The data source absorbs the whole burst, the renders compete for the same CPU, and every one of them but the last is thrown away when the next replaces it. That is the classic stampede, and it arrives exactly when a page is most popular. ## One render in flight per path The standard defence is to treat a refresh as a **per-path singleton**. The rule is simple to state: 1. A request finds the path due for refresh. 2. If no refresh for that path is running, one is started. 3. If one *is* running, nothing new is started. 4. In every case, the response is served from the stored copy. 5. When the render completes, the stored copy is replaced, and the next requests see the new page. Step 4 is the one people forget. The extra requests are not held, not queued behind the render, and not penalised for arriving during it. From their point of view the refresh does not exist — the same invariant that governs the single-request case. ## Where the load actually lands | Scenario | Renders under request-time rendering | Renders when refreshing out of band | |---|---|---| | 1,000 requests, one stale path | 1,000 | about 1 | | 1,000 requests spread over 40 stale paths | 1,000 | up to about 40 | | 1,000 requests, nothing due | 1,000 | 0 | The practical reading: **render cost is driven by how many distinct paths become due and how often, not by how many people look at them**. That is what makes the model affordable on a high-traffic site with expensive pages, and it is the sentence to say out loud in an interview. ## Where the coalescing stops helping Being precise about the limits is what separates a senior answer from a textbook one. - **It is per path, not per site.** A change that makes many paths due at once fans out into many renders. If those renders hit one data source, you have moved the stampede from the request path to the refresh path rather than removing it. - **It is usually per process.** Several server instances each keep their own idea of what is in flight, so the same path can be rendered once per instance around the same moment. Shared coordination is a property of the hosting arrangement, not of the rendering model. - **It does not bound render duration.** If a render takes longer than the freshness window, the path is due again as soon as it lands, and the system settles into refreshing continuously. - **It does not protect anything the render itself fans out to.** One render that issues fifty data requests still issues fifty. ## What to watch in production - **Renders started per path per minute.** It should look like a function of the freshness policy, not of traffic. If it tracks traffic, coalescing is not working the way you think. - **Render duration against the freshness window.** Duration approaching the window is the early warning for continuous refreshing. - **Concurrent renders across instances.** A number that scales with instance count tells you the in-flight state is not shared. - **Data-source load at refresh time.** Look for it to be flat and low relative to request volume; a spike shaped like your traffic curve means requests are still driving renders one-to-one. ## Where frameworks differ Frameworks implement the singleton differently and are not equally strict about it: some hold a lock for the duration of the render, some deduplicate on a key derived from the path, some cap the number of concurrent refreshes across all paths so a mass expiry cannot saturate the process. On a long-running server the in-flight state naturally lives in memory and covers everything that process serves; on a platform that spins up short-lived instances per request, there may be no shared memory to hold it in, and the deduplication has to come from wherever the output is stored. When you are asked the question, give the invariant — one render in flight per path, everyone else served from storage — and then say which of these variations applies to the system you are describing.

  • Are the requests that arrive during a refresh slowed down at all?
    No. They are answered from the stored copy on the normal fast path; they are not queued behind the render and they do not wait for it. Their only cost is that the content they receive predates the render in progress — which is the same cost every request inside a refresh window pays.
  • What happens if a render takes longer than the freshness window?
    The path is due again almost as soon as its refresh lands, so the system refreshes more or less continuously and the served copy is never near-current. It is a signal to lengthen the window, make the render cheaper, or move that surface to a different model — not something the coalescing rule can fix.
  • Why does a mass expiry across many paths still hurt?
    Coalescing is scoped to a path, so many due paths mean many concurrent renders, each doing its own data access. The stampede moves off the request path but still lands on the data source. Capping concurrent refreshes, or spreading when paths become due, is what limits it.

saying these in an interview costs you the question

  • Thinks each request that finds a copy due starts its own render.
  • Says requests during a refresh are queued behind it.
  • Assumes coalescing is global rather than scoped to one path.
  • Believes several instances automatically share in-flight state.
  • Claims refresh load scales with traffic to the page.