skip to content

As the lead, how do you decide which surfaces may serve a stale prerendered copy while a fresh render runs out of band?

level: principalimportance: nice to knowfreq 44%

answer

  1. eventually fresh, never fresh now
  2. cost of an old fact, per surface
  3. worst case, not typical age
  4. shared copy means no per-user content
  5. publish the age you serve

basics

~20 s

Classify content by the cost of being briefly wrong. Surfaces whose facts tolerate age take the stored copy; prices, availability, permissions and anything personalised do not. The model promises eventual freshness, never current data, so name a worst-case age and measure it.

solid answer

~50 s

Start from what the model actually guarantees: **eventually fresh**, with a worst case set by the refresh policy plus the render duration plus any period in which no refresh succeeds. Then sort surfaces by what an old fact costs. Editorial and marketing content, catalogues, documentation and aggregate figures tolerate minutes of age and belong here. Facts a user acts on irreversibly — a price at checkout, remaining stock, an entitlement — do not, and either move to rendering in the request or have the volatile value fetched by the client on top of the stored shell. Anything varying per user must not be a stored copy at all, since a stored copy is shared by whoever asks next. Finally, make the choice observable: agree a staleness ceiling per surface, publish the age you are actually serving, and alert on it — because this model fails by quietly serving old content rather than by going down.

go deeper

for a junior

The useful takeaway is that this choice is per page. Content that changes slowly and harms nobody by being a little old is what this model is for; anything a user acts on directly deserves a second look.

for a middle

Be able to justify a surface either way by naming the cost of an old fact on it, and remember that the served age includes render time and failed refreshes, not only the configured window.

for a senior

Show the operating side: a ceiling per surface, the age you actually serve published as a metric, an alert on age rather than status, and an escape hatch for the volatile value on an otherwise stale page.

for a principal

Own the contract and the consistency. A small set of named freshness tiers that product agrees to beats per-route numbers, and the difference between the configured window and the real guarantee has to be visible before an incident finds it.

## The guarantee you are buying The decision is easier once you stop treating the refresh policy as a maximum age. What this model provides is: *the served document is a complete render of this path from some earlier moment, and it will be replaced by a newer one in due course.* Nothing in that sentence promises any particular visitor a current page. So the question for each surface is not "how fresh can we make it" but **"what does it cost us when a user sees a page from some minutes ago?"** ## What sets the worst case The number to negotiate with stakeholders is a worst case, not a typical case, and it is assembled from several parts: - **The refresh policy**, which decides when a copy becomes eligible to be replaced. - **The render duration**, which extends the served age past that point for every refresh. - **Failing renders**, which hold the current copy in place indefinitely while nothing user-visible changes. - **Low-traffic paths**, whose copies can sit untouched for long stretches. - **Divergent instances**, if the refreshed output is not visible to every server that serves the path — in which case two users can be looking at different ages at the same instant. If you quote a single number, quote the sum of these, and be able to say which of them you measure. ## A classification you can defend | Content class | Cost of being a few minutes old | Verdict | |---|---|---| | Editorial, marketing, documentation | a reader sees a slightly older wording | good fit | | Catalogue and listing pages | ordering or membership lags reality | good fit, with a tighter window | | Aggregate figures and dashboards | a number trails the underlying data | fit, if the page says as of when | | Prices, availability, entitlements | a user acts on a fact the system will refuse | poor fit for the fact itself | | Anything specific to one user | the wrong person's data is served | never a shared stored copy | The last row is a correctness and privacy boundary rather than a tradeoff: a stored copy is, by construction, whatever the next requester receives. ## The escape hatches, in order of cost 1. **Split the page.** Keep the stored copy for the stable bulk and let the client fetch the volatile fact after load. The page stays cheap and the risky value is current. 2. **Shorten the window for that surface only.** Cheap to do, but it buys less than people expect: it moves the typical age, not the worst case, and it multiplies render load on the paths it applies to. 3. **Move the surface to rendering in the request.** Correct by construction, paid for on every view, and it reintroduces a dependency on the data source being up at request time. 4. **Say when it was produced.** For figures and listings, a visible "as of" line converts a silent inaccuracy into an honest one, and is often enough. ## What you owe a surface once you choose this model - **A named ceiling**, written where product people can see it, not inferred from a configuration value. - **Published age.** The age of what you are serving, per surface, as a metric with an alert. Request-level monitoring stays green through the entire failure mode, so it cannot be your detector. - **A decision for the too-old case**: keep serving and page someone, or stop trusting the copy on the surfaces where old is worse than slow. - **A review when the surface changes.** A page that was safely stale can stop being so the day a price is added to it, and nothing in the system will flag that. ## The organisational part The recurring failure is not technical. Someone sets a refresh policy as an implementation detail, it is communicated as "the page updates every minute", and months later an incident review discovers that no one owned the difference between that sentence and what the system guarantees. Two habits fix it: make the staleness ceiling a property of the surface that product and engineering agree on together, and demonstrate the first-request-after-expiry behaviour during design review so that nobody is surprised by it in production. Consistency helps too — a handful of named freshness tiers that teams choose from is far easier to reason about, and to audit, than a per-route number that each team picked independently.

  • A team wants a price on a stale-served page to always be current. What do you propose?
    Keep the stored copy for the page and fetch the price in the browser after load, or move that surface to rendering in the request if the price must be in the initial HTML. Shortening the refresh window is the tempting third option and the wrong one: it lowers typical age without removing the case where a user acts on an old number.
  • How do you set a staleness ceiling that survives contact with stakeholders?
    Express it as what a user may experience, not as a configuration value: "this page may show data up to ten minutes old, and we alert if it exceeds that." Derive the number from the cost of an old fact on that surface, and commit to measuring the age you actually serve rather than assuming the policy is the ceiling.
  • Why is per-user content excluded rather than merely discouraged?
    Because a stored copy is served to whoever requests that path next. Anything rendered into it that belonged to one user is then shown to another. That is a correctness and privacy failure rather than a freshness tradeoff, so it is a boundary on what may be stored at all, not a dial to tune.

saying these in an interview costs you the question

  • Treats the refresh window as a guaranteed maximum age.
  • Applies one freshness policy to every surface in the product.
  • Stores a page containing content specific to one user.
  • Relies on error rates to notice content going stale.
  • Assumes a shorter window removes the risk of an old fact.