How would you set the default staleness window for cached server data across an app, and who should own the exceptions?
answer
- the window is a written-down bet
- classify by consequence, not by feel
- one default, documented exceptions
- shrinking the window never reaches zero
- unreproducible support reports are a signal
basics
~20 sSet the default from the consequence of showing a slightly old value, not from a feeling about speed. Classify data by that consequence, give the app one modest non-zero default, and record every exception in a single reviewable table.
solid answer
~50 sStaleness is a product decision expressed as a number, so argue it from consequences. Classify server data by what it costs to show a value a minute old: money, stock, permissions and anything about to be acted on irreversibly at one end; labels, taxonomies and reference lists at the other. Pick one modest non-zero default, because zero quietly turns the cache into a request multiplier and a very long default hides wrong numbers in long-lived tabs. Allow exceptions per class, owned by the team that owns the data but recorded in one table naming the window, the reason and the triggers. Back it with evidence: requests per session, hit rate, and reports of the shape `wrong until I reloaded`. Where correctness needs more, invalidate at the moment of change or confirm as part of the action.
go deeper
Recall that how long cached data is trusted is a deliberate setting, not a technical constant, and that different kinds of data deserve different answers.
Explain both failure directions concretely: a near-zero window makes the cache a request multiplier, a very long one hides wrong values in long-lived sessions.
Argue from consequence per data class, know when to reach for invalidation or a server-confirmed action instead of a shorter window, and bring request and hit-rate numbers to the discussion.
Make it a reviewable policy with one owner for the default and named owners for exceptions, hold it to measured evidence, and re-classify surfaces when the product gives them irreversible actions.
## Why this is a policy question, not a tuning knob Every read of server-owned data is a bet: show what is in memory now, or spend a round trip to be sure. The staleness window is where that bet is written down. Set app-wide by whoever configured the data layer first, it becomes an accidental policy that every screen inherits — and the two failure modes it produces look nothing alike, which is why teams rarely connect them to the same number. - **A window of zero or near-zero** makes the cache a deduplicator and nothing more. Every entered screen, every return to the tab, every reconnect becomes a request. The symptoms are server cost, flaky behaviour on poor links, and a UI that spends its life in a subtly loading state. - **A very long window with no invalidation** gives an app that feels instant and is occasionally, invisibly wrong. The reports are unreproducible by definition: the value was right by the time anyone looked. ## Classify by consequence The only durable basis is what a slightly old value costs. A workable classification: | class | examples | window | extra measures | |---|---|---|---| | act-on-it data | amount to be charged, remaining stock, current permission, a limit being approached | very short, or treat every read as untrusted | confirm server-side as part of the action; never let the cached copy decide | | shared, fast-moving | a queue, an assignment, a status others change | short, with reconnect and focus refetching on that surface | consider invalidation at the moment of change, or a push channel | | personal, slow-moving | the user's own settings, their profile, their preferences | long | invalidate when the user changes it | | reference data | labels, categories, taxonomies, country lists, feature copy | very long | refresh on release or on explicit invalidation | The classification is the deliverable, not the individual numbers. Numbers get argued about once per class instead of once per screen, and a new screen inherits an answer rather than starting the discussion again. ## Ownership 1. **One default, owned centrally**, set deliberately and documented with its reasoning. Modest and non-zero is the right shape: long enough that a navigation does not always cost a request, short enough that nothing important hides behind it. 2. **Exceptions owned by the team that owns the data**, because only they know the consequence. A frontend guild cannot judge how bad a stale balance is. 3. **All exceptions in one visible table**, naming the class, the window, the refetch triggers and the reason. The point is reviewability: without a table the policy lives in scattered call sites and nobody can answer `how stale can this app be`. 4. **Escalation path for correctness.** When a class cannot tolerate any window, the answer is not an ever-shorter window — it is invalidating at the moment of change, pushing the change to the client, or re-reading and confirming as part of the action. A shrinking window is a tax, and it never reaches zero staleness anyway because the value can change while the response is on the wire. ## Evidence to hold the policy to - **Requests per session, split by key family and by trigger.** One family dominating is the usual finding, and it is usually fixed by its window rather than by disabling a trigger. - **Cache hit rate per family**, which tells you whether the window is doing anything at all. - **Support reports of the shape `wrong until I reloaded`**, mapped onto classes. Each one is either a misclassified class or a missing invalidation. - **Long-session behaviour.** A tab left open for a day is where a lax policy shows: what does it display, and what does it do on refocus and reconnect? ## The trade to state plainly There is no window that is both free and correct. Shortening it buys confidence with requests, latency and battery; lengthening it buys speed and calm with a risk measured in what the user might do with an old value. Naming which data may be old, for how long, and what happens instead where it may not, is a decision a lead should be able to defend in a sentence per class — and revisit when the product changes, because a screen that grows a `pay now` button has just changed classes.
- Why is a staleness window of zero not the safe default it appears to be?It removes the cache's value while keeping its cost, so every navigation, refocus and reconnect becomes a request, and the app degrades worst on the weak links where it matters most. It also does not make data correct: the value can change while the response is in flight, so anything genuinely correctness-critical still needs confirmation at the moment of action.
- What does this in-memory cache decide that an on-the-wire response cache does not?They answer different questions. The transport-level cache decides whether a request can be satisfied without a round trip. This cache decides which value the component tree renders right now, which readers share one copy, when to ask again, and what stays on screen while the answer is awaited. A transport hit is still a fetch the tree has to await.
- How would you notice that a staleness policy is set wrongly?Two signals. Requests per session dominated by one key family, with a hit rate near zero for it, means the window is too short for that class. Support reports that a value was wrong until reload, never reproducible afterwards, mean a class is too lax or is missing an invalidation at the moment the data changes.
saying these in an interview costs you the question
- Sets every window to zero and calls it correctness
- Chooses windows by feel rather than by consequence
- Lets each screen invent its own window with no record
- Assumes a short window removes the need for confirmation on write
- Judges the policy with no request or hit-rate numbers
- Never revisits a class when a screen gains a destructive action