skip to content

A document-QA cache with a five-minute window expires before almost every turn because users return after 40 minutes — longer-TTL tier, keep-warm traffic, or neither?

level: principalimportance: should knowfreq 34%

answer

  1. the average gap is the wrong statistic
  2. gap distribution, not mean, decides
  3. heartbeat cost scales with prefix count
  4. one shared prefix, many users: heartbeat wins
  5. bounded batch window favours the long tier

basics

~20 s

Start from the measured gap distribution, not the average. A longer window pays when a real share of follow-ups land inside it; keep-warm heartbeats pay only when one warm prefix serves many users; with per-user prefixes and 40-minute gaps, the honest answer is often neither.

solid answer

~60 s

First measure, then choose. The deciding number is the **share of requests arriving within each candidate window**, taken from the inter-request gap distribution for that prefix — a mean of 40 minutes hides a bimodal reality where half the follow-ups arrive in 30 seconds. Then weigh three levers. A **longer-TTL tier** costs a higher write premium once and wins when a meaningful fraction of sessions have a second turn inside the longer window. **Keep-warm traffic** — a heartbeat that touches the prefix before it lapses — only pays when few distinct prefixes serve many users, because its cost scales with the number of prefixes times the heartbeat rate, not with your traffic; per-user document prefixes make it strictly worse than doing nothing. **Neither** is a legitimate answer for single-turn or long-gap workloads, and saying so is better than instrumenting a cache that never hits. Traffic shape decides: bursty batch windows favour the long tier, steady trickle needs no help, idle-gapped per-user traffic favours restructuring the prompt so the shared part is what gets cached.

go deeper

for a junior

Know that a cache only helps when the next request arrives before the window closes, and that long gaps between user turns mean most requests start cold.

for a middle

Be ready to explain the two concrete levers — a longer window at a higher write premium, and a periodic request that refreshes the entry — and to name which traffic shapes each one suits.

for a senior

Show that you would measure the per-prefix gap distribution before choosing, gate any keep-warm traffic on operating hours, and recognise when the correct answer is to leave the path uncached.

for a principal

Own the framing: this is a workload-characterisation decision, not a prompt one. Argue for restructuring so shared material is what gets cached, state plainly when caching does not fit, and set a trigger for revisiting the decision as traffic shape drifts.

## The number that decides it The instinct is to compare the window length against the average gap between turns. That is the wrong statistic. What matters is the **cumulative distribution of inter-request gaps for a given prefix**: what fraction of requests arrive within five minutes of the previous one, within an hour, at all. A 40-minute mean is compatible with two very different realities — a uniform 40-minute spacing where no window under an hour ever helps, or a bimodal pattern where 60% of follow-ups arrive within a minute and a long tail drags the mean up. In the second case a short window is already capturing most of the value and the whole problem was a misread average. So the first move in the interview is to say what you would measure: per-prefix gap percentiles, the share of sessions with a second turn, and the size of the stable prefix (a 120k-token document rewrite hurts far more than a 2k system prompt). ## Lever one: a longer window Providers commonly offer an extended window at a higher write premium. It converts a recurring cost — writing the prefix once per gap — into a single more expensive write that covers a longer stretch. It wins when a real share of follow-ups fall between the short and long window, and it loses when sessions are genuinely one-shot, because you have then paid the premium for a write nobody reads. The clean case is a **scheduled batch window**: a job that runs 09:00–10:00 every weekday and nothing else. The entire run fits inside one long window, so one premium write serves the whole hour, and nothing needs to stay warm for the remaining 23 hours. Traffic that is dense inside a bounded window is the single best fit for the long tier. ## Lever two: keep-warm traffic A heartbeat — a small periodic request that touches the prefix just before it would lapse — pins an entry indefinitely by exploiting sliding expiry. Its economics are unusual because the cost does **not** scale with your traffic; it scales with the number of distinct prefixes multiplied by the heartbeat frequency, and it runs whether anyone is using the system or not. That gives a sharp rule. Heartbeats pay when one large, shared prefix serves many users, so a handful of keep-warm requests protects thousands of real ones. They stop paying when prefixes are per-user or per-document, because each one needs its own heartbeat and you are now writing continuously on behalf of users who may never return. They also stop paying overnight and at weekends, when the pinned entry serves nobody — so gate keep-warm on observed traffic, or on the schedule the workload actually runs to, rather than an unconditional cron. The honest check is a count, not a spreadsheet: does the number of keep-warm touches over a period come out below the number of writes those touches avoid? For one shared prefix under intermittent traffic it does, easily. For thousands of per-user prefixes it never does. ## Lever three: do not cache that path Sometimes the workload simply does not have the shape. Single-turn requests each carrying a different document, or genuinely uniform long gaps with per-user prefixes, produce a hit rate near zero however you tune the window. Caching then adds a write premium and operational surface for nothing. Recognising this and saying it is a stronger answer than defending a mechanism that does not fit. ## The fourth move: change the shape The most valuable answer is often none of the three. If per-user document prefixes are what makes caching useless, ask whether the genuinely shared material — the system prompt, the retrieval instructions, the shared policy or schema — can be the part that gets cached, so one warm entry serves everyone while the per-user document sits outside it. That converts a workload with thousands of cold prefixes into one with a single hot prefix, and it is the change with the largest effect on the whole problem. ## Putting it together Map shape to lever: **bursty inside a bounded window** → long tier, no heartbeats. **Steady trickle on a shared prefix** → short window is enough, natural traffic keeps it warm. **Intermittent traffic on one shared prefix** → heartbeat, gated on hours of operation. **Idle-gapped, per-user prefixes** → restructure so the shared part is cached, and accept misses on the rest. Then instrument write-versus-read tokens per route so the decision is revisited when traffic shape changes — because it will, and every one of these choices is a bet on a distribution rather than a permanent property of the system.

  • Your traffic is one prefix shared by all users, arriving in bursts every 20 minutes during business hours. What would you do?
    Heartbeat during business hours only. One prefix means one keep-warm stream, and a touch inside each short window pins the entry so every burst reads rather than writes. Gate it on the operating schedule so it stops overnight, when a pinned entry serves nobody. The long tier is unnecessary here because a single cheap touch already spans the gap.
  • How would you decide whether the extra write premium of a longer window is justified without doing full pricing analysis?
    Compare fractions, not currency. From the gap distribution, take the share of requests that would hit under the long window but miss under the short one; that is the read volume the premium buys. If it is a few percent of traffic, the premium is dead weight; if it is most of the follow-ups in a bounded batch window, it is clearly worth it. Detailed unit-cost work only matters once that fraction is materially large.
  • What changes your answer if the stable prefix is 2k tokens rather than 120k?
    Almost everything. The cost and latency of a miss scale with the prefix, so a small prefix makes the entire question low-stakes — accept the misses and do not add keep-warm machinery. Large prefixes are what justify a premium tier or a heartbeat, because one avoided rewrite is worth a great deal. Size the effort to the size of the prefix.
  • When would you revisit a keep-warm decision you made six months ago?
    When traffic shape moves: growth that turns intermittent bursts into steady traffic makes heartbeats redundant, and a shift toward per-user prefixes makes them actively wasteful. Watch the ratio of keep-warm writes to real reads on the same prefix, and treat a rising ratio as the signal to switch off rather than tune.

saying these in an interview costs you the question

  • Compares the average gap against the window instead of the distribution
  • Runs per-user heartbeats and calls it keep-warm
  • Leaves keep-warm traffic running overnight and at weekends
  • Assumes the longest available window is always the better choice
  • Never considers that this path should not be cached at all

context