What bounds the TTL of a cached risk flag whose inputs include a customer signal refreshed hourly?
answer
- the entry is not the only thing ageing
- staleness compounds
- tolerance first, then subtract
- the freshest input that can change it
- inputs inside the key impose no bound
basics
~20 sTwo quantities: the staleness the review workflow tolerates, and how old the inputs already were when the entry was written. A cached output ages on top of its inputs, so the TTL is the tolerance minus the worst input lag.
solid answer
~40 sA cached prediction freezes the feature values that produced it, so what a reviewer experiences is not the age of the entry but the age of the entry plus the age of those inputs when it was written. Work the bound backwards. State the staleness the workflow accepts - say six hours. Find the freshest input that can change the output and its worst-case lag: an hourly signal can be read moments before its next refresh, so one hour. Set the TTL to the difference, five hours. An entry written at 08:59 from a signal computed at 08:00 then expires at 13:59 having served an answer at most five hours fifty-nine minutes old. Inputs pinned in the key - the clause text, the scorer version - impose no bound at all.
code
pseudocode · 17 linestoleratedStaleness = 6 hours // agreed with the review workflow
signalRefreshPeriod = 1 hour // the customer signal is recomputed hourly
worstInputLag = signalRefreshPeriod // an entry can be written just before a refresh
ttl = toleratedStaleness - worstInputLag // 5 hours
on write(entry, prediction, signal):
entry.prediction = prediction
entry.featureAsOf = signal.computedAt
entry.expiresAt = now + ttl
on read(entry):
if now >= entry.expiresAt:
return miss // ordinary TTL expiry
if now - entry.featureAsOf > toleratedStaleness:
return miss // inputs were older than assumed
return entry.predictiongo deeper
Remember that a cached prediction is only as fresh as the data behind it, and that nothing in the system complains when a stale but well-formed answer is served.
Derive the number: subtract the worst-case age of the freshest changing input from the staleness the workflow tolerates, and set the TTL to what is left.
Separate the inputs that belong in the key from the ones the clock has to bound, state the tolerance as an agreement rather than a guess, and measure served input age to prove the bound holds.
Treat tolerated staleness as a product commitment: a longer TTL buys hit rate and capacity, and the cost is paid by reviewers acting on older evidence in a workflow with legal consequences.
## Staleness compounds A prediction cache stores an answer computed from inputs that already had an age. Serving it later adds a second age on top: `served staleness = age of the inputs when the entry was written + time since the entry was written` The second term is what the TTL controls. The first is set by the pipelines feeding the scorer and is completely invisible from inside the cache. A team that sets the TTL equal to the tolerated staleness has quietly promised something the system cannot deliver, and nothing will ever report the breach. ## Working the bound backwards 1. **State what the workflow tolerates.** In a legal-review queue, a risk flag that is a few hours behind is usually acceptable because a human reads it anyway; this number is a commitment to the workflow, not a guess. 2. **Find the freshest input that can change the output.** Not the slowest - the freshest, because one changed input is enough to change the label. An hourly recomputed customer signal has a worst-case lag of one hour, since an entry can be written moments before the next refresh lands. 3. **Subtract.** Six hours of tolerance minus one hour of input lag leaves a five-hour TTL, and the worst case lands exactly on the promise instead of an hour past it. | Input to the risk flag | How it changes | Bound it puts on the TTL | |---|---|---| | the clause text | never, for a given key | none - it is in the key | | the scorer version | on release | none - handled by the key, not the clock | | the customer's account tier | rarely, on contract change | the refresh interval, unless pinned in the key | | a rolling dispute count | recomputed hourly | the tolerance minus one hour | ## Not every input constrains the clock An input pinned in the key cannot change underneath an entry: a different value produces a different key, and the stale entry is simply never read again. That is why the clause text imposes no bound, and why the scorer version is handled by the key rather than by expiry. What remains - values fetched at scoring time and refreshed on their own schedule - is exactly the set the TTL has to bound. The lag figure itself comes from the feature pipelines, which publish the freshness bound they guarantee. How that bound is set, monitored and repaired after a backfill belongs to the feature-platform subject; here it is an input to an arithmetic problem. ## The failure that produces no error A stale cached prediction is a completely successful response. The lookup succeeds, the payload is well formed, no timeout fires, the error rate stays flat, and the reviewer acts on a flag computed from data that has since moved. Nothing in the serving path can tell the difference, because the path was never entered. So make it measurable. Stamp every entry with the timestamp of the inputs it was built from, and on every hit record the age of that stamp. The distribution of served input ages is the real freshness measurement, and its p99 compared against the stated tolerance is the only evidence that the bound you derived is the bound the system honours. A p99 above the tolerance means either the arithmetic was wrong or an input is lagging more than its owner claims. ## Trading freshness for hit rate The TTL is the single knob on this axis and it moves both quantities at once. A longer TTL serves older answers and lifts the hit rate, because more repeats fall inside the window; a shorter one does the reverse. The sensitivity depends on the gap distribution between repeats: where tail keys recur about twice a day, cutting a five-hour TTL to one hour converts most of the tail's repeats into fills and can cost a large share of the measured hit rate for freshness nobody asked for. Set the TTL at the longest value the tolerance allows, not the shortest the inputs permit. ## When a clock is the wrong instrument If an input changes rarely and its change is observable, two alternatives beat waiting out a TTL: - **Drop the affected entries when the change happens.** One account-tier change invalidates every entry for one customer, which is cheap if the namespace is organised so those entries can be addressed together. - **Pin the value in the key.** The change then produces new keys and the old entries age out unread, at the cost of multiplying the key space by the number of distinct values and cutting the hit rate accordingly. Which is cheaper depends on how many entries one change touches and how many values the input takes. A slowly changing input with two values is a good candidate for the key; one with a million values belongs behind the clock.
- The tolerance is six hours and the signal refreshes hourly, so why not just set the TTL to one hour?You can, and you pay for it in hit rate. Shortening the TTL turns every repeat arriving more than an hour after the last one into a fill, and in a corpus where tail keys recur a couple of times a day that is most of the tail. The TTL is the knob that trades hit rate against freshness, so set it at the longest value the tolerance permits rather than the shortest the inputs allow.
- What happens to the bound if the hourly signal's pipeline falls behind?The bound silently stops holding: entries are written from inputs older than the assumed one-hour lag, and the served staleness exceeds the tolerance without the TTL changing. That is why entries carry the input timestamp and hits record its age - the read-side check on that stamp is what turns a lagging upstream into a miss rather than a quietly wrong answer.
A cached prediction is a printed report: the TTL says how long you are willing to keep the printout, but the numbers on it were already a little old when it came off the printer. Age of the paper plus age of the figures is what the reader actually gets.
saying these in an interview costs you the question
- Sets the TTL from memory pressure rather than from input freshness.
- Assumes a freshly written entry was built from fresh inputs.
- Picks a round number like one hour with no tolerance stated.
- Counts inputs that are already part of the key as TTL constraints.
- Bounds the TTL by the slowest-changing input instead of the freshest.
- Expects a stale cached label to surface somewhere as an error.