In an online feature store, why does a value that stopped updating hours ago still serve a prediction without any error?
answer
- a lookup, not a subscription
- the read path consults no clock
- a plausible value, not a missing one
- last written wins, however old
- store writtenAt beside the value
basics
~20 sA feature read is a key lookup that returns whatever was last written, with no age check. The stale value is well-formed, so the model scores it and the wrong answer shows up only in the output.
solid answer
~50 sThe online tier keeps one row per entity key — here one row per home — and a read is a point lookup on that key. It returns the value the feature pipeline wrote last, and nothing on the read path compares that write time against the feature's declared age bound unless someone built that comparison. So when the uploader for a region stops, the indoor-temperature row keeps its 07:00 value and the comfort model scores a 19:00 request against it: the lookup hits, the number is inside a plausible range, the model returns a score, the request succeeds. Staleness is a *silent* failure because it produces a believable input rather than a missing one. It becomes visible only if the row carries a write timestamp and the serving path checks `now - writtenAt` against the bound.
go deeper
Be able to say that a feature read is a key lookup returning the value last written, and that an old value looks exactly like a fresh one to the model that scores it.
Explain the mechanics: the write timestamp has to be stored beside the value and compared at read time against that feature's bound, because no layer in the path compares clocks on its own.
Show how you would have caught it in production — age measured on values that were actually served, alarmed at a high percentile — and what the serving path does once the breach is visible.
Weigh where the default belongs: a platform that refuses by default is safe and noisy, one that serves flagged is quiet and dangerous, and someone has to own that choice across every consuming team.
## What a feature read actually is The **online tier** of a feature platform holds one row per **entity key**. In a smart-thermostat comfort predictor the entity is a home, and its row carries the values the model needs at request time: the current indoor temperature, the latest outdoor conditions, a thirty-day occupancy profile. A read is a **point lookup** — hand over the home id, get the row back. That is the whole mechanism, and it explains the question. A lookup is not a subscription. Nothing tells the reader that the value changed, that it failed to change, or how long ago it was written. The tier answers one question — *what is stored under this key* — and it answers it correctly even when the stored thing is twelve hours out of date. ## Why nothing in the path reports a problem Walk the request end to end and ask, at each layer, what would have to notice: | layer | what it would have to notice | why it does not | |---|---|---| | the online tier | that the row is older than the feature's bound | it stores a value; expiry, where a store offers it, is a storage policy rather than the feature's declared contract | | the serving code | that `now - writtenAt` exceeds the bound | it can only do this if the row carries a write timestamp **and** the code compares it | | the model | that this input is out of date | it receives numbers with no provenance attached | | the caller | that the prediction is degraded | it receives a well-formed response in the normal range | Every layer behaves correctly. There is no bug to find, because the guarantee nobody stated is the guarantee nobody enforced. ## What it looks like in the comfort platform Suppose the indoor-temperature feature declares a **sixty-second** age bound, and at 07:00 the ingestion path for one region stops accepting uploads. The rows for those homes freeze at their 07:00 values — cool rooms, early morning. At 19:00 a request arrives for one of those homes on a hot evening: - the lookup **hits**: the key exists, so there is no miss, no null, no handled error branch; - the value is **plausible**: 18.2 degrees is a legal indoor temperature at any hour; - the model **scores** it and returns a comfort score suggesting the home should be warmed; - the response is a number, returned inside the latency budget. Error rate flat. Latency flat. Every request-level dashboard green. The only damage is that thousands of homes are being heated against a twelve-hour-old reading, and the first report comes from a customer, not from the system. ## The distinction that matters: missing versus stale A **missing** value and a **stale** value are different failures, and only one of them is handled by default: - a missing value returns nothing, so the serving code hits its null branch — a default, a fallback, a refusal — because that branch had to exist for the code to compile and run at all; - a stale value returns something, so no branch fires; the request follows the happy path from end to end. This is why teams who have carefully handled missing features are still surprised by stale ones. The null path was forced on them; the age path was not. ## Making staleness loud Making the failure visible takes four deliberate steps, in this order: 1. **Store a write timestamp beside every value**, so a row carries not just *what* but *as of when*. 2. **Compare at read time** against the feature's declared age bound, on the serving path, per feature — the bound differs per feature, so one global check is not enough. 3. **Decide what a breach does**: serve the prediction with the staleness recorded, substitute a coarser feature that is still inside its own bound, or refuse to score. Choosing silently to serve is still a choice, and it is the one that hurt here. 4. **Measure the age of served values** and alarm on a high percentile against the bound, rather than trusting that a successful pipeline run implies fresh rows. None of this is exotic — it is a timestamp, a subtraction and a branch. It is skipped so often because the system without it looks perfectly healthy right up until someone reads the predictions.
- Should the age check live in the serving code or inside the feature store's read API?Putting it in the read API makes it uniform — every consumer gets the same check, and a new model cannot forget it. Putting it in the serving code lets each consumer choose its own response to a breach, which differs by caller. The common settlement is both: the read API returns the value together with its age, and the consumer decides what to do with a breach.
- If the uploader recovers after twelve hours, is the model correct again straight away?The serving path is, on the next write: the row is overwritten with a current value and requests go back to being scored on fresh input. History is not. The offline series still has a twelve-hour gap for those homes, so any training set assembled over that window is missing them or carries frozen values — which is repaired by a backfill, not by the uploader coming back.
A kitchen that plates every dish from a prep list nobody refreshed. The plates come out looking right, because the list never says when it was written.
saying these in an interview costs you the question
- Claims a stale feature read fails, times out, or returns an error
- Assumes every online tier evicts a value once it passes its age bound
- Thinks the model can tell a stale input from a fresh one
- Treats a successful pipeline run as proof the served values are fresh
- Says only fast-moving features need an age bound at all
- Assumes the missing-value branch also covers an out-of-date value