In a comfort platform, why is a p99 of served feature age a better freshness alarm than a materialization job's exit status?
answer
- measure the value, not the producer
- read time minus written-at
- sample stratified by region or partition
- a percentile maps to the bound
- green runs hide a stopped region
basics
~20 sA job's exit status says a run finished, not that homes hold fresh values. Age measured on what reads actually return catches a partial write, a stalled partition and a lagging source — all of which leave the job green.
solid answer
~40 sMeasure the thing you promised. The bound is a statement about the value a request receives, so the measurement is `readTime - writtenAt` on values that were actually served, gathered either from a sample of real reads or from a probe that reads a stratified sample of entity keys on a schedule. Then compare a high percentile against the bound, because the bound is a per-value promise: the reading you need is *what fraction of reads violate it*, which an average cannot express. A p99 of 3,600 seconds against a 60-second bound says at least one read in a hundred is an hour stale — across a million homes, ten thousand of them getting a wrong answer on every request — while the pipeline reports a week of clean runs.
go deeper
Know that freshness is measured on the value a request receives — read time minus the stored write time — and not inferred from whether the producing job finished.
Explain why a percentile is the right reading for a per-value promise, and what a p99 above the bound implies about the share of requests being served stale input.
Name the failures that leave a job green — a stopped region, a partial write, a late source, runs exceeding their interval — and design the stratified probe that surfaces each.
Decide what the platform measures by default for every feature, since an age signal added per team is a signal most teams will not add until after their first silent incident.
## What a green run actually asserts A successful materialization run asserts one thing: the process started, did not throw, and exited. It does not assert that every entity received a value, that the input it read was current, or that the write landed everywhere it was meant to. Between the job and the served value there are several places the promise can break while the exit status stays clean. ## Measure the value, not the producer The bound is defined on the value a request reads, so that is where the measurement belongs: - **age** = the read time minus the timestamp stored beside the value; - **sampled on served values** — tag a fraction of real reads, or run a probe that reads a random sample of entity keys on a schedule; - **stratified** by region, partition or ingestion path, because staleness is almost never uniform — it is usually shaped like whichever upstream path stopped; - **compared per feature**, against that feature's own declared bound, since a day-old occupancy profile is healthy and a day-old room temperature is not. ## Why a percentile rather than an average A bound is a promise about **each** value, so the question it poses is what share of reads violate it — and a mean cannot answer that. A distribution where almost everything is a few seconds old and a slice is hours old produces an average that belongs to neither population, and no threshold on it maps cleanly onto the contract. A high percentile maps directly. If the p99 of age is 3,600 seconds while the bound is 60 seconds, then at least 1% of reads are at least an hour past the bound. Across a million homes that is on the order of ten thousand homes receiving a prediction built on hour-old room state, every time they are scored. Which percentile to watch follows from the blast radius you care about: p99 surfaces a problem touching more than one read in a hundred, and catching a smaller slice requires looking further into the tail. ## Failures that leave the job green | failure | what the job reports | what age shows | |---|---|---| | one region's ingestion path stops | success — there was nothing to process for those homes | a distinct stale population in that region | | a run succeeds but writes only part of its output | success | ages rising for the entities that were skipped | | the upstream source starts publishing late | success — it processed what it received | every age shifted by the source's new delay | | runs now take longer than their interval | success, every time | ages growing steadily run after run | Every row is a real breach of the declared bound, and not one of them produces a failed run or a request error. ## Building the probe A workable measurement is small: - pick a stratified sample of entity keys, weighted so every partition and region is represented rather than only the busiest ones; - read each key's features on a schedule tighter than the tightest bound you are checking; - record the age distribution per feature and per stratum; - alarm when the percentile crosses the declared bound, and page on the rate of breach rather than on a single sample. The result is the only freshness signal that says what the contract says. Job health is still worth watching — it tells you *why* — but it can never tell you whether the promise is being kept.
- How do you measure served age without instrumenting every read?Sample. Tag a small fraction of real reads with their computed age, or run a probe that reads a stratified sample of entity keys on a schedule tighter than the tightest bound. Stratify across partitions and regions rather than sampling uniformly by traffic, because a stalled ingestion path affects a definable population that a traffic-weighted sample can miss entirely.
- Should pipeline health monitoring be dropped once served age is measured?No — they answer different questions. Served age tells you the contract is broken and roughly for whom; job health, lag and runtime tell you why, and often move before the age does. The distinction is which one pages: breach of the declared bound is the alarm, pipeline health is the diagnosis you reach for once it fires.
saying these in an interview costs you the question
- Treats a successful materialization run as proof the values are fresh
- Alarms on the mean age instead of the fraction past the bound
- Calls the job's wall-clock runtime the feature's age
- Samples one probe key and assumes it represents every entity
- Adds an age alarm only for continuously computed features