What does a per-feature age bound commit a comfort platform to, when its features range from daily to per-request?
answer
- a contract, not a schedule
- each feature states its own number
- consumers read the bound, not the job
- interval plus runtime equals served age
- cheapest mode that meets the number
basics
~20 sIt commits the platform to serving that feature no older than the stated age, and to noticing when it cannot. The bound is a contract on the value a request reads, not a description of how often the job runs.
solid answer
~50 sEach feature declares one number: the maximum age a served value may have. In a comfort predictor the thirty-day occupancy profile declares a day, outdoor conditions an hour, the current indoor temperature sixty seconds, and a setpoint delta that only exists with the request is computed there and so has no stored age at all. Bounds are per feature because the cost of holding one is not shared — pushing everything to sixty seconds means running continuous computation over data that changes daily, and pinning everything to a nightly job makes the indoor reading worthless. The critical distinction is that a **bound is a guarantee about the value**, while a **cadence is one implementation that may or may not satisfy it**: a job that starts every five minutes but takes twenty serves values around twenty-five minutes old, whatever the schedule says.
code
json · 17 lines[
{
"feature": "indoor_temperature_now",
"entity": "home_id",
"maxAgeSeconds": 60,
"computedBy": "stream",
"onBreach": "fallback",
"fallbackFeature": "indoor_temperature_hourly_mean"
},
{
"feature": "occupancy_profile_30d",
"entity": "home_id",
"maxAgeSeconds": 86400,
"computedBy": "batch",
"onBreach": "serve_flagged"
}
]go deeper
Recall that different features are allowed to be different ages, and that the allowed age is written down per feature rather than inferred from how often a job happens to run.
Explain the arithmetic behind a served value's age — waiting for the next run plus the run's own duration — and why that makes a schedule a weaker statement than a bound.
Demonstrate how you would pick the number: replay history with artificially aged values and take the age at which prediction error crosses what the product tolerates.
Consider the platform-wide effect of bounds: each tight one is a standing operational commitment, and a catalogue where every feature claims sixty seconds is a catalogue nobody can staff.
## A bound is a contract, not a schedule A **freshness bound** is a single declared number attached to one feature: *a value served for this feature will be at most this old*. It is written where consumers can read it — beside the feature's definition — because it is the thing a consuming model depends on. It is not a description of the pipeline. Two platforms can meet the same one-hour bound with completely different machinery, and a consumer should not have to know which. The distinction matters because it decides who is accountable. If the contract is *the job runs hourly*, the pipeline team has met it the moment the job runs, whatever the served rows look like. If the contract is *no served value is older than an hour*, the same team owes you detection and a response when the promise breaks. ## The features of a comfort predictor, and their bounds | feature | declared bound | why that number | typical computation | |---|---|---|---| | current indoor temperature | 60 seconds | it is the room state the prediction is about | continuous computation over the sensor stream | | outdoor conditions | 1 hour | the upstream source itself publishes hourly | scheduled batch | | thirty-day occupancy profile | 24 hours | a month-long profile barely moves in a day | nightly batch | | setpoint delta for this request | not stored | it only exists once the request names a target | computed at request time | Notice the last row. A feature computed at request time has no stored age, but it is not exempt from freshness — it reads **inputs**, and those inputs carry bounds of their own. Moving computation to the request moves the age question one hop upstream; it does not delete it. ## Why per feature rather than one number - **Cost is asymmetric.** Continuous computation of a feature that changes once a day buys nothing and bills constantly. - **Meaning is asymmetric.** A thirty-day occupancy profile computed an hour later is the same profile; an indoor temperature an hour later is a different room. - **Blast radius is asymmetric.** When a daily feature is a day late, one refresh is skipped. When a sixty-second feature is an hour late, every request in that hour is scored on the wrong room state. - **One global number is wrong twice.** Set at the tightest requirement it makes cheap features expensive; set at the loosest it makes the tight ones useless. ## Cadence is not the guarantee The most common mistake is publishing the schedule as though it were the bound. Work out what a schedule actually delivers: 1. A source event lands at time `t`. 2. The next run starts at most one **interval** later. 3. The run takes its **runtime** to finish and publish. 4. The value then sits in the online tier until the following publish. So a job that starts every 5 minutes and takes 20 minutes publishes a value roughly 25 minutes after the event that produced it — five minutes of waiting plus twenty of running. Declaring a five-minute bound for it is simply false. Worse, the obvious fix is the wrong one: once runtime exceeds the interval, starting runs more often does not lower the age, because each run still needs its full runtime before anything is published. What lowers it is doing less work per run (incremental materialization over the new window instead of a full recompute) or moving to continuous computation. ## What declaring a bound obliges the platform to do Declaring the number is the cheap part. The contract has three obligations: - **Publish it** alongside the feature definition, so a model author picks features by what they guarantee rather than by what they happen to do this month. - **Measure it** on values that were actually served, not on the pipeline's own health, and alarm on a high percentile against the declared number. - **Act on breach** with a defined response — serve with the staleness recorded, fall back to a coarser feature that is inside its own bound, or refuse to score. A bound with no defined breach behaviour is a comment, not a contract. ## Choosing the number honestly When nobody has stated a requirement, measure one. Replay a window of history, scoring each row first with values as of the request and then with values artificially aged by five, fifteen, sixty minutes, and watch the **mean absolute error** of the prediction climb. The bound is the age at which that error crosses what the product is willing to absorb — which is usually looser than the engineer's instinct for the fast-moving features and tighter for the slow ones than anyone expected.
- A feature job starts every five minutes and each run takes twenty — what bound can you honestly declare?About twenty-five minutes: up to five minutes waiting for the next run, plus twenty minutes of running before anything is published. Shortening the interval does not help, because the runtime still has to elapse before a value appears. To go lower you have to make each run cheaper — recompute only the new window instead of the whole feature — or move the feature onto continuous computation.
- Who should own a feature's bound, the team producing it or the models consuming it?The producer owns the number it can actually hold, and publishes it. Consumers choose features by that published number and are responsible for their own behaviour when it is breached. When a consumer needs something tighter than the producer offers, that is a negotiation about computation mode and cost — not a number a model author may quietly assume.
saying these in an interview costs you the question
- States the job's refresh interval as if it were the freshness guarantee
- Applies one global refresh cadence to every feature in the model
- Assumes a shorter job interval always lowers the served value's age
- Forgets that a request-time computation still reads bounded inputs
- Sets every bound to the tightest number the platform can currently hit