In a comfort platform, how does a feature's declared age bound decide whether it is computed by batch, by stream, or at request time?
answer
- the bound is a delay budget
- four terms: source, compute, publish, gap
- cheapest mode that fits, not freshest
- interval plus runtime caps a batch job
- request-time moves age one hop upstream
basics
~20 sThe bound is a ceiling on end-to-end delay, so you pick the cheapest mode that fits under it: batch for a daily bound, continuous computation for a sixty-second one, request-time for anything that depends on the request.
solid answer
~50 sBreak the bound into the delays that consume it: waiting for source data, computing, publishing, and then the gap until the next publish. A scheduled batch job spends minutes to hours on those and fits a bound measured in hours or days — the thirty-day occupancy profile rebuilt nightly. Continuous computation over an event stream spends seconds and is what a sixty-second indoor-temperature bound requires, at the price of a job that is always running and whose failure is an immediate breach rather than a late file. Request-time computation has effectively zero age but spends the model's own latency budget on every call, so it is reserved for values that cannot be precomputed at all — anything keyed by something the request itself supplies. You choose the cheapest mode that fits, not the freshest one available.
go deeper
Know that there are three places a feature value can be produced — on a schedule, continuously as data arrives, or during the request — and that the allowed age is what decides between them.
Explain the four delays that consume a bound and show why a schedule whose interval plus runtime exceeds the bound cannot meet it however it is tuned.
Argue for the cheapest mode that fits and defend the operational consequence: what breaks at 02:00 under each mode, and how long you have to respond before the bound is violated.
Take the view across the catalogue: every continuously computed feature is a standing on-call commitment, and the mode mix is a staffing decision as much as an architectural one.
## The bound as a delay budget A declared age bound is a ceiling, and a served value's age is a sum. Before choosing a computation mode, write down what consumes the budget: 1. **Source delay** — how long after the real-world event the data is available to compute on. 2. **Compute delay** — how long the computation takes once it starts. 3. **Publish delay** — writing the result into the tier the request reads. 4. **Publish gap** — how long the value then sits before the next one replaces it. For a sixty-second bound, a worked allocation might be 5 seconds of source delay, 10 seconds of computation, 2 seconds to publish and up to 20 seconds of gap: 37 seconds, comfortably under 60 with headroom for a slow moment. That sum is the whole decision. A mode is viable when its four terms fit under the bound with margin, and not otherwise. ## The three modes | mode | age it can hold | what it costs | where it fails | |---|---|---|---| | scheduled batch | hours to a day | cheapest per row; runs on a fixed schedule; a failed run is a late file | any bound shorter than one run's interval plus its runtime | | continuous computation over a stream | seconds to a couple of minutes | a job that is always up, always billing, and whose stall is an immediate breach | features needing a long historical window recomputed from scratch each time | | computed at request time | essentially zero stored age | spends the request's own latency budget, every call, per request | fan-out over many entities, or heavy computation inside a tight budget | ## Reading the table in the comfort platform - The **thirty-day occupancy profile** has a one-day bound and needs a month of history per home. A nightly batch pass over millions of homes is the correct and cheapest answer; putting it on a stream would mean maintaining a rolling thirty-day window in memory for every home to buy freshness the feature does not need. - The **current indoor temperature** has a sixty-second bound. No schedule can hold it — as soon as one run's interval plus its runtime exceeds a minute the bound is already broken — so it is computed continuously as readings arrive. - The **setpoint delta** exists only once the request names a target temperature. There is nothing to precompute, because the value is a function of a request parameter; it is computed on the call. ## The cheapest mode that fits, not the freshest The instinctive failure is to push everything onto continuous computation because "fresher is better". Three reasons it is not: - **Cost.** Continuous computation over millions of homes bills constantly, whether or not the feature changed. - **Fragility.** A batch job that fails at 02:00 is a late file you can rerun at 05:00 and still be inside a one-day bound. A continuous job that stalls is in breach within minutes, and it is in breach for every home at once. - **Backfill is still required.** Freshness says nothing about history. A feature computed continuously still needs its past rebuilt when its definition changes, which means the batch path exists anyway. ## When request-time is the only option, and when it is a trap It is the only option when the value depends on something that does not exist until the request arrives — a target the user just set, a reading attached to the call, a comparison against a parameter. Precomputing would mean enumerating every possible value of that parameter for every home. It becomes a trap in three situations: - the computation is heavy and the model's end-to-end budget is tight, so the feature's cost lands on the p99 of every call; - the feature must be computed for many entities per request, turning one computation into hundreds; - the inputs it reads are themselves stored values with their own bounds, so the "zero-age" feature is quietly built on a value that may be hours old — moving computation to the request moves the freshness question upstream rather than resolving it. ## One model, several modes The answer to a design round is almost never a single mode. A comfort predictor reads a nightly profile, an hourly outdoor snapshot, a continuously computed room state and a request-time delta in the same call, because the bounds differ per feature and each bound was priced separately. Choosing one mode for the whole feature set is the shape of the mistake, in both directions.
- What does moving a feature from a batch job to continuous computation cost beyond the compute bill?It changes the shape of failure. A late batch file is recoverable inside a loose bound by rerunning; a stalled continuous job is in breach within minutes, for every entity at once, and demands paging and a defined degraded path. You also still need the batch path, because rebuilding history after a definition change is a bulk job no stream performs.
- A feature has a one-hour bound but its upstream source only publishes hourly — what can you promise?Not one hour. The source delay alone consumes the whole budget, leaving nothing for computation, publishing or the gap before the next value. The honest bound is the source's own cadence plus your pipeline's delay — roughly an hour plus your processing time — and the only way to tighten it is to change the source, not the mode.
saying these in an interview costs you the question
- Puts every feature on continuous computation because fresher sounds better
- Assumes request-time computation can meet any bound for free
- Accepts a nightly batch job for a minute-level bound
- Sizes the computation mode by data volume alone, ignoring the bound
- Ignores the upstream source's own delay when promising a bound