A comfort feature's definition is corrected — how do you rebuild three years of history without disturbing what is live?
answer
- a definition change is a new feature
- new name, both series side by side
- chunk the time axis, record completion
- writes keyed by entity and timestamp
- verify on the overlap before cutover
basics
~20 sWrite the corrected values under a new definition name, rebuild history into it in re-runnable chunks, then move readers over once the new series is complete. Overwriting the old series in place destroys the record training was built on.
solid answer
~50 sTreat a definition change as a **new feature**, not an edit. Give it a new name or version so both series exist side by side, and leave every current reader on the old one. Run the rebuild as chunks — a month of history at a time — where each chunk's writes are keyed by entity and event timestamp, so re-running a chunk replaces exactly the rows it wrote before and nothing else; that makes a failed chunk restartable without double counting. Record each chunk's completion so a re-run skips finished work. Before cutting over, compare the two series on a window where both exist and confirm the differences are only the ones the correction intends. Then move training and serving to the new name in that order, and retire the old series once the models trained on it are gone.
code
pseudocode · 12 linesnewDefinition = "occupancy_profile_30d_v2"
for each chunk in months(from = today - 3 years, to = today):
if chunkState(newDefinition, chunk) == "complete":
continue
rows = compute(newDefinition, sourceEvents(chunk))
writeOffline(newDefinition, rows, replaceKey = (entityId, eventTimestamp))
markComplete(newDefinition, chunk)
// readers stay on occupancy_profile_30d_v1 until every chunk is completego deeper
Recall that correcting a feature's formula means producing a second series under a new name rather than editing the existing history, and that readers move only when the new one is complete.
Explain what makes a chunk re-runnable: writes keyed by entity and timestamp so a repeat replaces its own rows, plus recorded chunk state so a restart skips finished work.
Own the cutover order and the verification: overlap comparison first, training before serving, old series retired only when no production model depends on it, and the rebuild rate-limited against the live path.
Consider the catalogue-wide policy — how long two versions of a definition are allowed to coexist, who pays for the storage, and what forces the old one to be retired rather than lingering for years.
## A definition change is a new feature When the formula behind a feature changes — a corrected occupancy window, a fixed unit conversion, a bug in how a gap was handled — it is tempting to fix the code and rerun. That overwrites history in place, and it breaks three things at once: - the **models already in production** were trained on the old values, and the series that produced them no longer exists, so their behaviour can no longer be explained or reproduced; - any **evaluation** run against that history silently changes its result, with no version to point at; - if the rebuild fails part way, the series is left half old and half new, with nothing recording where the seam is. So the rule is to write the corrected values somewhere new. The old series stays exactly as it was until nothing reads it. ## Chunked and re-runnable Rebuilding three years for millions of homes is not one job. It is a sequence, and it must survive being interrupted: 1. **Pick a chunk boundary** on the time axis — a month of history per chunk is a common size, because it is large enough to amortise startup and small enough to redo. 2. **Make each chunk's writes replace-by-key.** Every row is keyed by entity and event timestamp, so writing the same chunk twice produces the same rows rather than duplicating them. 3. **Record chunk completion** in a small state table, so a re-run skips what already finished and picks up where the failure was. 4. **Rate-limit the rebuild** against the live path. A rebuild is a bulk workload, and pointed at the same resources the serving path uses it becomes a self-inflicted latency incident. Those four properties together are what "re-runnable" means concretely: interruption costs one chunk, not three years, and nobody has to reason about which half of history is which. ## Verifying before the cutover A finished job is not a correct rebuild. Verify on the overlap — the window where both the old and the new series exist: - **count rows per entity per period** in both, so a chunk that silently produced nothing is visible; - **compare values** and confirm that the differences are the ones the correction intends, in the direction it intends, and absent where it should not apply; - **recompute a handful of entities by hand** for a few dates and check both against the arithmetic, not against each other. ## The cutover, in order | step | what moves | what must be true first | |---|---|---| | 1 | nothing — rebuild runs | the new definition has its own name; readers untouched | | 2 | the training pipeline reads the new name | every chunk complete and verified | | 3 | a model is retrained on the new series | training data assembled from the new definition only | | 4 | the serving path materializes the new definition | the retrained model is the one being served | | 5 | the old series is retired | no model in production was trained on it | Step 4 before step 3 is the ordering mistake worth naming: serving the corrected values to a model that was trained on the old ones feeds it inputs whose meaning has changed, which is a worse state than the bug the correction was fixing. The definition and the model that consumes it move as a pair. ## Freshness during and after a backfill A backfill is about history, but it touches the live promise twice. While it runs, it competes for the same capacity the incremental materialization job needs, so the feature's ongoing age bound can be breached by the rebuild that was supposed to improve it. And after the cutover, the new definition's incremental job has to be running before readers move, or the new series is complete up to the cutover date and then stops. ## What goes wrong - **Overwriting in place** — history is replaced, old models become unexplainable, and a failed run leaves an undocumented seam. - **Appending under the same name** — two values now exist for the same entity and timestamp, and every reader silently picks one. - **One long job** — a failure at hour nine means starting again, and nothing records how far it got. - **Cutting over early** — readers see a series that is complete for recent months and empty for older ones, which looks like missing data rather than an unfinished rebuild. - **Declaring success on the job's exit status** — a run that completed while producing nothing for one region exits cleanly.
- How do you verify the rebuild before cutting readers over?Compare the two series on the window where both exist. Count rows per entity per period to catch a chunk that produced nothing, compare values to confirm the differences are the intended correction and appear only where it applies, and recompute a few entities by hand for a few dates. A clean exit status from the job is not verification.
- Which moves first after the rebuild — the training pipeline or the serving path?Training. Assemble training data from the new definition, retrain, and only then materialize the new definition into the online tier for the retrained model. Serving corrected values to a model trained on the old ones changes the meaning of its inputs underneath it, which is usually worse than the defect being corrected.
saying these in an interview costs you the question
- Overwrites the existing feature series in place with corrected values
- Runs the whole rebuild as one long job with no chunk boundaries
- Appends corrected rows beside the old ones under the same name
- Cuts readers over before the entire history range is rebuilt
- Materializes the new definition before retraining the model on it
- Declares the rebuild done on the job's exit status without comparing