skip to content

In a dispatch feature platform, what does it mean that the online key-value tier is a derived view of the offline history?

level: middleimportance: must knowfreq 58%

answer

  1. one provenance, not two sources
  2. the online row is produced, not received
  3. the history is what you rebuild from
  4. drop the tier and run the job again
  5. registry names key, type and tiers

basics

~20 s

It means every value in the online tier can be reproduced from the columnar history by a materialization job. The online row is an output of the platform, not an input to it, so losing the tier is an availability incident rather than data loss.

solid answer

~50 s

The columnar history is where feature values land and stay; the online row is produced from it by a **materialization job** that reads the history, reduces it to the current value per entity key, and writes one row per key into the low-latency tier. The practical test is reproducibility: if you dropped the online tier entirely, could you rebuild every row from the history and a feature definition alone? If yes, the tier is a derived view. If some values reached it by a path that wrote nowhere else, it has quietly become a second source of truth, and a value can now exist at dispatch that no training scan will ever see. A feature registry keeps that honest by naming each feature once — entity key, value type, source, and which tiers it is served from.

code

pseudocode · 17 lines
pseudocode
for each group in registry.groups_served_online():
    rows = offline_store.read_current(group.table)

    latest = {}
    for each row in rows:
        key = row[group.entity_key]
        if key not in latest or row.event_time > latest[key].event_time:
            latest[key] = row

    writes = []
    for each key, row in latest:
        writes.append(
            key   = group.name + ":" + key,
            value = project(row, group.features)
        )

    online_store.multi_put(writes)

go deeper

for a junior

Recall the direction: values land in the columnar history first, and the online row is produced from them. Nothing should be writing a feature value straight into the request-time tier.

for a middle

Explain the reproducibility test — could every online row be rebuilt from the history and a definition alone? — and describe the job that does it: read, reduce to one value per entity key, write one row per key.

for a senior

Demonstrate that you check the property rather than assume it. Name the accretions that break it, and say what the online tier's loss should cost: availability, not data. Correct values in the history, never in the row.

for a principal

The tradeoff a lead owns is how strictly to enforce single-provenance across many teams, given that a direct online write is always the faster fix in an incident and always the thing that makes the platform unreproducible a quarter later.

## Derived, not merely copied Saying the online tier is a **derived view** is a statement about provenance, not about mechanics. It asserts one property: every value a dispatch request can read is reproducible from the offline history plus the feature's definition. Nothing else needs to be true — not the schedule, not the engine, not the direction the bytes happened to travel. The usual shape is a **materialization job**: read the history for a feature group, reduce it to the current value per entity key, write one row per key into the low-latency tier. Architectures genuinely differ here — some platforms compute a value once from an event stream and land it in both tiers, others recompute on read from stored parts — and the derived-view property survives all of them as long as the history can reproduce what the request saw. What breaks it is a value that reaches the online tier by a path that writes nowhere else. ## What the property buys you 1. **Rebuildability.** Lose the online tier and you have lost availability, not data. The job runs again and the rows come back. This is the difference between a bad afternoon and an unrecoverable one. 2. **One place to correct.** A wrong feature value is fixed in the history and then re-materialized. Patching the online row instead leaves the training scan reading the wrong number forever. 3. **A single explanation for a decision.** Asked months later why a particular driver was ranked where they were, you can reconstruct the inputs, because the values the request saw have an offline ancestor rather than existing only in a hot store that has since expired them. 4. **A bounded blast radius for a definition change.** Redefining a feature changes what the job writes; it does not require hunting down writers scattered across services. ## The registry is the contract A feature registry is the list of definitions the platform runs on, not a place where values live. For each feature it names: - the **entity key** the feature hangs off — driver identity, zone identity, rider identity — which decides what a request must hold to read it; - the **value type and units**, so a count is not silently compared with a rate; - the **source** in the history the value is computed from; - **which tiers** the feature is served from, because plenty of features are training-only and never need a row. With that entry in place, the materialization job is generic: it iterates the features marked for online serving and writes them, rather than carrying a hand-written path per feature. ## How the property gets lost The failure is rarely a decision; it is an accretion. Watch for: - a service that computes something convenient and writes it straight into the online tier because it was already holding the value; - a backfill run against the online tier alone, to "fix production quickly", with the history left as it was; - a feature whose online value is produced by request-time code that stores no trace of what it produced; - an online tier with a retention policy shorter than the platform's memory, holding the only copy of anything. Each one creates a value at dispatch with no offline ancestor. The tell is operational: a rebuild of the online tier from the history changes what production reads. ## The dispatch case A driver-level aggregate — completed trips in 28 days, acceptance rate, minutes online today — is computed in the history where the events are, keyed by driver, and materialized to one driver row. The dispatch path never aggregates months of events itself; it reads a row. That is the whole point of the arrangement, and it is also why the answer to "where does this number come from?" is always the same: the history, through the job, into the row.

  • What does a feature registry have to name so one generic materialization job can serve every feature?
    The entity key the feature hangs off, its value type and units, the source in the history it is computed from, and whether it is served online at all. With those four, the job iterates definitions instead of carrying hand-written code per feature, and a request knows which key to present.
  • A driver aggregate over months of trips has to be readable at dispatch. What must exist for that?
    The aggregate is computed in the columnar history, where the events are, and keyed by driver; the materialization job then writes it as a field on that driver's online row. The dispatch path never aggregates months of events itself — it reads one row and moves on.

saying these in an interview costs you the question

  • The dispatch path writes back what it computed for next time.
  • Losing the online tier loses those feature values for good.
  • The registry stores the values, not just the definitions.
  • Each tier can be filled by its own independent writer.
  • Rebuilding the online tier means replaying upstream events.