skip to content

questions

5

In a ride-hailing dispatch platform, why is the same driver feature stored twice — in a columnar history and in a key-value row?

level: juniorimportance: must knowfreq 66%

answer

  1. two readers, two shapes
  2. scan a column vs fetch a row
  3. months of rows vs forty keys
  4. a latency bound on one side only
  5. one definition, one derived copy

basics

~20 s

Training and dispatch read the same feature with opposite access patterns: a training job scans months of one column across millions of rows, while dispatch fetches every feature for one driver in milliseconds. No single store serves both shapes well.

solid answer

~50 s

The feature has one definition but two readers. The nightly training job scans a handful of columns across every driver and every day in an 18-month window — billions of values, no one waiting, throughput is all that matters. The dispatch path shortlists roughly forty nearby drivers per request and needs every feature for each of them inside a few milliseconds. A `columnar` history stores each column contiguously so a long scan is cheap, but assembling one driver's row means touching every column segment. A low-latency key-value tier returns a whole row under a key in about a millisecond, but scanning eighteen months of it is billions of point reads. So you keep two physical copies of one logical feature, sized for the two reads, and require that the online copy be derivable from the history.

go deeper

for a junior

Recall that the same feature lives in two stores because two readers want it in two shapes, and be able to say what each read looks like: a long scan over columns, versus one row fetched by key.

for a middle

Explain the mechanics: contiguous columns make a scan cheap and a single row expensive, while a keyed row makes the lookup cheap and the sweep expensive. Tie each storage choice back to the read that motivated it.

for a senior

Show that you price the split. Two copies, a job between them, and a standing agreement question are real costs, and you should be able to say which features earn them and what the history still owes you after the online tier exists.

for a principal

The angle a lead owns is the platform bargain: every team gets two copies and a materialization path by default, and the organisation accepts the duplicated storage and the reconciliation burden in exchange for nobody hand-rolling a request-time aggregation.

## Two readers with opposite shapes A dispatch platform computes a feature once — say **trips completed by this driver in the last 28 days** — and then two very different readers want it. The **training read** happens after the day closes. The job assembles months of labelled dispatch decisions and, for each of them, the feature values as they stood. It touches a handful of columns across every driver and every day in an 18-month window: billions of values. Nobody is waiting on it; it is allowed to run for hours, and the only thing that matters is bytes scanned per second. The **dispatch read** happens while a rider watches a spinner. The platform shortlists roughly forty nearby drivers, scores each of them, and dispatches one. It needs **all** of a driver's features, for **forty** drivers, inside a slice of a request budget it shares with candidate retrieval, the model's forward pass and the dispatch decision itself. It happens thousands of times a second. | | the training read | the dispatch read | |---|---|---| | issued by | the nightly training job | every rider request | | entities touched | millions, across months | ~40 candidate drivers | | columns touched | one to a few | all of them | | latency bound | none — hours are fine | a few milliseconds | | rate | once a night | thousands per second | | values needed | every dated value in the window | one current value per key | ## Why one store cannot serve both The two layouts are not faster and slower versions of each other; they are organised along different axes. - A **columnar offline store** keeps each column contiguous and compressed, so reading one column over a year skips every other column entirely. The cost is that a single entity's row is scattered across every column's segments, and reading it means touching all of them plus the file's index structures — far outside a request budget. - A **low-latency key-value tier** keeps a whole row together under its key and answers a point lookup in about a millisecond. The cost is that there is no cheap way to sweep it: reading 18 months for two million drivers turns into billions of individual lookups against a tier sized for the working set, not the archive. - The **economics** differ too. The history is large, cold and grows forever, and is priced like bulk storage. The online tier is small, hot, often memory-resident and priced per byte per hour, which is exactly why you do not put 18 months in it. - The **retention** differs. The history must keep every dated value because a future training run will ask for a window that has not been chosen yet. The online tier needs only what a request can use. ## What the two copies must still share Two copies is a storage decision, not a licence to let the two sides drift apart: 1. **One definition.** A feature registry describes the feature once — its name, its entity key, its value type and units, and which tiers it is served from — and both the history and the online row are produced against that single entry. 2. **The same entity key.** If the history keys a feature by driver identity, the online row is keyed by driver identity too; a request that holds a driver id can then read it. 3. **A derived online copy.** The history is what the platform can rebuild from. Losing the online tier should be an availability incident, not data loss. ## What the split costs you Be honest about it in a design round: you are paying for two physical copies of every served feature, a job that keeps one in step with the other, and a new class of question — is what dispatch just read the same thing the training job scanned? That question is the reason the rest of the feature-platform design exists, and it is the price of the two reads being served well instead of both being served badly. ## Saying it in a design round - Name the two reads first, with numbers: months-by-millions scanned once a night, versus forty keys fetched thousands of times a second. - Derive the storage shapes from those reads rather than announcing them. - State the invariant that holds the two together: one definition, one entity key, and an online copy the history can reproduce.

  • Why does the training read tolerate hours while the dispatch read cannot tolerate tens of milliseconds?
    The training job is throughput-bound and runs once, unattended, after the day closes; its wall-clock time affects nobody's request. The dispatch read sits inside a rider-facing budget it shares with candidate retrieval, the model's forward pass and the dispatch decision, and it is paid thousands of times a second rather than once.
  • Once the online tier is populated, does the platform still need the columnar history?
    Yes, for two independent reasons. Every future training run scans it, over windows nobody has chosen yet, and it holds dated values the online tier does not keep. It is also what the online rows are rebuilt from, so discarding it turns a recoverable outage into permanent loss.

A restaurant keeps three years of supplier invoices in a filing cabinet and today's prices on a card taped by the till. Same numbers, two shapes: one is for the annual review, the other has to be readable while a customer waits.

saying these in an interview costs you the question

  • One store can serve both reads if you add enough replicas.
  • The online key-value tier is the platform's source of truth.
  • Training can just read the online tier key by key.
  • Each tier holds its own set of features, chosen per tier.
  • A columnar scan is fast enough to answer one request.
open as a page

In a dispatch feature platform, what does it mean that the online key-value tier is a derived view of the offline history?

level: middleimportance: must knowfreq 58%

basics

~20 s

It means every value in the online tier can be reproduced from the columnar history by a materialization job. The online row is an output of the platform, not an input to it, so losing the tier is an availability incident rather than data loss.

open as a page

A dispatch request scores 40 nearby drivers on 30 online features each, issuing one read per driver per feature — why does its p99 collapse?

level: seniorimportance: should knowfreq 51%

basics

~20 s

That shape issues 1,200 reads per request, and the request cannot finish until the slowest one returns. With a 1% per-read tail, almost every request contains at least one slow read, so request latency tracks the read tier's extreme tail rather than its median.

open as a page

In a dispatch model, a feature computed from trip history has no online row at request time — what does the scoring path receive?

level: seniorimportance: should knowfreq 47%

basics

~20 s

It receives nothing, and nothing is not an error: the read returns empty, the serving code fills the slot, and the model scores as if the filled value were real. The ranking degrades on every request with no failure anywhere to alert on.

open as a page

Why does a dispatch feature keyed on driver-and-zone pairs usually fail the online materialization decision?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Because the row count is the product of the two key cardinalities. Two million drivers and five thousand zones imply ten billion rows to write on every refresh, while a single request ever reads forty of them — the write side pays for a key space the read side never touches.

open as a page