Why does a dispatch feature keyed on driver-and-zone pairs usually fail the online materialization decision?
answer
- multiply the key cardinalities
- rows written against rows read
- ten billion against fifty million
- pairs observed, not pairs possible
- split the key or combine on request
basics
~20 sBecause the row count is the product of the two key cardinalities. Two million drivers and five thousand zones imply ten billion rows to write on every refresh, while a single request ever reads forty of them — the write side pays for a key space the read side never touches.
solid answer
~50 sMaterializing a feature means writing one row per value of its entity key, so a compound key multiplies. With 2,000,000 drivers and 5,000 zones the dense space is 10^10 rows; at roughly 200 bytes a row that is about 2 TB in a tier priced for hot data, and the job has to write all of it while one dispatch request reads forty pair rows. Three ways out, in order of preference: restrict the materialized rows to **observed** pairs — around 50 million if a driver works about 25 zones, more than a two-hundredfold cut; or key the row on the driver alone and carry a small map of that driver's recent zones inside it; or store the driver row and the zone row separately and combine them in the scoring path from parts already fetched.
code
sql · 8 linesSELECT COUNT(*) AS observed_pairs,
COUNT(DISTINCT driver_id) AS drivers,
COUNT(DISTINCT zone_id) AS zones
FROM (
SELECT DISTINCT driver_id, zone_id
FROM trips
WHERE trip_day >= CURRENT_DATE - 28
) AS pairsgo deeper
Recall that materializing a feature means one row per entity key, and that a key made of two things has as many rows as the two cardinalities multiplied together.
Explain the asymmetry: rows written per refresh scale with the key space, while rows read scale with traffic. Work the numbers for a compound key and name the observed-only and split-key alternatives.
Show that you measure the observed key space before committing, and that you can price each workaround — unserved cold pairs, a bounded map inside a row, or request-time combination — against the signal the feature actually carries.
The lead's angle is the platform rule: which key spaces the platform will materialize by default, who pays when a team wants a product-keyed feature anyway, and whether request-time combination is an approved pattern or a special case.
## The row count is a product A materialized feature costs one row per distinct value of its entity key. When the key is compound, the cost is the product of the cardinalities, and the product is what surprises people: - drivers: **2,000,000** - zones: **5,000** - dense pair space: **10,000,000,000 rows** At roughly 200 bytes per row that is about **2 TB** living in a tier that is priced, and often sized in memory, for hot data. The materialization job has to produce all of it on each refresh, which is a write volume completely unrelated to the traffic the feature serves. The read side, meanwhile, is tiny. One dispatch request holds one pickup zone and forty candidate drivers, so it reads **forty** pair rows. The vast majority of the ten billion rows are never read by anyone, because a driver works a handful of zones and never appears in the rest. ## Asymmetry is the diagnosis State it as a ratio in a design round: **rows written per refresh against rows read per unit of traffic.** A driver-keyed feature writes 2,000,000 rows and every one of them is reachable by some request. A pair-keyed feature writes 10^10 rows of which a tiny, strongly skewed subset is ever reachable. Same storage tier, same job, utterly different economics — and the difference came from one extra field in the key. | keying | rows materialized | read per request | verdict | |---|---|---|---| | driver | 2,000,000 | 40 | comfortable | | zone | 5,000 | 1 | trivial | | driver x zone, dense | 10,000,000,000 | 40 | not worth it | | driver x zone, observed only | ~50,000,000 | 40 | usually fine | ## Three ways out 1. **Materialize the observed key space only.** Restrict the rows to pairs actually seen in a recent window. If a driver works about 25 zones, that is 2,000,000 x 25 = **50,000,000** rows, roughly 10 GB at 200 bytes — a two-hundredfold cut. The price is that an unobserved pair has no row, so the serving path must have a defined behaviour for it, and the admissibility question comes back for the long tail. 2. **Fold one side into the other's row.** Key on the driver and store a small map of that driver's recent zones and their values inside the driver row. One key per driver, bounded row size, and the pair value arrives with the driver fetch rather than as an extra key. 3. **Combine on the request from stored parts.** Keep a driver row and a zone row, each cheap to materialize, and compute the interaction in the scoring path from the two values already fetched. This pays a little request-time work to avoid materializing a product, and it only works when the interaction is a function of the two sides rather than something that needs the pair's own history. ## What makes this a judgment call None of the three is free. Observed-only rows leave the cold tail unserved. Folding a map into a row bounds it arbitrarily and gets stale in its own way. Computing on the request moves cost into the latency budget and adds a second place where a feature's value is produced. The decision is a comparison — signal carried by the feature against the write volume, the storage bill and the risk each workaround introduces — and it is exactly the decision a design round is probing when it asks whether this feature belongs online at all. ## Check it before you build it Measure the observed key space before deciding anything. Count distinct pairs in a recent window, compare that with the product of the individual cardinalities, and look at the skew: if a small fraction of pairs carries almost all the traffic, the observed-only shape is a comfortable win. If the pairs are near-uniform, the cheap options evaporate and dropping or restructuring the feature is the honest answer.
- What is the alternative to a pair-keyed row when both sides genuinely matter to the score?Materialize the two sides separately — a driver row and a zone row, each with a small key space — and compute the interaction in the scoring path from the two values the request already fetched. It costs a little request-time work and works whenever the interaction is a function of the two sides rather than of the pair's own history.
- Why does the read side look cheap here while the write side does not?A request reads only the forty pair rows for its own zone and candidates, so read cost scales with traffic. The materialization job has to write every row in the key space on each refresh, so write cost scales with the product of the cardinalities — a number that is unrelated to how much traffic the feature actually serves.
saying these in an interview costs you the question
- Unread rows are free because bulk storage is cheap.
- Every feature in the model deserves its own online row.
- The pair count is the number of drivers, not the product.
- A sparse key space cannot be served online at all.
- Combining two stored rows at request time is always forbidden.