skip to content

In a dispatch model, a feature computed from trip history has no online row at request time — what does the scoring path receive?

level: seniorimportance: should knowfreq 47%

answer

  1. the read returns nothing, not an error
  2. a filled value is scored as real
  3. no alarm fires anywhere
  4. is the value knowable before the decision?
  5. measure online-row coverage per group

basics

~20 s

It receives nothing, and nothing is not an error: the read returns empty, the serving code fills the slot, and the model scores as if the filled value were real. The ranking degrades on every request with no failure anywhere to alert on.

solid answer

~50 s

A missing online row is not an exception — it is an empty result. The scoring path has to produce a ranking, so it fills the slot with something and scores. The model is perfectly happy to consume it, the request succeeds, latency is normal, error rate is flat, and the only symptom is that dispatch quietly gets worse for every request touching that feature. That is why availability at request time is an **admissibility rule** rather than an afterthought: a feature belongs in a served model only if its value is knowable before the prediction is made, hangs off an entity key the request actually holds, and has a materialization path into the online tier. A feature that trains beautifully and has no path to a row is a feature the model cannot have.

go deeper

for a junior

Recall that an online read for a missing key returns nothing rather than raising an error, so the request succeeds and the model scores anyway on whatever was put in the slot.

for a middle

Explain why this failure is silent: no exception, flat error rate, normal latency, only worse output. Then state the admissibility conditions a feature must meet before it can be part of a served model.

for a senior

Show that you check availability before promotion, not after an incident. Name the three options for an unservable feature, and defend dropping one that improved the offline metric because the serving path cannot supply it.

for a principal

The lead's call is where the veto sits: whether the platform refuses to promote a model whose features lack coverage, and who absorbs the offline metric a modeller loses when that veto fires.

## Nothing is not an error A lookup against the online tier for a key that has no row returns an **empty result**, which is a perfectly ordinary outcome for a key-value read. It is not a timeout and not an exception, so nothing in the request path treats it as a failure. The scoring path still owes the rider a ranked list, so it puts *something* in the slot and calls the model. The model then does what models do: it consumes a number without knowing where the number came from. The request returns 200, the p99 is unchanged, the error rate is flat, and the dispatch quality for every affected request is lower than the offline evaluation promised. This is the archetype of the silent ML failure — a degradation with no failure signal attached — and it is why an interviewer asks about it. ## The admissibility rule Before a feature enters a served model, three things must hold. They are cheap to check and expensive to skip. 1. **Knowable before the prediction.** The value must exist at the moment of the dispatch decision. Anything that only exists afterwards — the final fare, the trip duration, the rider's rating, whether the trip was cancelled — has no request-time value at all. Such features can be reconstructed for training and never served. 2. **Keyed by something the request holds.** The feature must hang off an entity the request carries: the candidate driver's identity, the pickup zone, the rider. A feature keyed by something the dispatch path cannot name is unreadable no matter how it is stored. 3. **A materialization path exists.** There has to be a route from the columnar history to a row in the online tier, and the feature's registry entry has to say it is served online. "We compute it in the training pipeline" is not that route. ## Coverage is a served-path measurement Admissibility is a design-time check; **coverage** is the running one. For each feature group, measure on live traffic the fraction of candidate entity keys whose online read returns a row, and read it per feature group rather than in aggregate. - Measure it **before** promoting a model that depends on a new feature, on shadow traffic if you have it. - Expect coverage to be below 100% for legitimate reasons — a driver who signed up an hour ago has no 28-day aggregate — and decide deliberately what the model does for them. - Alert on coverage **falling**, because that is the shape a broken materialization path takes when everything else looks healthy. ## The three options when a feature has no row | option | what it costs | when it is right | |---|---|---| | materialize it | storage and a job; only possible if the value is knowable and keyed | the feature carries real signal and the key is cheap | | replace it with a servable proxy | some offline metric, and a new definition to maintain | the original depends on data not present at request time | | drop it from the model | the offline gain it showed | the signal is small or the key space makes a row infeasible | The honest version of the third row deserves emphasis: **an offline gain is not by itself a reason to serve a feature.** A model is a contract with the serving path, and a feature the serving path cannot supply is not part of that contract, however good it looked in the evaluation. ## In a design round Say the failure mode first — empty read, filled slot, no error, worse ranking — then the rule that prevents it, then the measurement that catches the cases the rule missed. That sequence shows you have operated the thing rather than read about it.

  • What measurement on the served path tells you a feature's online rows are actually there?
    Online-row coverage: on live or shadow traffic, the fraction of candidate entity keys whose read returns a row, tracked per feature group rather than in aggregate. Check it before promoting any model that depends on a new feature, and alert when it falls, since a broken materialization path shows up there and nowhere else.
  • Which features can never have an online row at all, however much effort you spend?
    Those whose value only exists after the prediction is made — the completed fare, the trip duration, the cancellation flag, the rider's rating. They are knowable for a past trip and so reconstructable for training, but at the moment of dispatch they do not yet exist for the decision being made.

saying these in an interview costs you the question

  • A missing online row makes the request fail loudly.
  • The model refuses to score when an input is absent.
  • A strong offline gain alone qualifies a feature for serving.
  • Anything present in the training snapshot can be served.
  • A drift alarm will catch the filled-in value immediately.