skip to content

The per-user typing-history feature is missing for a keystroke the next-word scorer must score now — what are your options, and how do you choose?

level: seniorimportance: should knowfreq 56%

answer

  1. absent versus merely late
  2. match the training convention exactly
  3. a variant trained without the group
  4. zero is a real count, not unknown
  5. record which groups arrived

basics

~20 s

Three options: impute with the exact convention the training data used, route to a variant trained without that feature group, or skip the scorer and take the next fallback rung. Never substitute an in-range value the model will read as real evidence.

solid answer

~50 s

First I would separate absent from late. A stale value from the online store is still a value — scoring on a user state that is a few keystrokes old is usually fine, and it is a different problem from the feature not being there at all. If the group is genuinely absent, three options are honest: impute using the same convention the training pipeline used, so the model sees `missing` the way it was taught to see it; route the request to a variant trained without that feature group, which stays calibrated but costs a second head to maintain; or skip the scorer for this keystroke and take the cached or frequency-list rung. The dangerous option is a plausible sentinel such as zero, because a zero history count is a legitimate value meaning "this user has never typed this word". Whatever I pick, the response records which feature groups arrived.

go deeper

for a junior

Know that a scorer can run with an input absent and still return a number — nothing throws, so the wrongness is silent rather than visible.

for a middle

Explain the handling options and why an imputed value must match the convention the training data used, or the model reads the filler as real evidence.

for a senior

Diagnose it in production: separate absent from late, name the population the missingness concentrates on, and keep a per-request record of which feature groups arrived.

for a principal

Weigh a maintained fallback variant against one model trained to tolerate dropped groups, and decide how much degraded-path quality the product will pay for.

## Absent, late, and unusual are three different events Before choosing a handling strategy, name what actually happened. **Absent** means the online store returned nothing for this key: the user is new, the locale is uncovered, or the writer job never produced the row. **Late** means the store returned a value computed some time ago — the typing-history counters do not include the last few keystrokes. **Unusual** means the value arrived and is fine but sits far outside the range the training data contained. Only the first is a missing-feature problem. A late value is real data describing a slightly older user state, and the model handles it the way it handles any input: as evidence. Treating a stale feature as missing throws away a usable signal and inflates the degraded path for no reason. ## The three honest options 1. **Impute with the training convention.** If the training pipeline encoded absence as an explicit missing indicator, or filled it with a stated statistic computed on the training window, do exactly that at serving time. The rule is that the serving path must reproduce the training path's treatment of absence, not invent a reasonable-looking one. 2. **Route to a variant trained without the group.** A second head trained on the remaining inputs answers a request that has lost one feature group without pretending the group exists. Its scores stay calibrated because it was fitted on exactly the input set it is being given. 3. **Take the next fallback rung.** If the missing group carries most of the personalisation, a scorer without it may be no better than the global frequency list — in which case scoring at all is wasted work and the cheaper rung is the correct answer. | option | quality while degraded | what it costs | when it fits | |---|---|---|---| | impute by training convention | close to normal, if absence was represented in training | nothing extra at serving time | absence is rare and was trained for | | variant without the group | calibrated, visibly lower | a second head to train, ship and monitor | the group goes missing often enough to matter | | next fallback rung | the cheap rung's quality | nothing, but no personalisation at all | the missing group carried most of the signal | ## Why a sentinel wins the argument silently Passing zero for an absent count is attractive because nothing errors: the scorer accepts the vector, produces a score, and the strip renders. That is precisely the failure. Zero is a legal value in the feature's own domain and means "this user has never typed this word", which is a strong piece of evidence rather than the absence of evidence. The model applies the weight it learned for that evidence, and the strip shows a confident, wrong suggestion. There is no exception in the logs, no error-rate movement, and no way to find the affected requests later unless the response recorded that the group was absent. The same trap appears with a mean or median filled in on the fly. If the training pipeline never used that convention, the serving path is handing the model an input distribution it was not fitted on, and the resulting score is not calibrated in any direction you can predict. ## Training the model to expect the gap The most durable version of option 2 is not a separate model but one model trained with the feature group masked on a fraction of examples. The training pipeline drops the group deliberately; the serving path then selects the behaviour by which groups arrived. This keeps a single artefact, refreshes the degraded path on every retrain, and prevents the second head from quietly rotting between incidents because nobody exercises it. ## Missingness is never random A feature group goes missing for a reason, and the reason concentrates the damage: - **New users** have no typing history at all, so the degraded path is exactly the population whose first impression you are forming. - **A locale or keyboard layout** the feature pipeline does not cover fails permanently rather than transiently. - **A failed upstream job** takes out a whole shard or region at once. Because the affected population is specific, an aggregate acceptance rate will not reveal it: a small fraction of traffic with badly degraded suggestions moves the blended number by less than the noise band. Measure the imputed and variant paths against their own population, not against the whole. ## What the response must carry Record which feature groups arrived, which were absent, and which handling fired. That single record is what later lets you answer three otherwise unanswerable questions: how much traffic scored without personalisation, whether that traffic behaves worse, and whether an apparent model regression is really a feature-pipeline outage wearing a model's clothes.

  • How do you keep a no-history variant from becoming a second model nobody maintains?
    Train it from the same examples with the feature group masked on a fraction of them, so one pipeline produces both behaviours and both refresh together. The serving path then picks by which groups arrived. The degraded path gets exercised by every retrain instead of rotting between incidents, and there is one artefact to ship rather than two.
  • Missingness in production is rarely random — why does that matter here?
    A group goes missing for a reason: a new user with no history, an uncovered locale, or a failed upstream job. Each concentrates the degraded path on a specific population, so the quality loss is not spread evenly and a blended acceptance rate moves less than the noise band. Measure the degraded handling against the affected population, not the whole.

saying these in an interview costs you the question

  • Fills a missing count with zero and calls it safe
  • Treats a stale feature value as an absent one
  • Imputes with a statistic the training pipeline never used
  • Assumes a missing feature makes the scorer raise an error
  • Believes missingness is spread evenly across users