A scoring request arrives with a null device-age field — what value do you fill and where does it come from?
answer
- a fitted parameter, not a guess
- shipped inside the model artifact
- never recompute from today's traffic
- same request, same score, always
- watch the imputation rate per feature
basics
~20 sThe fill comes from a statistic computed once on the training data and shipped with the model as a fixed parameter. Never recompute it from live traffic: the same request would then score differently depending on its neighbours.
solid answer
~50 sThe fill value is a fitted parameter of the model, not something the serving code invents. You compute it once on the training data — say the median device age — freeze it into the model artifact alongside the weights, and apply that exact number to any request whose device age is null, plus the matching was-missing flag. What you must not do is recompute it from today's traffic or from the current request batch: that makes a single request's score depend on its neighbours, and it lets the fill drift from the level the model trained around, silently. A null also means two different things — genuinely absent, or a failed upstream lookup — and both quietly become the median. So I monitor the imputation rate per feature and alert when it jumps, because a broken upstream shows up as a rising fill rate long before a metric drop.
code
python · 17 linesfrom statistics import median
# --- training time: fit the fill once, then freeze it -------------
train_device_age = [12, 30, 7, 45, None, 22, None]
observed = [v for v in train_device_age if v is not None]
fill_value = median(observed) # 22 -> stored in the model artifact
# --- shared by training and serving: one row in, features out ----
def featurize(device_age):
return {
"device_age": fill_value if device_age is None else device_age,
"device_age_was_missing": 1 if device_age is None else 0,
}
print(fill_value) # 22
print(featurize(None)) # {'device_age': 22, 'device_age_was_missing': 1}
print(featurize(3)) # {'device_age': 3, 'device_age_was_missing': 0}go deeper
Remember that the number used to fill a gap is decided during training and reused unchanged afterwards. Do not compute it from whatever data happens to be in front of you at prediction time.
Explain why the fill belongs in the model artifact next to the weights, and what goes wrong if the serving code recomputes it: batch-dependent predictions and an input distribution that slides away from what the model learned.
Demonstrate operational judgement. Guarantee an identical transformation on both paths with a pinned test, know that single-row serving rules out population-based strategies unless reference data ships too, and monitor the imputation rate because a fill turns an upstream failure into a plausible number.
Decide the policy: which features may be silently imputed at all, which must fail the request loudly instead of scoring on a fabricated input, and what imputation-rate threshold is allowed to page someone or hold a deployment.
## The fill value is a model parameter The key idea is that an imputation rule is *fitted*, exactly like a coefficient. "Fill device age with 14 months" is not a rule of thumb the serving code makes up; it is a number learned from the training data and stored with the model. The same is true of the more elaborate rules: a per-segment median is a small table of learned numbers, and a model-based imputer is a fitted model in its own right. Whatever the form, the artifact you deploy must carry it. That reframing answers most of the question. Where does the value come from? From the training data, computed once, at training time, on the training split alone. What applies it? The same feature-building code that ran during training, executed on one row. ## Why it must not be recomputed at serving There is a tempting shortcut: compute the median of the incoming batch, or of the last hour of traffic, and fill with that. It is wrong in three separate ways. **It makes predictions non-deterministic in the wrong variable.** If the fill comes from the current batch, an identical request scores differently depending on which other requests were batched with it. The same customer, retried a minute later, gets a different answer. That is very hard to debug and impossible to explain to a caller. **It breaks the correspondence with training.** The model learned a relationship that treats the filled group as sitting at a particular level. If live traffic skews younger and the recomputed median moves to 6 months, the filled rows land in a region the model associates with something else entirely. The weights are unchanged; the input distribution has moved under them. **It hides drift instead of surfacing it.** A frozen fill plus a monitored imputation rate makes a shift in the population visible. A rolling fill silently tracks the shift and reports nothing. For a single-row request there is also a plain mechanical problem: the population statistics are not available. Serving sees one row, not a column. Any strategy that needs neighbours — a k-nearest-neighbour fill, for instance — requires shipping reference rows with the model and paying for a lookup on every request. That is a real design decision with latency and memory consequences, not a detail. ## Training and serving must run the same code path The most common production defect here is not a wrong statistic but a *divergent* one: the training notebook fills with the median and adds an indicator; the serving service fills with zero because that is what its default deserialisation does, and never builds the indicator. Nothing errors. Predictions are simply wrong for the subset of rows that matter most. The defences are structural: build features with the same code in both paths, ship the fitted fill values inside the model artifact rather than as constants copied into the service, and add a test that scores a fixed row with a null field and asserts the exact prediction the training-time pipeline produced for it. ## Two very different nulls A null device age can mean: 1. **Genuinely absent** — a new device, an unregistered one, a customer who never provided it. This is the case the training data represents, and filling is the right response. 2. **Upstream failure** — the device registry timed out, the join key was wrong, the schema changed. The value exists; you just could not fetch it. The imputation layer cannot tell these apart, and it will happily convert both into the median. That is what makes a broken upstream so dangerous: it produces no error, no exception, no failed request. Predictions degrade gradually toward whatever the model outputs for the average device, and the first signal is often a business metric weeks later. The operational answer is to instrument the imputation itself. Emit, per feature, the fraction of scored requests that had to be filled. Compare it against the rate seen in training. Alert on a step change. A feature that was filled 4% of the time in training and is being filled 60% of the time this morning is a broken dependency, whatever the latency dashboards say. For features where the model's output is critical, you may go further and fail the request loudly rather than score it on a fabricated input. ## What a good answer contains Name the fill as a fitted, frozen artifact; refuse to recompute it from live data and say why; insist that the serving transformation reproduce training exactly, including the missingness flag; acknowledge that single-row serving constrains which imputation strategies are even available; and finish with the monitoring, because the fill is the one transformation in the pipeline that turns a failure into a plausible-looking number.
- Why can't you fill with the median of the current scoring batch?Because a request's score would then depend on which other requests arrived with it — the same input retried later gets a different answer. It also lets the fill wander away from the level the model learned around, and it hides population drift instead of exposing it. Batch composition is not a property of the customer you are scoring.
- How would you catch a serving path that fills differently from training?Build features with the same code on both paths and ship the fitted fill values inside the artifact rather than copying constants into the service. Then add a regression test that scores a fixed row with a null field and asserts the exact prediction the training pipeline produced. Silent divergence here throws no exception, so only an assertion catches it.
- The upstream device registry starts timing out — what do you see?Nothing, in the obvious places. Every request still succeeds and every prediction still looks reasonable, because nulls are quietly converted into the stored median. The visible signal is the imputation rate for that feature jumping from a few percent to most requests, which is why it should be emitted as a metric and alerted on independently of latency and error rates.
- Does a k-nearest-neighbour fill work at serving time?Only if you ship reference rows with the model and query them per request. That turns a cheap constant lookup into a search with real latency, memory and staleness costs, and it means the deployed artifact now contains training records. It is doable, but it is a deliberate architectural choice rather than a free upgrade over a stored median.
saying these in an interview costs you the question
- Recomputes the fill from live traffic or the current batch
- Treats the fill as a constant hardcoded in the serving code
- Assumes training and serving fill the same way without a test
- Cannot distinguish a genuinely absent value from a failed lookup
- Never monitors how often a feature is being imputed