In a checkout wait-time model, a new kitchen gets the imputed training median offline but a zero from the serving fallback - what breaks?
answer
- the missing case, decided twice
- zero is a value, not an absence
- below every split the model learnt
- new kitchens quoted the fastest wait
- put the default inside the definition
basics
~20 sThe model meets a value its training column never held. Zero sits below every learnt prep-time split, so brand-new kitchens are scored like the fastest ones and the estimate comes out far too short for the cohort least able to absorb it.
solid answer
~40 sOffline, the definition filled the empty average with the training set's median of 12 minutes, so the model learnt splits over a column whose smallest values were realistic prep times. At serving, the fallback writes `0`. Zero is below everything the model ever saw, so the row extrapolates off the bottom of the learnt range and the customer is quoted something like 4 minutes against a 20-minute reality. The damage is concentrated: it hits exactly the kitchens being onboarded, and nothing errors. The fix is to stop treating the default as each path's private business - put it inside the transform artifact both paths execute, and better still give the model an explicit no-history indicator so the missing case has learnt behaviour instead of a smuggled sentinel.
go deeper
Remember that a missing feature still has to become a number, and that the number chosen at serving may differ from the one chosen during training. Zero is a decision, not the absence of one.
Explain why an unseen sentinel is worse than a plausible constant: the model extrapolates below every learnt split, so the estimate moves in a predictable direction rather than becoming random noise.
Show that you would enumerate the missing case for every aggregate in the served vector, place the rule inside the shared definition, and assert it in a replay audit rather than trusting a code review to keep two fallbacks aligned.
The judgment call is whether missingness is modelled or imputed platform-wide. Imputation keeps definitions simple and hides a state from the model; an explicit indicator is more honest and costs a column, a contract and a retrain on every feature you adopt it for.
## Two paths, two answers to the same question Every aggregate feature has a missing case. `kitchen_recent_prep` is the mean preparation time of the last twenty finished orders, and a kitchen that opened this morning has none. The nightly training job resolved that case deliberately: fill it with the training set's median, 12 minutes, so that the column contains a plausible number and the model learns splits over realistic values. The request path resolved it by accident: an empty average became `0`, which is what an empty sum over an empty count tends to produce when nobody makes a decision. Neither path errors. Both are defensible in isolation. Together they are a skew, and it is the most concentrated kind: it fires for a single cohort, and it fires hard. ## What zero does at scoring time The model never saw `0` in that column. Its lowest learnt split on `kitchen_recent_prep` sits somewhere among genuine fast kitchens, so a zero routes below all of them: | path | value supplied | model's reading | estimate shown | reality | |---|---|---|---|---| | training | 12 (imputed median) | an average kitchen | calibrated | matches | | serving | 0 (fallback) | faster than the fastest kitchen seen | far too short | a long wait for a new kitchen | The direction matters and is worth stating explicitly in an interview: fewer minutes in means a shorter estimate out. A new kitchen, which is usually *slower* than average while its staff learn the flow, is quoted the shortest wait in the catalogue. Other features shift the overall level, but this feature's contribution is pinned at its low extreme on every affected request, so the bias has a consistent direction rather than being noise. ## Who it hits and how hard - **The cohort with the least slack.** Newly onboarded kitchens have no history precisely because they are new, and a badly wrong first estimate is a poor introduction for both the customer and the merchant. - **A small share of traffic, so aggregate dashboards hide it.** If new kitchens are two per cent of orders, an eight-minute error on them moves overall mean error by a fraction of a minute - visible in the 2.1-against-5.4 gap only as a contributor, not as a spike. - **Retraining does not heal it.** Tomorrow's training job imputes the median again, so the offline number stays good and the loop looks healthy. ## Why zero is a worse choice than a wrong constant Two separate problems ride on a zero sentinel: 1. **It is outside the learnt range**, so the model extrapolates rather than interpolating. A wrong-but-plausible constant at least lands in territory the model has learnt behaviour for. 2. **It is ambiguous.** For a feature where zero genuinely occurs - an open-order count, say - a sentinel zero is indistinguishable from a true zero, and no downstream audit can separate the two. ## Where the default belongs The missing-value rule is part of the feature definition, not part of either path's code. Two consequences follow: - Express it inside the transform artifact both paths execute, so the fallback cannot be reinvented by whoever writes the request path next. - Assert it in the replay audit: a recomputation of a no-history kitchen must produce the same value the serving path emitted, and a mismatch there is the cheapest possible detection. ## A better shape than a sentinel Stronger still is to stop smuggling the missing case into a numeric column. Add an explicit indicator - a `has_prep_history` flag - present in both the training rows and the served vector, and let the model learn what to do when it is false. The imputed number then stops carrying two meanings at once, and the missing case becomes a modelled state rather than a value the model must guess about. This only works if the indicator itself is produced by the shared definition; an indicator computed independently on each path is simply the same skew with an extra column. ## What to check before shipping For every aggregate feature in the served vector, answer three questions and write the answers down: what does the training job put here when the entity has no rows, what does the request path put here, and does the model see a flag telling it which case it is in. A design round that reaches this level of specificity has separated the candidate who has run one of these systems from the candidate who has read about them.
- Why is zero a worse fallback than a wrong non-zero constant?Two reasons. It sits outside the learnt range, so the model extrapolates instead of interpolating, and the error is larger than a plausible constant would produce. It is also ambiguous for any feature where zero genuinely occurs, such as an open-order count, so no later audit can tell a sentinel apart from a real measurement.
- How do you make it impossible for the two defaults to disagree again?Move the rule into the transform artifact both paths execute, so there is one place where the missing case is decided, and add an assertion to the replay audit that a no-history entity recomputes to the same value the serving path emitted. Documentation and code review do not survive the next person writing a fallback under time pressure.
- Does adding a no-history indicator column remove the need to agree on the numeric default?No. The indicator gives the model learnt behaviour for the missing state, but the numeric column still carries some value on those rows, and if the two paths fill it differently the skew persists underneath the flag. The indicator helps only when it and the imputed value both come from the shared definition.
saying these in an interview costs you the question
- Says zero is a safe neutral fallback for a missing numeric feature
- Assumes the model treats an unseen sentinel as missing data
- Believes both paths inherited one default because neither team picked one
- Claims the error is negligible because only a few kitchens lack history
- Thinks a missing-value indicator fixes a mismatched sentinel on its own