A checkout wait-time model gets a feature computed differently in the serving path than in training - why is there no error?
answer
- nothing checks what the number means
- wrong value, right shape
- models saturate, they do not throw
- the outcome arrives twenty minutes later
- 2.1 minutes offline, 5.4 live
basics
~20 sA model validates nothing about what its inputs mean: any number in the right slot is scored. Training-serving skew therefore lands as a confidently wrong estimate rather than an exception, and shows up only when real outcomes arrive.
solid answer
~40 sThe serving path assembles a vector and hands it to the model; nothing in that sequence asserts units, window length or missing-value handling. A value the nightly training job produced in whole minutes, recomputed on the request path in seconds, is simply a larger number in the same slot, so the call succeeds and the latency budget is met. A type-and-range check passes too, because `420` is a valid non-negative number. The model degrades instead of failing: values far outside the trained range fall past every learnt split and land in one extreme leaf, and a sentinel the column never contained is extrapolated rather than recognised as missing. The evidence is a gap between two error measurements - 2.1 minutes offline against 5.4 minutes measured on finished orders - not an alarm.
go deeper
Remember the shape of the failure: the prediction endpoint stays healthy while the answer is wrong. A model scores whatever number it is handed, so a feature built differently at serving produces bad output with no exception anywhere.
Explain why each guard fails - a schema pins types not meaning, a name pins a slot not a unit, an offline score pins the training pairing not the serving one - and describe how a model degrades rather than raising when a value leaves the trained range.
Show that you would measure the live error against real outcomes and compare it with the held-out figure as a standing practice, because that gap is the only early signal a healthy-looking serving tier ever gives you.
The tradeoff is where the correctness contract lives. Pushing it into the platform costs a shared artifact and per-request logging; leaving it with each team costs a class of silent product harm that no dashboard you own today would surface.
## The failure has no error path A deployed model is a function from a vector of numbers to a number. At checkout the serving path assembles that vector - the kitchen's recent prep time, the count of orders ahead of this one, the hour of the week - hands it to the model and returns an estimate. Nothing in that sequence asserts what the numbers *mean*. A feature the nightly training job produced in whole minutes, recomputed on the request path in seconds, is to the model just a larger number in the same slot. The call succeeds, the latency budget is met, every health check is green, and the customer is shown a wrong wait. That is what separates **training-serving skew** - one feature definition computed differently on the two paths - from an ordinary outage. There is no exception, no status code to alarm on and no unusual log line. The only witness is the outcome, and for a wait-time estimate the outcome arrives twenty minutes later, when the order is actually handed over. ## Why nothing upstream objects Three checks that engineers expect to catch this do not: - **Type and range validation** at the serving boundary checks that the field is a number, and perhaps that it is non-negative. `420` seconds is a perfectly good non-negative number; so is `0` for a kitchen with no history. Neither violates a schema. - **Name matching** between the offline definition and the online one proves that two things are called `kitchen_recent_prep`. It proves nothing about units, rounding, window length, time zone or missing-value handling. - **A green offline metric.** The held-out score measures the model against features the *training job* produced. It cannot observe the serving path at all, so it certifies a pairing that never occurs in production. ## What the model does instead of failing Models degrade in different ways, and none of them raise: | input the model receives | what happens at serving | |---|---| | a value 60x larger than anything in training | it falls past every learnt split, so every row served by that branch lands in one extreme leaf and gets a nearly constant estimate | | a sentinel such as `0` that the training column never held | it extrapolates off the bottom of the learnt range, so a brand-new kitchen is scored like the fastest one | | a value half a minute off, on every request | a small systematic bias that looks unremarkable in any single prediction | The magnitudes differ; the visibility does not. All three return a confident number. ## How it surfaces It surfaces as a gap between two measurements, not as an incident. A wait-time model that scores a mean absolute error of 2.1 minutes on a held-out day can measure 5.4 minutes against finished orders once it is live, with a flat error rate on the endpoint and unchanged latency. Because the two numbers usually live on different dashboards, weeks can pass before anyone lines them up. The consequences compound quietly: 1. **Product harm lands first.** Customers act on the estimate - they wait, they cancel, they complain - long before anyone notices the gap. 2. **The retraining loop confirms the wrong thing.** The next nightly run trains on the same offline features and reproduces the same good offline score, so the loop looks healthy. 3. **Triage goes to the wrong component.** Teams hunt for a model problem - a bad checkpoint, stale weights, a data change - when the model is fine and the input assembly is not. ## The line against two neighbouring failures - **This is not leakage.** A transform fitted on the whole dataset before the split is a training-time correctness bug: it inflates the offline number, and the honest offline number would have been worse. Skew has the opposite shape - the transform is fitted correctly, the offline number is honest, and the *execution at serving* differs. - **This is not the live input distribution moving.** Skew is a disagreement between two computations of the same quantity at the same instant; it is present on day one, before anything in the world changes. A model whose inputs shift away from training over weeks is a separate subject with separate instrumentation. ## What to say in a design round Name the missing property: nothing asserts that the vector scored in production is the vector the training job would have produced for that entity at that moment. Every mechanism this subject covers - one shared transform artifact, the logged served input vector, a replay audit, a skew rate with a stated tolerance - exists to turn that missing contract into something a system can check, because the model itself never will.
- How is this different from a transform fitted on the whole dataset before the train-test split?That is leakage: a training-time bug that inflates the offline score, so the honest offline number would already have been worse. Skew is the mirror image - the transform is fitted correctly on the training fold and the offline score is honest, but the serving path executes a different computation, so the gap opens only in production.
- What does a gradient-boosted tree model do with a value 60 times larger than anything it saw in training?It routes past the topmost split learnt for that feature, so every such row reaches the same extreme leaf and receives a near-constant contribution. There is no error and no null - the prediction simply stops responding to that feature, which is why the symptom is a flat, confidently wrong estimate for the affected rows.
- Would validating the input vector against the training data's schema have caught it?Not by itself. A schema pins types and often a range, and a mis-scaled or defaulted value usually satisfies both. Catching it requires comparing the served value against what the training definition would have produced for the same entity at the same instant - a comparison, not a schema.
A kitchen that plates every order from yesterday's prep list serves a full plate on time and never rings an alarm. The food is wrong, and only the diner finds out.
saying these in an interview costs you the question
- Assumes a type-and-range check on the input vector catches a unit mismatch
- Claims the model returns an error or a null for an unseen value
- Believes a green health check on the serving path means the features are right
- Thinks matching feature names across the two paths guarantees matching values
- Says a good held-out score proves the serving path assembles inputs correctly