skip to content

A checkout wait-time model gets a feature computed differently in the serving path than in training - why is there no error?

level: juniorimportance: must knowfreq 72%

answer

  1. nothing checks what the number means
  2. wrong value, right shape
  3. models saturate, they do not throw
  4. the outcome arrives twenty minutes later
  5. 2.1 minutes offline, 5.4 live

basics

~20 s

A model validates nothing about what its inputs mean: any number in the right slot is scored. Training-serving skew therefore lands as a confidently wrong estimate rather than an exception, and shows up only when real outcomes arrive.

solid answer

~40 s

The serving path assembles a vector and hands it to the model; nothing in that sequence asserts units, window length or missing-value handling. A value the nightly training job produced in whole minutes, recomputed on the request path in seconds, is simply a larger number in the same slot, so the call succeeds and the latency budget is met. A type-and-range check passes too, because `420` is a valid non-negative number. The model degrades instead of failing: values far outside the trained range fall past every learnt split and land in one extreme leaf, and a sentinel the column never contained is extrapolated rather than recognised as missing. The evidence is a gap between two error measurements - 2.1 minutes offline against 5.4 minutes measured on finished orders - not an alarm.

go deeper

for a junior

Remember the shape of the failure: the prediction endpoint stays healthy while the answer is wrong. A model scores whatever number it is handed, so a feature built differently at serving produces bad output with no exception anywhere.

for a middle

Explain why each guard fails - a schema pins types not meaning, a name pins a slot not a unit, an offline score pins the training pairing not the serving one - and describe how a model degrades rather than raising when a value leaves the trained range.

for a senior

Show that you would measure the live error against real outcomes and compare it with the held-out figure as a standing practice, because that gap is the only early signal a healthy-looking serving tier ever gives you.

for a principal

The tradeoff is where the correctness contract lives. Pushing it into the platform costs a shared artifact and per-request logging; leaving it with each team costs a class of silent product harm that no dashboard you own today would surface.

## The failure has no error path A deployed model is a function from a vector of numbers to a number. At checkout the serving path assembles that vector - the kitchen's recent prep time, the count of orders ahead of this one, the hour of the week - hands it to the model and returns an estimate. Nothing in that sequence asserts what the numbers *mean*. A feature the nightly training job produced in whole minutes, recomputed on the request path in seconds, is to the model just a larger number in the same slot. The call succeeds, the latency budget is met, every health check is green, and the customer is shown a wrong wait. That is what separates **training-serving skew** - one feature definition computed differently on the two paths - from an ordinary outage. There is no exception, no status code to alarm on and no unusual log line. The only witness is the outcome, and for a wait-time estimate the outcome arrives twenty minutes later, when the order is actually handed over. ## Why nothing upstream objects Three checks that engineers expect to catch this do not: - **Type and range validation** at the serving boundary checks that the field is a number, and perhaps that it is non-negative. `420` seconds is a perfectly good non-negative number; so is `0` for a kitchen with no history. Neither violates a schema. - **Name matching** between the offline definition and the online one proves that two things are called `kitchen_recent_prep`. It proves nothing about units, rounding, window length, time zone or missing-value handling. - **A green offline metric.** The held-out score measures the model against features the *training job* produced. It cannot observe the serving path at all, so it certifies a pairing that never occurs in production. ## What the model does instead of failing Models degrade in different ways, and none of them raise: | input the model receives | what happens at serving | |---|---| | a value 60x larger than anything in training | it falls past every learnt split, so every row served by that branch lands in one extreme leaf and gets a nearly constant estimate | | a sentinel such as `0` that the training column never held | it extrapolates off the bottom of the learnt range, so a brand-new kitchen is scored like the fastest one | | a value half a minute off, on every request | a small systematic bias that looks unremarkable in any single prediction | The magnitudes differ; the visibility does not. All three return a confident number. ## How it surfaces It surfaces as a gap between two measurements, not as an incident. A wait-time model that scores a mean absolute error of 2.1 minutes on a held-out day can measure 5.4 minutes against finished orders once it is live, with a flat error rate on the endpoint and unchanged latency. Because the two numbers usually live on different dashboards, weeks can pass before anyone lines them up. The consequences compound quietly: 1. **Product harm lands first.** Customers act on the estimate - they wait, they cancel, they complain - long before anyone notices the gap. 2. **The retraining loop confirms the wrong thing.** The next nightly run trains on the same offline features and reproduces the same good offline score, so the loop looks healthy. 3. **Triage goes to the wrong component.** Teams hunt for a model problem - a bad checkpoint, stale weights, a data change - when the model is fine and the input assembly is not. ## The line against two neighbouring failures - **This is not leakage.** A transform fitted on the whole dataset before the split is a training-time correctness bug: it inflates the offline number, and the honest offline number would have been worse. Skew has the opposite shape - the transform is fitted correctly, the offline number is honest, and the *execution at serving* differs. - **This is not the live input distribution moving.** Skew is a disagreement between two computations of the same quantity at the same instant; it is present on day one, before anything in the world changes. A model whose inputs shift away from training over weeks is a separate subject with separate instrumentation. ## What to say in a design round Name the missing property: nothing asserts that the vector scored in production is the vector the training job would have produced for that entity at that moment. Every mechanism this subject covers - one shared transform artifact, the logged served input vector, a replay audit, a skew rate with a stated tolerance - exists to turn that missing contract into something a system can check, because the model itself never will.

  • How is this different from a transform fitted on the whole dataset before the train-test split?
    That is leakage: a training-time bug that inflates the offline score, so the honest offline number would already have been worse. Skew is the mirror image - the transform is fitted correctly on the training fold and the offline score is honest, but the serving path executes a different computation, so the gap opens only in production.
  • What does a gradient-boosted tree model do with a value 60 times larger than anything it saw in training?
    It routes past the topmost split learnt for that feature, so every such row reaches the same extreme leaf and receives a near-constant contribution. There is no error and no null - the prediction simply stops responding to that feature, which is why the symptom is a flat, confidently wrong estimate for the affected rows.
  • Would validating the input vector against the training data's schema have caught it?
    Not by itself. A schema pins types and often a range, and a mis-scaled or defaulted value usually satisfies both. Catching it requires comparing the served value against what the training definition would have produced for the same entity at the same instant - a comparison, not a schema.

A kitchen that plates every order from yesterday's prep list serves a full plate on time and never rings an alarm. The food is wrong, and only the diner finds out.

saying these in an interview costs you the question

  • Assumes a type-and-range check on the input vector catches a unit mismatch
  • Claims the model returns an error or a null for an unseen value
  • Believes a green health check on the serving path means the features are right
  • Thinks matching feature names across the two paths guarantees matching values
  • Says a good held-out score proves the serving path assembles inputs correctly