A checkout wait-time model scores 2.1 minutes offline but 5.4 live with no errors - how do you prove skew caused it?
answer
- compare, do not reason about it
- log what was served, not the request
- recompute for the same entity and instant
- diff per field, not per prediction
- a tolerance and a skew rate
basics
~20 sLog the exact input vector the model was scored on with the entity key, the serving timestamp and the definition version, then recompute each field from history for that same entity and instant and diff them field by field.
solid answer
~50 sReasoning about the two implementations does not settle it; only a comparison does. Instrument the serving path to log the assembled vector - not the request payload - stamped with the entity key, the time the estimate was served, and the feature-definition and model versions. Then replay a sample: for each logged record recompute every field as the nightly definition would have for that kitchen at that instant, and diff. Test against a per-feature tolerance rather than exact equality, because floating-point results and sub-second timing differ harmlessly. Report a **skew rate** per feature - the share of requests outside tolerance - and the signed magnitude, so the shape of the diff names the source: a constant offset means the implementations differ, an all-or-nothing gap on a cohort means the defaults do, a gap that tracks load means a stale read.
code
json · 13 lines{
"request_id": "r-91f3c2",
"served_at": "2026-05-14T19:31:07.412Z",
"entity": { "kitchen_id": "k-8842", "order_id": "o-77310" },
"feature_definition_version": "prep-v4",
"model_version": "wait-v12",
"features": {
"kitchen_recent_prep": 7.0,
"orders_ahead": 3,
"hour_of_week": 115
},
"prediction_minutes": 9.4
}go deeper
The takeaway is that a gap between an offline score and a live one is evidence of a question, not an answer. Proving the cause needs a record of the numbers the model actually received.
Be able to describe the loop: log the served vector with keys, timestamp and versions, recompute the same fields from history, diff within a tolerance, and report a rate per feature.
Show the diagnosis, not just the tooling: read the shape of the diff to name the source, check the sign against the observed symptom, and say plainly that the audit shows disagreement without saying which side is wrong.
Decide how much of this is a platform obligation. Per-request input logging costs write volume, retention and a privacy review; not having it costs every team a class of failure that only customers can detect.
## The only evidence is a comparison A 2.1-minute held-out error and a 5.4-minute live error are two measurements of two different pairings: the model against features the nightly job produced, and the model against features the request path produced. Nothing about either number identifies which field diverged, and no amount of reading both implementations settles it, because the disagreement is usually in an edge case neither author had in mind. The instrument that closes the question is a per-request comparison between what was actually served and what the definition says should have been served. ## Step one - log the vector that was scored Log at the point where the vector is handed to the model, after every fallback and cache has had its say. Logging the request payload instead is the classic mistake: rebuilding the vector from the payload re-runs the very code under suspicion, so the reconstruction agrees with itself by construction. The record needs five things: - the **entity keys** the features were fetched for; - the **serving timestamp**, precise enough to recompute against; - the **feature values exactly as scored**, at full precision; - the **feature-definition version**, so a replay months later recomputes with the definition that was live rather than today's; - the **model version**, so the diff can be attributed to a deployment. Sampling is fine. A few per cent of requests, plus every request for a cohort you suspect, gives a stable rate without doubling the write volume of the serving tier. ## Step two - recompute and diff For each logged record, run the nightly definition over history for that entity as of the logged serving timestamp, and compare field by field. The unit of comparison is the field, not the prediction: a prediction diff tells you something is wrong, a field diff tells you what. ## Step three - a tolerance, not equality Exact equality is the wrong test and produces an alarm on nearly every record. Two harmless sources of difference exist even when both paths are correct: floating-point accumulation differs with summation order, and the replay's view of history can include an order that landed a few hundred milliseconds after the estimate was served. Set a per-feature tolerance from the resolution the model can actually distinguish - roughly the narrowest split gap it learnt on that feature - and treat anything beyond it as skew. Then publish two numbers per feature: the **skew rate** (share of sampled requests outside tolerance) and the **signed median magnitude** (how far, and in which direction). ## Reading the shape of the diff | what the diff looks like | which source it is | |---|---| | a near-constant offset on almost every request | two implementations differing in rounding or units | | exact agreement on most rows, a large gap on a small cohort | a missing-entity default that disagrees between paths | | agreement at night, growing gaps during peaks | a stale read - the online row is older than the request | | a step change dated to a deploy | one path's definition changed and the other did not | The signed direction is as informative as the size. A served value consistently below the recomputation on `kitchen_recent_prep` explains an estimate that is consistently too short, which is exactly the live symptom to look for. ## What the audit does not tell you - **It does not say which side is wrong.** A diff means the two disagree. The nightly definition may be the one that changed, and the request path may be right. - **It is not a comparison of distributions.** Two paths can produce identically distributed values over a day and still disagree on every individual request, so a marginal comparison can look clean while the per-row skew rate is high. The pairing is what carries the information. - **It does not cover a field the request path never had.** If the serving path cannot fetch an input inside its budget, the replay will show a permanent gap that no code change closes; that is a design decision to make, not a bug to fix. ## Turning the investigation into a standing check The same instrument that answers the question once is the monitor that keeps answering it. Run the replay continuously on a sample, alarm on the per-feature skew rate crossing its tolerance, and treat a new feature as unlaunched until its rate has been observed for a few days. That converts a silent, outcome-delayed failure into an alert that fires within a sampling window, which is the only real defence when the model itself will never raise.
- What does a constant offset in every diff tell you, versus an all-or-nothing one?A near-constant offset on almost every request points at the two implementations - a rounding or unit difference applied uniformly. An all-or-nothing gap confined to a cohort points at the missing-entity default, because it fires only where one path takes its fallback branch. A gap that grows with traffic points at a stale read rather than at code.
- How do you choose the tolerance for a feature?From the resolution the model can distinguish - roughly the narrowest split gap learnt on that feature - and from the harmless measurement noise between serving and replay. State it per feature, publish it beside the definition, and alarm on the share of requests outside it rather than on any single record.
- Why stamp the record with the feature-definition version?So a replay run weeks later recomputes with the definition that actually served the request. Without it, a legitimate definition change looks like a fleet-wide skew event, and a real skew that predates the change becomes impossible to reproduce. The version also lets the diff be attributed to a specific deploy.
saying these in an interview costs you the question
- Reruns the offline evaluation to prove the model itself is fine
- Logs the request payload instead of the assembled feature vector
- Tests served and recomputed values for exact equality
- Compares aggregate feature distributions instead of per-request pairs
- Assumes any diff found by replay must be the serving path's bug
- Waits for outcome labels before investigating a suspected skew