A feature table matched each label row to the nearest sensor reading by stamp, and the model scored far better offline than live - what did the match do?
answer
- three directions, not one
- nearest points both ways
- the sign of the gap, not its size
- what was knowable at that stamp
- offline and live read different tables
basics
~20 sNearest takes whichever reading is closer, earlier or later. Every row where the later reading was closer carries a value that did not exist at the label's stamp, so the table contains information from the future. Only the latest-at-or-before direction cannot.
solid answer
~50 sAn inexact ordered match has a **direction**: it may look only earlier, only later, or take whichever is nearer. Nearest sounds the most accurate and is the one that leaks, because on every row where the later reading happens to be closer - roughly half of them for an evenly sampled feed - the attached value was recorded after the moment the row describes. Nothing errors, the row count is right, and the column looks plausible. Offline the model reads a value it could never have at serving time, so its score is inflated; live, that value is simply not there yet. The repair is the earlier-only direction - the latest reading at or before each stamp - plus a check you can run on any output: carry the matched record's own stamp into the result and assert that driving stamp minus matched stamp is never negative.
go deeper
Recall that matching by stamp can look earlier, later, or at whichever is nearer, and that only the earlier-only choice guarantees the attached value already existed at the moment being described.
Explain why nearest silently behaves as later-only on a large share of rows, and why nothing in the output - row count, types, ranges, absence count - reveals it.
Show the diagnosis and the assertion: carry the matched stamp through, check the signed gap is never negative, and accept the extra absences the earlier-only direction produces at the start of the span.
The standing rule to argue for is that any table built by matching feeds ships with the matched stamp and a signed-gap assertion, so that no future author can reintroduce this by choosing a direction that reads as more accurate.
## Three directions, and only one of them is safe An inexact ordered match - each record paired with a record on another feed by stamp rather than by an equal key, an *as-of match* - always makes a choice about which way it may look. Designs that offer the operation expose all three; designs without it force you to build one of the three by hand. They are not interchangeable. | direction | which record it attaches | can it use a record that did not exist yet? | |---|---|---| | latest at or before | the most recent record whose stamp is not later than the driving stamp | no - by construction | | next at or after | the first record whose stamp is not earlier | yes, always | | nearest | whichever of those two is closer in time | yes, on every row where the later one is closer | The third row is the one that ships to production. Nearest is chosen because it reads as the most accurate: it minimises the time difference, and minimising an error sounds obviously right. It does minimise the gap. It also silently becomes the second row on a large share of the data. ## Why nearest is the trap For a feed sampled at roughly even intervals, a randomly placed driving stamp falls in the second half of its interval about as often as the first, so the later reading is closer about as often as the earlier one. That means a large fraction of the rows in the feature table carry a measurement taken **after** the event they are supposed to describe. The mix is not even visible by inspection - the values are all plausible sensor readings, in range, of the right type. An important detail: this is not a mistake about *time zones*, *rounding* or *sorting*. The match did exactly what it was asked. The defect is in the requirement: a row that describes a moment may only carry values that existed at that moment. ## Why it does not show up until production The measurement loop conceals it perfectly: - The row count is unchanged, so the cheap structural assertion passes. - Absent values do not increase - nearest actually produces fewer absences than earlier-only, because it can match at the start of the span from a later record. - Every offline evaluation reads the same table, so the inflated signal is present in the training rows and the evaluation rows alike, and the two agree with each other. - Only at serving time does the environment change: at the moment a prediction is required, the later reading has not arrived. The feature computed live is a different feature from the one the model was fitted on, and the score falls. A model that is much better offline than live has a short list of causes, and a match that could see later records is high on it. ## The check that catches it This is cheap and you should run it on any table built by matching two feeds: 1. Carry the **matched record's own stamp** into the output, not just its values. 2. Compute the signed gap: driving stamp minus matched stamp. 3. Assert the minimum of that column is not negative. A single negative row means something later was used. 4. While you are there, look at the maximum too - that is the staleness question, and the same column answers both. The sign is the point. Checking only the magnitude of the gap - "all matches are within 30 seconds" - passes happily on a table built with nearest, because the leaking rows are the *closest* ones. ## The repair, and the edge case worth naming Switch the direction to latest at or before and re-run. Expect two consequences and be ready to defend them: rows at the very start of the span now have no match, because nothing earlier exists; and the average gap grows, because you gave up the closer half of the matches. That second consequence is the point, not a regression - the closer half was the unusable half. The edge case is an exactly equal stamp. Whether a reading stamped at precisely the driving stamp counts as knowable depends on what the stamps mean: an event time and a measurement published at the same instant may or may not have been available to a consumer then. Some designs let you include or exclude exact equality explicitly. Decide it once, write it down, and do not let an unstated default decide it - and never assume an unnamed direction means the safe one.
- Why does a check that all matches are within 30 seconds fail to catch this?Because it tests magnitude, not direction. The rows that used a later reading are the ones with the smallest gaps, so they sit comfortably inside any magnitude threshold. The assertion has to be on the signed difference: driving stamp minus matched stamp must never be negative.
- After switching to the earlier-only direction, more rows have absent features. Is that a regression?No - it is the honest state. Rows at the start of the span have nothing earlier to match, and rows whose feed had a hole now show it. You have traded invented values for visible absence, which is the trade you wanted.
- Should a reading stamped at exactly the driving record's stamp be matched?It depends what the stamps mean - whether a value published at that instant was actually available to a consumer then. Some designs let you include or exclude exact equality. Decide it deliberately from the feed's semantics and record the decision; do not inherit it from an unstated default.
Asked what the speed limit is, a driver answers from the last sign passed, not from the nearest sign - which may be the one just ahead, not yet read. The nearest sign is closer and is exactly the one they are not allowed to use.
saying these in an interview costs you the question
- Says nearest is safest because it minimises the time difference
- Believes leakage would show up as a changed row count
- Assumes an unnamed match direction means earlier-only
- Treats a value recorded one second later as knowable at the stamp
- Checks the size of the gap and never its sign
- Believes the offline score was real and production is the anomaly