A match to the latest earlier reading always succeeds while any earlier reading exists, so what is a staleness tolerance for?
answer
- the match cannot fail
- success is not the same as current
- the age of the matched value
- an honest absence instead
- keep the matched record's stamp
basics
~20 sA staleness tolerance caps how old a matched value may be. Beyond that age the match is abandoned and an absent value is produced instead, so a reading from hours ago is never silently presented as the value in effect now.
solid answer
~50 sMatching each record to the latest reading at or before its stamp - an *as-of match*, matched as of that instant - has no notion of failure: as long as one earlier reading exists anywhere in the feed, every record gets a match, even if that reading is three days old because the sensor died. A **staleness tolerance** is the cap on that age. Where the gap between the driving stamp and the matched stamp exceeds it, the match is abandoned and the matched columns come back absent. That turns a silently wrong number into a visibly missing one, which is the only version a downstream check can catch. Some designs take the cap as an argument on the match; others express it as a limit on how far a carried value may travel; where neither exists, keep the matched record's own stamp in the output and invalidate the row afterwards.
go deeper
Recall that matching to the latest earlier reading always succeeds once any earlier reading exists, and that a cap on the matched value's age is what turns a very old match into an absent one instead.
Explain the age arithmetic - driving stamp minus matched stamp - and that exceeding the cap blanks the matched columns without removing the driving row, so the result's row count is unaffected.
Demonstrate that you set the cap from the feed's measured arrival spacing and the volatility of the quantity, and that you treat the count of abandoned matches as a monitoring signal on the upstream feed.
The call to own is where the cap lives: as an argument to whichever tool ran the match, or as a stamp-and-age column in your own output, which costs a little width and buys portability, auditability and tuning without a re-run.
## A match that cannot fail is not a match that is right An inexact ordered match - each record paired with the latest record on the other feed at or before its stamp - has no failure mode once the feeds overlap at all. The first driving records may have nothing earlier to match, and after that every single record matches something. The operation cannot tell the difference between a price recorded 40 milliseconds ago and the last price anyone published before the venue went down on Friday. Both are "the latest value at or before this stamp", and both are attached with identical confidence. That is the gap a **staleness tolerance** closes. It is a cap on the **age of the match**: the difference between the driving record's stamp and the stamp of the record that was attached to it. Where that age exceeds the cap, the match is thrown away and the matched columns come back absent. ## What the cap actually buys It does not improve any match. It converts one class of wrong answers into a class of honest absences: | feed situation | with no age cap | with one | |---|---|---| | feed is healthy, sub-second gaps | correct value attached | identical result | | feed stopped two hours ago | a two-hour-old value attached and treated as current | matched columns absent from the moment the cap is passed | | feed has a backfill hole in the middle of the day | the last value before the hole is stretched silently across it | the hole appears as absent rows, with a count you can alert on | | driving records before the feed's first record | already absent | already absent, unchanged | Note what the cap does **not** do: it does not drop the driving record. The output still carries one row per driving record - only the matched columns become absent. A candidate who expects the row count to shrink has confused the cap with a filter on the driving side. ## Choosing the number The cap is a statement about the quantity being matched, not a round number picked for comfort. Two inputs decide it: 1. **The feed's own sampling rate.** If readings normally arrive every five seconds, a matched value 90 seconds old means roughly eighteen readings went missing; the cap belongs somewhere above normal jitter and well below that. 2. **How fast the quantity moves.** A price can be meaningfully wrong within a second. An ambient-temperature reading is usually fine at ten minutes. The same feed spacing justifies very different caps depending on what the number is used for. A cap set so wide it can never fire is decoration, and the usual sign of one is that nobody can say what the unmatched count was last week. ## Where no such argument exists This differs by design, and it is worth saying so rather than assuming. Some tools take the cap directly on the match. Some express only a limit on how far a value may be carried forward into positions where nothing was measured. Some offer neither. The portable construction works everywhere and is worth preferring even where the argument exists: - Carry the **matched record's own stamp** into the output alongside its values. - Compute the signed age: driving stamp minus matched stamp. - Blank the matched columns where that age exceeds the cap - and keep the age column, because it is the evidence. That construction has a second benefit: the cap becomes a parameter of your pipeline rather than a parameter of whichever tool you used, so it survives being ported and can be tuned without re-running the match. ## The unmatched count is a monitoring signal Once the cap exists, the number of driving records whose match was abandoned becomes a measurement of the upstream feed, available for free from a job you were running anyway. A step change in that count is a feed that slowed, stopped or started arriving late - visible before anyone notices the numbers downstream look odd. Discarding the absent rows without counting them throws that signal away and leaves you with a table that looks complete. What you then *do* with an absent matched value - leave it, fill it, drop the row - is a separate decision made downstream on its own merits, and it should be made deliberately rather than by whatever the match produced.
- Does exceeding the staleness tolerance remove the driving record from the result?No. The output keeps one row per driving record; only the matched columns become absent. The cap governs the match, not the driving side. If your row count changed after adding a cap, something else changed too.
- How would you pick the cap for a sensor that reports every five seconds?Start from the normal arrival spacing plus realistic jitter, then check it against how fast the measured quantity can move and how wrong a stale reading would make the downstream number. Then measure: plot the distribution of match ages over a week of real data and see where the tail begins.
- Your tool offers no age cap on the match at all. What do you do?Carry the matched record's stamp into the output, compute driving stamp minus matched stamp, and blank the matched values where that exceeds your cap. Keep the age column - it is both the evidence and the input to tuning the cap later.
saying these in an interview costs you the question
- Thinks a successful match means the value was current
- Picks a round number without looking at the feed's arrival spacing
- Believes an absent result means the match is broken rather than honest
- Sets the cap so wide it can never fire
- Expects the cap to drop rows from the driving side
- Assumes every design offers an age cap as an argument